A network and system for athlete action recognition based on motion and timing enhancement

By designing an athlete's movement recognition network based on sports and timing enhancement, using the sports feature excitation module and timing enhancement module to extract the sports information and long-term characteristics of players in sports videos, the problem that existing models cannot fully extract sports information and perform long-term modeling is solved, and the accurate identification of player's movements in sports videos is achieved.

CN116092197BActive Publication Date: 2025-05-16CHONGQING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310175366.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-28
Publication Date
2025-05-16
Estimated Expiration
2043-02-28

AI Technical Summary

Technical Problem

Existing models cannot directly extract motion information from RGB frame sequences, making it difficult to perform long-term modeling.

Method used

A sportsman's action recognition network based on motion and timing enhancement is designed. Through the sequentially connected convolutional layer and fully connected layer, the motion feature excitation module and timing enhancement module are used to extract local motion features and global timing information, and perform spatiotemporal feature extraction and action recognition.

Benefits of technology

Effectively extract the players' sports information and long-term characteristics of the sports video, and can identify the players' movements in the sports video end-to-end, solving the problems of insufficient extraction of sports information and difficulty in extracting long-term characteristics.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116092197B_ABST
    Figure CN116092197B_ABST
Patent Text Reader

Abstract

The present invention relates to the fields of behavior recognition and computer vision, and discloses a network and system for athlete action recognition based on motion and timing enhancement. The network architecture includes a first convolutional layer, a second convolutional layer, a third convolutional layer, a fourth convolutional layer, a fifth convolutional layer and a fully connected layer connected in sequence; the first convolutional layer is used to perform convolution operations and maximum pooling operations on the input motion video frames; the second convolutional layer to the fifth convolutional layer are all used to perform dimensionality reduction processing, motion feature excitation, timing enhancement and spatiotemporal feature extraction on the received feature graphs; the fully connected layer is used to perform classification and recognition based on the feature information output by the fifth convolutional layer, and output the athlete action recognition results. The present invention can complete the action recognition task of players in sports basketball videos end-to-end, and can solve the problems of insufficient extraction of motion information of players in sports videos and difficulty in extracting long time series features.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of behavior recognition and computer vision, and in particular to an athlete action recognition network and system based on motion and timing enhancement. Background Art

[0002] In recent years, computer vision has been increasingly widely used in the intelligent analysis of sports videos. In multi-player collaborative sports events, automatic and accurate recognition of the movements of each athlete on the field can not only liberate coaches and enable them to focus on formulating techniques and tactics, but also provide them with technical movement details and guidance that ordinary people cannot notice, further providing a basis for analyzing game tactics and strategies. Therefore, the application of deep learning and computer vision technology to accurately recognize the movements of athletes in sports videos has broad and important practical significance. However, in sports videos with similar backgrounds, accurately identifying the movements of athletes is still a challenging task.

[0003] In sports videos, similar background information cannot provide effective features and is not very helpful for action recognition models. However, dynamic motion information in videos is particularly critical for action recognition models. At present, the commonly used method for extracting motion information is the optical flow method. However, calculating optical flow will bring a lot of computational overhead and cannot meet real-time needs. In addition to extracting motion information in videos, the correlation between different time series frames is also very important. However, it is difficult to extract long time series features in videos.

[0004] In view of this, this application is hereby filed. Summary of the invention

[0005] The purpose of the present invention is to provide an athlete action recognition network and system based on motion and timing enhancement to solve the problem that the existing model cannot fully extract motion information directly from RGB frame sequences and perform long-term temporal modeling.

[0006] The present invention is achieved through the following technical solutions:

[0007] On the one hand, a sportsman action recognition network based on motion and timing enhancement is provided, comprising a first convolutional layer, a second convolutional layer, a third convolutional layer, a fourth convolutional layer, a fifth convolutional layer and a fully connected layer connected in sequence; the first convolutional layer is used to perform convolution operations and maximum pooling operations on the input motion video frame, and output a feature map to the second convolutional layer; the second convolutional layer to the fifth convolutional layer are all used to extract local motion features and explicitly enhance motion sensitive channels on the received feature map, and then globally enhance the key sequence of the explicitly enhanced feature map, and finally extract spatiotemporal features from the feature map with enhanced motion information and timing information;

[0008] The fully connected layer is used to perform classification and recognition based on the feature information output by the fifth convolutional layer, and output the athlete action recognition result.

[0009] Among them, the second convolutional layer to the fifth convolutional layer all include a plurality of convolutional blocks combined in the form of residuals, each of which is used to perform dimensionality reduction processing on the channels of the received feature maps using 1×1×1 three-dimensional convolution, and to perform feature extraction and explicit enhancement of motion-sensitive channels on the local mixed motion of the feature maps after dimensionality reduction processing by embedding motion feature excitation modules, and to perform global temporal enhancement on the key sequence of the explicitly enhanced feature maps by embedding timing enhancement modules, and to perform spatiotemporal feature extraction on the feature maps with enhanced motion information and timing information by using 3×3×3 three-dimensional convolution. .

[0010] Furthermore, the second convolution layer includes 3 convolution blocks, the third convolution layer includes 4 convolution blocks, the fourth convolution layer includes 6 convolution blocks, and the fifth convolution layer includes 3 convolution blocks.

[0011] Furthermore, the motion feature excitation module includes a local motion feature extraction unit, which is used to extract dynamic local motion features from the input feature map after 1×1×1 three-dimensional convolution; and a motion-sensitive channel excitation unit, which is used to explicitly enhance the channel that captures the local motion features to obtain a feature map after motion feature excitation.

[0012] Furthermore, the local motion feature extraction unit includes a first dimensional transformation subunit, which is used to perform dimensionality reduction processing on the channel of the feature map after the 1×1×1 three-dimensional convolution using a 1×1 two-dimensional convolution; a time series division subunit, which is used to divide the feature map after the 1×1 two-dimensional convolution according to the time series to obtain the feature map X of each frame R (t), t represents the timestamp of the feature map of each frame, t = 1, 2, ..., T, T represents the number of frames of the input feature map; the feature map transformation subunit is used to transform the feature map X R (t) Perform 3×3 channel-by-channel convolution to obtain the transformed feature map Y R (t); adjacent frame motion feature calculation subunit, used to calculate the feature map X R (t) is the motion feature of the adjacent frame, and the motion feature of the adjacent frame is the feature map X R (t) and the transformed feature map Y R (t+1); the cross-frame motion feature calculation subunit is used to calculate the feature map X R (t) is a cross-frame motion feature, wherein the cross-frame motion feature is a feature map X R (t) and the transformed feature map Y R (t+2); feature fusion subunit, used to calculate the feature map XR (t) The fused motion features at the corresponding moment, where the fused motion features are the sum of the adjacent frame motion features and the cross-frame motion features; a feature splicing subunit, used to splice the fused motion features corresponding to each moment in time sequence to obtain the local motion features of the input feature map.

[0013] Furthermore, the motion-sensitive channel excitation unit includes a second dimensional transformation subunit, which is used to interact and transform the channel dimension of the local motion feature using a 1×1 two-dimensional convolution to obtain the transformed local motion feature; a third dimensional transformation subunit, which is used to transform the channel dimension of the local motion feature into 1 using a 1×1 two-dimensional convolution to obtain the global spatial feature; a feature enhancement subunit, which is used to enhance the global spatial feature to obtain the enhanced global spatial feature; a channel descriptor acquisition subunit, which is used to perform a dot multiplication operation on the transformed local motion feature and the enhanced global spatial feature to obtain the channel descriptor of the local motion feature; a motion information weight acquisition subunit, which is used to restore the channel dimension to the channel dimension of the input feature map and obtain the motion information weight of each channel; and a motion feature excitation subunit, which is used to multiply the motion information weight of each channel with the input feature map and output the feature map after motion feature excitation.

[0014] Furthermore, the motion feature excitation module includes a query feature map acquisition unit, which is used to perform a convolution operation with a convolution kernel of 3 on the feature map after the motion sensitive channel is enhanced in time sequence to obtain a query feature map The key feature map acquisition unit is used to perform a convolution operation with a convolution kernel of 3 on the feature map after the motion sensitive channel enhancement in time sequence. For the feature map after the convolution, the channel dimension and the time sequence dimension are transposed to obtain the key feature map. The value feature map acquisition unit is used to perform a convolution operation with a convolution kernel of 1 and a dimension transformation on the feature map after the motion sensitive channel enhancement in time sequence to obtain a value feature map that retains the original features. The temporal weight allocation unit is used to obtain the association weights of each time series by using the temporal association function to obtain the association weights of each time series with the value feature map. The frame sequence enhancement unit is used to perform corresponding multiplication operations on the association weights of each time series with the value feature map in time series, transform the dimension of the value feature map after the multiplication operation to the dimension of the feature map after explicit enhancement, add the value feature map after the dimension transformation and the feature map after explicit enhancement to obtain the feature map after temporal enhancement, that is, the key partial sequence in the video data is enhanced.

[0015] Furthermore, the association weight calculation is expressed as: Among them, σ(·) represents the Softmax function, p(·) represents the global pooling operation, ⊙ represents the dot product operation, Att T Indicates the associated weight of each time series.

[0016] On the other hand, a system for identifying athlete actions based on motion and timing enhancement is provided, comprising an action recognition module, a video processing module and a model training module; the action recognition module comprises an athlete action recognition network model based on motion and timing enhancement as described in any one of claims 1 to 8; the video processing module is used to capture video clips containing athlete actions from sports videos, and annotate the video with action category labels to obtain a data set; the model training module is used to train the network architecture model using the data set, and optimize the parameters in the network architecture model through back propagation, and iterate the network architecture model based on the loss function until a preset number of iterations is reached.

[0017] Furthermore, the expression of the loss function is: Among them, C represents the total number of action categories, represents the predicted probability that the sample belongs to the i-th category, Represents the probability distribution of the sample.

[0018] Compared with the prior art, the present invention has the following advantages and beneficial effects: the present invention fully characterizes the local motion information by fusing the feature-level differences between short-distance video frames, and explicitly stimulates the motion channel while suppressing the background channel. At the same time, by constructing a temporal correlation function for long-distance video frames to capture the global dependency between the temporal sequences, the temporal global information extracted is used to reasonably allocate the temporal weights, and enhance the important temporal sequences that are critical to action recognition. Compared with the existing models, the present invention captures and enhances the motion information and temporal information in the motion video frames from a local and global perspective, and can complete the action recognition task of the players in the sports basketball video end-to-end, which can solve the problems of insufficient extraction of the motion information of the players in the sports video and difficulty in extracting long temporal features. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the technical solutions of the exemplary embodiments of the present invention, the drawings required for use in the embodiments will be briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.

[0020] Figure 1 A schematic diagram of a network for identifying athlete movements based on motion and timing enhancement provided in Example 1 of the present invention;

[0021] Figure 2 A schematic diagram of the working principle of the motion feature excitation module provided in Example 1 of the present invention;

[0022] Figure 3 A schematic diagram of the working principle of the timing enhancement module provided in Example 1 of the present invention. DETAILED DESCRIPTION

[0023] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with embodiments and drawings. The exemplary embodiments of the present invention and their description are only used to explain the present invention and are not intended to limit the present invention.

[0024] Example 1

[0025] This embodiment provides a Figure 1 The athlete action recognition network based on motion and timing enhancement shown in the figure includes the first convolution layer Conv1, the second convolution layer Conv2_x, the third convolution layer Conv3_x, the fourth convolution layer Conv4_x, the fifth convolution layer Conv5_x and the fully connected layer Fc connected in sequence. Among them, the first convolution layer Conv1 is used to perform 1×7×7 convolution operation and maximum pooling operation on the input motion video frame, and output the feature map to the second convolution layer Conv2_x. The second convolution layer Conv2_x contains 3 convolution blocks Block, the third convolution layer Conv3_x contains 4 convolution blocks Block, the fourth convolution layer Conv4_x contains 6 convolution blocks Block, and the fifth convolution layer Conv5_x contains 3 convolution blocks Block. The following operations are performed in each convolution block: 1. The channels of the received feature map are reduced in dimension using a 1×1×1 three-dimensional convolution; 2. The mixed motion features are extracted from the feature map after dimensionality reduction by embedding a motion feature excitation module, and the motion sensitive channels are excited; 3. The feature map after explicit enhancement is temporally enhanced by embedding a timing enhancement module; 4. The spatiotemporal features of the feature map after temporal enhancement are extracted using a 3×3×3 three-dimensional convolution.

[0026] The motion feature excitation module extracts dynamic local motion information from short-distance video frames, and explicitly enhances the channels of captured motion information and suppresses the channels of captured background information. T represents the number of frames of the input feature map, C, H and W represent the number of channels, length and width of the input feature map respectively. The working principle of the motion feature excitation module is as follows Figure 2 As shown, in terms of module composition, the motion feature excitation module has the following functional units:

[0027] The first dimension transformation unit is used to reduce the dimension of the channel of the feature map after the 1×1×1 three-dimensional convolution using a 1×1 two-dimensional convolution, reducing the channel dimension of the feature map to 1 / 16 of the original, and obtaining the feature map after dimensionality reduction.

[0028] The time division unit is used to divide the feature map after 1×1 two-dimensional convolution Divide by time sequence to get the feature map X of each frame R (t), t represents the timestamp of the feature map of each frame, t = 1, 2, ..., T, T represents the frame number of the input feature map.

[0029] Feature map transformation unit, used to transform feature map X R (t) Perform 3×3 channel-by-channel convolution to obtain the transformed feature map Y R (t);

[0030] Adjacent frame motion feature calculation unit, used to calculate the feature map X R (t) is the motion feature of the adjacent frame, and the motion feature of the adjacent frame is the feature map X R (t) and the transformed feature map Y R The characteristic difference between (t+1)

[0031] Cross-frame motion feature calculation unit, used to calculate the feature map X R (t) is a cross-frame motion feature, wherein the cross-frame motion feature is a feature map X R (t) and the transformed feature map Y R The characteristic difference between (t+2)

[0032] Feature fusion unit, used to calculate the feature map X R (t) The fused motion features at the corresponding time t The fused motion features are the sum of adjacent frame motion features and cross-frame motion features;

[0033] The feature splicing unit is used to splice the fused motion features corresponding to each moment in time sequence to obtain the local motion features of the input feature map

[0034] The second dimension transformation unit is used to interact and transform the dimension of the channel dimension of the local motion feature using a 1×1 two-dimensional convolution, transform the dimension of the feature map into T×C / 16×HW, and obtain the transformed local motion feature;

[0035] The third dimension transformation unit is used to transform the channel dimension of the local motion feature into 1 using a 1×1 two-dimensional convolution to obtain the global spatial feature and transform its dimension into T×HW×1;

[0036] A feature enhancement unit is used to enhance the global spatial features using a Softmax function to obtain enhanced global spatial features;

[0037] The channel descriptor acquisition unit is used to perform a dot multiplication operation on the transformed local motion feature and the enhanced global space feature to obtain the channel descriptor of the local motion feature.

[0038] The motion information weight acquisition unit is used to restore the two-dimensional convolution of the channel dimension 1×1 to the channel dimension C of the input feature map, and obtain the motion information weight of each channel through the Sigmoid function;

[0039] The motion feature excitation subunit is used to multiply the motion information weight of each channel with the input feature map and output the feature map after motion feature excitation.

[0040] Furthermore, the working principle of realizing timing enhancement is as follows Figure 3 As shown in the figure, the global timing information in the long-distance video frame is extracted through the timing correlation function, the timing weight is reasonably allocated, and the important timing that is critical to action recognition is enhanced. In terms of module composition, the timing enhancement module includes the following functional units:

[0041] The query feature map acquisition unit is used to perform a convolution operation with a convolution kernel of 3 on the feature map after explicit enhancement in time sequence to obtain a query feature map

[0042] The key feature map acquisition unit is used to perform a convolution operation with a convolution kernel of 3 on the feature map after explicit enhancement in time sequence. For the feature map after convolution, the channel dimension and the time sequence dimension are transposed to obtain the key feature map.

[0043] The value feature map acquisition unit is used to perform a convolution operation with a convolution kernel of 1 and a dimension transformation on the feature map after explicit enhancement in time sequence to obtain a value feature map that retains the original features.

[0044] The temporal weight allocation unit is used to allocate the query feature graph X Q and key feature graph Through the time series correlation function g θ The association weights of each time series are obtained. The association weight calculation is expressed as: Among them, σ(·) represents the Softmax function, p(·) represents the global pooling operation, ⊙ represents the dot product operation, Att T Represents the associated weight of each time series. The mth row represents the feature information in the mth frame. The nth column of represents the feature information in the nth frame (1≤m≤T,1≤n≤T), and The element in the mth row and nth column of represents the association weight of the mth frame to the nth frame. The larger the value, the greater the correlation between the time series and the global one and the more important it is.

[0045] The frame sequence enhancement unit is used to perform corresponding multiplication operations on the associated weights of each time sequence and the value feature map in time sequence, transform the dimension of the value feature map after the multiplication operation to the dimension of the feature map after explicit enhancement, add the value feature map after the dimension transformation and the feature map after explicit enhancement to obtain the feature map after time sequence enhancement, and complete the enhancement of important frame sequences in the full time domain video frame.

[0046] In addition, the above-mentioned network architecture also includes a fully connected layer, which is used to perform classification and recognition based on the feature information output by the fifth convolutional layer. When the video frame is input into the network architecture, it is processed in sequence through the first to fifth convolutional layers, and finally the fully connected layer outputs the athlete's action recognition result.

[0047] Example 2

[0048] Based on the network architecture provided in Example 1, this embodiment provides an athlete action recognition system based on motion and timing enhancement. The system includes an action recognition module, a video processing module and a model training module.

[0049] Among them, the action recognition module includes the network architecture model described in Example 1.

[0050] The video processing module is used to collect game video data (such as basketball game videos), capture player action clips and mark action category labels to create a data set. The video data can be downloaded from the Internet or collected by shooting with a camera. The acquisition method is reasonable and reliable. If the task requirements are met, video clips containing athlete actions can be captured from sports videos, and action category labels can be marked on the video clips to obtain a data set.

[0051] The model training module is used to train the network architecture model using the data set output by the video processing module, and optimize the parameters in the network architecture model through back propagation, and iterate the network architecture model based on the loss function until the preset number of iterations is reached. The expression of the loss function used is Among them, C represents the total number of action categories, represents the predicted probability that the sample belongs to the i-th category, Represents the probability distribution of the sample.

[0052] The specific implementation methods described above further illustrate the objectives, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above description is only a specific implementation method of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A method for constructing an athlete action recognition network based on motion and timing enhancement, characterized in that: The action recognition network includes a first convolutional layer, a second convolutional layer, a third convolutional layer, a fourth convolutional layer, a fifth convolutional layer and a fully connected layer connected in sequence; The first convolutional layer is used to perform convolution operations and maximum pooling operations on the input motion video frame, and output feature maps to the second convolutional layer; The second convolutional layer to the fifth convolutional layer are all used to extract local motion features and explicitly enhance motion-sensitive channels on the received feature map, then globally enhance the key sequence of the explicitly enhanced feature map, and finally extract spatiotemporal features on the feature map with enhanced motion information and timing information; wherein, Local motion feature extraction includes: The first dimension transformation subunit is used to perform dimensionality reduction processing on the channels of the feature map after the 1×1×1 3D convolution using 1×1 2D convolution; The time division subunit is used to divide the feature map after 1×1 two-dimensional convolution into time series to obtain the feature map X of each frame. R (t), t represents the timestamp of the feature map of each frame, t = 1, 2, ..., T, T represents the frame number of the input feature map; Feature map transformation subunit, used to transform the feature map X R (t+1), X R (t+2) Perform 3×3 channel-by-channel convolution to obtain the transformed feature map Y R (t+1), Y R (t+2); The adjacent frame motion feature calculation subunit is used to calculate the feature map X R (t) is the motion feature of the adjacent frame, and the motion feature of the adjacent frame is the feature map X R (t) and the transformed feature map Y R The characteristic difference between (t+1); The cross-frame motion feature calculation subunit is used to calculate the feature map X R (t) is a cross-frame motion feature, wherein the cross-frame motion feature is a feature map X R (t) and the transformed feature map Y R The characteristic difference between (t+2); Feature fusion subunit, used to calculate the feature map X R (t) The fused motion features at the corresponding moment, where the fused motion features are the sum of the adjacent frame motion features and the cross-frame motion features; The feature splicing subunit is used to splice the fused motion features corresponding to each moment in time sequence to obtain the local motion features of the input feature map; Explicit enhancements of motion-sensitive pathways include: The second dimension transformation subunit is used to perform dimension transformation on the channel dimension of the local motion feature using a 1×1 two-dimensional convolution to obtain the transformed local motion feature; The third dimension transformation subunit is used to perform dimension transformation on the channel of the local motion feature using a 1×1 two-dimensional convolution, so that the channel dimension of the local motion feature is transformed to 1, thereby obtaining the global spatial feature; The feature enhancement subunit is used to enhance the global spatial features through the softmax function to obtain enhanced global spatial features; A channel descriptor acquisition subunit is used to perform a dot multiplication operation on the transformed local motion feature and the enhanced global space feature to obtain a channel descriptor of the local motion feature; The motion information weight acquisition subunit is used to restore the channel dimension to the channel dimension of the input feature map and obtain the motion information weight of each channel through the Sigmoid function; The motion feature excitation subunit is used to multiply the motion information weight of each channel with the input feature map and output the feature map after motion feature excitation; The fully connected layer is used to perform classification and recognition according to the feature information output by the fifth convolutional layer, and output the athlete action recognition result.

2. A method for constructing a network for identifying athlete movements based on motion and timing enhancement according to claim 1, characterized in that: The second convolutional layer to the fifth convolutional layer all include a plurality of convolutional blocks combined in a residual form, each of the convolutional blocks being used to perform dimensionality reduction processing on the channels of the received feature maps using a 1×1×1 three-dimensional convolution, and to perform feature extraction and explicit enhancement of motion-sensitive channels on the local mixed motion of the feature maps after dimensionality reduction processing by embedding a motion feature excitation module, and to perform global timing enhancement on the key sequence of the explicitly enhanced feature maps by embedding a timing enhancement module, and to perform spatiotemporal feature extraction on the feature maps with enhanced motion information and timing information using a 3×3×3 three-dimensional convolution.

3. A method for constructing a network for identifying athlete movements based on motion and timing enhancement according to claim 2, characterized in that: The second convolution layer includes 3 convolution blocks, the third convolution layer includes 4 convolution blocks, the fourth convolution layer includes 6 convolution blocks, and the fifth convolution layer includes 3 convolution blocks.

4. A method for constructing a sportsman action recognition network based on motion and timing enhancement according to claim 2, characterized in that: The timing enhancement module comprises: The query feature map acquisition unit is used to perform a convolution operation with a convolution kernel of 3 on the feature map after the motion sensitive channel enhancement in time sequence to obtain the query feature map The key feature map acquisition unit is used to perform a convolution operation with a convolution kernel of 3 on the feature map after the motion sensitive channel enhancement in time sequence. For the feature map after the convolution, the channel dimension and the time sequence dimension are transposed to obtain the key feature map. The value feature map acquisition unit is used to perform a convolution operation with a convolution kernel of 1 and a dimension transformation on the feature map after the motion sensitive channel enhancement in time sequence to obtain a value feature map that retains the original features. A time series weight allocation unit, used to obtain the association weights of each time series by using the time series association function to associate the query feature graph and the key feature graph; The frame sequence enhancement unit is used to perform corresponding multiplication operations on the associated weights of each time sequence and the value feature map in time sequence, transform the dimension of the value feature map after the multiplication operation to the dimension of the feature map after explicit enhancement, add the value feature map after the dimension transformation and the feature map after explicit enhancement, and obtain the feature map after time sequence enhancement.

5. A method for constructing a network for identifying athlete movements based on motion and timing enhancement according to claim 1, characterized in that: The association weight calculation is expressed as: Among them, σ(·) represents the Softmax function, p(·) represents the global pooling operation, ⊙ represents the dot product operation, Att t Indicates the associated weight of each time series.

6. A sportsman action recognition system based on motion and timing enhancement, characterized in that: Includes action recognition module, video processing module and model training module; The action recognition module is obtained by a method for constructing an athlete action recognition network based on motion and timing enhancement as described in any one of claims 1 to 5; The video processing module is used to capture video clips containing athlete movements from sports videos, and annotate the video with action category labels to obtain a data set; The model training module is used to train the network model using the data set, optimize the parameters in the network model through back propagation, and iterate the network model based on the loss function until a preset number of iterations is reached.

7. A sportsman action recognition system based on motion and timing enhancement according to claim 6, characterized in that: The expression of the loss function is Among them, C represents the total number of action categories, represents the predicted probability that the sample belongs to the i-th category, Represents the probability distribution of the sample.