Standing Long Jump Stage Classification Method Based on Feature Adaptive Fusion

By using a video classification network with feature adaptive fusion and double-layer pooling timing attention module in the standing long jump movement, the problem of inaccurate classification of the existing technology is solved, and more efficient and accurate classification of the motion stage is achieved.

CN115359292BActive Publication Date: 2025-06-20SHAANXI NORMAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211005577.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-22
Publication Date
2025-06-20
Estimated Expiration
2042-08-22

AI Technical Summary

Technical Problem

It is difficult for the prior art to achieve accurate classification and personalized description of the opposite fixed jump motion, and the existing methods are inefficient and inaccurate, which are greatly affected by human factors and optical flow extraction efficiency.

Method used

The standing long jump stage classification method based on feature adaptive fusion is adopted. By building a video classification network, the motion information adaptive fusion and the double-layer pooling timing attention module are used to extract motion feature information and perform additive feature adaptive fusion to improve the accuracy of classification results.

Benefits of technology

The refinement of the acquisition of motion feature information and the accuracy of classification results has been achieved. Compared with other mainstream networks, the classification accuracy rate has been increased by 14.1%, which can effectively perform stage classification of standing long jump movement.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115359292B_ABST
    Figure CN115359292B_ABST
Patent Text Reader

Abstract

A standing long jump stage classification method based on feature adaptive fusion consists of constructing a video classification dataset, adaptively fusing motion information, constructing a video classification network, training the video classification network, and detecting test set videos. The present invention adopts an adaptive fusion method of motion information to change the input modality to focus on motion boundary information, and constructs a double-layer pooling temporal attention module that focuses on local and global feature information in the feature extraction backbone network, obtaining a standing long jump stage classification network with feature adaptive fusion. Compared with the prior art, the present invention has the advantages of more refined acquisition of motion feature information and more accurate classification results. The comparative simulation experiment results on the self-built dataset of standing long jump show that, compared with the existing mainstream methods, the classification accuracy is increased by 14.1%, and it can be used for the stage classification of this specific sport of standing long jump.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer technology, and particularly relates to sports classification. Background Art

[0002] As a sport that integrates physical qualities such as bounce, explosive power, body coordination, and technique, standing long jump is listed as a mandatory test item for annual physical fitness tests in junior high schools, senior high schools, and universities by the National Student Physical Health Standard, and is one of the important physical fitness test items to measure students' physical fitness. The main problems existing in the existing test methods are as follows: (1) Each link is easily affected by human factors, and the operation process is cumbersome; (2) The test method has low efficiency and inaccurate test results; (3) The evaluation result depends on the teacher's personal experience in observing the long jump process, lacking a quantitative objective standard. Therefore, applying advanced computer vision technology to the current standing long jump physical fitness test and giving qualitative motion analysis to achieve the purpose of intelligent guidance and scientific analysis has become a research hotspot and challenging topic in the field of artificial intelligence in recent years. At the present stage, there is no "AI + assisted training" technology for students' standing long jump, and there is also little research on the fine-grained classification of this sport, which cannot meet the actual needs of intelligent guidance.

[0003] The action classification technology based on deep learning can be classified into two categories: based on 3D convolutional neural network and based on 2D convolutional neural network structure. The three-dimensional convolutional network helps to capture time information by adding a time dimension, and can effectively extract spatio-temporal features. However, the deployment cost based on 3D convolutional neural network is too high, and the 3D convolutional kernel has a high computational cost. Subsequently, some variant methods based on 3D convolution have emerged, decomposing 3D convolution into spatial 2D convolution and temporal 1D convolution to learn spatio-temporal features, but still have a higher computational cost than 2D convolution. The methods based on 2D convolutional neural network can be mainly divided into two categories. One is the two-stream network idea: using two 2D convolutional neural networks to process RGB frames (spatial stream) and pre-computed optical flow (temporal stream) respectively, and finally applying a late fusion strategy to obtain spatio-temporal semantics. This type of method can improve the recognition accuracy, but the relatively independent two networks also hinder the information exchange, making it difficult to effectively combine information in each dimension, and the two-stream convolution heavily relies on optical flow as the motion representation, and the extraction of optical flow is time-consuming and storage-demanding, and the efficiency is also not satisfactory. The other is to add an attention mechanism to the network: by adding an attention module, modeling the channels and time series to strengthen important features to achieve the effect of improving the accuracy rate.

[0004] Considering the characteristics of the standing long jump, from the perspective of the duration of the collected video, compared with the existing experimental data sets, the movements in the standing long jump stage change instantaneously, and the durations of different movement instances are extremely uneven. There are also significant differences in the frame times of different samples under the same movement instance. From the background and movements of the standing long jump, the movement processes of students in the same class have similar backgrounds, and all students perform the same movement program. The subtle differences in these movement programs are mainly reflected in the changes in spatial semantic information. Existing methods are difficult to accurately classify this sport and effectively make personalized descriptions of the movement process of the standing long jump. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to overcome the shortcomings of the prior art and provide a standing long jump stage classification method based on feature adaptive fusion with fine acquisition of motion feature information and accurate classification results.

[0006] The technical solution adopted to solve the above technical problem consists of the following steps:

[0007] (1) Construct a standing long jump data set

[0008] Collect videos of students' standing long jump movements in the natural environment of the playground. Divide each complete movement video into 10 stage clips to construct a standing long jump stage data set {S0, S1, S2, S3, S4, S5, S6, S7, S8, S9}. S0 represents the stage from the start of the video to the highest point of the swing arm. S1 represents the stage from the highest point of the swing arm to the moment when the sole of the foot starts to leave the ground. S2 represents the stage from the sole of the foot starting to leave the ground to the stage of full sole leaving the ground and taking off. S3 represents the stage from full sole leaving the ground and taking off to the stage of pulling the back arch. S4 represents the stage from pulling the back arch to the stage when the thigh is perpendicular to the ground. S5 represents the stage from the thigh being perpendicular to the ground to the highest point of the lifted knee. S6 represents the stage from the highest point of the lifted knee to the moment when the heel starts to land. S7 represents the stage from the heel starting to land to the stage of full sole landing. S8 represents the stage from full sole landing to the lowest point of the hip bone. S9 represents the stage from the lowest point of the hip bone to the end of the video. Divide the data set into a training set, a validation set, and a test set according to a ratio of 3:1:1.

[0009] (2) Adaptive fusion of motion information

[0010] Divide the video input sequence V into M video segments of equal length. V is {S1, S2,..., S M}, and select adjacent f frames from each segment to stack as {I1, I2,..., I M}. The f frames in the adjacent frame set are Input them into the motion enhancement module to obtain motion information The motion information features of each segment form {ME1, ME2,..., ME M}, and the first frame of each f-frame stack is Fed to the video classification network in an additive feature adaptive fusion manner, where M is a finite positive integer and f takes the value of 4.

[0011] (3) Construct the video classification network

[0012] The video classification network is sequentially composed of a 7×7 convolutional layer and a max pooling layer, a second-stage residual block, a third-stage residual block, a fourth-stage residual block, and a fifth-stage residual block. The second-stage residual block is composed of 3 residual basic blocks in series, the third-stage residual block is composed of 4 residual basic blocks in series, the fourth-stage residual block is composed of 6 residual basic blocks in series, and the fifth-stage residual block is composed of 3 residual basic blocks in series. Each residual basic block is sequentially composed of a 1×1 convolutional layer and a double-layer pooling temporal attention module, a 3×3 convolutional layer, and a 1×1 convolutional layer.

[0013] (4) Train the video classification network

[0014] 1) Initialize the video classification network

[0015] Initialize the parameters of the video classification network using the Xavier method.

[0016] 2) Set the hyperparameters of the video classification network

[0017] Adjust the video frame size of the training set to 224×224; during training, the data batch size is 8, the learning rate is 0.0025, the learning rate decays to 0.00025 after 36 epochs, and the learning rate decays to 0.000025 in the last 10 epochs.

[0018] 3) Train the video classification network

[0019] Input all the videos in the training set into the video classification network for forward propagation, and determine the loss function L according to the following formula:

[0020]

[0021] The loss function L is the negative log-likelihood loss; use the stochastic gradient descent method to reduce the loss value for backpropagation, repeatedly cycle forward propagation and backpropagation, and update the weights and biases of the video classification network until the loss function converges, the training ends, and the trained video classification network is obtained.

[0022] (5) Detect the test set videos

[0023] Input the test set into the trained video classification network and output the video classification results.

[0024] The (2) motion information adaptive fusion step of the present invention is: divide the video input sequence V into M video segments of equal length, where V is {S1, S2,..., S M}, select adjacent f frames in each segment and stack them as {I1, I2,..., I M} The f frames in the adjacent frame set are Input it into the motion enhancement module to obtain motion information The motion information features of each segment constitute {ME1, ME2,..., ME M} The first frame stacked with every f frames is Feed it into the video classification network in an additive feature adaptive fusion manner, and M takes values from 3 to 8.

[0025] In the (2) motion information adaptive fusion of the present invention, the additive feature adaptive fusion method is as follows:

[0026] Determine the feature F according to the following formula IM :

[0027] F IM = I + ME

[0028] Where I is the RGB feature of the first frame set stacked with every f frames, and ME is the mapping feature obtained by the adjacent frame pairs stacked with every f frames through the motion enhancement module. Determine the feature fusion result F according to the following formula ADP :

[0029] F ADP = α1 × F IIM + α2 × F MEIM

[0030]

[0031] ∑ i α i = 1

[0032] Where the feature F IIM is obtained by concatenating the feature I and the feature F IM in the channel dimension column by column. The feature F MEIM is obtained by concatenating the feature ME and the feature F IM in the channel dimension column by column. α i is the normalized weight, ω i is the initial weight coefficient, ω j is the feature weight, i ∈ {1, 2}, j ∈ {1, 2}, and an adaptive learnable weight coefficient is used to assign weights to different features.

[0033] In the (3) construction of the video classification network of the present invention, the construction method of the double-layer pooling temporal attention module is as follows: The input feature of this module is X0 ∈ [N, T, C, H, W], where N is the batch size, T is the time dimension of the feature, C is the number of channels, H is the length in the spatial dimension, and W is the width in the spatial dimension. Transpose the tensor dimension to X1 ∈ [N, C, T, H, W], and determine the feature according to the following formula

[0034]

[0035]

[0036] Among them, n, c, t, h, and w represent all the values of their corresponding dimensions, and Conv1d() is a one-dimensional convolution.

[0037] Determine the temporal attention weight F according to the following formula t :

[0038]

[0039] Determine the feature X2 according to the following formula:

[0040]

[0041]

[0042] Among them, represents element-wise addition, and ⊙ represents element-wise multiplication.

[0043] Determine the importance weight F sensitive to temporal positions according to the following formula s :

[0044]

[0045] Among them, Relu() is the rectified linear activation function, Sigmoid() is the S-shaped activation function, K is the convolution kernel size, K takes the value of 3, and β is a hyperparameter, taking the value of 2 to 8.

[0046] Determine the feature X3 according to the following formula:

[0047]

[0048] Keep the feature output with the same size as the original input, and perform tensor dimension transposition to X4 ∈ [N, T, C, H, W].

[0049] The present invention uses an adaptive fusion method of motion information to focus on motion boundary information, and constructs a two-layer pooling temporal attention module that focuses on local and global feature information in the feature extraction backbone network, obtaining a stage classification network for standing long jump with feature adaptive fusion. Compared with the prior art, it has the advantages of more refined acquisition of motion feature information and more accurate classification results. The comparison simulation experiment results show that the classification accuracy of the present invention is improved by 14.1% compared with other mainstream networks, and it can be used for stage classification of standing long jump motion. Description of the Drawings

[0050] Figure 1It is the flow chart of the present invention.

[0051] Figure 2 is Figure 1 Schematic diagram of the standing long jump stage classification network structure for feature adaptive fusion in

[0052] Figure 3 It is the result diagram of the 1st stage of the standing long jump.

[0053] Figure 4 It is the result diagram of the 2nd stage of the standing long jump.

[0054] Figure 5 It is the result diagram of the 3rd stage of the standing long jump.

[0055] Figure 6 It is the result diagram of the 7th stage of the standing long jump. Detailed implementation manners

[0056] The present invention will be further described below in conjunction with the accompanying drawings and examples, but the present invention is not limited to the following embodiments.

[0057] Embodiment 1

[0058] The method for classifying the stages of the standing long jump based on feature adaptive fusion in this embodiment consists of the following steps (see Figure 1 ).

[0059] (1) Construct a standing long jump data set

[0060] Collect the standing long jump motion videos of students in the natural environment scene of the playground, divide each complete motion video into 10 stages, and construct a standing long jump stage data set {S0, S1, S2, S3, S4, S5, S6, S7, S8, S9}, where S0 represents the stage from the start of the video to the highest point of the swing arm, S1 represents the stage from the highest point of the swing arm to the moment when the sole of the foot starts to leave the ground, S2 represents the stage from the sole of the foot starting to leave the ground to the stage of taking off with the whole sole off the ground, S3 represents the stage from taking off with the whole sole off the ground to the stage of pulling the back arch, S4 represents the stage from pulling the back arch to the stage when the thigh is perpendicular to the ground, S5 represents the stage from the thigh being perpendicular to the ground to the highest point of the lifted knee, S6 represents the stage from the highest point of the lifted knee to the moment when the heel starts to land, S7 represents the stage from the heel starting to land to the stage of the whole sole landing, S8 represents the stage from the whole sole landing to the lowest point of the hip bone, and S9 represents the stage from the lowest point of the hip bone to the end of the video. Divide the data set into a training set, a validation set, and a test set according to a ratio of 3:1:1;

[0061] (2) Adaptive fusion of motion information

[0062] Divide the video input sequence V into M video segments of equal length. V is {S1, S2,..., S M} and stack adjacent f frames for each segment to form {I1, I2,..., I M}, the f frames in the adjacent frame set are Input it into the motion enhancement module to obtain motion information The motion information features of each segment form {ME1, ME2,..., ME M}, and the first frame stacked with each f frames is Feed it into the video classification network in an additive feature adaptive fusion manner. M is a finite positive integer, and the value of M in this embodiment is 3, and the value of f in this embodiment is 4.

[0063] The additive feature adaptive fusion method in this embodiment is:

[0064] Determine the feature F according to the following formula IM :

[0065] F IM = I + ME

[0066] where I is the RGB feature of the first frame set stacked with each f frames, and ME is the mapped feature obtained by the adjacent frame pairs stacked with each f frames through the motion enhancement module. Determine the feature fusion result F according to the following formula ADP :

[0067] F ADP = α1 × F IIM + α2 × F MEIM

[0068]

[0069] ∑ i α i = 1

[0070] where the feature F IIM is obtained by concatenating the feature I and the feature F IM by columns in the channel dimension. The feature F MEIM is obtained by concatenating the feature ME and the feature F IM by columns in the channel dimension. α i is the normalized weight, ω i is the initialized weight coefficient, ω j is the feature weight, i ∈ {1, 2}, j ∈ {1, 2}, and the adaptive learnable weight coefficient is used to assign weights to different features.

[0071] (3) Construct the video classification network

[0072] In Figure 2In this case, the video classification network of this embodiment is successively formed by connecting in series a 7×7 convolutional layer 1, a max pooling layer 2, a second-stage residual block 3, a third-stage residual block 4, a fourth-stage residual block 5, and a fifth-stage residual block 6. The second-stage residual block 3 is formed by connecting 3 residual basic blocks in series, the third-stage residual block 4 is formed by connecting 4 residual basic blocks in series, the fourth-stage residual block 5 is formed by connecting 6 residual basic blocks in series, and the fifth-stage residual block 6 is formed by connecting 3 residual basic blocks in series. Each residual basic block of this embodiment is successively formed by connecting in series a 1×1 convolutional layer, a double pooling temporal attention module, a 3×3 convolutional layer, and a 1×1 convolutional layer.

[0073] The construction method of the double pooling temporal attention module of this embodiment is as follows: The input feature of this module is X0 ∈ [N, T, C, H, W], where N is the batch size, T is the time dimension of the feature, C is the number of channels, H is the length in the spatial dimension, and W is the width in the spatial dimension. The tensor dimension is transposed to X1 ∈ [N, C, T, H, W], and the feature is determined according to the following formula

[0074]

[0075]

[0076] where n, c, t, h, and w represent all values of their corresponding dimensions, and Conv1d() is a one-dimensional convolution;

[0077] The temporal attention weight F is determined according to the following formula t :

[0078]

[0079] The feature X2 is determined according to the following formula:

[0080]

[0081]

[0082] where denotes element-wise addition, and ⊙ denotes element-wise multiplication;

[0083] The importance weight F sensitive to temporal positions is determined according to the following formula s :

[0084]

[0085] where Relu() is a rectified linear activation function, Sigmoid() is an S-shaped activation function, K is the convolution kernel size, K takes a value of 3, β is a hyperparameter, β takes a value from 2 to 8, and β of this embodiment takes a value of 4.

[0086] Determine the feature X3 according to the following formula:

[0087]

[0088] Keep the feature output with the same size as the original input, and perform tensor dimension transposition to X4 ∈ [N, T, C, H, W].

[0089] (4) Train the video classification network

[0090] 1) Initialize the video classification network

[0091] Initialize the parameters of the video classification network using the Xavier method.

[0092] 2) Set the hyperparameters of the video classification network

[0093] Adjust the video frame size of the training set to 224×224; during training, the data batch is 8, the learning rate is 0.0025, the learning rate decays to 0.00025 after 36 rounds, and the learning rate decays to 0.000025 in the last 10 times.

[0094] 3) Train the video classification network

[0095] Input all the videos in the training set into the video classification network for forward propagation, and determine the loss function L according to the following formula:

[0096]

[0097] The loss function L is the negative log-likelihood loss; use the stochastic gradient descent method to reduce the loss value for backpropagation, repeatedly cycle forward propagation and backpropagation, and update the weights and biases of the video classification network until the loss function converges, the training ends, and the trained video classification network is obtained.

[0098] (5) Detect the test set videos

[0099] Input the test set into the trained video classification network and output the video classification results.

[0100] Complete the standing long jump stage classification method based on feature adaptive fusion.

[0101] Example 2

[0102] The standing long jump stage classification method based on feature adaptive fusion in this example consists of the following steps.

[0103] (1) Construct the standing long jump data set

[0104] This step is the same as that in Example 1.

[0105] (2) Adaptive fusion of motion information

[0106] Divide the video input sequence V into M video segments of equal length V as {S1, S2,..., S M}, and select adjacent f frames from each segment and stack them as {I1, I2,..., I M}, and the f frames in the adjacent frame set are Input it into the motion enhancement module to obtain motion information The motion information features of each segment constitute {ME1, ME2,..., ME M}, and the first frame of each stack of f frames is Feed it into the video classification network in an additive feature adaptive fusion manner. M is a finite positive integer, and the value of M in this embodiment is 5, and the value of f in this embodiment is 4.

[0107] The additive feature adaptive fusion method in this embodiment is:

[0108] Determine the feature F according to the following formula IM :

[0109] F IM = I + ME

[0110] where I is the RGB feature of the first frame set of each stack of f frames, and ME is the mapping feature obtained by the adjacent frame pairs of each stack of f frames through the motion enhancement module. Determine the feature fusion result F according to the following formula ADP :

[0111] F ADP = α1 × F IIM + α2 × F MEIM

[0112]

[0113] ∑ i α i = 1

[0114] where the feature F IIM is obtained by concatenating the feature I and the feature F IM in columns along the channel dimension. The feature F MEIM is obtained by concatenating the feature ME and the feature F IM in columns along the channel dimension. α i is the normalized weight, ω i is the initialized weight coefficient, ω j is the feature weight, i ∈ {1, 2}, j ∈ {1, 2}, and an adaptive learnable weight coefficient is used to assign weights to different features.

[0115] (3) Construct a video classification network

[0116] The video classification network of this embodiment is composed of a 7×7 convolutional layer 1, a max pooling layer 2, a second-stage residual block 3, a third-stage residual block 4, a fourth-stage residual block 5, and a fifth-stage residual block 6 connected in series in sequence. The second-stage residual block 3 is composed of 3 residual basic blocks connected in series, the third-stage residual block 4 is composed of 4 residual basic blocks connected in series, the fourth-stage residual block 5 is composed of 6 residual basic blocks connected in series, and the fifth-stage residual block 6 is composed of 3 residual basic blocks connected in series. Each residual basic block of this embodiment is composed of a 1×1 convolutional layer, a double pooling temporal attention module, a 3×3 convolutional layer, and a 1×1 convolutional layer connected in series in sequence.

[0117] The construction method of the double pooling temporal attention module of this embodiment is as follows: The input feature of this module is X0∈[N,T,C,H,W], where N is the batch size, T is the time dimension of the feature, C is the number of channels, H is the length in the spatial dimension, and W is the width in the spatial dimension. The tensor dimension is transposed to X1∈[N,C,T,H,W], and the feature is determined according to the following formula

[0118]

[0119]

[0120] where n, c, t, h, w represent all values of their corresponding dimensions, and Conv1d() is a one-dimensional convolution;

[0121] The temporal attention weight F is determined according to the following formula t :

[0122]

[0123] The feature X2 is determined according to the following formula:

[0124]

[0125]

[0126] where denotes element-wise addition, and ⊙ denotes element-wise multiplication;

[0127] The importance weight F sensitive to temporal positions is determined according to the following formula s :

[0128]

[0129] where Relu() is a rectified linear activation function, Sigmoid() is an S-shaped activation function, K is the convolution kernel size, K takes the value of 3, β is a hyperparameter, β takes values from 2 to 8, and β takes the value of 2 in this embodiment.

[0130] Determine the feature X3 according to the following formula:

[0131]

[0132] Maintain the feature output with the same size as the original input, and perform tensor dimension transposition to X4 ∈ [N, T, C, H, W].

[0133] Other steps are the same as those in Embodiment 1. Complete the standing long jump stage classification method based on feature adaptive fusion.

[0134] Embodiment 3

[0135] The standing long jump stage classification method based on feature adaptive fusion in this embodiment consists of the following steps.

[0136] (1) Construct a standing long jump data set

[0137] This step is the same as that in Embodiment 1.

[0138] (2) Adaptive fusion of motion information

[0139] Divide the video input sequence V into M video segments V of equal length {S1, S2,..., S M}}, and select adjacent f frames in each segment to stack as {I1, I2,..., I M}, and the f frames in the adjacent frame set are Input it into the motion enhancement module to obtain motion information The motion information features of each segment form {ME1, ME2,..., ME M}}, and the first frame of each f-frame stack is Feed it into the video classification network in an additive feature adaptive fusion manner. M is a finite positive integer, and the value of M in this embodiment is 8, and the value of f in this embodiment is 4.

[0140] The additive feature adaptive fusion method in this embodiment is as follows:

[0141] Determine the feature F according to the following formula IM :

[0142] F IM = I + ME

[0143] where I is the RGB feature of the first frame set of each f-frame stack, and ME is the mapping feature obtained by the adjacent frame pairs of each f-frame stack through the motion enhancement module. Determine the feature fusion result F according to the following formula ADP :

[0144] F ADP = α1 × F IIM + α2 × F MEIM

[0145]

[0146] ∑ i α i = 1

[0147] Among them, feature F IIM is obtained by concatenating feature I and feature F IM in the column direction of the channel dimension. Feature F MEIM is obtained by concatenating feature ME and feature F IM in the column direction of the channel dimension. α i is the normalized weight, ω i is the initialized weight coefficient, ω j is the feature weight. i ∈ {1, 2}, j ∈ {1, 2}, and adaptive learnable weight coefficients are used to assign weights to different features.

[0148] (3) Construct a video classification network

[0149] The video classification network of this embodiment is sequentially composed of a 7×7 convolutional layer 1, a max pooling layer 2, a second-stage residual block 3, a third-stage residual block 4, a fourth-stage residual block 5, and a fifth-stage residual block 6 in series. The second-stage residual block 3 is composed of 3 residual basic blocks in series, the third-stage residual block 4 is composed of 4 residual basic blocks in series, the fourth-stage residual block 5 is composed of 6 residual basic blocks in series, and the fifth-stage residual block 6 is composed of 3 residual basic blocks in series. Each residual basic block of this embodiment is sequentially composed of a 1×1 convolutional layer, a double pooling temporal attention module, a 3×3 convolutional layer, and a 1×1 convolutional layer.

[0150] The construction method of the double pooling temporal attention module in this embodiment is as follows: The input feature of this module is X0 ∈ [N, T, C, H, W], where N is the batch size, T is the time dimension of the feature, C is the number of channels, H is the length in the spatial dimension, and W is the width in the spatial dimension. The tensor dimension is transposed to X1 ∈ [N, C, T, H, W], and the feature is determined according to the following formula

[0151]

[0152]

[0153] where n, c, t, h, w represent all values of their corresponding dimensions, and Conv1d() is a one-dimensional convolution;

[0154] The temporal attention weight F is determined according to the following formula t :

[0155]

[0156] The feature X2 is determined according to the following formula:

[0157]

[0158]

[0159] Among them, denotes element-wise addition, and ⊙ denotes element-wise multiplication;

[0160] Determine the importance weight F sensitive to the timing position according to the following formula s :

[0161]

[0162] Among them, Relu() is the rectified linear activation function, Sigmoid() is the S-shaped activation function, K is the convolution kernel size, K takes the value of 3, β is a hyperparameter, and β takes the value of 2 to 8. In this embodiment, β takes the value of 8.

[0163] Determine the feature X3 according to the following formula:

[0164]

[0165] Keep the feature output the same size as the original input, and perform tensor dimension transposition to X4 ∈ [N, T, C, H, W].

[0166] Other steps are the same as those in Embodiment 1. Complete the standing long jump stage classification method based on feature adaptive fusion.

[0167] To verify the beneficial effects of the present invention, the inventor used the standing long jump stage classification method based on feature adaptive fusion in Embodiment 1 of the present invention (hereinafter referred to as Embodiment 1) and compared it with "Wang L, Xiong Y, Wang Z, et al. Temporal segment networks for action recognition in videos [J]. IEEE transactions on pattern analysis and machine intelligence, 2018, 41(11): 2740-2755." (hereinafter referred to as Comparative Experiment 1), "Lin J, Gan C, Han S. Tsm: Temporal shift module for efficient video understanding [C] Proceedings of the IEEE International Conference on Computer Vision. 2019: 7083-7093." (hereinafter referred to as Comparative Experiment 2), "Liu Z, Wang L, Wu W, et al. Tam: Temporal adaptive module for video recognition [C] Proceedings of the IEEE International Conference on Computer Vision. 2021: 13708-13718." (hereinafter referred to as Comparative Experiment 3), "Yang C, Xu Y, Shi J, et al. Temporal pyramid network for action recognition [C] Proceedings of the IEEE conference on computer vision and pattern recognition. 2020: 591-600." (hereinafter referred to as Comparative Experiment 4), "Zhou, B., Andonian, A., Oliva, A., & Torralba, A. Temporal relational reasoning in videos. In Proceedings of the European conference on computer vision. 2018, 803-818." (hereinafter referred to as Comparative Experiment 5) for comparative experiments, and evaluated the video classification results by the highest accuracy rate and the top-five accuracy rate. The experimental results are shown in Table 1.

[0168] Table 1 Experimental results of the method in Example 1 and Comparative Experiments 1-5

[0169] Experimental group Highest accuracy rate Top-five accuracy rate Comparative experiment 1 0.7379 0.9517 Comparative experiment 2 0.8034 0.9862 Comparative experiment 3 0.8196 0.9508 Comparative experiment 4 0.8379 0.9897 Comparative experiment 5 0.8517 0.9931 Example 1 0.9448 0.9965

[0170] As can be seen from Table 1, compared with Comparative Experiments 1-5, the scores of various evaluation indicators in Example 1 have been greatly improved. The highest accuracy rate and top-five accuracy rate in Example 1 are increased by 20.7% and 4.4% respectively compared with Comparative Experiment 1, by 14.1% and 1.0% respectively compared with Comparative Experiment 2, by 12.5% and 4.6% respectively compared with Comparative Experiment 3, by 10.6% and 0.6% respectively compared with Comparative Experiment 4, and by 9.3% and 0.3% respectively compared with Comparative Experiment 5.

[0171] The above experiments show that compared with the comparative experiments, all indicators of the present invention are superior to the comparative experiments. The test results of the present invention can achieve correct stage classification. The results from the highest point of the swing arm to the stage when the sole of the foot starts to leave the ground are as Figure 3 shown, the results from the sole of the foot starting to leave the ground to the stage of taking off with the whole sole of the foot off the ground are as Figure 4 shown, the results from taking off with the whole sole of the foot off the ground to the stage of pulling the back arch are as Figure 5 shown, and the results from the heel starting to land to the stage of the whole sole of the foot landing are as Figure 6 shown. The experimental results and visualization diagrams further prove that the method of the present invention can effectively classify the standing long jump stages.

Claims

1. A standing long jump stage classification method based on feature adaptive fusion, characterized in that It consists of the following steps: (1) Construct a standing long jump dataset Collect the standing long jump motion videos of students in the natural environment scene of the playground, divide each complete motion video into 10 stage clips, and construct a standing long jump stage dataset {S0, S1, S2, S3, S4, S5, S6, S7, S8, S9}. S0 represents the stage from the start of the video to the highest point of arm swing, S1 represents the stage from the highest point of arm swing to the moment when the sole of the foot begins to leave the ground, S2 represents the stage from the sole of the foot beginning to leave the ground to the stage of taking off with the whole sole off the ground, S3 represents the stage from taking off with the whole sole off the ground to the stage of pulling the back arch, S4 represents the stage from pulling the back arch to the stage where the thigh is perpendicular to the ground, S5 represents the stage from the thigh being perpendicular to the ground to the highest point of the lifted knee, S6 represents the stage from the highest point of the lifted knee to the moment when the heel begins to land, S7 represents the stage from the heel beginning to land to the stage of the whole sole landing, S8 represents the stage from the whole sole landing to the lowest point of the hip bone, and S9 represents the stage from the lowest point of the hip bone to the end of the video. Divide the dataset into a training set, a validation set, and a test set according to 3:1:1; (2) Adaptive fusion of motion information Divide the video input sequence V into M video segments V of equal length {S1, S2,..., S M}, and select adjacent f frames from each segment and stack them as {I1, I2,..., I M}, and the f frames in the adjacent frame set are Input them into the motion enhancement module to obtain motion information The motion information features of each segment form {ME1, ME2,..., ME M}, and the first frame stacked with every f frames is Feed them into the video classification network in an additive feature adaptive fusion manner. M is a finite positive integer, and f takes the value of 4; (3) Construct a video classification network The video classification network is sequentially composed of a 7×7 convolutional layer (1), a max pooling layer (2), a second-stage residual block (3), a third-stage residual block (4), a fourth-stage residual block (5), and a fifth-stage residual block (6). The second-stage residual block (3) is composed of 3 residual basic blocks in series. The third-stage residual block (4) is composed of 4 residual basic blocks in series. The fourth-stage residual block (5) is composed of 6 residual basic blocks in series. The fifth-stage residual block (6) is composed of 3 residual basic blocks in series. Each residual basic block is sequentially composed of a 1×1 convolutional layer, a double-layer pooling temporal attention module, a 3×3 convolutional layer, and a 1×1 convolutional layer; (4) Train the video classification network 1) Initialization of the video classification network Initialize the parameters of the video classification network using the Xavier method; 2) Set the hyperparameters of the video classification network Adjust the video frame size of the training set to 224×224; during training, the data batch size is 8, the learning rate is 0.0025, the learning rate decays to 0.00025 after 36 epochs, and the learning rate decays to 0.000025 in the last 10 epochs; 3) Train the video classification network Input all the videos in the training set into the video classification network for forward propagation, and determine the loss function L according to the following formula: The loss function L is the negative log-likelihood loss; use the stochastic gradient descent method to reduce the loss value for backpropagation, repeatedly cycle forward propagation and backpropagation, and update the weights and biases of the video classification network until the loss function converges and the training ends to obtain a trained video classification network; (5) Detect the test set videos Input the test set into the trained video classification network and output the video classification results.

2. The standing long jump stage classification method based on feature adaptive fusion according to claim 1, characterized in that The described step (2) of adaptive fusion of motion information is as follows: Divide the video input sequence V into M video segments of equal length. V is {S1, S2,..., S M}, and for each segment, stack adjacent f frames to form {I1, I2,..., I M}. The f frames in the adjacent frame set are Input them into the motion enhancement module to obtain motion information The motion information features of each segment form {ME1, ME2,..., ME M}, and the first frame of each stack of f frames is Feed them into the video classification network in an additive feature adaptive fusion manner, where M ranges from 3 to 8.

3. The standing long jump stage classification method based on feature adaptive fusion according to claim 1, characterized in that: The additive feature adaptive fusion method in the (2) adaptive fusion of motion information is as follows: Determine the feature F according to the following formula IM : F IM = I + ME Among them, I is the RGB feature of the first frame set stacked every f frames, and ME is the mapped feature obtained by the adjacent frame pairs stacked every f frames through the motion enhancement module. The feature fusion result F is determined by the following formula ADP :[[]]END]] F ADP = α1 × F IIM + α2 × F MEIM ∑ i α i =1 Among them, feature F IIM is obtained by concatenating feature I and feature F IM in the channel dimension by column. Feature F MEIM is obtained by concatenating feature ME and feature F IM in the channel dimension by column. α i is the normalization weight, ω i is the initial weight coefficient, ω j is the feature weight. i ∈ {1, 2}, j ∈ {1, 2}, and adaptive learnable weight coefficients are used to assign weights to different features.

4. The standing long jump stage classification method based on feature adaptive fusion according to claim 1, characterized in that: In the construction of the video classification network in (3), the construction method of the double-layer pooling temporal attention module is as follows: the input feature of this module is X0 ∈ [N, T, C, H, W], where N is the batch size, T is the time dimension of the feature, C is the number of channels, H is the length in the spatial dimension, and W is the width in the spatial dimension. The tensor dimension is transposed to X1 ∈ [N, C, T, H, W], and the feature is determined according to the following formula Among them, n, c, t, h, w represent all the values of their corresponding dimensions, and Conv1d() is a one-dimensional convolution; Determine the temporal attention weight F according to the following formula t :[[]]END]] Determine the feature X2 according to the following formula: wherein, represents element-by-element addition, and ⊙ represents element-by-element multiplication; Determine the importance weight F sensitive to the timing position according to the following formula s : Among them, Relu() is the rectified linear activation function, Sigmoid() is the sigmoid activation function, K is the convolution kernel size, K takes the value of 3, and β is a hyperparameter, taking values from 2 to 8; Determine the feature X3 according to the following formula: Keep the feature output the same as the original input size, and perform tensor dimension transposition to X4 ∈ [N, T, C, H, W].