Method and device for counting repetitive sports actions based on multi-scale transformation network

The multi-scale transformation network accurately counts repetitive sports motions by analyzing skeletal key points, addressing inefficiencies and inaccuracies in existing methods, and providing precise, interpretable results across varying video lengths.

CN116129528BActive Publication Date: 2025-07-15XIAN YUNYING YITONG TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310166784.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-24
Publication Date
2025-07-15
Estimated Expiration
2043-02-24

AI Technical Summary

Technical Problem

The existing repetitive action counting methods consume manpower, insufficient refinement, low accuracy and poor versatility. Especially in fitness and exercise scenarios, traditional methods consume manpower and have large errors, high cost of sensor methods and affect movement. Deep learning-based methods have serious background interference during video processing and lack local periodic feature recognition capabilities.

Method used

Using a method based on a multi-scale transformation network, we use the method to detect bone key points, perform multi-scale sampling and bone-by-bone joint node coding, and periodic feature encoding using the time autocorrelation matrix, combining multi-scale periodic fusion and pulse graph regression to realize repeated action counting.

Benefits of technology

It improves the accuracy and generalization of counting, reduces the cost of label annotation, is suitable for videos of different lengths, has interpretability and high anti-interference ability, and is suitable for fitness exercises and physical testing scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116129528B_ABST
    Figure CN116129528B_ABST
Patent Text Reader

Abstract

The present invention relates to a method and device for counting repetitive sports actions based on a multi-scale transformation network. The method includes: detecting skeletal key points in an original video to obtain an original skeletal sequence, sampling the original skeletal sequence at three scales to obtain three down-sampled skeletal sequences; encoding the temporal periodicity features of each down-sampled skeletal sequence for each skeletal joint point to obtain a temporal autocorrelation matrix for each skeletal joint; performing a coarse-to-fine periodic selection on the temporal autocorrelation matrix of each skeletal joint to obtain three weighted fusion period cubes; performing multi-scale periodic fusion on the three weighted fusion period cubes to obtain a multi-scale period cube; performing impulse graph regression on the multi-scale period cube based on a multi-scale period transformation network to obtain an impulse graph; adding all elements in the impulse graph to obtain a repetitive action count result. The embodiments of the present invention are accurately tested, and the count result is interpretable.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer vision, and particularly relates to a method and device for counting repetitive sports actions based on a multi-scale transformation network. Background Art

[0002] Repetitive action counting is a task of estimating the number of repetitive actions performed by humans over a period of time. Intelligent sports action analysis is crucial in the era of national fitness. It can monitor and record human movements in real time, generate reasonable and personalized exercise plans, and help avoid sports injuries. In many fitness assessments and sports tests, such as squats, pull-ups, and sit-ups, the number of times the same action is completed is an indicator reflecting exercise intensity and physical fitness level. Therefore, repetitive sports action counting has broad application scenarios. The intelligent repetitive action counting method can not only record people's physical exercise process in detail, formulate personalized exercise plans for reference, but also help evaluate people's exercise status to provide real-time feedback.

[0003] Existing repetitive action counting methods are mainly divided into two categories: traditional methods and computer vision-based methods.

[0004] Traditional methods mainly include manual counting and sensor-assisted counting. Manual counting requires a dedicated recorder, which is labor-consuming. Moreover, it is often difficult to accurately count some actions with a relatively high frequency, there are counting errors caused by reaction delays, and there may also be counting errors due to the fatigue of the recorder. Sensor-assisted counting methods generally install infrared sensors, pressure sensors, etc. in the sports venue, or let the athletes wear corresponding sensors, and then analyze the data information of the sensors to achieve repetitive action counting. Although this method has high accuracy, the equipment installation is complex, the sensors used for different actions are not the same, and the layout cost is relatively high. In addition, wearing sensors is very likely to affect performance or cause safety accidents.

[0005] Computer vision-based methods provide new solutions to overcome the disadvantages of low efficiency and contact characteristics based on traditional methods. Early repetition counting methods assume that repetitions are periodic and predict rough counting results by estimating the period length. These methods may fail in real-world scenarios, especially in fitness exercise scenarios, because physical exercises do not have a fixed periodicity, and due to the physical limitations of humans themselves or the interference of others, abnormal movements such as pauses may occur within one movement or between two movements. Recently, emerging deep learning-based methods solve the above problems in a data-driven manner through context awareness or temporal correlation modeling, thus achieving repetition counting in general scenarios. Although representative works based on deep learning have achieved good performance, their counting accuracy far from meets the actual application requirements in the physical fitness test scenario. Most existing counting methods are directly executed on videos, and in fitness action videos, people and human movements are the objects of concern. When directly processing videos, background information may interfere with human movement features, resulting in inaccurate feature descriptions and large counting errors. In addition, existing methods focus on global spatial information by taking each frame of the video as a whole, lacking the ability to distinguish local region features with periodic movements, thus making it difficult to identify fine-grained local periodic movements and causing repetition counting failures.

[0006] In summary, the existing repetitive action counting methods have disadvantages such as consuming manpower, insufficient refinement, low accuracy, and weak generality. Summary of the Invention

[0007] To solve the above problems existing in the prior art, the present invention provides a method and device for counting repetitive sports actions based on a multi-scale transformation network. The technical problems to be solved by the present invention are realized through the following technical solutions:

[0008] An embodiment of the present invention provides a method for counting repetitive sports actions based on a multi-scale transformation network, including the steps of:

[0009] S1. Detect the skeletal key points in the original video to obtain the original skeletal sequence, and sample the original skeletal sequence at three scales to obtain three downsampled skeletal sequences;

[0010] S2. Encode the temporal periodicity features of each downsampled skeletal sequence for each skeletal joint point to obtain the temporal autocorrelation matrix of each skeletal joint;

[0011] S3. Perform periodic selection from coarse to fine on the temporal autocorrelation matrix of each skeletal joint to estimate the periodic weight of each skeletal joint, and obtain three weighted fusion period cubes;

[0012] S4. Perform multi-scale periodic fusion on the three weighted fusion period cubes to obtain a multi-scale period cube;

[0013] S5. Based on a multi-scale period transformation network, perform impulse map regression on the multi-scale period cube to obtain an impulse map, where a random frame is labeled for each repeated action in the impulse map, and the true value of the random frame is 1, and the true values of other frames are 0;

[0014] S6. Add all elements in the impulse map to obtain a repeated action count result.

[0015] In an embodiment of the present invention, step S1 includes:

[0016] S11. Use a pose estimation algorithm to perform pose estimation on the athletes in the original video, detect the bone key points, and obtain the original bone sequence;

[0017] S12. Perform equally spaced sampling on the original bone sequence at three scales to obtain the three downsampled bone sequences.

[0018] In an embodiment of the present invention, step S2 includes:

[0019] S21. For each downsampled bone sequence, add local time motion information to the bone joint features of each frame, and form an encoded feature for each bone joint point from the bone joint features of each frame after adding the local time motion information:

[0020]

[0021]

[0022] where F i represents the encoded feature for each bone joint point, represents the bone joint feature of the i-th joint at the t-th frame, represents the 2D position coordinate of the i-th bone joint at the t-th frame, represents the motion offset of the i-th bone joint along the x-axis direction at the t-th frame, represents the motion offset of the i-th bone joint along the y-axis direction at the t-th frame; T k represents the number of frames;

[0023] S22. Calculate the temporal autocorrelation matrix of the i-th bone joint according to the encoded feature for each bone joint point:

[0024]

[0025] Among them, represents the value of the p-th row and q-th column of the matrix M i of M i represents the time autocorrelation matrix of the i-th skeletal joint, represents a real number, represents a similarity metric function, represents the skeletal joint feature of the i-th skeletal joint at the p-th frame, represents the joint feature of the i-th skeletal joint at the q-th frame.

[0026] In an embodiment of the present invention, step S3 includes:

[0027] S31. Based on the prior knowledge of the element numerical size relationship in the time autocorrelation matrix of non-periodic joints, exclude the time autocorrelation matrices of non-periodic joints in the time autocorrelation matrices of all skeletal joints to obtain a rough selection fusion cube;

[0028] S32. Use channel attention to estimate the periodic weights of the time autocorrelation matrices retained in the rough selection fusion cube to obtain the three weighted fusion periodic cubes.

[0029] In an embodiment of the present invention, step S31 includes:

[0030] Stack the time autocorrelation matrices of all skeletal joints layer by layer to obtain a time autocorrelation cube;

[0031] Input the time autocorrelation cube into a two-dimensional average pooling layer and a global max pooling layer connected in sequence to obtain a periodic metric vector;

[0032] Sort the elements in the periodic metric vector from large to small to generate a periodic mask;

[0033] Multiply the time autocorrelation cube by the periodic mask and remove all zero channels in the multiplication result to obtain the rough selection fusion cube.

[0034] In an embodiment of the present invention, step S32 includes:

[0035] Unfold the rough selection fusion cube layer by layer and perform two-dimensional discrete cosine transform on the features of each channel after unfolding to obtain a two-dimensional discrete cosine transform result;

[0036] Merge the frequency components of the two-dimensional discrete cosine transform result and input the merged result into a one-dimensional convolutional layer, and then perform Sigmoid activation on the output of the one-dimensional convolutional layer to obtain channel attention weights;

[0037] Multiply the rough selection fusion cube by the channel attention weight to obtain the three weighted fusion period cubes, where the three weighted fusion period cubes are arranged in ascending order of scale as the first weighted fusion period cube, the second weighted fusion period cube, and the third weighted fusion period cube.

[0038] In an embodiment of the present invention, step S4 includes:

[0039] Downsample the third weighted fusion period cube to the same scale size as the second weighted fusion period cube, then input the first downsampling result into the first two-dimensional convolutional layer, and add the output of the first two-dimensional convolutional layer to the second weighted fusion period cube to obtain an intermediate periodic cube;

[0040] Downsample the intermediate periodic cube to the same scale size as the first weighted fusion period cube, then input the second downsampling result into the second two-dimensional convolutional layer, and add the output of the second two-dimensional convolutional layer to the first weighted fusion period cube to obtain the multi-scale periodic cube.

[0041] In an embodiment of the present invention, step S5 includes:

[0042] Input the multi-scale periodic cube into a convolutional layer to obtain local temporal context information, resulting in a feature cube;

[0043] Perform dimensional transformation on the feature cube and input the dimensional transformation result into a linear projection layer to obtain a feature embedding;

[0044] Add the position encoding to the feature embedding and send the added result into the encoding layer of the Transformer network to obtain an encoding result;

[0045] Input the encoding result into two consecutive fully connected layers to obtain the impulse diagram.

[0046] In an embodiment of the present invention, after step S6, there is further a step of inferring the start time and end time of each action based on the average interval between two adjacent local extrema in the impulse diagram.

[0047] Another embodiment of the present invention provides a repetitive sports action counting device based on a multi-scale transformation network, including: a multi-scale sampling module, a per-skeleton joint point encoding module, a skeleton joint weight estimation module, a multi-scale periodic fusion module, an impulse diagram regression module, and a counting result calculation module, where

[0048] The multi-scale sampling module is used to detect the skeleton key points in the original video to obtain the original skeleton sequence, and sample the original skeleton sequence at three scales to obtain three down-sampled skeleton sequences;

[0049] The per-skeleton joint encoding module is used to perform per-skeleton joint encoding on the time periodic features of each of the down-sampled skeleton sequences to obtain the time autocorrelation matrix of each skeleton joint;

[0050] The skeleton joint weight estimation module is used to perform coarse-to-fine periodic selection on the time autocorrelation matrix of each skeleton joint to estimate the weight of each skeleton joint, obtaining three weighted fusion period cubes;

[0051] The multi-scale periodic fusion module is used to perform multi-scale periodic fusion on the three weighted fusion period cubes to obtain a multi-scale period cube;

[0052] The impulse map regression module is used to perform impulse map regression on the multi-scale period cube based on a multi-scale period transformation network to obtain an impulse map, wherein a random frame is labeled for each repeated action in the impulse map, and the true value of the random frame is 1, and the true values of other frames are 0;

[0053] The counting result calculation module is used to add up all the elements in the impulse map to obtain the repeated action counting result.

[0054] Compared with the prior art, the beneficial effects of the present invention are:

[0055] 1. The method for counting repetitive sports actions of the present invention first performs per-skeleton joint encoding on the time periodic features of the down-sampled skeleton sequence, and then performs coarse-to-fine periodic selection on the time autocorrelation matrix of each skeleton joint, so that it can focus on the time changes occurring in the local space, especially those places where repeated movements have occurred, enhancing the ability to capture motion information and having a high prediction accuracy. Therefore, this method has the advantages of focusing on fine-grained actions in space and high counting accuracy.

[0056] 2. The method of the present invention performs impulse map regression on the multi-scale period cube. The impulse map is used to predict the probability of an action occurring in a time series, maintaining the interpretability of the model. At the same time, using the random frame of each repeated action as a label and avoiding using the start frame and end frame as labels, thus reducing the cost of data annotation; therefore, the prediction result of this method has interpretability while having a low label annotation cost.

[0057] 3. The method of the present invention samples the original bone sequence at three scales and performs multi-scale periodic fusion on three weighted fusion period cubes, which can better handle the sports repetitive action counting of various lengths of videos, be applicable to actions with large changes in time length, and have a wide range of usage scenarios.

[0058] 4. The method of the present invention encodes the periodic patterns contained in bone joints using a temporal autocorrelation matrix. The use of the temporal autocorrelation matrix can achieve periodic encoding of action classes. Since subsequent processing is performed in the self-similarity space rather than in the high-dimensional action feature space, it has good generality for unseen action classes and has the advantages of strong generalization and being applicable to different actions.

[0059] 5. The method of the present invention uses the human body motion bone sequence as input. The human body bone is a high-level representation with relatively low complexity, which can accurately reflect the human body motion in the fitness action scene. In addition, the human body bone can be easily obtained through advanced pose estimation algorithms, and the bone sequence is a lightweight data modality relative to videos, which makes the model smaller and the inference speed faster when used in the model.

[0060] 6. The method of the present invention is a non-contact detection method that counts repetitive actions through computer vision, with high accuracy, strong anti-interference ability, and high safety. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] Figure 1 It is a schematic flowchart of a method for counting repetitive sports actions based on a multi-scale transformation network provided by an embodiment of the present invention;

[0062] Figure 2 It is a network structure diagram of a method for counting repetitive sports actions based on a multi-scale transformation network provided by an embodiment of the present invention;

[0063] Figures 3a - 3d It is the test result of the present invention on the FitnessRep dataset;

[0064] Figures 4a - 4d It is the test result of the present invention on the RepCount dataset;

[0065] Figures 5a - 5d It is the test result of the present invention on the UCFRep dataset. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0066] The following further describes the present invention in detail with reference to specific embodiments, but the embodiments of the present invention are not limited thereto.

[0067] Embodiment 1

[0068] Please refer to Figure 1 and Figure 2, Figure 1 It is a flowchart of a method for counting repetitive sports actions based on a multi-scale transformation network provided by an embodiment of the present invention. Figure 2 It is a network structure diagram of a method for counting repetitive sports actions based on a multi-scale transformation network provided by an embodiment of the present invention.

[0069] As Figure 2 shown, the method for counting repetitive sports actions based on a multi-scale transformation network includes five parts: multi-scale sampling, joint-by-joint encoding and selection (JES), multi-scale periodic fusion, impulse map regression, and calculating the counting result. Among them, joint-by-joint encoding and selection include two parts: 1) the temporal autocorrelation matrix of each skeletal joint point, that is, encoding the temporal periodic features for each skeletal joint point; 2) coarse-to-fine periodic selection for estimating the periodic weight of each joint.

[0070] The method for counting repetitive sports actions based on a multi-scale transformation network specifically includes the following steps:

[0071] S1. Detect the skeletal key points in the original video to obtain the original skeletal sequence, and perform sampling at three scales on the original skeletal sequence to obtain three downsampled skeletal sequences. Specifically, it includes the following steps:

[0072] S11. Use a pose estimation algorithm to perform pose estimation on the athlete in the original video, detect the skeletal key points, and obtain the original skeletal sequence.

[0073] Specifically, for a video with T frames, use the MediaPipe pose estimation algorithm to perform pose estimation on the athlete in the video, detect the skeletal key points, and obtain the original skeletal sequence S = {s1, s2,..., s T}, where the skeletal model of each frame has N joint points.

[0074] S12. Perform equidistant sampling on the original skeletal sequence at three scales to obtain the three downsampled skeletal sequences.

[0075] Specifically, perform equidistant sampling on the original skeletal sequence S at three scales to obtain three downsampled skeletal sequences S k (k = 1, 2, 3). Finally, use the MinMax method to scale the data of the downsampled skeletal sequence so that its value is between 0 and 1.

[0076] In this embodiment, the three scales refer to three different numbers of frames. As Figure 2As shown by the multi-scale sampling in, in this embodiment, the MediaPipe pose estimation algorithm is used to detect the original skeleton sequence S of N×T×C, where N represents the number of skeleton points, T represents the time scale, and C represents the number of channels; then, the original skeleton sequence S is sampled at three scales to obtain the downsampled skeleton sequences of N×64×C, N×96×C, and N×128×C.

[0077] This embodiment uses the human motion skeleton sequence as the input. The human skeleton is a high-level representation with relatively low complexity, which can accurately reflect the human motion in the fitness action scene; in addition, the human skeleton is easily obtained through advanced pose estimation algorithms, and the skeleton sequence is a lightweight data modality relative to the video, making the model smaller and the inference speed faster when used in the model.

[0078] S2. Encode the time periodicity features of each of the downsampled skeleton sequences for each skeleton joint point to obtain the time autocorrelation matrix of each skeleton joint. As Figure 2 shown by the per-joint point time autocorrelation matrix in, step S2 specifically includes the steps:

[0079] S21. For each downsampled skeleton sequence, add local time motion information to the skeleton joint features in each frame, and form the encoded features for each skeleton joint point from the skeleton joint features after adding the local time motion information.

[0080] Specifically, for each downsampled skeleton sequence S with a fixed number of frames T k (k = 1, 2, 3), each skeleton joint sequence can be expressed as k (k = 1, 2, 3), each skeleton joint sequence can be expressed as where is the 2D position coordinate of the i-th joint at the t-th frame. In addition to the spatial position information, add local time motion information to the skeleton joint features in each frame, and then the joint features in each frame after adding the local time motion information are represented as follows:

[0081]

[0082] where, represents the skeleton joint feature of the i-th joint at the t-th frame, represents the 2D position coordinate of the i-th skeleton joint at the t-th frame, represents the motion offset of the i-th skeleton joint along the x-axis direction at the t-th frame, represents the motion offset of the i-th skeleton joint along the y-axis direction at the t-th frame, and are calculated as follows:

[0083]

[0084]

[0085] Among them, T k represents the number of frames.

[0086] Then, the encoded features of each bone joint point are formed by the joint features of each frame after adding local time motion information:

[0087]

[0088] Among them, F i represents the encoded features of each bone joint point.

[0089] S22. Calculate the temporal autocorrelation matrix of the i-th bone joint according to the encoded features of each bone joint point.

[0090] Specifically, after obtaining the encoded features F i of each bone joint point, the temporal autocorrelation matrix M i of the i-th bone joint is calculated as follows:

[0091]

[0092] Among them, represents the value of the p-th row and q-th column of the matrix M i , M i represents the temporal autocorrelation matrix of the i-th bone joint, represents a real number, represents a similarity metric function, represents the bone joint feature of the i-th bone joint at the p-th frame, represents the joint feature of the i-th bone joint at the q-th frame.

[0093] The temporal autocorrelation matrix serves as the information bottleneck of the network. Subsequent processing is performed in the self-similarity space rather than in the high-dimensional action feature space. The use of the temporal autocorrelation matrix realizes the periodic encoding of actions of unknown classes.

[0094] In this embodiment, the temporal autocorrelation matrix is used to encode the periodic patterns contained in the bone joints. The use of the temporal autocorrelation matrix can realize the periodic encoding of action classes. Since subsequent processing is performed in the self-similarity space rather than in the high-dimensional action feature space, it has good universality for unseen action classes and has the advantages of strong generalization and applicability to different actions.

[0095] S3. Perform coarse-to-fine periodic selection on the temporal autocorrelation matrix of each bone joint to estimate the periodic weight of each bone joint, and obtain three weighted fusion period cubes.

[0096] Generally speaking, not all joints are involved in repetitive movements. The time autocorrelation matrix of the joints that remain stationary does not contribute to the subsequent inference of the network and may even become interfering noise. Therefore, in this embodiment, joint selection is performed according to its periodicity, and comprehensive strong periodic features are generated, which cover the periodic features of all skeletal joints. The coarse-to-fine periodic selection mainly estimates the periodic weights of each joint, which includes two parts: coarse selection and fine selection. Therefore, as shown in the coarse-to-fine periodic selection in Figure 2 Step S3 specifically includes the following steps:

[0097] S31. Based on the prior knowledge of the magnitude relationship of the elements in the time autocorrelation matrix of non-periodic joints, exclude the time autocorrelation matrix of non-periodic joints in the time autocorrelation matrices of all skeletal joints to obtain a coarsely selected fusion cube.

[0098] Specifically, this step is the coarse selection part. Based on the prior knowledge that most of the elements in the time autocorrelation matrix of non-periodic joints are small values, the time autocorrelation matrix of non-periodic joints can be roughly excluded.

[0099] The specific implementation of the coarse selection part is as follows:

[0100] Stack the time autocorrelation matrices M of all skeletal joints i layer by layer to obtain a time autocorrelation cube P o . Input the time autocorrelation cube P o into a two-dimensional average pooling layer (AP) and a global max pooling layer (GMP) connected in sequence to obtain a periodicity metric vector D p . Sort the elements in the periodicity metric vector D p from largest to smallest to generate a periodicity mask D m ; among them, in the periodicity mask D m , the elements corresponding to the top half of the periodicity metric rankings are 1, and the other elements are 0. Multiply the time autocorrelation cube P o by the periodicity mask D m , and remove all zero channels in the multiplication result to obtain a coarsely selected fusion cube P c . The number of channels of the coarsely selected fusion cube P c is half of the original cube P o .

[0101] It should be noted that each skeletal joint corresponds to a time autocorrelation matrix, and each downsampled skeletal sequence corresponds to multiple skeletal joints. Therefore, stacking the time autocorrelation matrices M of all skeletal joints i layer by layer means stacking multiple skeletal joint time autocorrelation matrices in order (for example, from bottom to top).

[0102] S32. Estimate the periodic weights of the time autocorrelation matrix retained in the coarsely selected fusion cube using channel attention to obtain the three weighted fusion period cubes.

[0103] Specifically, this part is the fine selection part. The fine selection part uses channel attention to estimate the periodic weights of the time autocorrelation matrix retained in the coarsely selected part.

[0104] The specific implementation of the fine selection part is as follows:

[0105] First, layer-by-layer expand the coarsely selected fusion cube P c obtained from the coarsely selected part, and perform two-dimensional discrete cosine transform (DCT) on the features of each channel after expansion to obtain the two-dimensional discrete cosine transform result. Then, merge the frequency components of the two-dimensional discrete cosine transform result, input the merged result into a one-dimensional convolutional layer, and then perform Sigmoid activation on the output of the one-dimensional convolutional layer to obtain the channel attention weight W p ; where a one-dimensional convolutional layer is used to replace the fully connected layer to better fuse the information of the input vector and complete cross-channel information interaction. Finally, multiply the coarsely selected fusion cube P c by the channel attention weight W p to obtain the weighted fusion period cube.

[0106] Furthermore, after the joint-by-joint encoding and selection part, three weighted fusion period cubes C k (k = 1, 2, 3) corresponding to the three downsampled skeleton sequences S k (k = 1, 2, 3) are obtained, and they are arranged in ascending order of scale as the first weighted fusion period cube, the second weighted fusion period cube, and the third weighted fusion period cube.

[0107] It should be noted that layer-by-layer expanding the coarsely selected fusion cube P c obtained from the coarsely selected part means sequentially expanding the coarsely selected fusion cube P c to obtain the features of each skeleton joint point. Therefore, the features of each channel after expansion refer to the features of each skeleton joint point.

[0108] In the human fitness scenario, there are usually many fine repetitive actions, such as wrist rotation and dumbbell shrugging. These repetitive actions only involve local body parts. To solve this problem, this embodiment proposes the joint point-by-joint encoding and selection part described in steps S2 and S3, that is, extracting periodic features in units of human skeletal joints rather than single-frame images, and generating weighted strong periodic features according to joint point-by-joint periodic selection. Based on this, the method first encodes the time period features of the downsampled skeletal sequence at each skeletal joint point, and then performs coarse-to-fine periodic selection on the time autocorrelation matrix of each skeletal joint, so as to focus on the time changes occurring in the local space, especially those places where repetitive movements have occurred, enhancing the ability to capture motion information and having a high prediction accuracy. Therefore, the method has the advantages of focusing on fine-grained actions in space and having a high counting accuracy.

[0109] S4. Perform multi-scale periodic fusion on the three weighted fusion period cubes to obtain a multi-scale period cube.

[0110] Considering the challenges brought by long videos in the real scenario, this embodiment designs a multi-scale periodic fusion module to better handle the counting of sports repetitive actions in videos of different lengths.

[0111] Specifically, first, downsample the third weighted fusion period cube C 3 to the same scale size as the second weighted fusion period cube C 2 . Then input the first downsampling result into the first two-dimensional convolutional layer, and add the output of the first two-dimensional convolutional layer to the second weighted fusion period cube C 2 to obtain an intermediate periodic cube C t .

[0112] Then, downsample the intermediate periodic cube C t to the same scale size as the first weighted fusion period cube C 1 . Then input the second downsampling result into the second two-dimensional convolutional layer, and add the output of the second two-dimensional convolutional layer to the first weighted fusion period cube C 1 to obtain a multi-scale period cube C m .

[0113] In this embodiment, the first two-dimensional convolutional layer and the second two-dimensional convolutional layer before element-wise summation are designed to eliminate the confounding effects caused by downsampling and enhance the learning ability of the network.

[0114] In this embodiment, sampling the original skeletal sequence at three scales and performing multi-scale periodic fusion on the three weighted fusion period cubes can better handle the counting of sports repetitive actions in videos of various lengths, be applicable to actions with large time length variations, and have a wide range of usage scenarios.

[0115] S5. Perform pulse map regression on the multi-scale periodic cube based on the multi-scale periodic transformation network to obtain a pulse map, where the pulse map annotates a random frame for each repeated action, and the true value of the random frame is 1, and the true values of other frames are 0.

[0116] Specifically, to make the counting result interpretable, the method proposed in this embodiment can not only obtain the accurate number of counts, but also estimate the position of each counting unit. The pulse map regression part mainly predicts the probability map of action occurrence in the time series, reduces the cost of data annotation, and at the same time retains interpretability. The pulse map only needs to annotate a random frame of each repeated action, set the true value of the annotated random frame to 1, and the true values of other frames to 0. The pulse map only encodes one moment when the action occurs. This is a rough action probability map, but it can be used for repeated action counting.

[0117] In this embodiment, a multi-scale periodic transformation network is used for pulse map regression, where the multi-scale periodic transformation network is constructed based on the Transformer network. The specific implementation steps of using the multi-scale periodic transformation network for pulse map regression are as follows: First, input the multi-scale periodic cube C m into a convolutional layer to obtain local temporal context information and get a feature cube; then, perform a dimension transformation on the feature cube, and input the dimension transformation result into a linear projection layer to obtain a feature embedding; add a position encoding to the feature embedding, and input the added result into the encoding layer of the Transformer network to obtain an encoding result; input the encoding result into two fully connected layers connected in sequence to obtain a pulse map.

[0118] Most existing repeated counting methods only output the number of repeated actions, which lacks interpretability because it may happen that the total number of repeated actions is correct, but the position of each counting unit may be wrong. In this embodiment, by performing pulse map regression on the multi-scale periodic cube, the pulse map is used to predict the probability of action occurrence in the time series, maintaining the interpretability of the model. At the same time, using the random frame of each repeated action as a label and avoiding using the start frame and end frame as labels reduces the cost of data annotation; therefore, the prediction result of this method has interpretability while having a low label annotation cost.

[0119] In addition, the local maximum value in the pulse map indicates that an action has occurred nearby. The start and end positions of each action can be estimated through post-processing methods, providing more refined information about the entire motion process, and providing richer information for subsequent personalized exercise plan formulation or human motion state assessment.

[0120] S6. Add all the elements in the pulse diagram to obtain the repetitive action count result.

[0121] Specifically, the local maximum in the pulse diagram indicates that an action has occurred nearby. Therefore, by adding all the elements in the pulse diagram, the repetitive action count result can be obtained.

[0122] Furthermore, the average interval between two adjacent local extrema in the pulse diagram gives a clue to an approximate period. Therefore, the start time and end time of each action can be inferred based on the average interval between two adjacent local extrema.

[0123] The present invention is a non-contact detection method that counts repetitive actions through computer vision, with high accuracy, strong anti-interference ability, and high safety.

[0124] In summary, the method for counting repetitive sports actions based on the multi-scale transformation network in this embodiment can avoid manpower consumption and complex sensor installation, improve the accuracy of repetitive action counting, and can be applied to daily fitness exercises and physical fitness tests, such as the intelligent counting and standardized evaluation of pull-ups with a wide range of applicable scenarios.

[0125] In this embodiment, the following test experiments are further carried out to illustrate the effect of the method for counting repetitive sports actions based on the multi-scale transformation network.

[0126] I. Test conditions:

[0127] The test data are the FitnessRep dataset, the RepCount dataset, and the UCFRep dataset. The FitnessRep dataset is a collection of 1662 physical fitness exercise videos of tested personnel on a university campus, including 12 types of actions, including 2 types of physical fitness tests (pull-ups and sit-ups) and 10 types of fitness actions.

[0128] The simulation test platform is a PC with a 12th Gen Intel(R) Core(TM) i7-12700KF × 20 CPU at 3.6 GHz, 64 GB of memory, and an Nvidia RTX3090Ti graphics card. The test platform is the Ubuntu20.04.4 LTS operating system, using the Pytorch deep learning framework and implemented in the Python language.

[0129] Evaluation metrics: Five metrics are used to evaluate the performance of the algorithm as follows:

[0130] Out-of-Bounds One (OBO) error less than 1: If the predicted count is within one count of the true value, the predicted value is considered correctly classified; otherwise, it is considered misclassified. It represents the repetitive counting error rate of the entire dataset.

[0131]

[0132] Among them, OBO represents the counting error less than 1, N represents the number of data, represents the true counting value of the i-th data, c i represents the predicted counting value of the i-th data.

[0133] Mean Absolute Error (MAE) of counting: This error metric measures the absolute difference between the true count and the predicted count, and then normalizes it by dividing by the true count. The MAE error is the average of the normalized absolute differences for the entire dataset.

[0134]

[0135] To better evaluate the error situation of the entire dataset, in this embodiment, the OBO counting error is generalized. OBX represents that if the predicted count is within X counts of the true value, the predicted value is considered correctly classified; otherwise, it is considered misclassified. In this embodiment, OB3 and OB5 errors are used as extensions of the counting error.

[0136]

[0137] Among them, X represents the number of X counts.

[0138] Absolute Mean Error (AME) of counting: This error metric measures the absolute difference between the true count and the predicted count. The AME error is the average of the absolute differences for the entire dataset.

[0139]

[0140] II. Test Content and Results:

[0141] Under the above test conditions, the method of this embodiment is used to process the FitnessRep dataset, RepCount dataset, and UCFRep dataset to obtain all predicted values, calculate the number of actions for each sample, and obtain OBO, MAE, OB3, OB5, and AME errors. The test results are shown in Table 1, where the * symbol in RepCount(part-A*) and UCFRep(*) represents the data remaining after removing video examples with failed pose estimation in the dataset.

[0142] Table 1 Error Statistics of the Method of this Embodiment

[0143] MAE↓ OBO↑ OB3↑ OB5↑ AME↓ FitnessRep 0.051 0.930 0.987 0.996 0.599 RepCount(part - A*) 0.178 0.619 0.836 0.891 3.228 UCFRep(*) 0.285 0.618 0.763 0.909 2.090

[0144] As can be seen from Table 1, the AME error of the method in this embodiment in the real-world collected sports fitness dataset FitnessRep is less than 1, within the error range of practical application, indicating that the counting method in this embodiment can be applied to actual sports fitness scenarios.

[0145] See Figures 3a - 3d , Figures 3a - 3d which is the test result of the present invention in the FitnessRep dataset. Figures 3a - 3d In it, the first row represents the video; the second row represents the true value of the labeled action pulse diagram, with a random frame labeled in each repeated action; the third row represents the action pulse diagram predicted in this embodiment. From Figures 3a - 3d it can be concluded that whether it is an action with periodic changes (such as Figure 3c sit-ups, Figure 3d pull-ups), an action with a large number of actions (such as Figure 3a backward squeeze, Figure 3c sit-ups), or an action with local changes (such as Figure 3b left stretch), the prediction results of this embodiment can well fit the local maximum values of the pulse diagram. The method of this embodiment produces accurate results while maintaining interpretability.

[0146] See Figures 4a - 4d , Figures 4a - 4d which is the test result of the present invention in the RepCount dataset. Figures 4a - 4d In it, the first row represents the video; the second row represents the true value of the labeled action pulse diagram, with a random frame labeled in each repeated action; the third row represents the action pulse diagram predicted in this embodiment. From Figures 4a - 4d it can be concluded that the prediction results of this embodiment well fit the true values of the pulse diagram and produce accurate counts in sports actions, as shown in Figure 4a and Figure 4b . For actions where the human skeleton cannot be fully extracted, the prediction results of this embodiment are basically consistent with the true values, and the counting errors are within an acceptable range, as shown in Figure 4c . The prediction results of the method of this embodiment for the "playing the cello" action are slightly deviated, probably because the repetition of this action is not significant, as shown in Figure 4d .

[0147] See Figures 5a - 5d , Figures 5a - 5d which is the test result of the present invention in the UCFRep dataset. Figures 5a - 5d In it, the first row represents the video; the second row represents the true value of the labeled action pulse diagram, with a random frame labeled in each repeated action; the third row represents the action pulse diagram predicted by the present invention. As shown in Figure 5aAs shown, for actions with obvious periodic changes, the pulse diagram predicted in this embodiment well fits each peak of the true value, resulting in correct counting results. For Figure 5b , Figure 5c and Figure 5d in the action samples, the prediction of this embodiment has a slight deviation from the true value near the peak, but the counting result is within the error range.

[0148] In summary, the repetitive sports action counting method based on the multi-scale transformation network described in this embodiment is accurately tested, and the counting result is interpretable.

[0149] Embodiment 2

[0150] On the basis of Embodiment 1, this embodiment further provides a repetitive sports action counting device based on a multi-scale transformation network. The device includes a multi-scale sampling module, a per-bone joint point encoding module, a bone joint weight estimation module, a multi-scale periodic fusion module, a pulse diagram regression module, and a counting result calculation module.

[0151] Specifically, the multi-scale sampling module is used to detect the bone key points in the original video to obtain the original bone sequence, and perform sampling on the original bone sequence at three scales to obtain three downsampled bone sequences. The per-bone joint point encoding module is used to encode the time periodic features of each downsampled bone sequence for each bone joint to obtain the time autocorrelation matrix of each bone joint. The bone joint weight estimation module is used to perform coarse-to-fine periodic selection on the time autocorrelation matrix of each bone joint to estimate the weight of each bone joint, obtaining three weighted fusion period cubes. The multi-scale periodic fusion module is used to perform multi-scale periodic fusion on the three weighted fusion period cubes to obtain a multi-scale period cube. The pulse diagram regression module is used to perform pulse diagram regression on the multi-scale period cube based on the multi-scale period transformation network to obtain a pulse diagram, where a random frame is annotated for each repeated action in the pulse diagram, and the true value of the random frame is 1, and the true value of other frames is 0. The counting result calculation module is used to add all the elements in the pulse diagram to obtain the repeated action counting result.

[0152] Furthermore, the bone joint weight estimation module includes a coarse selection module and a fine selection module. Among them, the coarse selection module is used to exclude the time autocorrelation matrix of non-periodic joints in the time autocorrelation matrix of all bone joints based on the prior knowledge of the element numerical size relationship in the time autocorrelation matrix of non-periodic joints, obtaining a coarse selection fusion cube. The fine selection module is used to estimate the periodic weight of the time autocorrelation matrix retained in the coarse selection fusion cube using channel attention to obtain three weighted fusion period cubes.

[0153] For the specific implementation steps and beneficial effects achieved by each of the above modules, please refer to Embodiment 1, which will not be elaborated herein.

[0154] The above content is a further detailed description of the present invention in combination with specific preferred embodiments. It cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention pertains, without departing from the concept of the present invention, several simple deductions or substitutions can be made, which should all be regarded as falling within the protection scope of the present invention.

Claims

1. A method for counting repetitive sports actions based on a multi-scale transformation network, characterized in that, Including the steps: S1. Detect the skeletal key points in the original video to obtain the original skeletal sequence, and perform sampling at three scales on the original skeletal sequence to obtain three downsampled skeletal sequences; S2. Encode the temporal periodicity features of each of the downsampled skeletal sequences for each skeletal joint point to obtain the temporal autocorrelation matrix of each skeletal joint; Step S2 includes: S21. For each downsampled skeletal sequence, add local temporal motion information to the skeletal joint features of each frame, and form the encoded features for each skeletal joint point from the skeletal joint features of each frame after adding the local temporal motion information; Among them, F i represents the encoded feature of each skeletal joint point, and f t (i) represents the skeletal joint feature of the i-th joint at the t-th frame, represents the 2D position coordinate of the i-th skeletal joint at the t-th frame, represents the motion offset of the i-th skeletal joint in the x-axis direction at the t-th frame, represents the motion offset of the i-th skeletal joint in the y-axis direction at the t-th frame; T k represents the number of frames; S22. Calculate the temporal autocorrelation matrix of the i-th skeletal joint according to the encoded features for each skeletal joint point; Among them, represents the value of the p-th row and q-th column of the matrix M i and M i represents the temporal autocorrelation matrix of the i-th skeletal joint, represents a real number, represents a similarity metric function, represents the skeletal joint feature of the i-th skeletal joint at the p-th frame, represents the joint feature of the i-th skeletal joint at the q-th frame; S3. Perform a coarse-to-fine periodic selection on the temporal autocorrelation matrix of each skeletal joint to estimate the periodic weight of each skeletal joint, and obtain three weighted fusion periodic cubes; S4. Perform multi-scale periodic fusion on the three weighted fusion periodic cubes to obtain a multi-scale periodic cube; S5. Perform impulse map regression on the multi-scale periodic cube based on a multi-scale periodic transformation network to obtain an impulse map, wherein a random frame is labeled for each repeated action in the impulse map, and the true value of the random frame is 1, and the true values of other frames are 0; S6. Add all the elements in the impulse map to obtain the repeated action count result.

2. The method for counting repetitive sports actions based on a multi-scale transformation network according to claim 1, characterized in that Step S1 includes: S11. Use a pose estimation algorithm to perform pose estimation on the athletes in the original video, detect the skeletal key points, and obtain the original skeletal sequence; S12. Perform equidistant sampling at three scales on the original skeletal sequence to obtain the three downsampled skeletal sequences.

3. The method for counting repetitive sports actions based on a multi-scale transformation network according to claim 1, characterized in that, Step S3 includes: S31. Based on the prior knowledge of the magnitude relationship of the elements in the temporal autocorrelation matrix of non-periodic joints, exclude the temporal autocorrelation matrices of non-periodic joints in the temporal autocorrelation matrices of all skeletal joints to obtain a coarse selection fusion cube; S32. Use channel attention to estimate the periodic weight of the remaining temporal autocorrelation matrices in the coarse selection fusion cube to obtain the three weighted fusion periodic cubes.

4. The method for counting repetitive sports actions based on a multi-scale transformation network according to claim 3, wherein, Step S31 includes: Stack the temporal autocorrelation matrices of all skeletal joints layer by layer to obtain a temporal autocorrelation cube; Input the temporal autocorrelation cube into a two-dimensional average pooling layer and a global maximum pooling layer connected in sequence to obtain a periodicity metric vector; Sort the elements in the periodicity metric vector from large to small to generate a periodicity mask; Multiply the temporal autocorrelation cube by the periodicity mask, and remove all channels with zero results in the multiplication result to obtain the coarse selection fusion cube.

5. The method for counting repetitive sports actions based on a multi-scale transformation network according to claim 3, wherein, Step S32 includes: Unfold the coarse selection fusion cube layer by layer, and perform two-dimensional discrete cosine transform on the features of each channel after unfolding to obtain a two-dimensional discrete cosine transform result; Merge the frequency components of the two-dimensional discrete cosine transform result, input the merged result into a one-dimensional convolutional layer, and then perform Sigmoid activation on the output of the one-dimensional convolutional layer to obtain the channel attention weight; Multiply the rough selection fusion cube by the channel attention weights to obtain the three weighted fusion period cubes, where the three weighted fusion period cubes are arranged in ascending order of scale as the first weighted fusion period cube, the second weighted fusion period cube, and the third weighted fusion period cube.

6. The method for counting repetitive sports actions based on a multi-scale transformation network according to claim 5, characterized in that Step S4 includes: Downsample the third weighted fusion period cube to the same scale size as the second weighted fusion period cube, then input the first downsampling result into the first two-dimensional convolutional layer, and add the output of the first two-dimensional convolutional layer to the second weighted fusion period cube to obtain an intermediate periodic cube; Downsample the intermediate periodic cube to the same scale size as the first weighted fusion period cube, then input the second downsampling result into the second two-dimensional convolutional layer, and add the output of the second two-dimensional convolutional layer to the first weighted fusion period cube to obtain the multi-scale period cube.

7. The method for counting repetitive sports actions based on a multi-scale transformation network according to claim 1, characterized in that Step S5 includes: Input the multi-scale period cube into a convolutional layer to obtain local temporal context information, resulting in a feature cube; Perform dimensional transformation on the feature cube and input the dimensional transformation result into a linear projection layer to obtain a feature embedding; Add the position encoding to the feature embedding and send the addition result into the encoding layer of the Transformer network to obtain an encoding result; Input the encoding result into two consecutive fully connected layers to obtain the impulse graph.

8. The method for counting repetitive sports actions based on a multi-scale transformation network according to claim 1, characterized in that After step S6, there is also a step: infer the start time and end time of each action based on the average interval between two adjacent local extrema in the impulse graph.

9. A repetitive sports action counting device based on a multi-scale transformation network, characterized in that, Includes: A multi-scale sampling module, a per-bone joint encoding module, a bone joint weight estimation module, a multi-scale periodic fusion module, an impulse graph regression module, and a counting result calculation module, where The multi-scale sampling module is used to detect the bone key points in the original video to obtain the original bone sequence, and perform sampling on the original bone sequence at three scales to obtain three downsampled bone sequences; The per-bone joint encoding module is used to perform per-bone joint encoding on the temporal periodic features of each downsampled bone sequence to obtain the temporal autocorrelation matrix of each bone joint; includes: For each downsampled bone sequence, add local temporal motion information to the bone joint features of each frame, and form the encoding features of per-bone joint from the bone joint features of each frame after adding the local temporal motion information; Among them, F i represents the encoded feature of each bone joint point, and f t (i) represents the bone joint feature of the i-th joint at the t-th frame, represents the 2D position coordinate of the i-th bone joint at the t-th frame, represents the motion offset of the i-th bone joint along the x-axis direction at the t-th frame, represents the motion offset of the i-th bone joint along the y-axis direction at the t-th frame; T k represents the number of frames; Calculate the temporal autocorrelation matrix of the i-th bone joint according to the encoding features of per-bone joint; Among them, represents the value of the p-th row and q-th column of the matrix M i and M i represents the temporal autocorrelation matrix of the i-th skeletal joint, represents a real number, represents a similarity metric function, represents the skeletal joint feature of the i-th skeletal joint at the p-th frame, represents the joint feature of the i-th skeletal joint at the q-th frame; The bone joint weight estimation module is used to perform coarse-to-fine periodic selection on the temporal autocorrelation matrix of each bone joint to estimate the weight of each bone joint, obtaining three weighted fusion period cubes; The multi-scale periodic fusion module is used to perform multi-scale periodic fusion on the three weighted fusion period cubes to obtain a multi-scale period cube; The pulse diagram regression module is used to perform pulse diagram regression on the multi-scale periodic cube based on the multi-scale periodic transformation network to obtain a pulse diagram, wherein a random frame is labeled for each repeated action in the pulse diagram, and the true value of the random frame is 1, and the true values of other frames are 0; The counting result calculation module is used to add all elements in the pulse diagram to obtain a repeated action counting result.

Citation Information

Patent Citations

  • Dynamic gesture recognition method based on multi-view three-dimensional skeleton information fusion

    CN114612938A

  • Human skeleton behavior recognition method and system based on multi-scale residual image convolutional network

    CN114743273A