A barbell detection and trajectory tracking method and system

Through the improved barbell detection network and CoTracker model, the problem of low barbell object detection accuracy in the prior art is solved, high-precision object detection and tracking are achieved, and barbell motion parameters are accurately calculated.

CN119579652BActive Publication Date: 2025-06-13CHINA INST OF SPORT SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411740614.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-29
Publication Date
2025-06-13
Estimated Expiration
2044-11-29

AI Technical Summary

Technical Problem

The prior art has problems such as complex background, foreground occlusion, and rapid movement in barbell target detection and tracking, resulting in low detection accuracy and prone to missed detection and missed detection.

Method used

A barbell detection network based on YOLOv8n is adopted, combined with the CBFormer backbone network and the TriFPN neck, multi-directional cross-scale features of left and right barbell pieces are fusion, and the target tracking is used to calculate the barbell motion parameters.

Benefits of technology

It improves the accuracy and speed of barbell target detection, reduces missed detection and misdetection, can maintain high-precision target tracking performance under complex backgrounds and rapid movement, and accurately calculates barbell motion parameters.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119579652B_ABST
    Figure CN119579652B_ABST
Patent Text Reader

Abstract

The present invention discloses a barbell detection and trajectory tracking method and system, including: sampling the original video frames I with a frame rate of f rate to obtain a set of sampled images I'; inputting the barbell detection network improved based on YOLOv8n to detect the left and right barbell plates, and obtaining the center point coordinates (b x , b y ) and the length and width dimensions (b h , b w ) of the target bounding box. The barbell detection network includes a backbone network based on CBFormer, a neck based on TriFPN, and a detection head; generating a query vector Q tr for the tracking point according to the start frame τ of the first detected target; using the CoTracker model to obtain the position vector Tr of the tracking point frame by frame in the video according to the generated query vector Q tr ; calculating the barbell motion parameters according to the position vector Tr and the frame rate f rate of the original video frames. The present invention can detect and track the trajectory of the barbell.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of barbell target detection and target tracking based on computer vision, and particularly to a barbell detection and trajectory tracking method and system. Background Art

[0002] In weightlifting, accurately detecting the barbell target from the video stream and tracking its movement trajectory can effectively assist athletes or fitness enthusiasts in optimizing their training movements and prevent sports injuries. The barbell target detection and tracking task has problems such as complex backgrounds, foreground occlusion, and fast movement. Previous technical solutions mainly rely on object detection algorithms such as YOLO, and there may be certain accuracy bottlenecks when detecting barbells, resulting in phenomena such as missed detections and misdetections. Summary of the Invention

[0003] The present invention provides a barbell detection and trajectory tracking method and system to solve the problems existing in the above-mentioned prior art. The technical solutions are as follows:

[0004] On the one hand, a barbell detection and trajectory tracking method is provided, including:

[0005] S1. Perform a sampling operation on the original video frame I with a frame rate of f rate to obtain a set of multiple sampled images I';

[0006] S2. Input the set of multiple images I' into a barbell detection network improved based on YOLOv8n to detect the left and right barbell plates, and obtain the center point coordinates (b x , b y ) and the length and width dimensions (b h , b w ) of the target bounding box. The barbell detection network includes a backbone network based on CBFormer, a neck based on the ternary feature pyramid network TriFPN, and a detection head;

[0007] S3. According to the start frame τ of the first detected target, generate a query vector Q tr for the tracking points. The query vector Q tr contains the initial positions and start frames of the center points of the left and right barbell plates to be tracked;

[0008] S4. Use the CoTracker model to obtain the position vector Tr of the tracking points frame by frame in the video according to the generated query vector Q tr ;

[0009] S5. Calculate the barbell movement parameters according to the position vector Tr and the frame rate f rate of the original video frame.

[0010] Optionally, the CBFormer-based backbone network is used to extract image features. It adopts a four-layer pyramid structure. An image patch embedding module is used in the first layer, and image patch merging modules are used in the second to fourth layers to reduce the input spatial resolution and increase the number of channels. Then, in each layer, after the image patch embedding module and the image patch merging module, the CBFormer block is used for feature transformation.

[0011] Optionally, for the CBFormer block, first, a 3×3 depthwise convolution is used to implicitly encode relative position information, and the output is skip-connected through a residual connection. Subsequently, layer normalization is performed. Then, a cascaded bi-level routing attention (CBRA) module is applied for cross-position relationship modeling, and the output is added through a residual connection. Subsequently, layer normalization is performed again. Finally, a two-layer multi-layer perceptron (MLP) is used for position embedding, and the output result is also added through a residual connection.

[0012] Optionally, the cascaded bi-level routing attention (CBRA) module obtains a more complete feature representation through the cascaded operation of multiple bi-level routing attention (BRA). Each bi-level routing attention (BRA) can achieve a sparse attention pattern in an adaptive query dynamic manner. The key idea is to perform attention calculation sequentially at two levels in the coarse-grained region and the fine-grained region. At the coarse-grained region level, the most relevant key-value pairs are selected, which is achieved by constructing and pruning a directed graph at the region level. Then, token-to-token attention is applied in the combination of the fine-grained region. From the perspective of single-input single-head self-attention, the calculation process of BRA attention is as follows:

[0013] Given a 2D input feature map The width and height of the feature map are W and H, and the number of channels is C. It is divided into S×S non-overlapping coarse-grained regions, such that each region contains feature vectors. Subsequently, query, key, and value tensors are calculated through linear projection

[0014] Q = X r W q , K = X r W k , V = X r W v

[0015] where are the projection weights of the query, key, and value respectively;

[0016] Then, by constructing a directed graph, the participation relationship of each region is found, and the region-level query and key are derived by applying the average value of each region to Q and K respectively

[0017] Then through Qr and K r The matrix multiplication between and the transpose of K calculates the adjacency matrix of the region-to-region directed graph

[0018] A r = Q r (K r ) T

[0019] Subsequently, the directed graph is pruned by retaining only the top k connections for each region, thereby deriving a routing index matrix

[0020] I r = topkIndex(A r )

[0021] Using the region-to-region routing index matrix I r , apply fine-grained token-to-token attention. For each query token in region i, it will attend to all key-value pairs in the union of the k regions indexed by I r (i,1) , I r (i,2) , …, I r (i,k) , and collect the key and value tensors

[0022] K g = gather(K, I r ), V g = gather(V, I r )

[0023] Subsequently, apply the attention operation to the collected key-value pairs to obtain the output. Here, a local context enhancement term LCE(V) is introduced, which is parameterized using depthwise convolution with a kernel size set to 5, as follows:

[0024] O = Attention(Q, K g , V g ) + LCE(V)j.

[0025] Optionally, the neck based on the TriFPN (Triple Feature Pyramid Network) is used for further feature fusion. It adds triple feature fusion to the traditional feature pyramid network that fuses features from top to bottom. Through three-branch feature aggregation and diffusion connections, it allows feature information to propagate between different scale levels, and while improving the feature fusion effect, it does not bring excessive additional computational overhead;

[0026] And since different input feature information has different resolutions and contributes differently to the output feature information, TriFPN adds a learnable additional weight to each input and uses fast normalization fusion to perform weighted fusion on the input features as follows:

[0027]

[0028] Among them, In i is the input feature, Out is the output feature, ω i , ω j are the added additional weights, and ∈ = 0.0001 is a small value to avoid numerical instability.

[0029] Optionally, the loss function for training the barbell detection network is divided into a classification branch and a bounding box regression branch;

[0030] The classification branch uses binary cross-entropy loss BCE, and the calculation of the classification loss is as follows:

[0031]

[0032] Among them, N is the number of target categories, is the category to which the i-th sample belongs, and p i is the category probability predicted for the i-th sample;

[0033] The bounding box regression branch includes distribution focal loss DFL and MPSDIoU loss, and the calculation is as follows:

[0034]

[0035] Among them, y i and y i+1 are the two values closest to the label y in the label discrete distribution such that y i ≤y≤y i+1 , S i and S i+1 are the Sigmoid outputs of the corresponding probabilities respectively;

[0036] For the MPSDIoU loss, it not only considers the overlapping or non-overlapping regions of the target box, but also considers the center point distance, aspect ratio, corner square distance, and scale change effects. Let (x 1 , y 1 ), (x 2 , y 2 ) represent the coordinates of the upper left and lower right points of the bounding box respectively, then the true bounding box is expressed as The predicted bounding box is expressed as If the width and height of the original input image are w and h, the calculation of the MPSDIoU loss is as follows:

[0037]

[0038]

[0039] where IoU is the IoU loss function;

[0040] Therefore, the calculation of the total loss function in the training of the barbell detection network is as follows:

[0041]

[0042] where λ Cls , λ DFL , λ Box are hyperparameters during training and are the weights corresponding to the respective losses.

[0043] Optionally, the S5 specifically includes:

[0044] S51: Perform conversion from pixel units to physical length units;

[0045] Since the physical size of the barbell plate itself is fixed, the calculation of the length unit conversion coefficient α is as follows:

[0046]

[0047] where l real is the true length of the barbell plate in meters, and l pixel is the pixel length of the barbell plate bounding box;

[0048] S52: Calculate the speed v during barbell movement;

[0049] The barbell speed at the i-th frame is obtained by subtracting the position vector at the (i - 1)-th frame from the position vector at the i-th frame. Due to possible noise, the Savitzky-Golay filter is further used to smooth the barbell speed vector. The specific calculation process is as follows:

[0050]

[0051] where l window is the moving window length of the Savitzky-Golay filter, and p is the order of polynomial fitting;

[0052] In addition, to eliminate the small speed fluctuations caused by interference or detection errors, set v min as the speed threshold, and if it is lower than this, it is regarded as 0. The specific formula is as follows:

[0053]

[0054] S53. Calculation of the height h in barbell exercise;

[0055] The height of the barbell at the i-th frame is calculated from the difference between the position vectors at the i-th frame and the start-tracking moment and the actual initial height. Due to possible noise, the Savitzky-Golay filter is further used to smooth the barbell height vector. The specific calculation process is as follows:

[0056]

[0057] where h 0 is the initial height of the barbell, l window is the moving window length of the Savitzky-Golay filter, and p is the order of polynomial fitting;

[0058] In addition, to eliminate the height error caused by interference or detection, h 0 is set as the height threshold, and values below this are regarded as h 0 . The specific formula is as follows:

[0059]

[0060] S54. Calculation of the acceleration a in barbell exercise;

[0061] The acceleration of the barbell at the i-th frame is calculated from the difference between the height vectors at the i-th frame and the (i - 1)-th frame and the video frame rate. Due to possible noise, the Savitzky-Golay filter is further used to smooth the barbell acceleration vector. The specific calculation process is as follows:

[0062]

[0063] where l window is the moving window length of the Savitzky-Golay filter, and p is the order of polynomial fitting;

[0064] In addition, to ensure that the acceleration is within a reasonable range, a max is set as the acceleration threshold to constrain the occurrence of unreasonable accelerations. The specific formula is as follows:

[0065]

[0066] Finally, is smoothed:

[0067]

[0068] On the other hand, a barbell detection and trajectory tracking system is provided, the system comprising:

[0069] A sampling module for sampling the original video frames I with a frame rate of f rate to obtain a set of sampled images I';

[0070] A detection module for inputting the set of images I' into a barbell detection network improved based on YOLOv8n to perform left and right barbell plate detection, and obtaining the center point coordinates (b x , b y ) and the length and width dimensions (b h , b w ) of the target bounding box, the barbell detection network comprising a backbone network based on CBFormer, a neck based on the triple feature pyramid network TriFPN, and a detection head;

[0071] A generation module for generating a query vector Q of the tracking points according to the start frame τ of the first detected target tr , the query vector Q tr including the initial positions and the start frame of the center points of the left and right barbell plates to be tracked;

[0072] A tracking module for using the CoTracker model to obtain the position vector Tr of the tracking points frame by frame in the video according to the generated query vector Q tr ;

[0073] A calculation module for calculating barbell motion parameters according to the position vector Tr and the frame rate f of the original video frames rate .

[0074] On the other hand, an electronic device is provided, the electronic device comprising a processor and a memory, and at least one instruction is stored in the memory, and the at least one instruction is loaded and executed by the processor to implement the above barbell detection and trajectory tracking method.

[0075] On the other hand, a computer-readable storage medium is provided, and at least one instruction is stored in the storage medium, and the at least one instruction is loaded and executed by a processor to implement the above barbell detection and trajectory tracking method.

[0076] The beneficial effects brought by the technical solution provided by the present invention at least include:

[0077] 1) The position and category information of barbell plates are identified from multiple pictures sampled from the weightlifting action video stream using a barbell detection network improved based on YOLOv8n. A backbone network based on CBFormer is adopted, and to adapt to barbell plates of different sizes, a TriFPN structure is introduced for multi-directional cross-scale feature fusion. In terms of technical effects, the backbone network based on CBFormer can well capture the global dependencies in the image and has higher computational efficiency. TriFPN can effectively process multi-scale features, ensuring that the model can still accurately identify the barbell when faced with barbell plates of different sizes. Therefore, the barbell target detection network can accurately and quickly identify the target position and provide key point queries for barbell movement tracking.

[0078] 2) To address the problem that the prediction of barbell target position is still not precise enough, the IoU loss function for training the barbell detection network is improved, and the MPSDIoU loss function is used to further improve the detection accuracy. In terms of technical effects, the MPSDIoU loss function comprehensively considers the overlapping or non-overlapping parts of the target box, the distance between the center points, the aspect ratio, the squared distance of the corner points, and the influence of scale changes, significantly improving the accuracy of barbell target position prediction.

[0079] 3) The CoTracker model is used to track the movement trajectories of the barbell plates on the left and right sides, and motion parameters such as the height, speed, and acceleration during the barbell movement are calculated based on the video context information and the results of the tracking algorithm. In terms of technical effects, the model can maintain high-precision target tracking performance in complex backgrounds and fast movements, better reflecting the trajectory and position information of the barbell movement. Moreover, based on real-world physical modeling, accurate calculation of barbell motion parameters is achieved, thus truly reflecting the technical characteristics and related quantitative indicators of the weightlifting action. BRIEF DESCRIPTION OF THE DRAWINGS

[0080] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0081] Figure 1 is a flowchart of a barbell detection and trajectory tracking method provided by an embodiment of the present invention;

[0082] Figure 2 is an overall flowchart of a barbell detection and trajectory tracking method provided by an embodiment of the present invention;

[0083] Figure 3 is a schematic diagram of the barbell detection network structure provided by an embodiment of the present invention;

[0084] Figure 4 It is a schematic diagram of the CBFormer block structure provided by an embodiment of the present invention;

[0085] Figure 5 It is a schematic diagram of the TriFPN structure provided by an embodiment of the present invention;

[0086] Figure 6 It is a block diagram of a barbell detection and trajectory tracking system provided by an embodiment of the present invention;

[0087] Figure 7 It is a schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. Detailed implementation manners

[0088] To make the technical problems, technical solutions and advantages to be solved by the present invention clearer, the following will be described in detail with reference to the accompanying drawings and specific embodiments.

[0089] An embodiment of the present invention provides a barbell detection and trajectory tracking method, which can be implemented by an electronic device, and the electronic device can be a terminal or a server. As Figure 1 shown in a flowchart of a barbell detection and trajectory tracking method, the processing flow of the method may include the following steps:

[0090] S1. Perform a sampling operation on the original video frame I with a frame rate of f rate to obtain a set of multiple sampled pictures I';

[0091] For the original video to be processed, basic information such as the video frame rate f rate (25fps in the embodiment of the present invention), the frame height, etc. are obtained from the meta-information to ensure the consistency and accuracy of the video in subsequent processing. The original video frame I = {I 1 , I 2 , …, I o} will be subjected to a sampling operation, and the sampling operation is performed in the time dimension at a sampling rate r sample to reduce the computational complexity and improve the processing efficiency. Finally, a set of multiple sampled pictures I' = {I' 1 , I' 2 , …, I' s} is obtained, as Figure 2 shown.

[0092] S2. Input the set of multiple pictures I' into a barbell detection network improved based on YOLOv8n to perform left and right barbell plate detection, and obtain the center point coordinates (b x , b y ) and the length and width dimensions (b h , b w), the barbell detection network includes a backbone network based on CBFormer, a neck based on the triadic feature pyramid network TriFPN, and a detection head;

[0093] Optionally, as Figure 3 shown, the backbone network based on CBFormer is used to extract image features, adopts a four-layer pyramid structure, uses an image patch embedding module in the first layer ( Figure 3 the C 2 layer), and uses an image patch merging module in the second to fourth layers ( Figure 3 the C 3 to C 5 layer) to reduce the input spatial resolution and increase the number of channels, and then each layer uses a CBFormer block for feature transformation after the image patch embedding module and the image patch merging module.

[0094] Optionally, as Figure 4 shown, the CBFormer block first uses a 3×3 depthwise convolution to implicitly encode relative position information, the output is skip-connected through a residual connection, then layer normalization is performed, then a cascaded bi-level routing attention CBRA module is applied for cross-position relationship modeling and the output is added through a residual connection, then layer normalization is performed again, and finally a 2-layer multi-layer perceptron MLP is used for position embedding, and the output result is also added through a residual connection.

[0095] Optionally, the cascaded bi-level routing attention CBRA module obtains a more complete feature representation through the cascaded operation of multiple bi-level routing attention (BRA). Each bi-level routing attention BRA can achieve a sparse attention pattern in an adaptive query dynamic manner. The key idea is to divide the attention calculation into two levels in the coarse-grained region and the fine-grained region successively. The most relevant key-value pairs are selected at the coarse-grained region level, which is achieved by constructing and pruning a directed graph at the region level, and then the token-to-token attention is applied in the combination of the fine-grained region. From the perspective of single-input single-head self-attention, the calculation process of BRA attention is as follows:

[0096] Given a 2D input feature map The width and height of the feature map are W and H, and the number of channels is C. It is divided into S×S non-overlapping coarse-grained regions, so that each region contains feature vectors, and then the query, key, and value tensors are calculated through linear projection

[0097] Q = X r W q , K = X r Wk , V = X r W v

[0098] where are the projection weights for query, key, and value respectively;

[0099] Then, by constructing a directed graph to find the participation relationships of each region, the region-level queries and keys are derived by applying the average value of each region to Q and K respectively

[0100] Then through Q r and the transpose of K r , the adjacency matrix of the region-to-region directed graph is calculated through matrix multiplication

[0101] A r = Q r (K r ) T

[0102] Subsequently, the directed graph is pruned by keeping only the top k connections for each region, thereby deriving a routing index matrix

[0103] I r = topkIndex(A r )

[0104] Using the region-to-region routing index matrix I r , apply fine-grained token-to-token attention. For each query token in region i, it will attend to all key-value pairs in the union of the k regions indexed by I r (i,1) , I r (i,2) , …, I r (i,k) , and collect the key and value tensors

[0105] K g = gather(K, I r ), V g = gather(V, I r )

[0106] Subsequently, the attention operation is applied to the collected key-value pairs to obtain the output. Here, a local context enhancement term LCE(V) is introduced, which is parameterized using depthwise convolution with a kernel size set to 5, as follows:

[0107] O = Attention(Q, K g , V g) + LCE(V)j。

[0108] Optionally, the neck based on the TriFPN (Triple Feature Pyramid Network) is used for further feature fusion. It adds triple feature fusion to the traditional Feature Pyramid Network that fuses features from top to bottom. Through three-branch feature aggregation and diffusion connections, it allows feature information to be propagated between different scale levels, and while improving the feature fusion effect, it does not bring excessive additional computational overhead.

[0109] And since different input feature information has different resolutions and different contributions to the output feature information, TriFPN adds a learnable additional weight to each input and uses fast normalization fusion to perform weighted fusion on the input features as follows:

[0110]

[0111] where, In i is the input feature, Out is the output feature, ω i , ω j are the added additional weights, and ∈ = 0.0001 is a small value to avoid numerical instability.

[0112] The TriFPN proposed in the embodiment of the present invention can capture multi-scale information and perform feature fusion more efficiently, better match the requirements of barbell detection tasks of different sizes in different scenarios, thereby reducing the phenomenon of missed detection and false detection of barbell plates. Its specific structure is as Figure 5 shown. For the bottom-up P3 - P7 features, feature propagation is performed through multiple three-branch aggregation and diffusion modules. Triple feature fusion enables higher-level features to be transmitted to lower levels to enhance the global information of low-resolution features, and at the same time, lower-level features can be transmitted to higher levels to enrich the detailed information of high-resolution features. Feature fusion is performed at the weighted fusion nodes. Each node will fuse feature maps from different scales, and TriFPN introduces a weighted feature fusion mechanism, enabling the network to adaptively learn the weights of each input feature, thereby enhancing the feature fusion effect. Through this multi-directional feature transfer and weighted fusion mechanism, TriFPN can efficiently combine multi-level features, thereby improving the performance of the model when processing multi-scale targets.

[0113] Optionally, the loss function for training the barbell detection network is divided into a classification branch and a bounding box regression branch;

[0114] The classification branch uses Binary Cross-Entropy (BCE). The calculation of the classification loss is as follows:

[0115]

[0116] where N is the number of target categories, is the category to which the i-th sample belongs, and p i is the predicted category probability of the i-th sample;

[0117] The bounding box regression branch includes Distribution Focal Loss (DFL) and MPSDIoU loss, and the calculation is as follows:

[0118]

[0119]

[0120] where y i and y i+1 are the two values closest to the label y in the label discrete distribution such that y i ≤y≤y i+1 , and S i and S i+1 are the Sigmoid outputs of the corresponding probabilities respectively;

[0121] The MPSDIoU loss not only considers the overlapping or non-overlapping regions of the target boxes, but also considers the center point distance, aspect ratio, corner square distance, and scale change effects. Let (x 1 ,y 1 ),(x 2 ,y 2 ) represent the coordinates of the upper left point and the lower right point of the bounding box respectively, then the ground truth bounding box is represented as The predicted bounding box is represented as If the width and height of the original input image are w and h, then the calculation of the MPSDIoU loss is as follows:

[0122]

[0123] where IoU is the IoU loss function;

[0124] The MPSDIoU loss solves the problem that the existing bounding box regression loss function cannot be optimized when the predicted box and the ground truth box have the same aspect ratio. Therefore, it can show better performance in the barbell detection task where the target has a specific aspect ratio, significantly improving the accuracy of barbell position prediction and providing a more reliable basis for subsequent barbell tracking.

[0125] Therefore, the total loss function in the training of the barbell detection network is calculated as follows:

[0126]

[0127] where λCls and λ DFL and λ Box are hyperparameters during training and are the weights corresponding to the respective losses.

[0128] S3. Generate a query vector Q for the tracking points based on the start frame τ when the target is first detected tr , where the query vector Q tr contains the initial positions and the start frame of the center points of the left and right barbell plates to be tracked;

[0129] S4. Use the CoTracker model to obtain the position vector Tr of the tracking points frame by frame in the video according to the generated query vector Q tr ;

[0130] The CoTracker model is a prior art and will not be elaborated here.

[0131] S5. Calculate the barbell movement parameters according to the position vector Tr and the frame rate f of the original video frame rate .

[0132] Optionally, S5 specifically includes:

[0133] S51. Perform the conversion from pixel units to physical length units (since the foregoing steps are all processed based on video input, the barbell target size, barbell trajectory position vector, etc. obtained are all measured in pixel units. Therefore, in order to calculate the barbell movement parameters in the real physical world, it is necessary to perform the conversion from pixel units to physical length units);

[0134] Since the physical size of the barbell plate itself remains fixed, the calculation of the length unit conversion coefficient α is as follows:

[0135]

[0136] where l real is the actual length of the barbell plate, in meters, and l pixel is the pixel length of the barbell plate bounding box;

[0137] S52. Calculate the speed v during barbell movement;

[0138] The barbell speed at the i-th frame is obtained by subtracting the position vector at the (i - 1)-th frame from the position vector at the i-th frame. Since there may be noise, the Savitzky-Golay filter is further used to smooth the barbell speed vector. The specific calculation process is as follows:

[0139]

[0140] where l windowis the moving window length of the Savitzky-Golay filter, and p is the order of polynomial fitting;

[0141] In addition, to eliminate the small velocity fluctuations caused by interference or detection errors, set v min as the velocity threshold. If it is lower than this value, it is regarded as 0. The specific formula is as follows:

[0142]

[0143] S53. Calculation of the height h in barbell movement;

[0144] The height of the barbell at the i-th frame is calculated from the difference between the position vectors at the i-th frame and the start tracking time, and the actual initial height. Due to possible noise, the Savitzky-Golay filter is further used to smooth the barbell height vector. The specific calculation process is as follows:

[0145]

[0146] where h 0 is the initial height of the barbell, l window is the moving window length of the Savitzky-Golay filter, and p is the order of polynomial fitting;

[0147] In addition, to eliminate the height error caused by interference or detection, set h 0 as the height threshold. If it is lower than this value, it is regarded as h 0 , and the specific formula is as follows:

[0148]

[0149] S54. Calculation of the acceleration a in barbell movement;

[0150] The acceleration of the barbell at the i-th frame is calculated from the difference between the height vectors at the i-th frame and the (i - 1)-th frame, and the video frame rate. Due to possible noise, the Savitzky-Golay filter is further used to smooth the barbell acceleration vector. The specific calculation process is as follows:

[0151]

[0152] where l window is the moving window length of the Savitzky-Golay filter, and p is the order of polynomial fitting;

[0153] In addition, to ensure that the acceleration is within a reasonable range, set a max as the acceleration threshold to constrain the occurrence of unreasonable acceleration. The specific formula is as follows:

[0154]

[0155] Finally, perform smoothing processing on :

[0156]

[0157] As Figure 6 shown, an embodiment of the present invention also provides a barbell detection and trajectory tracking system, and the system includes:

[0158] A sampling module 610, configured to perform sampling operations on the original video frames I with a frame rate of f rate to obtain a set of sampled multiple pictures I';

[0159] A detection module 620, configured to input the set of multiple pictures I' into a barbell detection network improved based on YOLOv8n to perform left and right barbell plate detection, and obtain the center point coordinates (b x , b y ) and the length and width dimensions (b h , b w ) of the target bounding box. The barbell detection network includes a backbone network based on CBFormer, a neck based on the ternary feature pyramid network TriFPN, and a detection head;

[0160] A generation module 630, configured to generate a query vector Q tr of the tracking point according to the start frame τ when the target is first detected. The query vector Q tr contains the initial positions and the start frame of the center points of the left and right barbell plates to be tracked;

[0161] A tracking module 640, configured to use the CoTracker model to obtain the position vector Tr of the tracking point frame by frame in the video according to the generated query vector Q tr ;

[0162] A calculation module 650, configured to calculate barbell motion parameters according to the position vector Tr and the frame rate f rate of the original video frames.

[0163] A barbell detection and trajectory tracking system provided by an embodiment of the present invention has a functional structure corresponding to a barbell detection and trajectory tracking method provided by an embodiment of the present invention, and will not be elaborated here.

[0164] Figure 7It is a schematic structural diagram of an electronic device 700 provided by an embodiment of the present invention. The electronic device 700 may vary greatly due to different configurations or performances, and may include one or more central processing units (CPUs) 701 and one or more memories 702. Among them, at least one instruction is stored in the memory 702, and the at least one instruction is loaded and executed by the processor 701 to implement the steps of the above barbell detection and trajectory tracking method.

[0165] In an exemplary embodiment, a computer-readable storage medium is also provided, such as a memory including instructions, and the above instructions can be executed by a processor in a terminal to complete the above barbell detection and trajectory tracking method. For example, the computer-readable storage medium may be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.

[0166] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above embodiments can be completed by hardware, or can be completed by instructing relevant hardware through a program. The program can be stored in a computer-readable storage medium, and the above-mentioned storage medium can be a read-only memory, a magnetic disk, or an optical disc, etc.

[0167] The above are only the preferred embodiments of the present invention, and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A barbell detection and trajectory tracking method, characterized in that: The method comprises: S1, frame rate is f rate The original video frame I is sampled to obtain a set of multiple sampled pictures I'; S2, input the plurality of picture sets I' into the barbell detection network improved based on YOLOv8n, perform left and right barbell piece detection, and obtain the center point coordinates of the target bounding box (b x ,b y ) and length and width (b h ,b w ), the barbell detection network includes a backbone network based on CBFormer, a neck and a detection head based on a ternary feature pyramid network TriFPN; S3. Generate the query vector Q of the tracking point based on the starting frame τ where the target is first detected tr , the query vector Q tr Contains the initial position and start frame of the center points of the left and right barbell plates to be tracked; S4. Use the CoTracker model to generate the query vector Q tr , obtain the position vector Tr of the tracking point frame by frame in the video; S5, according to the position vector Tr and the frame rate f of the original video frame rate , calculate the barbell motion parameters; The CBFormer-based backbone network is used to extract image features and adopts a four-layer pyramid structure. An image block embedding module is used in the first layer, and an image block merging module is used in the second to fourth layers to reduce the input spatial resolution and increase the number of channels. Then, each layer uses a CBFormer block for feature transformation after the image block embedding module and the image block merging module. The CBFormer block first uses 3×3 depth-wise convolution to implicitly encode relative position information, and the output is skip-connected through residual connections, followed by layer normalization, and then a cascaded two-stage routing attention CBRA module is applied to model cross-position relationships and the outputs are added through residual connections, followed by layer normalization again, and finally a 2-layer multi-layer perceptron MLP is used for position embedding, and the output results are also added through residual connections; The cascaded two-level routing attention CBRA module obtains a more complete feature representation through multiple two-level routing attention BRA cascade operations. Each two-level routing attention BRA can realize sparse attention mode in a dynamic way of adaptive query. The key idea is to divide the attention calculation into two levels, coarse-grained area and fine-grained area, and perform it in sequence. The most relevant key-value pairs are screened out at the coarse-grained area level by building and pruning the directed graph at the area level, and then the mark-to-mark attention is applied in the union of the fine-grained areas. From the perspective of single-head self-attention with a single input, the calculation process of BRA attention is: Given a 2D input feature map The feature map has width W and height H and channel C, which is divided into S×S non-overlapping coarse-grained regions, so that each region contains feature vectors, and then the query, key, and value tensors are calculated through linear projection Q=X r W q ,K=X r W k ,V=X r W v in are the projection weights for query, key, and value, respectively; Then, we construct a directed graph to find the participation relationship of each region, and apply the average value of each region to Q and K to derive the region-level query and key Then through Q r and K r The adjacency matrix of the region-to-region directed graph is calculated by matrix multiplication between the transposes of A r =Q r (K r ) T The directed graph is then pruned by retaining only the top k connections for each region, resulting in a routing index matrix I r =topkIndex(A r ) Using the area-to-area routing index matrix I r , applies fine-grained token-to-token attention, where for each query token in region i, it focuses on the nodes located in region i. r (i,1) ,I r (i,2) ,…,I r (i,k) All key-value pairs in the union of k regions of the index, collecting key and value tensors K g =gather(K,I r ),V g =gather(V,I r ) The attention operation is then applied to the collected key-value pairs to obtain the output. A local context enhancement term LCE(V) is introduced here, which is parameterized using depth-wise convolution with the kernel size set to 5, as shown below: O=Attention(Q,K g ,V g )+LCE(V)j; The neck based on the ternary feature pyramid network TriFPN is used for further feature fusion. It adds ternary feature fusion to the traditional top-down feature pyramid network. It allows feature information to be propagated between different scale levels through three-branch feature aggregation and diffusion connection, and does not bring too much additional computational overhead while improving the feature fusion effect. And because different input feature information has different resolutions and contributes differently to the output feature information, TriFPN adds an additional learnable weight to each input and uses fast normalization fusion to perform weighted fusion on the input features, as shown below: Among them, i is the input feature, Out is the output feature, ω i ,ω j is the additional weight added, ∈ = 0.0001 is a small value to avoid numerical instability; The loss function of the barbell detection network training is divided into a classification branch and a bounding box regression branch; The classification branch uses the binary cross entropy loss BCE, and the classification loss is calculated as follows: Where N is the number of target categories, is the category to which the i-th sample belongs, p i The class probability predicted for the i-th sample; The bounding box regression branch includes distribution focus loss DFL and MPSDIoU loss, which are calculated as follows: where y i and i+1 are the two values ​​in the labeled discrete distribution closest to the label y such that y i ≤y≤y i+1 , S i and S i+1 They are the Sigmoid outputs of the corresponding probabilities respectively; The MPSDIoU loss not only considers the overlapping or non-overlapping areas of the target box, but also considers the center point distance, aspect ratio, corner point square distance and scale change. Let (x1, y1) and (x2, y2) represent the coordinates of the upper left and lower right points of the bounding box respectively. Then the true bounding box is expressed as The predicted bounding box is represented as The width and height of the original input image are w and h, then the MPSDIoU loss is calculated as follows: Where IoU is the IoU loss function; Therefore, the total loss function in the barbell detection network training is calculated as follows: where λ Cls , DFL , Box are the hyperparameters during training, and are the weights corresponding to the corresponding losses.

2. The method according to claim 1, characterized in that: The S5 specifically includes: S51, converting from pixel units to physical length units; Since the physical dimensions of the barbell plates themselves are fixed, the length unit conversion factor α is calculated as follows: Among them l real is the actual length of the barbell plate in meters, l pixel is the pixel length of the barbell plate bounding box; S52. Calculation of velocity v in barbell motion; The barbell velocity at the i-th frame is obtained by subtracting the position vector at the i-1-th frame from the position vector at the i-th frame. Due to the possible existence of noise, the barbell velocity vector is further smoothed using a Savitzky-Golay filter. The specific calculation process is as follows: Among them l window is the moving window length of the Savitzky-Golay filter, and p is the order of the polynomial fit; In addition, in order to eliminate the small speed fluctuations caused by interference or detection errors, v min The speed threshold is below which it is regarded as 0. The specific formula is as follows: S53. Calculation of height h in barbell exercise; The height of the barbell at the i-th frame is calculated by the difference between the position vector at the i-th frame and the time when tracking starts, and the actual initial height. Due to the possible existence of noise, the Savitzky-Golay filter is further used to smooth the barbell height vector. The specific calculation process is as follows: Where h0 is the initial height of the barbell, l window is the moving window length of the Savitzky-Golay filter, and p is the order of the polynomial fit; In addition, in order to eliminate the height error caused by interference or detection, h0 is set as the height threshold. If it is lower than this, it is regarded as h0. The specific formula is as follows: S54. Calculation of acceleration a in barbell motion; The barbell acceleration at the i-th frame is calculated by the difference between the height vectors at the i-th frame and the i-1-th frame, with the same frame rate as the video. Due to the possible presence of noise, the barbell acceleration vector is further smoothed using a Savitzky-Golay filter. The specific calculation process is as follows: Among them l window is the moving window length of the Savitzky-Golay filter, and p is the order of the polynomial fit; In addition, in order to ensure that the acceleration is within a reasonable range, set a max is the acceleration threshold to constrain the occurrence of unreasonable acceleration. The specific formula is as follows: Finally To perform smoothing:

3. A barbell detection and trajectory tracking system, characterized in that: The system comprises: Sampling module, used to sample frame rate f rate The original video frame I is sampled to obtain a set of multiple sampled pictures I'; The detection module is used to input the plurality of image sets I' into the barbell detection network improved based on YOLOv8n, perform left and right barbell piece detection, and obtain the center point coordinates of the target bounding box (b x ,b y ) and length and width (b h ,b w ), the barbell detection network includes a backbone network based on CBFormer, a neck and a detection head based on a ternary feature pyramid network TriFPN; The generation module is used to generate the query vector Q of the tracking point according to the starting frame τ where the target is first detected tr , the query vector Q tr Contains the initial position and start frame of the center points of the left and right barbell plates to be tracked; The tracking module is used to generate the query vector Q using the CoTracker model. tr , obtain the position vector Tr of the tracking point frame by frame in the video; A calculation module is used to calculate the position vector Tr and the frame rate f of the original video frame according to the position vector Tr and the frame rate f of the original video frame. rate , calculate the barbell motion parameters; The CBFormer-based backbone network is used to extract image features and adopts a four-layer pyramid structure. An image block embedding module is used in the first layer, and an image block merging module is used in the second to fourth layers to reduce the input spatial resolution and increase the number of channels. Then, each layer uses a CBFormer block for feature transformation after the image block embedding module and the image block merging module. The CBFormer block first uses 3×3 depth-wise convolution to implicitly encode relative position information, and the output is skip-connected through residual connections, followed by layer normalization, and then a cascaded two-stage routing attention CBRA module is applied to model cross-position relationships and the outputs are added through residual connections, followed by layer normalization again, and finally a 2-layer multi-layer perceptron MLP is used for position embedding, and the output results are also added through residual connections; The cascaded two-level routing attention CBRA module obtains a more complete feature representation through multiple two-level routing attention BRA cascade operations. Each two-level routing attention BRA can realize sparse attention mode in a dynamic way of adaptive query. The key idea is to divide the attention calculation into two levels, coarse-grained area and fine-grained area, and perform it in sequence. The most relevant key-value pairs are screened out at the coarse-grained area level by building and pruning the directed graph at the area level, and then the mark-to-mark attention is applied in the union of the fine-grained areas. From the perspective of single-head self-attention with a single input, the calculation process of BRA attention is: Given a 2D input feature map The feature map has width W and height H and channel C, which is divided into S×S non-overlapping coarse-grained regions, so that each region contains feature vectors, and then the query, key, and value tensors are calculated through linear projection Q=X r W q ,K=X r W k ,V=X r W v in are the projection weights for query, key, and value, respectively; Then, we construct a directed graph to find the participation relationship of each region, and apply the average value of each region to Q and K to derive the region-level query and key Then through Q r and K r The adjacency matrix of the region-to-region directed graph is calculated by matrix multiplication between the transposes of A r =Q r (K r ) T The directed graph is then pruned by retaining only the top k connections for each region, resulting in a routing index matrix I r =topkIndex(A r ) Using the area-to-area routing index matrix I r , applies fine-grained token-to-token attention, where for each query token in region i, it focuses on the nodes located in region i. r (i,1) ,I r (i,2) ,…,I r (i,k) All key-value pairs in the union of k regions of the index, collecting key and value tensors K g =gather(K,I r ),V g =gather(V,I r ) The attention operation is then applied to the collected key-value pairs to obtain the output. A local context enhancement term LCE(V) is introduced here, which is parameterized using depth-wise convolution with the kernel size set to 5, as shown below: O=Attention(Q,K g ,V g )+LCE(V)j; The neck based on the ternary feature pyramid network TriFPN is used for further feature fusion. It adds ternary feature fusion to the traditional top-down feature pyramid network. It allows feature information to be propagated between different scale levels through three-branch feature aggregation and diffusion connection, and does not bring too much additional computational overhead while improving the feature fusion effect. And because different input feature information has different resolutions and contributes differently to the output feature information, TriFPN adds an additional learnable weight to each input and uses fast normalization fusion to perform weighted fusion on the input features, as shown below: Among them, i is the input feature, Out is the output feature, ω i ,ω j is the additional weight added, ∈ = 0.0001 is a small value to avoid numerical instability; The loss function of the barbell detection network training is divided into a classification branch and a bounding box regression branch; The classification branch uses the binary cross entropy loss BCE, and the classification loss is calculated as follows: Where N is the number of target categories, is the category to which the i-th sample belongs, p i The class probability predicted for the i-th sample; The bounding box regression branch includes distribution focus loss DFL and MPSDIoU loss, which are calculated as follows: where y i and i+1 are the two values ​​in the labeled discrete distribution closest to the label y such that y i ≤y≤y i+1 , S i and S i+1 They are the Sigmoid outputs of the corresponding probabilities respectively; The MPSDIoU loss not only considers the overlapping or non-overlapping areas of the target box, but also considers the center point distance, aspect ratio, corner point square distance and scale change. Let (x1, y1) and (x2, y2) represent the coordinates of the upper left and lower right points of the bounding box respectively. Then the true bounding box is expressed as The predicted bounding box is represented as The width and height of the original input image are w and h, then the MPSDIoU loss is calculated as follows: Where IoU is the IoU loss function; Therefore, the total loss function in the barbell detection network training is calculated as follows: where λ Cls , DFL , Box are the hyperparameters during training, and are the weights corresponding to the corresponding losses.

4. An electronic device, comprising a processor and a memory, wherein at least one instruction is stored in the memory, wherein: The at least one instruction is loaded and executed by the processor to implement the barbell detection and trajectory tracking method as claimed in claim 1 or 2.

5. A computer-readable storage medium, wherein at least one instruction is stored in the storage medium, characterized in that: The at least one instruction is loaded and executed by the processor to implement the barbell detection and trajectory tracking method as claimed in claim 1 or 2.

Citation Information

Patent Citations

  • Barbell recognition and tracking control method based on YOLO and improved template matching

    CN114743125A

  • Small target detection method based on improved YOLOv8n

    CN116895007A