Joint point detection dynamic model algorithm for racing boat
By adding posture estimation branches to the YOLOv5 model, the YoloBoat model was developed, which solved the problem of athletes' posture detection in complex sports scenarios, and achieved high-precision and real-time node detection, which was suitable for rowing and other scenarios.
Patent Information
- Application Number
- CN202510186635.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-20
- Publication Date
- 2025-06-10
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The prior art is difficult to achieve high-precision athlete posture detection in complex sports scenarios, especially in rowing. Traditional methods are limited by high equipment costs, complex operation and low athlete comfort, and are difficult to widely use.
A dynamic recognition model based on YOLO is proposed. Combined with the real-time object detection capability of the YOLO series and advanced pose estimation technology, the accurate tracking of key points of the human body is achieved by adding pose estimation branches based on the YOLOv5 model.
It realizes real-time and accurate detection of athletes' joints in rowing, provides high-precision motion analysis, improves detection accuracy and robustness, and maintains stable detection performance in complex environments.
Smart Images

Figure CN120125962A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision and motion pose detection, and more specifically, particularly relates to a dynamic model algorithm for joint point detection for rowing sports. Background Art
[0002] With the popularization of water sports, especially in sports such as rowing and sculling, how to accurately monitor the motion postures and action performances of athletes has become the key to improving training effects and preventing sports injuries. Traditional motion posture assessment methods mostly rely on complex sensor devices or high-precision visual tracking systems. These methods are often restricted by factors such as high equipment costs, complex operations, and low athlete comfort, and are difficult to be widely applied in actual training. Therefore, it is particularly important to develop an efficient and convenient posture detection technology that can evaluate the action performances of athletes in real time and accurately.
[0003] In the field of pose estimation, traditional computer vision methods mostly rely on the detection of human key points in static images. However, these methods usually cannot effectively handle pose changes in fast-moving or dynamic scenes. In recent years, deep learning-based models, especially the YOLO series of object detection algorithms, have gradually become a new trend in human pose estimation due to their excellent real-time performance and high-precision detection capabilities. However, traditional YOLO models are mainly used for static object detection, and how to simultaneously perform object detection and pose estimation in complex motion scenes is still a technical challenge.
[0004] To address this challenge, we propose a dynamic recognition model YoloBoat based on YOLO. This model combines the real-time object detection capabilities of the YOLO series and advanced pose estimation techniques. By adding a pose estimation branch on the basis of the YOLOv5 model, it realizes the accurate tracking of human key points. For rowing sports videos, the model can, during the rowing training process of athletes, in real time identify and track the key parts of the human body, including 7 key points: the head, heel, knee, hip, shoulder, elbow, and wrist side, so as to accurately analyze the motion trajectories and pose performances of athletes. The model can not only track the postures of athletes in real time in dynamic motion scenes, but also analyze the motion fluency and coordination of athletes according to the changes of key points, and provide high-precision motion feedback. This enables this project to no longer be limited to traditional visual tracking methods in sports training, but can provide more accurate and efficient data support for training through deep learning models.
[0005] Traditional pose estimation methods often rely on highly accurate sensors or complex device configurations. They not only require a large amount of training data but also cause significant interference to athletes' movements. In contrast, the YoloBoat model has strong adaptability and real-time feedback capabilities. It can accurately provide athletes' pose analysis results in complex water training environments without relying on high-cost equipment, helping coaches and athletes understand their movement performance in real-time and make adjustments. Summary of the Invention
[0006] To solve the above technical problems, the present invention provides a dynamic model algorithm for joint point detection for rowing sports to solve the above problems.
[0007] The dynamic model algorithm for joint point detection for rowing sports includes the following steps:
[0008] Select the YoloBoat model, preprocess the input rowing sports image data, normalize it to the interval [0,1], and simulate diverse scenarios through data augmentation methods such as rotation, scaling, and cropping. The preprocessed image is fixed at a size of 640×640.
[0009] The backbone network of the model uses CSPDarknet53 for feature extraction, which consists of multiple convolutional layers, batch normalization, and LeakyReLU activation functions. The convolutional layers perform convolutional operations on the input image through kernels.
[0010] The feature fusion layer uses PathAggregationNetwork (PANet) to pass rich semantic information through a bottom-up path, enhance the context information of the features, and perform feature merging and optimization.
[0011] The output layer generates detection results through a fully connected layer, including bounding box coordinates, object categories, and key point coordinates. The bounding box predicts the position of the target through regression.
[0012] Perform spline interpolation on the key points to smooth the trajectory.
[0013] Represent the human body structure through skeleton connection, and at the same time perform optimization and improvement of confidence filtering, historical position update, and trajectory smoothing.
[0014] Under the training of its own data with different granularities of fineness and coarseness, the model basically achieves stable prediction for long-distance and short-distance scenes, with an average missed frame rate of 1 / 800 frames.
[0015] Preferably, in the preprocessing of the input image, the data augmentation method includes randomly adding noise or blur to simulate low-light or complex background scenes.
[0016] Preferably, the convolution formula is: F = I * K, where * represents the convolution operation, I is the input image, K is the convolution kernel, and F is the output feature map.
[0017] Preferably, the formula for the activation function LeakyReLU is: where α is a small constant, usually taking the value of 0.1.
[0018] Preferably, batch normalization normalizes the input of each layer, and the formula for batch normalization is: where μ is the mean of the batch, σ is the standard deviation, and x is the input.
[0019] Preferably, the feature aggregation formula is: F aggregated = F bottom-up + F top-down , where F bottom-up is the underlying feature, F top-down is the top feature, and they are combined through an addition operation.
[0020] Preferably, the bounding box regression formula is: box = (x, y, w, h), where (x, y) is the center of the bounding box, w and h are the width and height of the bounding box, and the key points predict their coordinates through regression. The key point regression formula is: (x k , y k ) = YOLO predict (I), where (x k , y k ) are the coordinates of the key points, and the model predicts these coordinates through regression.
[0021] Preferably, the spline interpolation formula is: S(t) = a 3 t 3 + b 3 t 2 + c 3 t + d 3 , where t is the interpolation parameter, and a3, b3, c3, d3 are coefficients calculated from existing data.
[0022] Preferably, the model is trained using its own dataset. The dataset contains approximately 1000 images, with the annotation format being JSON. It is divided into a training set, a test set, and a validation set in a ratio of 7:2:1. Data annotation is performed using an artificial annotation tool, and the annotated key points are concentrated on the main parts of the human skeleton, such as the shoulders, elbows, knees, ankles, etc.
[0023] Preferably, by comparing with the OpenPose, YOLOv8n-pose, and YOLOv8x-pose models, and using BoxPrecision, PosePrecision, mAP@50, and mAP@50-95 as evaluation metrics, the final selection plan of the model includes data preprocessing and annotation, model training, inference optimization, testing, and verification.
[0024] Compared with the prior art, the present invention has the following beneficial effects:
[0025] 1. Improve detection accuracy: By selecting the YoloBoat model and conducting elaborate architecture design and optimization, high-precision keypoint detection can be achieved in the close-range and long-range scenes of rowing sports, providing accurate and reliable data for athletes' motion analysis.
[0026] 2. Enhance robustness: It can effectively cope with various interference situations such as complex backgrounds, low light, and multi-person occlusion, maintaining stable detection performance in complex and changeable actual environments, and ensuring that the detection results are not significantly affected by changes in the external environment.
[0027] 3. Achieve fast inference: With high inference speed, it can be applied to rowing training and competition scenes in real time, quickly providing timely feedback to coaches and athletes, which helps improve training efficiency and the timeliness of competition decision-making.
[0028] 4. Comprehensively evaluate performance: Comprehensively evaluate the model performance through multi-dimensional evaluation metrics to ensure that the model performs excellently in all key aspects and meets the high standards and diverse needs of rowing sports keypoint detection.
[0029] 5. Great practical application value: Provide a scientific basis for rowing training, help coaches formulate more accurate training plans, assist athletes in improving technical movements, thus promoting the development of rowing sports and improving athletes' competition results.
[0030] 6. Optimize data processing: Effectively preprocess the input rowing sports image data, including normalization and data augmentation, improving data quality and the generalization ability of the model.
[0031] 7. Optimize key point trajectories: Use spline interpolation to smooth the key point trajectories, and at the same time conduct optimization and improvement of confidence filtering, historical position update, and trajectory smoothing, making the detection results more accurate, stable, and reliable.
[0032] 8. Efficient model training: Use about 1000 self-owned manually annotated pictures to construct a dataset, reasonably divide the training set, test set, and validation set, and adopt an adaptive learning rate and various data augmentation methods to improve the training effect and generalization ability of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 It is a schematic diagram of the average frame leakage rate in the present invention;
[0034] Figure 2 It is a training example of the present invention;
[0035] Figure 3 It is a schematic diagram of the training example when the ratio of the training set, test set, and validation set in the present invention is 7:2:1;
[0036] Figure 4 It is a schematic diagram of the dynamic environment adaptability in the present invention;
[0037] Figure 5 It is the YoloBoat model diagram of the present invention. Detailed implementation manners
[0038] The following further describes the implementation manners of the present invention in detail with reference to the drawings and embodiments. The following embodiments are used to illustrate the present invention, but cannot be used to limit the scope of the present invention.
[0039] The present invention aims to provide an efficient and accurate dynamic model algorithm for joint point detection for rowing sports to meet the needs of athlete posture evaluation in rowing sports and provide valuable data support for training and competitions.
[0040] 1. Model selection and basic architecture:
[0041] Select the YoloBoat model. This model consists of an input layer, a backbone network, a feature fusion layer, and an output layer. Under the training of its own data with different granularities, the model basically achieves stable prediction for long-distance and short-distance scenes, with an average frame leakage rate of 1 / 800 frames. See Figure 1 ; YoloBoat is the latest version of the YOLO series of models, with efficient and general object detection capabilities, and at the same time supports extended tasks, including pose estimation. In pose estimation, YoloBoat can detect the bone points of the target and generate a human skeleton structure, providing strong support for tasks such as sports analysis, behavior recognition, and human pose estimation.
[0042] Input layer:
[0043] Preprocess the input rowing sports image, and normalize the pixel values to the range of [0, 1].
[0044] Adopt data augmentation methods such as rotation, scaling, and cropping to simulate diverse scenes. The preprocessed images are uniformly sized at 640×640.
[0045] Backbone network (CSPDarknet53):
[0046] It consists of multiple convolutional layers, batch normalization, and LeakyReLU activation functions.
[0047] The convolutional layer performs a convolution operation on the input image through a kernel, and extracts the spatial features of the image according to the convolution formula.
[0048] The formula of the LeakyReLU activation function is, where the value is usually 0.1.
[0049] Batch normalization normalizes the input of each layer according to the formula, where is the mean of the batch, is the standard deviation, and is the input.
[0050] Feature Fusion Layer (PathAggregationNetwork - PANet):
[0051] It passes rich semantic information through the bottom - up path, enhancing the context information of the features.
[0052] Feature merging and optimization are carried out according to the feature aggregation formula, where is the bottom - layer feature, is the top - layer feature, and they are merged through an addition operation.
[0053] Output Layer:
[0054] The detection results are generated through a fully - connected layer, including bounding box coordinates, object categories, and key - point coordinates. The bounding box predicts the target position according to the regression formula, where is the center of the bounding box, and and are the width and height of the bounding box.
[0055] The key - point coordinates are predicted through regression according to the formula.
[0056] 2. Improvement in Key - point Smoothing and Optimization:
[0057] Spline interpolation is used to smooth the key - point trajectory according to the spline interpolation formula, where is the interpolation parameter, and is the coefficient calculated from the existing data.
[0058] The human body structure is represented by skeleton connection, and at the same time, optimization and improvement of confidence filtering, historical position update, and trajectory smoothing are carried out. Only when the confidence of the key - point reaches the set threshold, the key - point is drawn and connected.
[0059] 3. Dataset:
[0060] A dataset is constructed using approximately 1000 self - owned images, and the annotation format is JSON.
[0061] The images in the dataset are manually annotated, and the annotated key - points are concentrated on the main parts of the human body skeleton, such as shoulders, elbows, knees, ankles, etc.
[0062] The dataset is divided into a training set, a test set, and a validation set according to the ratio of 7:2:1.
[0063] Training process:
[0064] Train on the in-house dataset with the goal of improving the model's accuracy and robustness in identifying skeletal points in close-range and long-range scenarios.
[0065] Adopt adaptive learning rate and data augmentation methods such as random cropping, rotation, scaling, adding noise or blurring, etc. to address key point fluctuations and occlusion problems in dynamic scenarios, object scale and feature fusion problems in long-range scenarios, multi-person scenarios and occlusion problems, low-light and complex background scenarios, as well as scene variations.
[0066] Optimize the inference speed and resolution adaptability so that the model can maintain efficient inference performance under different hardware configurations and video resolutions.
[0067] 4. Model comparison and evaluation metrics:
[0068] Compare with OpenPose, YOLOv8n-pose, and YOLOv8x-pose models.
[0069] The evaluation metrics include BoxPrecision, PosePrecision, mAP@50, mAP@50-95, etc.
[0070] 5. Final selected solution and performance:
[0071] The final selected solution covers data preprocessing and annotation, model training, inference optimization, testing and verification, etc.
[0072] On a Lenovo Legion laptop (CPU i9, GPU RTX4060), the inference time for each image is 21.4ms, and the average frame drop rate is 1 / 800 frames. At different resolutions, such as the inference time for 1280×720 is 20.45ms, and the inference time for 1920×1080 is 37.05ms.
[0073] The feature fusion layer is used to merge and optimize feature maps from different scales, enhancing the model's ability to identify objects of different sizes. YoloBoat uses the Path Aggregation Network (PANet), which enhances the context information of features, transmits rich semantic information through the bottom-up path, and improves the detection ability of small objects. On the GPU RTX 4060, the inference speed is further improved. By optimizing the inference engine of YoloBoat, the inference speed remains efficient under different video sizes and resolutions. In long-range scenarios and dynamically changing scenarios, YoloBoat maintains a stable inference speed and basically has no frame drop phenomenon, with an average frame drop rate of 1 / 800 frames. The accuracy comparison is as follows:
[0074]
[0075] The Precision of YoloBoat and YoloBoat is significantly higher than that of YOLOv8n-pose and YOLOv8x-pose. YOLOv11 provides a higher Recall value, especially reaching the highest value of 0.978 in the large-sized model (YoloBoat). The small model (n-pose) of YOLOv8 has the fastest inference speed (15.2ms), but YoloBoat provides a better balance (YoloBoat: 21.4ms). The inference speed of YoloBoat (21.4ms) is about 14% slower than that of YOLOv11n-pose (18.7ms), but the improvement in performance is usually more important for practical applications.
[0076] Examples:
[0077] Data collection and preprocessing:
[0078] To obtain high-quality rowing motion image data, we set up multiple high-definition cameras at rowing training venues and competition sites to capture the training and competition processes of athletes from different angles and distances. These cameras have high frame rates and high resolutions, capable of capturing the subtle movement changes of athletes.
[0079] The collected raw image data is transmitted to the data processing center. First, preliminary screening is carried out to remove images with poor quality such as blurring, too dark or too bright lighting. Then, the screened images are preprocessed by normalizing their pixel values to the range [0, 1]. Data diversity is increased through operations such as rotation, scaling, and cropping to simulate various possible scenarios. At the same time, noise or blur is randomly added to simulate low-light or complex background scenarios, making the model more robust. The preprocessed images are uniformly fixed to a size of 640×640 for subsequent model input.
[0080] Model architecture and training:
[0081] The YoloBoat model is selected as the basic architecture. The backbone network of the model uses CSPDarknet53 for feature extraction. Multiple convolutional layers, batch normalization, and LeakyReLU activation functions cooperate with each other to extract valuable features from the input image according to the convolution formula.
[0082] The feature fusion layer uses the Path Aggregation Network (PANet) to merge and optimize features from different levels according to the feature aggregation formula, enhancing the context information of the features.
[0083] The output layer generates detection results through a fully connected layer, including bounding box coordinates, object categories, and keypoint coordinates. The bounding box predicts the target position according to the formula through regression, and the keypoints predict their coordinates according to the formula through regression.
[0084] A dataset is constructed using approximately 1000 self-owned images with manual annotations, and the annotation format is JSON. The annotated keypoints are concentrated on the main parts of the human skeleton, such as shoulders, elbows, knees, ankles, etc. The dataset is divided into a training set, a test set, and a validation set according to the ratio of 7:2:1.
[0085] During the training process, an adaptive learning rate adjustment strategy is adopted, and the initial learning rate is set to 0.001. The AdamW optimizer is used to continuously optimize the model's parameters through stochastic gradient descent. At the same time, data augmentation methods such as random cropping, rotation, and scaling are adopted to increase the diversity of data and the generalization ability of the model.
[0086] Key point smoothing and optimization improvement:
[0087] Spline interpolation smoothing is performed on the predicted key point trajectory. According to the spline interpolation formula, the jitter and discontinuity in the key point trajectory are eliminated to make it smoother and more natural.
[0088] At the same time, confidence filtering is performed, and only the key points with a confidence level reaching the set threshold are drawn and connected to improve the accuracy of the results. Moreover, through the optimization of historical position update and trajectory smoothing, the stability and reliability of key point detection are further improved.
[0089] Model comparison and evaluation:
[0090] The YoloBoat model of the present invention is compared with the OpenPose, YOLOv8n-pose, and YOLOv8x-pose models. Box Precision, Pose Precision, mAP@50, and mAP@50-95 are used as evaluation indicators.
[0091] In the experiment, different models are used to detect and analyze the same rowing sports image dataset respectively. The results show that the YoloBoat model performs excellently in Box Precision and Pose Precision, reaching 0.97 and 0.965 respectively, which are significantly higher than other comparison models. In terms of the mAP@50 and mAP@50-95 indicators, the YoloBoat model also achieves excellent results of 0.96 and 0.887, outperforming other models.
[0092] Inference optimization and practical application:
[0093] In terms of inference optimization, the inference speed of the model has been improved by compressing the model structure and optimizing the computational graph. On a Lenovo Legion notebook (CPU i9, GPU RTX 4060), the inference time for each image is only 21.4 ms. At different resolutions, such as 1280×720, the inference time is 20.45 ms, and at 1920×1080, the inference time is 37.05 ms. This shows that the model consumes more computing resources at higher resolutions but still maintains a relatively fast inference speed.
[0094] The trained model is applied to actual rowing training and competition analysis. Coaches and athletes can intuitively understand the athletes' movement postures, discover potential problems and deficiencies, and thus conduct targeted training and adjustments through the real-time obtained joint detection results. For example, by analyzing the joint trajectories of the shoulders and elbows when the athlete is rowing, it can be judged whether the force is uniform and the movement is standard; by observing the changes in the joints of the knees and ankles, the coordination and stability of the athlete's leg movements can be evaluated.
[0095] The joint detection dynamic model algorithm for rowing proposed in the present invention has the following remarkable advantages:
[0096] High-precision detection: Through a carefully designed model architecture and optimization strategy, it can accurately detect the joints of athletes in both close-range and long-range scenarios, providing a reliable data basis for motion analysis.
[0097] Strong robustness: In the face of various interference factors such as complex backgrounds, low light, and multi-person occlusion, it can still maintain stable detection performance and is not affected by changes in the external environment.
[0098] Fast inference: The high-efficiency inference speed enables the model to be applied to rowing training and competitions in real time, providing timely feedback to coaches and athletes, which helps improve training efficiency and competition results.
[0099] Comprehensive evaluation: Through multi-dimensional evaluation metrics such as Box Precision, Pose Precision, mAP@50, mAP@50-95, etc., the performance of the model is comprehensively evaluated to ensure that the model has excellent performance in different aspects.
[0100] High practical application value: In actual rowing training and competitions, it can provide strong support for coaches to formulate scientific training plans and athletes to improve their technical movements, promoting the development and progress of rowing.
[0101] In summary, the joint point detection dynamic model algorithm for rowing sports of the present invention has important application value and broad development prospects in the field of rowing sports. Through continuous optimization and improvement, this algorithm will provide stronger technical support for the scientific training and precise analysis of rowing sports, helping athletes achieve better results on the field.
[0102] The embodiments of the present invention are given for purposes of illustration and description, and are not exhaustive or limit the invention to the disclosed form. Many modifications and variations are obvious to those of ordinary skill in the art. The embodiments are chosen and described in order to best explain the principles of the invention and its practical application, and to enable those of ordinary skill in the art to understand the invention and design various embodiments with various modifications suitable for a particular purpose.
Claims
1. A dynamic model algorithm for joint point detection for rowing sports, characterized in that: The following steps are involved: The YoloBoat model is used to preprocess the input rowing sports image data, normalize it to the [0,1] interval, and simulate diverse scenes. The preprocessed image is fixed to a size of 640×640. The backbone network of the model uses CSPDarknet53 for feature extraction, which consists of multiple convolutional layers, batch normalization, and LeakyReLU activation functions. The convolutional layer performs convolution operations with the input image through the kernel; The feature fusion layer uses Path Aggregation Network (PANet) to transfer rich semantic information through a bottom-up path to merge and optimize features; The output layer generates detection results through the fully connected layer, including bounding box coordinates, object categories, and key point coordinates. The bounding box predicts the location of the target through regression. Perform spline interpolation to smooth the trajectory of key points; The human body structure is represented by skeleton connection, while confidence filtering, historical position update and trajectory smoothing are optimized and improved; When trained with our own data of different granularities, the average frame loss rate is 1 / 800 frames.
2. The dynamic model algorithm for joint point detection for rowing sports as claimed in claim 1, characterized in that: In the preprocessing of the input image, the data enhancement method includes randomly adding noise or blur to simulate low-light or complex background scenes.
3. The dynamic model algorithm for joint point detection for rowing as claimed in claim 1, characterized in that: The convolution formula is: F=I*K, where * represents the convolution operation, I is the input image, K is the convolution kernel, and F is the output feature map.
4. The dynamic model algorithm for joint point detection for rowing as claimed in claim 1, characterized in that: The formula for the activation function LeakyReLU is: Here α is a small constant, usually taken as 0.
1.
5. The dynamic model algorithm for joint point detection for rowing as claimed in claim 1, characterized in that: Batch normalization normalizes the input of each layer. The formula for batch normalization is: Where μ is the mean of the batch, σ is the standard deviation, and x is the input.
6. The dynamic model algorithm for joint point detection for rowing sports as claimed in claim 1, characterized in that: The feature aggregation formula is: F aggregated =F bottom-up +F top-down , where F bottom-up is the underlying feature, F top-down are the top features, merged by the addition operation.
7. The dynamic model algorithm for joint point detection for rowing as claimed in claim 1, characterized in that: The bounding box regression formula is: box = (x, y, w, h), where (x, y) is the center of the bounding box, w and h are the width and height of the bounding box, and the key points are predicted by regression. The key point regression formula is: (x k ,y k )=YOLO predict (I), where (x k ,y k ) are the coordinates of the key points, and the model predicts these coordinates through regression.
8. The dynamic model algorithm for joint point detection for rowing as claimed in claim 1, characterized in that: The spline interpolation formula is: S(t) = a3t 3 +b3t 2 +c3t+d3, where t is the interpolation parameter, and a3, b3, c3, d3a_3, b_3, c_3, d_3a3, b3, c3, d3 are coefficients calculated from existing data.
9. The dynamic model algorithm for joint point detection for rowing as claimed in claim 1, characterized in that: The model is trained using its own dataset in JSON format and divided into training set, test set and validation set in a ratio of 7:2:
1. Data is annotated using manual annotation tools, and the key annotation points are concentrated on the main parts of the human skeleton.
10. The dynamic model algorithm for joint point detection for rowing sports as claimed in claim 1, characterized in that: By comparing with OpenPose, YOLOv8n-pose and YOLOv8x-pose models, and taking BoxPrecision, PosePrecision, mAP@50 and mAP@50-95 as evaluation indicators, the final model selection plan includes data preprocessing and labeling, model training, inference optimization, testing and verification.
Citation Information
Cited By
Lightweight visible light ship target detection method based on edge feature guidance
CN120656032A
Joint point detection dynamic model algorithm for racing boat
CN120877383A
Application method of a joint point dynamic detection model for racing boat sports
CN120877383B