Water athlete detection and tracking method based on improved YOLOv10

By improving the YOLOv10 target detection network, combined with TGS-SPPF, GSConv, VoV-GSCSP, SPDConv and SimAM modules, the problems of easy loss of targets detected and tracked by water athletes in complex surface environments are solved, and efficient and real-time detection and tracking effects are achieved.

CN120108041AInactive Publication Date: 2025-06-06DONGHUA UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510584914.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-08
Publication Date
2025-06-06
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

When the prior art detects and tracks water athletes in complex water surface environments, there are problems such as targets being easily lost, poor real-time detection and poor robustness.

Method used

The detection and tracking method of water athletes based on improved YOLOv10 is adopted, and multi-scale feature extraction is performed by embedding the TGS-SPPF module in the backbone network, GSConv convolution and VoV-GSCSP modules are introduced to the neck network for multi-scale feature fusion, and SPDConv module and SimAM module are embedded in the head network to form detection heads with different receptive fields. Combined with small sample training and inference networks with cross-domain transfer learning, the generalization ability and robustness of the model are improved.

Benefits of technology

It effectively improves detection accuracy and real-time performance, solves the problems of easy loss of targets and poor real-time performance, and can take into account detection efficiency and accuracy in complex water surface environments, realizes real-time detection of rapidly changing targets, and provides frame-by-frame detection results for water athletes' tracking.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120108041A_ABST
    Figure CN120108041A_ABST
Patent Text Reader

Abstract

The invention relates to a water athlete detection and tracking method based on improved YOLOv10. The method comprises the following steps: acquiring an image frame sequence of a water athlete and preprocessing the image frame sequence; the preprocessed image frames are sequentially input into a target detection model based on improved YOLOv10, and all the water athletes in each image frame and corresponding detection frames of the water athletes are obtained; the improved YOLOv10 comprises the following steps: embedding a TGS-SPPF module in a backbone network to carry out multi-scale feature extraction; a GSConv convolution module and a VoV-GSCSP module are introduced into the neck network to carry out multi-scale feature fusion; an SPDConv module and a SimAM module are embedded in the head network, and three groups of detection heads with different receptive fields are formed; and performing multi-target tracking on the water athletes based on a detection result of the target detection model. According to the invention, on the basis of keeping a network lightweight architecture, the water athletes can be rapidly and accurately detected and tracked in real time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of water sports detection, and in particular to a water sports athlete detection and tracking method based on improved YOLOv10. Background Art

[0002] With the widespread development of water sports, traditional video surveillance methods have poor target recognition and tracking effects on athletes due to the influence of environmental factors such as water surface fluctuations, lighting changes, wind direction and speed. Existing technologies mainly use drone photography and portable smart positioning to identify and track athletes. The inconvenience of wearing portable smart devices will interfere with the performance of athletes. In recent years, target tracking algorithms based on deep learning have made significant progress in processing complex scenes and multi-target tracking tasks. However, when using target detection methods to identify and track the motion trajectories of water athletes, the following problems exist: 1) Lightweight and accurate target recognition: When using existing target detection methods, there is always a problem of target (water athlete) morphology detection delay due to wind speed and wave conditions, limited computing power of drones, etc. Therefore, an algorithm is needed to efficiently add and fuse the data of different targets and different shapes under different wind and wave conditions to reduce the computational complexity and inference time of the detector; 2) Real-time and accuracy issues of trajectory tracking: In water sports scenarios with target collision, target disappearance (falling into the water) and complex motion trajectories (curves, acceleration), existing target tracking methods are often unable to accurately and stably track the athlete's trajectory; 3) Robustness issue: There are relatively few research results in the field of water sports events, and the data is relatively scarce, resulting in poor model robustness under small sample conditions.

[0003] Therefore, a detection and tracking method is needed to solve the above problems. Summary of the invention

[0004] The technical problem to be solved by the present invention is to provide a method for detecting and tracking water athletes based on improved YOLOv10, which can solve the problems of easy target loss and poor detection real-time performance when detecting and tracking high-speed changing targets in complex water environments.

[0005] The technical solution adopted by the present invention to solve the technical problem is: to provide a method for detecting and tracking water athletes based on improved YOLOv10, comprising the following steps: Obtain image frame sequences of water athletes and perform preprocessing; The pre-processed image frames are sequentially input into the target detection model based on the improved YOLOv10 to detect all the aquatics athletes and their corresponding detection frames in each image frame; the improved YOLOv10 includes: Embed the TGS-SPPF module in the backbone network to extract multi-scale features; Introduce GSConv convolution and VoV-GSCSP modules in the neck network for multi-scale feature fusion; The SPDConv module and SimAM module are embedded in the head network to form three sets of detection heads with different receptive fields; Based on the detection results of the target detection model, multi-target tracking is performed on the aquatic athletes to obtain the real-time action trajectory of each aquatic athlete.

[0006] Furthermore, the TGS-SPPF module is placed between the attention mechanism module of the backbone network and the neck network.

[0007] Furthermore, after the TGS-SPPF module processes the input features using the improved spatial pyramid pooling layer, the output features are sent to the first convolution branch and the second convolution branch for parallel processing, and the processed features are sent to the neck network after feature fusion; wherein, the first convolution branch includes a convolution processing module, and the second convolution branch structure includes a convolution processing module and a bottleneck Transformer connected in sequence.

[0008] Furthermore, the first convolution branch is activated by a Mish function.

[0009] Furthermore, the improved spatial pyramid pooling layer sends the input features to the GSConv convolution and the SPPF module based on the GSConv convolution for parallel processing, and the generated features are output after feature fusion.

[0010] Furthermore, the SPPF module based on GSConv convolution is obtained by replacing the convolution module in the SPPF module with GSConv convolution.

[0011] Furthermore, the introduction of GSConv convolution and VoV-GSCSP modules in the neck network for multi-scale feature fusion is achieved by replacing the convolution module and the cross-stage partial fusion module of the neck network with GSConv convolution and VoV-GSCSP modules respectively.

[0012] Furthermore, embedding the SPDConv module and the SimAM module in the head network to form three groups of detection heads with different receptive fields is achieved by respectively embedding a group of SPDConv modules and SimAM modules connected in sequence between the neck network and each group of detection heads.

[0013] Furthermore, the target detection model is trained by the following method: Taking water sports as the target domain, determine the corresponding source domain; Collect image samples from the target domain and the source domain respectively, and divide the image samples from the target domain into training samples and test samples; Pre-training the object detection model using image samples from a source domain; The pre-trained target detection model is fine-tuned using training samples in the target domain.

[0014] Furthermore, when fine-tuning the target detection model, the number of categories of each group of detection heads is adjusted to a set value, and the set parameters of the backbone network are frozen, and no more than The learning rate.

[0015] Beneficial Effects

[0016] Due to the adoption of the above technical solution, the present invention has the following advantages and positive effects compared with the prior art: (1) The present invention effectively improves detection accuracy and real-time performance by introducing modules such as TGS-SPPF, Slim-Neck and composite detection head based on the YOLOv10 target detection network, solves the problems of easy target loss and poor real-time detection, and can balance detection efficiency and accuracy in complex water environments, realize real-time detection of rapidly changing targets, and provide frame-by-frame detection results for tracking of water athletes; (2) The present invention improves the real-time tracking accuracy of athletes under conditions of water surface fluctuations and high wind speeds through the synergy of anti-interference detection design (TGS-SPPF module in the improved YOLOv10 model), motion model optimization (Kalman filter algorithm in the DeepSORT tracking network) and appearance feature enhancement (cross-domain transfer learning), and can effectively prevent the mistracking phenomenon that may occur during the tracking process of the athlete's forward movement, cornering, acceleration, center of gravity fluctuations, etc. (3) The present invention proposes a small sample training and inference network based on cross-domain transfer learning, which can continuously optimize the generalization ability of the detection and tracking modules by collecting small samples; the visual feature migration from land sports (source domain) to water sports (target domain), through pre-training the target detection model, the grouping feature decoupling of GSConv in the model and the multi-scale context fusion of the TGS-SPPF module force the network to learn the common features that are independent of illumination and posture, realize domain-invariant feature extraction, and solve the distribution offset problems such as illumination difference, target posture difference, and background interference between the two domains; the cross-stage reuse mechanism of VoV-GSCSP reuses the common features of source domain pre-training, accelerates the adjustment of target domain parameters, and realizes rapid model adaptation under small sample data; the sub-pixel feature retention of SPDConv and the SimAM attention mechanism suppress noise and generate more stable appearance embedding for DeepSORT to reduce ID switching when calculating similarity.

[0017] Together, these innovations improve our technology’s performance for target recognition and motion tracking in water sports, making it more practical and adaptable. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 It is a diagram of the structure of the YOLOv10 lightweight detection network improved by the present invention; Figure 2 It is a structural diagram of TGS-SPPF introduced in the improved YOLOv10 lightweight detection network of the present invention; Figure 3 It is a GSConv structure diagram introduced in the improved YOLOv10 lightweight detection network of the present invention; Figure 4 It is a Vov-GSCSP structure diagram introduced in the improved YOLOv10 lightweight detection network of the present invention; Figure 5 It is a structural diagram of SPDConv introduced in the improved YOLOv10 lightweight detection network of the present invention; Figure 6 It is a structural diagram of SimAM introduced in the improved YOLOv10 lightweight detection network of the present invention; Figure 7 It is a schematic diagram of a small sample training and reasoning network based on cross-domain transfer learning of the present invention; Figure 8 It is a model block diagram of the DeepSORT target tracking network of the present invention; Fig. 9 This is a diagram showing the detection results of the water athlete image using the network model proposed by the present invention; Fig.10 It is a comparison chart of the F1 score curves of YOLOv8s, YOLOv8n, YOLOv10s and the model of the present invention; Fig.11 It is a comparison chart of the accuracy-confidence curves of YOLOv8s, YOLOv8n, YOLOv10s and the model of the present invention; Fig.12 It is a comparison chart of the precision-recall curves of YOLOv8s, YOLOv8n, YOLOv10s and the model of the present invention; Fig.13 It is a comparison chart of the recall-confidence curves of YOLOv8s, YOLOv8n, YOLOv10s and the model of the present invention; Fig.14 It is a comparison chart of the normalized confusion matrix of YOLOv8s, YOLOv8n, YOLOv10s and the model of the present invention. DETAILED DESCRIPTION

[0019] The present invention will be further described below in conjunction with specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and are not intended to limit the scope of the present invention. In addition, it should be understood that after reading the content taught by the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent forms fall within the scope limited by the appended claims of the application equally.

[0020] An embodiment of the present invention relates to a method for real-time detection and tracking of water athletes based on cross-domain migration and lightweight YOLOv10, comprising the following steps: Obtain image frame sequences of water athletes and perform preprocessing; The preprocessed image frames are sequentially input into the target detection model based on the improved YOLOv10 to detect all the aquatics athletes and their corresponding detection frames in each image frame; wherein the improved YOLOv10 includes: Embed the TGS-SPPF module in the backbone network to extract multi-scale features; Introduce GSConv convolution and VoV-GSCSP modules in the neck network for multi-scale feature fusion; The SPDConv module and SimAM module are embedded in the head network to form three sets of detection heads with different receptive fields; Based on the detection results of the above target detection model, multi-target tracking is performed on the aquatic athletes to obtain the real-time action trajectory of each aquatic athlete.

[0021] More specifically, the target detection model based on the improved YOLOv10 can be used Figure 1 The structure in includes: In the backbone network, the group spatial convolution (GSConv) and bottleneck transformer (i.e., bottleneck transformer) models are introduced to form the TGS-SPPF architecture through operations such as convolution, pooling, and feature fusion. Introduce GSConv, VoV cross-stage partial network (VoV-GSCSP) and organically combine upsampling, feature fusion and other operations to form a Slim-Neck component; The spatial-to-depth convolution (SPDConv) and the parameter-free attention module (SimAM) are introduced as a new head component.

[0022] This implementation retains several main modules in the original YOLOv10 including C2fCIB, CIB, PSA and SCDown.

[0023] The structure of the TGS-SPPF module is as follows Figure 2As shown in the figure, the Conv in the fast-spatial pyramid pooling layer (SPPF) of the YOLOv10 backbone network is replaced with GSconv, and the generated features are fused with the input features processed by the GSConv convolution and then output. The replaced SPPF module, together with the GSConv that processes the input features and the feature fusion module, constitute the GS-SPPF module. This module avoids excessive smoothing of global convolution by channel grouping. The GS-SPPF terminal adopts a dual convolution branch structure. The first branch consists of a 1×1 convolution and a bottleneck Transformer, and the second branch uses a 1×1 convolution. By applying The activation function performs nonlinear transformation. A feature fusion module connected in series with GS-SPPF is added to the structure to form the TGS-SPPF module.

[0024] The Transformer branch captures long-term posture changes by calculating global spatial relationships. The branch strengthens the action details of a single frame by retaining more gradient information in the negative range, forming a "global-local" feature complementarity. The activation function is expressed as

[0025] In the formula, is the output of the 1×1 convolution.

[0026] The structure of the GSConv module is as follows Figure 3 As shown in the figure, the number of channels of the input feature map is C1. After convolution processing, the number of channels of the output feature map is adjusted to C2 / 2. The depthwise separable convolution further converts the feature map into a feature map with C2 / 2 channels, concatenates the previous and next feature maps, and applies the channel shuffle operation to generate an output feature map with C2 channels.

[0027] Using GSConv instead of Conv in the bottleneck network can reduce the amount of calculation, but it will lead to the problem of information isolation between channels. Therefore, by introducing GS Bottleneck and adopting a one-time aggregation strategy (avoiding redundant calculations of multi-branch splicing and reducing the number of memory reads and writes), the cross-stage partial network (CSP) is improved to obtain the VoV-GSCSP module, whose structure is as follows: Figure 4 This module cascades and fuses shallow detail features (such as subtle features such as object edges and textures, which help to accurately locate the target) with deep semantic features (which are helpful in determining which category the target belongs to) to restore cross-group information interaction.

[0028] The multi-branch splicing operation of the cross-stage partial fusion module (C2f) in the neck network of the original YOLOv10 model requires frequent memory read and write, and the feature interaction between different branches depends on residual connections and splicing operations, resulting in weakened cross-channel semantic associations. Using VoV-GSCSP to replace all C2f modules in the neck module layer can avoid redundant calculations of multi-branch splicing and reduce the number of memory read and write times. The neck network formed by combining GSConv and VoV-GSCSP has a smaller model size. The network structure is optimized by reducing parameters through GSConv and optimizing feature flow through VoV-GSCSP.

[0029] In the detection head part, the SPDConv module is combined with the cascaded SimAM attention module as a new head component. SPDConv effectively retains high-frequency detail features through the strided convolution decomposition strategy, and retains the key limb details of long-distance athletes (small targets) in water sports scenes through feature reorganization across spatial dimensions. SimAM dynamically allocates feature weights through energy functions, has better feature information focusing and capturing capabilities, and complements the fine-grained features extracted by SPDConv - SPDConv provides high-resolution feature maps, and SimAM suppresses background interference.

[0030] The structure of the SPDConv module is as follows Figure 5 As shown in the figure, the input feature map is divided into 4 sub-regions through the spatial pyramid decomposition strategy, and the input feature map of size S×S×C1 is sampled at intervals, where S×S represents the size of the image block and C1 represents the number of channels. After the interval sampling, the size of each part is S / 2×S / 2×C1, and the four parts are spliced ​​in the channel dimension to form a single feature map of size S / 2×S / 2×(4×C1), and a feature map of S / 2×S / 2×C2 is generated by 1x1 convolution.

[0031] The structure of SimAM attention module is as follows Figure 6 As shown, its processing logic can be expressed as:

[0032] In the formula, is the input feature map In the channel , spatial location The eigenvalues ​​at is the number of channels, is the spatial dimension; is the output feature map, is the energy function value, is a hyperparameter.

[0033] The above target detection model can use small sample training and inference network of cross-domain transfer learning to continuously optimize the generalization ability of detection and tracking modules by collecting small samples. Specifically, it includes: Obtain training sample sets of the source domain (land sports domain) and the target domain (water sports domain); Input the training sample set of the source domain (land surface motion domain) for training, and save the trained model weights as the initialization parameters of the target domain task; The training sample set of the target domain (water sports field) is input into the network trained with source domain samples, the output layer structure is adjusted (the number of classification categories is changed to 4), the backbone components of the model are selectively frozen, a smaller learning rate is used for fine-tuning, and the trained model parameters are exported to the ONNX model file; the 4-layer classification corresponds to water athletes, rowing, paddleboarding and drowning, respectively, and a layered unfreezing strategy is adopted for the backbone components. In the first 50 rounds, only the Head and Slim-Neck components are trained and detected, 50% of the backbone layers are unfrozen in rounds 51-100, and all parameters are trained in rounds 101-200; the initial learning rate is set to , use cosine annealing method for adjustment, reduce the batch size to 16 to prevent overfitting, and export the ONNX model; Use the inference framework to load the ONNX model file parameters, input the test sample set of the target domain (water sports field), infer each frame of the image, and obtain the corresponding detection box result; the output result includes the target category, ID, normalized coordinates, confidence score, and the data is encapsulated in JSON format.

[0034] Among them, Figure 7 As shown, the training sample set can be obtained in the following ways: Collect athlete images and video data from the original source domain (land sports) and the target domain (water sports) (public datasets can also be used); Perform data preprocessing on the collected image and video data. The preprocessing process includes but is not limited to the steps of video decoding, data cleaning, pixel restoration, filtering and denoising, and data selection. The processing can be carried out in sequence according to the above steps; The preprocessed image data is format converted, that is, the image data is labeled to obtain a training sample set for model training, which includes high-quality image data and txt file labels corresponding to the image data.

[0035] More specifically, the data preprocessing process includes: Video decoding, restoring the collected video data to a displayable image sequence, and cropping it together with the collected picture data to obtain a cropped image of the athlete with a size of 448×448; Data cleaning: While keeping the preprocessing of the source domain and the target domain consistent, perform data cleaning on the cropped images to remove abnormal data such as errors, duplications, and missing values; Pixel repair: repair image pixels, video freeze frames, and garbled text for the cleaned data; Filtering and denoising: the repaired data is denoised using filtering and signal processing technology to eliminate image noise, video noise, screen jitter and text redundancy; Data selection, when screening image data, can include as many sports scenes and different environmental factors as possible, including but not limited to water sports scenes such as target collision, target disappearance (falling into the water) and complex motion trajectories (curves, acceleration), as well as environmental factors such as water surface fluctuations, lighting changes, wind direction and speed.

[0036] In some preferred implementations, the denoised data may also be enhanced, including but not limited to data enhancement techniques such as Mosaic, random flipping, color jittering, etc., to expand the training data and enhance the model's ability to learn target domain features.

[0037] In this implementation, the DeepSORT target tracking network is used for trajectory tracking. Figure 8 As shown, the specific steps include: Use the person re-identification (ReID) model to extract the appearance feature vector of the detection box: Kalman filtering is used to predict the motion state of each ID in the next frame; The detection box and the tracker are associated through the Hungarian algorithm, the state of the successfully matched tracker is updated, and the maximum and minimum values ​​of the objective function are solved using the coefficient matrix of the conditional function; The overlap and size ratio of the boxes are calculated using the distance-assisted intersection-over-union (IoU) measure.

[0038] Among them, the maximum number of lost frames is set to 100 when DeepSORT is initialized, and the minimum number of consecutive matches for confirming tracking is set to 3 according to the trajectory state confirmation mechanism.

[0039] The Kalman filter can be expressed as:

[0040] In the formula, is the mean of the observation vector, is the observation vector covariance matrix, For the Kalman prediction results, It is The covariance matrix of the sub-Kalman prediction results, is the observation matrix.

[0041] The Hungarian algorithm can be expressed as:

[0042] In the formula, is a binary function. For multi-target association tasks, if Indicates The first goal and A matching is completed with the corresponding trajectories. It is composed of The weight matrix , for multi-target association tasks, It means the The first goal and The correlation measure between the trajectories; The distance-assisted intersection-over-union measurement calculates the overlap and size ratio of the frames while weighing the distance between the center points of the frames, further enhancing the model's ability to accurately predict the position of the target frame.

[0043] In the formula, is the traditional intersection-union ratio measure, , , and They are the horizontal and vertical coordinates of the upper left corner of the predicted frame and the horizontal and vertical coordinates of the upper left corner of the real frame, respectively. is the Euclidean distance between the center point of the predicted bounding box and the true bounding box, is the minimum diagonal length of the enclosing rectangle of the predicted box and the real box, , , and They are the maximum width, minimum width, maximum height and minimum height of the smallest enclosing rectangle respectively.

[0044] The present implementation mode is further described below by taking the image data of water athletes autonomously photographed by drones in professional water sports events held in Shanghai as an example.

[0045] Fig. 9 This is the detection result of the acquired water athlete image by the embodiment of the present invention. It should be noted that in order to protect personal privacy, some faces in the image are blurred and covered. This processing is only to meet the desensitization requirements and does not belong to the technical content of this embodiment.

[0046] The specific implementation example calculation steps are as follows: Step 1: Crop the 100 images of water athletes and some images of athletes from land sports events with a resolution of 4096×4096 to obtain cropped images of 448×448. The athlete images of the original source domain (land sports) and the target domain (water sports) are subjected to Mosaic enhancement (4-image stitching), random rotation (±45°), HSV color perturbation (hue ±0.1, saturation ±0.7, brightness ±0.4), and denoising using non-local mean filtering (NL-Means). Divide the processed source domain and target domain image sets into two parts: training set and test set, respectively, with a ratio of 8 to 2; Step 2: Build an improved network model based on YOLOv10. The construction of the network model is exactly the same as that of the above specific implementation method. Step 3. Send the training sample set of the source domain (land sports) divided in step 1 into the network training, adopt the PyTorch deep learning framework to train and test our model, and save the model weights as the initialization parameters of the target domain task. The experimental equipment used includes Python 3.8, CUDA 11.1, 15 vCPU AMD EPYC 7543 32-core processor and RTX 3090 GPU with 24GB memory. The stochastic gradient descent (SGD) optimizer is used for training, the momentum size is set to 0.9, and the weight decay is 0.0001. The batch size is 16 and the number of training rounds is 200. The learning rate is initially set to 0.001 and adjusted using cosine annealing. The input image size is fixed to 448×448; Step 4: Input the training sample set of the target domain (water sports) into the network, and change the number of classification categories to 4 (corresponding to the four categories of water athletes, rowing, paddling, and drowning). Adopt a layered unfreezing strategy, and only train the detection of Head and Slim-Neck components in the first 50 rounds, unfreeze 50% of the Backbone layer in rounds 51-100, and train all parameters in rounds 101-200. Use a learning rate of Fine-tune and export the trained model parameters to the ONNX model file; Step 5. Use the inference framework to load the ONNX model file parameters, input the test sample set of the target domain (water sports), infer each frame of the image, and obtain the corresponding detection box result. The output result includes the target category, ID, normalized coordinates, and confidence score. The data is encapsulated in JSON format. Step 6: Initialize the DeepSORT target tracking network, set the maximum number of lost frames to 100, set the minimum number of consecutive matches to confirm tracking to 3 according to the trajectory state confirmation mechanism, set the Kalman filter parameters to process noise covariance Q = 0.01, and observation noise covariance R = 0.1. Import the detection box of the first frame input image into the DeepSORT target tracking network for trajectory tracking; Step 7: The DeepSORT target tracking network outputs the tracking trajectory of the water athlete with identity information (ID), and determines the athlete's victory or defeat based on the athlete's trajectory and time.

[0047] The detection and tracking effect of this embodiment is as follows Figure 10-Figure 14 As shown, in the order of upper left, upper right, lower left, and lower right, they are YOLOv8s, YOLOv8n, YOLOv10s, and the target detection model described in this embodiment, and the target categories are water athletes (person), rowing (boat), paddleboard (surfboard), drowning people (drowner), and all targets (allclasses).

[0048] See Fig.10 The F1 score curve in Figure 1 shows how the F1 score of different categories changes with the confidence threshold. From the dark blue curve, the highest F1 score of 0.93 is achieved at a confidence threshold of 0.723, indicating that the model performs best at this confidence level among all categories.

[0049] See Fig.11 The Precision-Confidence curve in reflects the precision performance of the model at different confidence thresholds. When the confidence threshold reaches 0.957, the precision of all categories reaches 1.0. The model has a high annotation quality for high-confidence samples at this threshold, but sacrifices a certain detection coverage (low recall rate). The curve rises steeply to 1.00 and then maintains a plateau, so the model chooses a threshold slightly lower than 0.957 (0.95) when it is actually deployed.

[0050] See Fig.12The Precision-Recall curve in Figure 1 shows that in water sports events, the number of samples of athletes falling into the water is relatively small, and the data set is relatively unbalanced. This curve can reflect the performance of the model in distinguishing athletes from athletes falling into the water. AP (average precision) represents the area under the PR curve, which measures the average precision of the model under all thresholds. In the figure, mAP50 reaches 0.971, and the curve is close to the upper right as a whole. There is no "biased" phenomenon in which some categories have high recall but low precision. When the recall rate reaches 0.9, the precision rate remains above 0.7, indicating that the model can not only capture most of the positive samples (high recall rate), but also a considerable proportion of the results predicted as positive samples are indeed correct (high precision rate). This means that this model has found a good balance point in classification.

[0051] See Fig.13 The Recall-Confidence curve in reflects the recall performance of the model under different confidence thresholds. At the lowest threshold, the recall rate is very high (1.0), indicating that the model can capture most of the real targets under loose conditions.

[0052] See Fig.14 The normalized confusion matrix in the figure shows the performance of the model in identifying the categories of "water athlete", "paddleboard", "rowing", and "drowner". Each cell represents the probability that a specific category is misclassified as another category, which is used to evaluate the recognition accuracy of the model in different categories. It can be seen that the model of the present invention is relatively accurate in predicting the categories of person (water athlete), boat (boat), and drowner (drowner).

[0053] In order to further verify the effectiveness of this implementation, mainstream models YOLOv8s, YOLOv8n, and YOLOv10s were selected to conduct a horizontal comparison test with the target detection model described in this embodiment.

[0054] Table 1, Table 2, and Table 3 respectively list the evaluation indicators P (precision), R (recall), mAP50 (average accuracy of the model when the IoU threshold is 0.5), and mAP50-95 (average accuracy in the range of IoU thresholds from 0.5 to 0.95) for the categories person (water athletes), drowner (drowned person), and all (all) obtained through comparative tests. The table clearly shows that the model proposed in the present invention performs better than other models in detecting water athletes and drowned persons, and can more effectively improve the detection accuracy of water sports targets. In addition, although the number of parameters is not clearly provided in the table, because the present technology introduces lightweight modules such as GSConv when improving YOLOv10, it has fewer parameters while showing excellent detection capabilities.

[0055] Table 1. Evaluation indicators of different models for detecting water athletes on private datasets Model P (%) R (%) mAP50 (%) mAP50-95 (%) YOLOv8s 93.3 87.7 97.5 82.2 YOLOv8n 92.6 85.8 97.6 82.8 YOLOv10s 91.7 87.1 95.0 76.7 Proposed model 99.9 94.9 99.3 84.1 Table 2. Evaluation indicators of different models for detecting drowning people on private datasets Model P (%) R (%) mAP50 (%) mAP50-95 (%) YOLOv8s 79.2 50.0 70.8 61.6 YOLOv8n 83.0 50.0 65.7 59.0 YOLOv10s 96.7 66.7 76.7 70.3 Proposed model 97.7 99.9 99.5 83.6 Table 3. Evaluation indicators of different models for detecting all categories on private datasets Model P (%) R (%) mAP50 (%) mAP50-95 (%) YOLOv8s 85.8 68.9 85.6 65.5 YOLOv8n 84.4 67.5 83.7 65.4 YOLOv10s 94.3 72.1 87.4 68.8 Proposed model 95.6 90.8 97.1 77.1 Although the present invention has been disclosed as a preferred embodiment, it is not intended to limit the present invention. Any person skilled in the art can make possible changes and modifications to the technical solution of the present invention by using the methods and technical contents disclosed above without departing from the spirit and scope of the present invention. Therefore, any simple modification, equivalent change and modification made to the above embodiments according to the technical essence of the present invention without departing from the content of the technical solution of the present invention shall fall within the protection scope of the technical solution of the present invention.

Claims

1. A method for detecting and tracking water athletes based on improved YOLOv10, characterized in that: The following steps are involved: Obtain image frame sequences of water athletes and perform preprocessing; The pre-processed image frames are sequentially input into the target detection model based on the improved YOLOv10 to detect all the aquatics athletes and their corresponding detection frames in each image frame; The improved YOLOv10 includes: Embed the TGS-SPPF module in the backbone network to extract multi-scale features; Introduce GSConv convolution and VoV-GSCSP modules in the neck network for multi-scale feature fusion; The SPDConv module and SimAM module are embedded in the head network to form three sets of detection heads with different receptive fields; Based on the detection results of the target detection model, multi-target tracking is performed on the aquatic athletes to obtain the real-time action trajectory of each aquatic athlete.

2. The method according to claim 1, characterized in that The TGS-SPPF module is placed between the attention mechanism module of the backbone network and the neck network.

3. The method according to claim 2, characterized in that After the TGS-SPPF module processes the input features using the improved spatial pyramid pooling layer, the output features are respectively sent to the first convolution branch and the second convolution branch for parallel processing, and the processed features are sent to the neck network after feature fusion; wherein, the first convolution branch includes a convolution processing module, and the second convolution branch structure includes a convolution processing module and a Bottleneck Transformer connected in sequence.

4. The method according to claim 3, characterized in that The first convolution branch is activated by the Mish function.

5. The method according to claim 3, characterized in that: The improved spatial pyramid pooling layer sends the input features to the GSConv convolution and the SPPF module based on the GSConv convolution for parallel processing, and the generated features are output after feature fusion.

6. The method according to claim 5, characterized in that The GSConv convolution-based SPPF module is obtained by replacing the convolution module in the SPPF module with the GSConv convolution.

7. The method according to claim 1, characterized in that The introduction of GSConv convolution and VoV-GSCSP modules in the neck network for multi-scale feature fusion is achieved by replacing the convolution module and the cross-stage partial fusion module of the neck network with GSConv convolution and VoV-GSCSP modules respectively.

8. The method according to claim 1, characterized in that The method of embedding the SPDConv module and the SimAM module in the head network and forming three groups of detection heads with different receptive fields is achieved by respectively embedding a group of SPDConv modules and SimAM modules connected in sequence between the neck network and each group of detection heads.

9. The method according to claim 1, characterized in that: The target detection model is trained by the following method: Taking water sports as the target domain, determine the corresponding source domain; Collect image samples from the target domain and the source domain respectively, and divide the image samples from the target domain into training samples and test samples; Pre-training the object detection model using image samples from a source domain; The pre-trained target detection model is fine-tuned using training samples in the target domain.

10. The method according to claim 9, characterized in that When fine-tuning the target detection model, the number of categories of each group of detection heads is adjusted to the set value, and the set parameters of the backbone network are frozen, and no more than The learning rate.