Water column detection tracking algorithm based on ByteTrack
By using the ByteTrack-based water column detection and tracking algorithm in offshore water column detection, combined with the static feature extraction network and multi-objective tracker, the problems of low detection accuracy and low efficiency in the existing technology are solved, and accurate detection and stable tracking of water columns are achieved, meeting the needs of offshore landing accuracy assessment tasks.
Patent Information
- Application Number
- CN202411888221.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-20
- Publication Date
- 2025-05-13
AI Technical Summary
The prior art has problems such as low measurement accuracy, low efficiency, complex equipment, long preparation time, limited measurement range, missed inspection, false inspection, repeated inspection and slow detection speed caused by sea environment complexity and water column morphology variability, which cannot meet the needs of the sea landing accuracy assessment task.
ByteTrack-based water column detection and tracking algorithm is adopted, combined with the static feature extraction network and ByteTrack multi-objective tracker, through static feature extraction and dynamic tracking, the static features and dynamic characteristics of the water column are comprehensively considered to improve detection accuracy and efficiency.
Accurate detection and stable tracking of water columns are achieved, detection accuracy and efficiency are improved, and the requirements of sea landing accuracy assessment tasks are met.
Smart Images

Figure CN119992401A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of target detection, and in particular to a water column detection and tracking algorithm based on ByteTrack. Background Art
[0002] The effective detection and tracking of water columns at sea is an indispensable core link in the task of assessing the accuracy of landing points at sea. At present, there are three main means of detecting water columns at sea. The first is the naked eye observation method, but this method has problems such as low measurement accuracy and low efficiency, and cannot meet the needs of high-precision assessment. The second method is to integrate shipborne radar and optoelectronic information for detection. Although it improves accuracy and efficiency, it has limitations such as complex equipment, long preparation time, and limited measurement range. The third method is to use airborne optoelectronic information for detection. This method has the advantages of wide field of view, flexible deployment, and variable observation positions. However, the complexity of the marine environment and the variability of the water column morphology bring huge challenges to detection.
[0003] The uncertainty of the shape and duration of the water column, the dynamic changes of the camera angle, and the limitation of the terminal computing power may lead to problems such as missed detection, false detection, repeated detection, and slow detection speed. These problems make the online detection results unreliable and unable to meet the needs of practical applications.
[0004] Therefore, it is necessary to develop more efficient and accurate water column detection methods to improve the accuracy and efficiency of detection and meet the requirements of water column detection technology for offshore drop point accuracy assessment tasks. Summary of the invention
[0005] In order to avoid the shortcomings of the prior art, the present invention provides a water column detection and tracking algorithm based on ByteTrack, which comprehensively considers the static characteristics and dynamic characteristics of the water column through a static feature extraction network and a ByteTrack multi-target tracker to improve the accuracy and efficiency of water column detection.
[0006] The present invention provides a water column detection and tracking algorithm based on ByteTrack, comprising: Step 1, obtaining video frame data of the object to be detected; Step 2: Input the video frame data into the static feature extraction network, use the static feature extraction network to extract and fuse the features of the video frame data, and output the detection frame corresponding to each detected water column target and its corresponding confidence; Step 3: Input the content to be detected into the ByteTrack multi-target tracker, predict the trajectory of the previous video frame through the Kalman filter model, obtain the predicted trajectory of the current frame, and output the predicted box corresponding to the predicted trajectory; Step 4: Use the Hungarian algorithm to match the detection box and the prediction box based on the confidence level, and output the water column target trajectory that successfully matches and the water column target trajectory that failed to match.
[0007] Wherein, step 2 further comprises: Step 21, input the content to be detected into the Backbone feature extraction network of the YOLOv8 model, and extract feature maps of different scales containing static features of the water column through convolution and deconvolution layers; Step 22, inputting the feature maps of different scales into the Neck feature fusion network of the YOLOv8 model, performing feature fusion and enhancement based on the feature maps of different scales, and obtaining a multi-scale feature map; In step 23, the multi-scale feature map is transferred to the Head detection head in the YOLOv8 model. The Head detection head obtains the detection frame and corresponding confidence of the water column target through convolution, pooling and post-processing operations.
[0008] Wherein, step 22 further comprises: Step 221, performing preliminary feature extraction on the input feature map through the Conv layer; Step 222, effectively aggregating the initially extracted features through the C2f module to achieve model compression; Step 223, using upsampling technology, improves the resolution of the feature map after feature aggregation and compression.
[0009] Wherein, step 2 further comprises: According to the confidence level, the detection frames are divided into high-score frames and low-score frames.
[0010] Wherein, step 4 further comprises: Step 41, performing a matching operation on the high-resolution frame and the predicted trajectory through the Hungarian algorithm, and obtaining the matched trajectory and the high-resolution frame, the unsuccessfully matched trajectory, and the unsuccessfully matched high-resolution frame; Step 42, performing a matching operation on the low-resolution frame and the unsuccessfully matched trajectory through the Hungarian algorithm, to obtain the matched trajectory and the low-resolution frame, the unsuccessfully matched trajectory, and the unsuccessfully matched low-resolution frame; Step 43: for the unsuccessfully matched high-resolution frame, if its confidence is higher than a preset confidence threshold, a new tracking track is created for the unsuccessfully matched high-resolution frame; Step 44, through the Kalman filter prediction model, the state update operation is performed on the matched trajectory, the unsuccessfully matched trajectory and the newly-created tracking trajectory to obtain an updated water column target trajectory.
[0011] Wherein, step 41 further comprises: Step 411, calculating the intersection-over-union ratio between each high-score frame and the predicted frame; Step 412, using the Hungarian algorithm, matching the high-resolution frame with the predicted frame to obtain matched tracks and high-resolution frames, unsuccessfully matched tracks and unsuccessfully matched high-resolution frames; Step 413: for the matched track, update its position and state, and update the frame in the track to the matched high-score detection frame.
[0012] Wherein, step 42 further comprises: Step 421, calculating the intersection-over-union ratio between each low-scoring frame and the unsuccessfully matched prediction frame; Step 422, using the Hungarian algorithm, matching the low-resolution frame with the unsuccessfully matched prediction frame to obtain the matched trajectory and the low-resolution frame, the unsuccessfully matched trajectory, and the unsuccessfully matched low-resolution frame; Step 423: for the matched track, update its position and state, and update the box in the track to the matched detection box.
[0013] Wherein, step 43 further includes: setting a retention frame number threshold for the newly created tracking trajectory, and if the trajectory is not detected again within the retention frame number, marking it as deleted or discarded; wherein, the retention frame number threshold is 30 frames.
[0014] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, wherein the computer program implements any of the above-mentioned water column detection and tracking algorithms when executed by a processor.
[0015] The present invention also provides a computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements any of the above-mentioned water column detection and tracking algorithms.
[0016] A water column detection and tracking algorithm based on ByteTrack is provided in the present invention. The YOLOv8 model is used as the water column static feature extraction network to output the detection frame and the corresponding confidence corresponding to the water column target. Then, the detection result is input into the ByteTrack multi-target tracker, and the water column target is tracked by associating the upper and lower frame information to obtain the water column trajectory of successful matching and the water column trajectory of failed matching.
[0017] Specifically, the method of the present invention proposes a water column detection and tracking framework including a static feature extraction network and a multi-target tracker. The framework can comprehensively consider the static and dynamic characteristics of the water column, thereby providing accurate water column detection and tracking results, and providing more scientific and accurate technical support for the task of offshore drop point accuracy assessment. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The accompanying drawings herein are incorporated into the specification and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the specification are used to explain the principles of the present disclosure. Obviously, the accompanying drawings described below are only some embodiments of the present disclosure, and for ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without creative work.
[0019] Figure 1 It is a schematic diagram of a water column detection and tracking framework provided by the present invention; Figure 2 It is a schematic diagram of the YOLOv8 model structure provided by the present invention; Figure 3 It is a flow chart of the water column detection and tracking algorithm provided by the present invention; Figure 4 This is a diagram showing the water column detection and tracking effect provided by the present invention. DETAILED DESCRIPTION
[0020] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0021] Water column detection and tracking technology refers to the detection of water column targets to achieve the accuracy assessment of the landing point at sea. At present, the mainstream methods of water column detection and identification at sea include simple judgment based on empirical rules, fusion of shipborne radar and optoelectronic information, and detection using airborne optoelectronic information. Water column detection and tracking is a key technology for the accuracy assessment of the landing point at sea. However, the current water column detection and tracking algorithm cannot adapt to the complexity of the marine environment and the variability of the water column morphology, which limits the comprehensiveness and accuracy of the prediction results, and a more comprehensive evaluation method is urgently needed.
[0022] In general, the method of the present invention first uses the YOLOv8 model to efficiently extract the static features of the water column, which include key information such as the shape, size, color, and texture of the water column. Subsequently, these extracted static features are input into the ByteTrack tracker for subsequent dynamic tracking of the water column. In the ByteTrack tracker, the Kalman filter algorithm is used to predict the motion state of the water column, and the Hungarian algorithm is combined to achieve accurate matching of the water column trajectory. Through a series of algorithmic processing, the present invention can achieve accurate detection and stable tracking of the water column, providing strong technical support for the task of accurately assessing the landing point at sea.
[0023] Figure 1Schematic diagram of the water column detection and tracking framework of the present invention. Figure 1 As shown, the present invention provides a water column detection and tracking algorithm based on ByteTrack, including: step 1, obtaining video frame data of the object to be detected; step 2, inputting the video frame data into a static feature extraction network, using the static feature extraction network to perform feature extraction and feature fusion on the video frame data, and outputting a detection frame corresponding to each detected water column target and its corresponding confidence; step 3, inputting the content to be detected into a ByteTrack multi-target tracker, predicting the trajectory of the previous video frame through a Kalman filter model, obtaining a predicted trajectory of the current frame, and outputting a prediction frame corresponding to the predicted trajectory; step 4, matching the detection frame and the prediction frame based on the confidence through the Hungarian algorithm, and outputting a successfully matched water column target trajectory and a failed matched water column target trajectory.
[0024] Among them, Figure 2 As shown, step 2 further includes: The content to be detected is input into the Backbone feature extraction network of the YOLOv8 model. Through the convolution and deconvolution layers, feature maps of different scales containing static features of the water column are extracted. Residual connections and bottleneck structures are introduced into the Backbone feature extraction network to improve the performance of feature extraction.
[0025] The feature maps of different scales are input into the Neck feature fusion network of the YOLOv8 model for feature fusion and enhancement to obtain multi-scale feature maps.
[0026] In the Neck feature fusion network, the Conv layer is used for the preliminary extraction of feature maps and the adjustment of the number of channels. The Conv layer moves on the input image through a sliding window (convolution kernel), calculates the dot product between the convolution kernel and the local area of the image, and generates feature maps. These feature maps capture the local features of the input data, such as edges, textures, etc.; then, the C2f module is used to achieve cross-stage partial aggregation, which can effectively aggregate multi-scale information and compress feature maps through convolution operations, reducing the amount of calculation while maintaining or enhancing the expressiveness of the model.
[0027] In addition, upsampling operations are also applied to the Neck feature fusion network to enlarge the low-resolution feature map from Backbone so that it can be fused with the higher-resolution feature map. Upsampling is achieved through interpolation algorithms or transposed convolution (deconvolution). In the upsampling process, the Concat module is combined to concatenate the upsampled feature map with the feature map from the previous stage to achieve multi-scale feature fusion.
[0028] The multi-scale feature map is transferred to the Head detection head in the YOLOv8 model. Through convolution, pooling and post-processing operations, the detection box and corresponding confidence of the water column target are obtained. The detection head consists of a series of convolutional layers and deconvolutional layers, which are responsible for generating the final detection results. The Head detection head of the YOLOv8 model has a bounding box prediction branch and a category prediction branch. The bounding box prediction branch is responsible for determining the location of the water column target, while the category prediction branch is responsible for determining the category of the target. Both branches calculate the bounding box loss and category loss through convolution blocks and Conv2d layers, and finally output the detection box and corresponding confidence of the water column target.
[0029] Specifically, static features are those that do not change much or remain unchanged over time. This includes properties such as water column color, texture, and shape. The data representation of static features is usually an image dataset, which contains images and their associated labels for training and testing models. Image datasets are a common data representation in deep learning. They contain a large number of image files and corresponding annotation information, such as bounding boxes, category labels, etc.
[0030] Wherein, step 2 further includes: dividing the detection frame into a high-score frame and a low-score frame according to the confidence of the detection frame.
[0031] Among them, Figure 3 As shown, step 4 further includes: First, the intersection-over-union ratio between each high-score frame and the predicted frame is calculated. The intersection-over-union ratio is a key indicator to measure the degree of overlap between two frames. It is obtained by calculating the ratio of the intersection area of the two frames to the intersection area. Then, the Hungarian algorithm is used to match the high-score frame with the predicted frame one by one, and the high-score frame and the predicted frame are optimally matched to minimize the overall matching cost. Finally, the successfully matched trajectory and high-score frame, the unsuccessfully matched trajectory, and the unsuccessfully matched high-score frame are obtained. For the successfully matched trajectory, according to the information carried by the newly matched high-score detection frame, its position and state are updated, and the frame in the trajectory is updated to the matched high-score detection frame.
[0032] Calculate the intersection-over-union ratio between each low-score frame and the unsuccessfully matched prediction frame. Use the Hungarian algorithm to match the low-score frame with the unsuccessfully matched prediction frame to obtain the matched track and low-score frame, the unsuccessfully matched track and the unsuccessfully matched low-score frame; for the matched track, update its position and state, and update the box in the track to the matched detection box.
[0033] For the unmatched high-scoring detection frames, match them with the inactive tracks to obtain three results: match, unmatched track, and unmatched detection frame. For the matching update status, the unmatched track is marked as deleted, and the unmatched detection frame is discarded if the confidence is low. However, if the confidence is high enough, a new tracking track is created and retained for 30 frames, and matched when it appears again.
[0034] Through the Kalman filter prediction model, the state update operation is performed on the matched trajectories, the unsuccessfully matched trajectories and the newly-created tracking trajectories to obtain the updated water column target trajectory.
[0035] It can be seen that the ByteTrack-based water column detection and tracking algorithm proposed in the present invention effectively utilizes the static feature information and dynamic characteristics of the water column. Figure 4 The tracking effect of the test video is shown. The test video was shot from the water surface perspective, and the perspective changes slightly, resulting in the water column and the background being similar in color and difficult to distinguish. However, as can be seen from the figure, the ByteTrack-based water column detection and tracking algorithm proposed in the present invention can output a stable ID number for the same water column, improving the stability of water column detection and tracking and reducing the probability of repeated detection or missed detection.
[0036] On the other hand, the present invention also provides a computer program product, which includes a computer program, and the computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute a water column detection and tracking algorithm provided by the above methods, and the method includes: step 1, obtaining video frame data of the object to be detected; step 2, inputting the video frame data into a static feature extraction network, using the static feature extraction network to extract features and fuse features on the video frame data, and outputting a detection frame corresponding to each detected water column target and its corresponding confidence; step 3, inputting the content to be detected into a ByteTrack multi-target tracker, predicting the trajectory of the previous video frame through a Kalman filter model, obtaining a predicted trajectory of the current frame, and outputting a prediction frame corresponding to the predicted trajectory; step 4, matching the detection frame and the prediction frame based on the confidence through the Hungarian algorithm, and outputting a successfully matched water column target trajectory and a failed matched water column target trajectory.
[0037] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute a water column detection and tracking algorithm provided by the above methods, the method comprising: step 1, obtaining video frame data of the object to be detected; step 2, inputting the video frame data into a static feature extraction network, using the static feature extraction network to perform feature extraction and feature fusion on the video frame data, and outputting a detection frame corresponding to each detected water column target and its corresponding confidence; step 3, inputting the content to be detected into a ByteTrack multi-target tracker, predicting the trajectory of the previous video frame through a Kalman filter model, obtaining a predicted trajectory of the current frame, and outputting a prediction frame corresponding to the predicted trajectory; step 4, matching the detection frame and the prediction frame based on the confidence through the Hungarian algorithm, and outputting a successfully matched water column target trajectory and an unmatched water column target trajectory.
[0038] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.
[0039] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0040] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A water column detection and tracking algorithm, characterized in that: The following steps are involved: Step 1, obtaining video frame data of the object to be detected; Step 2, inputting the video frame data into a static feature extraction network, using the static feature extraction network to perform feature extraction and feature fusion on the video frame data, and outputting a detection frame corresponding to each detected water column target and its corresponding confidence; Step 3: Input the content to be detected into the ByteTrack multi-target tracker, predict the trajectory of the previous video frame through the Kalman filter model, obtain the predicted trajectory of the current frame, and output the predicted box corresponding to the predicted trajectory; Step 4: Match the detection frame and the prediction frame based on the confidence level through the Hungarian algorithm, and output the water column target trajectory of successful matching and the water column target trajectory of failed matching.
2. The water column detection and tracking algorithm according to claim 1, characterized in that: Step 2 also includes: Step 21, input the content to be detected into the Backbone feature extraction network of the YOLOv8 model, and extract feature maps of different scales containing static features of the water column through convolution and deconvolution layers; Step 22, inputting the feature maps of different scales into the Neck feature fusion network of the YOLOv8 model, performing feature fusion and enhancement based on the feature maps of different scales to obtain a multi-scale feature map; Step 23, transmitting the multi-scale feature map to the Head detection head in the YOLOv8 model, and the Head detection head obtains the detection frame and corresponding confidence of the water column target through convolution, pooling and post-processing operations.
3. The water column detection and tracking algorithm according to claim 2, characterized in that: Step 22 further comprises: Step 221, performing preliminary feature extraction on the input feature map through the Conv layer; Step 222, effectively aggregating the initially extracted features through the C2f module to achieve model compression; Step 223, using upsampling technology, improves the resolution of the feature map after feature aggregation and compression.
4. The water column detection and tracking algorithm according to claim 1, characterized in that: Step 2 also includes: According to the confidence level, the detection frame is divided into a high-scoring frame and a low-scoring frame.
5. The water column detection and tracking algorithm according to claim 4, characterized in that: Step 4 also includes: Step 41, performing a matching operation on the high-resolution frame and the predicted trajectory through the Hungarian algorithm to obtain a matched trajectory and a high-resolution frame, an unsuccessfully matched trajectory, and an unsuccessfully matched high-resolution frame; Step 42, performing a matching operation on the low-resolution frame and the unsuccessfully matched trajectory by using the Hungarian algorithm, to obtain a matched trajectory and a low-resolution frame, an unsuccessfully matched trajectory, and an unsuccessfully matched low-resolution frame; Step 43: for the unsuccessfully matched high-resolution frame, if its confidence is higher than a preset confidence threshold, create a new tracking track for the unsuccessfully matched high-resolution frame; Step 44, using the Kalman filter prediction model, a state update operation is performed on the matched trajectory, the unsuccessfully matched trajectory and the newly-created tracking trajectory to obtain an updated water column target trajectory.
6. The water column detection and tracking algorithm according to claim 5, characterized in that: Step 41 further comprises: Step 411, calculating the intersection-over-union ratio between each high-score frame and the predicted frame; Step 412, using the Hungarian algorithm, matching the high-resolution frame with the predicted frame to obtain matched tracks and high-resolution frames, unsuccessfully matched tracks and unsuccessfully matched high-resolution frames; Step 413: for the matched track, update its position and state, and update the frame in the track to the matched high-score detection frame.
7. The water column detection and tracking algorithm according to claim 5, characterized in that: Step 42 further comprises: Step 421, calculating the intersection-over-union ratio between each low-scoring frame and the unsuccessfully matched prediction frame; Step 422, using the Hungarian algorithm, matching the low-resolution frame with the unsuccessfully matched prediction frame to obtain a matched trajectory and a low-resolution frame, an unsuccessfully matched trajectory, and an unsuccessfully matched low-resolution frame; Step 423: for the matched track, update its position and state, and update the box in the track to the matched detection box.
8. The water column detection and tracking algorithm according to claim 5, characterized in that: Step 43 further includes: setting a retention frame number threshold for the newly created tracking trajectory, and if the trajectory is not detected again within the retention frame number, marking it as deleted or discarded; wherein the retention frame number threshold is 30 frames.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, a water column detection and tracking algorithm as claimed in any one of claims 1 to 8 is implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, a water column detection and tracking algorithm as claimed in any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Impact point water column detection method based on improved YOLOv4 algorithm
CN115909072A
Traffic tracking detection system for view angle of unmanned aerial vehicle
CN118918148A
Cited By
Ground target statistical method and system based on YOLOV10 model
CN120198656A
Multi-target tracking identity recovery method based on lightweight feature embedding
CN121661358A
Sonar image continuous frame target detection method and system based on tracker fusion
CN122131311A