Self-adaptive identification and tracking method and system for floating objects on water surface
Through the improved SSD model and Swin Transformer network combined with LSTM and Kalman filtering algorithm, the floating object detection and tracking problems in complex water surface scenarios are solved, and high-precision and low-complexity monitoring of water surface float objects is achieved.
Patent Information
- Application Number
- CN202510820155.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-06-19
AI Technical Summary
The prior art cannot achieve high-precision detection and tracking in complex scenarios such as dynamic floating objects on the water surface, small-sized target floating objects on the water surface, fluctuations on the water surface and strong light reflection.
The improved enhanced SSD model is fused with the Swin Transformer network, combined with LSTM and Kalman filtering algorithms, and an adaptive water surface floating object recognition and tracking system is built, target detection is performed through multi-scale default boxes and self-attention mechanisms, and the Kalman filtering parameters are dynamically adjusted for state estimation.
High-precision floating object detection and tracking is achieved in complex scenarios, with low computational complexity, fast identification and strong consistency of multi-scene adaptability, suitable for dynamic water surface floating object monitoring.
Smart Images

Figure CN120355753A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image recognition, and particularly to a method and system for adaptively recognizing and tracking floating objects on the water surface. Background Art
[0002] Water environment monitoring and protection are important components of global environmental monitoring, and among them, the monitoring of floating objects on the water surface is a key area of water environment monitoring. Floating objects on the water surface refer to solid waste floating in water bodies such as rivers, lakes, and seas. Floating objects on the water surface affect the water quality environment and the ecosystem, causing serious negative impacts on human life, especially leading to problems such as river blockages and urban waterlogging. Therefore, the monitoring and treatment of floating objects on the water surface have become an important part of urban environmental monitoring and protection.
[0003] Traditional monitoring of floating objects on the water surface mainly relies on manual inspections or image processing techniques. The monitoring data of floating objects on the water surface is collected by fixed or mobile image acquisition devices, and then through image processing techniques, such as image target recognition algorithms based on deep learning, the recognition accuracy and efficiency of floating objects on the water surface are improved. However, most of the floating object recognition based on image processing techniques basically achieves high recognition accuracy in the scenario of a calm water surface. Under natural conditions, complex scenarios such as water surface fluctuations, floating objects with water flow, strong light reflection on the water surface, and foggy water surfaces are common, and the use of traditional image processing techniques can no longer meet the high-precision recognition and tracking of floating objects on the water surface in complex scenarios.
[0004] The patent with the publication number CN118918486B discloses a method and system for identifying targets in a water environment to solve the technical problem that the detection accuracy is significantly reduced by water surface reflection and floating object reflection. It proposes a yolov8-OMCIL network model, which is based on the yolov8 network model, a hybrid network reverse residual attention mechanism, and a cross-modal attention mechanism, enhancing the ability to extract features of water environment targets under different lighting conditions, but it does not involve the detection of floating objects on the water surface in scenarios such as dynamically floating objects on the water surface, small-sized target floating objects, and water surface fluctuations.
[0005] The patent with the publication number CN118887572B discloses a method for identifying the density of floating objects on the water surface based on image recognition, focusing on solving the monitoring spatial area of floating objects on the water surface and the water level calibration points, and performing abnormal verification of the monitoring position on the image, achieving an effective quantitative evaluation of the density of floating objects, and not involving issues such as the recognition of floating objects on the water surface in multiple scenarios and the accuracy guarantee.
[0006] The patent with the publication number CN216899018U discloses an integrated real-time online monitoring and early warning system for water surface floating objects, focusing on the construction of an integrated real-time data acquisition platform for water surface floating objects and a remote data analysis and early warning platform. It realizes the monitoring of floating objects by using fixed observation point data acquisition and remote data analysis, but does not involve the detection of water surface floating objects in scenarios such as dynamic water surface floating objects, small-size target floating objects, and water surface fluctuations.
[0007] The patent with the publication number CN114220044B discloses a method for detecting river floating objects based on AI algorithms, which solves the problem of high-precision identification and tracking of water surface floating objects in existing complex scenarios.
[0008] Therefore, the existing methods or systems fail to solve the problems of detecting and tracking water surface floating objects in scenarios such as dynamic water surface floating objects, small-size target floating objects, water surface fluctuations, and strong light reflection. Summary of the Invention
[0009] Aiming at the above-mentioned existing technical deficiencies, the technical problem to be solved by the present invention is to provide a method and system for adaptively identifying and tracking water surface floating objects, aiming to solve the problem that the image processing technology is not accurate enough in detecting and tracking water surface floating objects in scenarios such as dynamic water surface floating objects, small-size target floating objects, water surface fluctuations, and strong light reflection.
[0010] To solve the above technical problems, the present invention adopts the following technical solutions: The present invention provides a method for adaptively identifying and tracking water surface floating objects, including: by adjusting the structures of the deep low-resolution detection layer and the shallow high-resolution detection layer of the SSD model, and dividing the detection layer input to the SSD model into several scales for processing, and at the same time removing the set spatial resolution checking unit, constructing an improved enhanced SSD model to detect water surface floating objects; Dividing the open-source water surface floating object dataset into a training set and a validation set for training and validating the improved enhanced SSD model, and applying the improved enhanced SSD model to actual image data for testing to detect the performance and generalization ability of the improved enhanced SSD model; Among them, the improved enhanced SSD model also fuses the Swin Transformer local window self-attention mechanism to process the image layer by layer; Inputting the historical coordinates of the detected water surface floating objects into the LSTM to predict the Kalman filter parameters, ensuring that the Kalman filter dynamically performs the optimal state estimation of the water surface floating objects, and realizing the tracking of the dynamic water surface floating object targets.
[0011] Further, the SSD generates default boxes of different scales on different feature layers. In the SSD300, the high-level feature map is 3×3, and the default box scale is relatively large, with the maximum scale = 0.85; The low-level feature map is 76×76, and the default box scale is relatively small, with the smallest scale = 0.15. At the same time, the aspect ratio of the default box is automatically adjusted according to the extreme aspect ratios in the training set.
[0012] Furthermore, the 6 detection layers input to SSD300 are divided into three scales for processing.
[0013] Furthermore, the 76×76 and 38×38 spatial resolutions are regarded as large sizes, 19×19 and 10×10 as medium sizes, and 5×5 and 3×3 spatial resolutions as small sizes.
[0014] Furthermore, the 1×1 spatial resolution checking unit is removed.
[0015] Furthermore, an LSTM is used to dynamically adjust the process noise covariance Q and the observation noise covariance R of the Kalman filter.
[0016] Furthermore, the steps to construct an adaptive target tracking model for floating objects on the water surface are as follows: 1) Initialization: Set the initial process noise covariance Q0 and the initial observation noise covariance R0, and initialize the LSTM hidden state; 2) Perform loop tracking at each time step k: Input historical data: Input the errors and residuals of the past T steps into the LSTM; Generate adjustment factors: The LSTM outputs α k and β k , and ensure positive numbers through the activation function Softplus(); Adjust the noise covariance: Calculate the process noise covariance at the current time step Q k and the observation noise covariance at the current time step R k as follows: ; ; where, is the process noise adjustment factor, is the observation noise adjustment factor; The adaptive noise adjustment is as follows: , ; where, is the adjusted process noise covariance at the current time step, is the adjusted observation noise covariance at the current time step; Kalman filter prediction and update: Use dynamic and Perform state estimation; Record errors and residuals: for input at the next time step; Minimize the tracking error as follows: ; where, is the true target state at time k, is the updated state by Kalman filter; Backpropagation Through Time (BPTT) optimizes the LSTM parameters while fixing the Kalman filter parameters.
[0017] An adaptive recognition and tracking system for floating objects on the water surface, comprising: An identification module for constructing an improved enhanced SSD model to identify floating object targets on the water surface; A tracking module for dynamically adjusting the Kalman filter parameters through LSTM, constructing an adaptive tracking model for floating object targets on the water surface, and tracking the floating object targets on the water surface.
[0018] The beneficial effects of the present invention are as follows: 1. The present invention adopts a fusion mode of an improved enhanced SSD and a Swin Transformer network, which solves the problem of high-precision detection of floating objects on the water surface under foggy and fluctuating water conditions; moreover, it uses an LSTM and Kalman filter floating object target tracking algorithm to ensure that the Kalman filter dynamically performs optimal state estimation on the floating objects, ensuring the tracking and monitoring of moving floating objects on the water surface; 2. The present invention has the advantages of low computational complexity, fast floating object recognition, high accuracy, accurate dynamic tracking and monitoring, and strong accuracy consistency in multi-scene applications, and is suitable for multi-scene dynamic monitoring of floating objects on the water surface. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0020] Figure 1 It is a schematic flowchart of the adaptive recognition and tracking method for floating objects on the water surface provided in Embodiment 1 of the present invention.
[0021] Figure 2 It is a schematic diagram of the detection results of the improved enhanced SSD model in the open-source dataset in Embodiment 1 of the present invention.
[0022] Figure 3Schematic diagram of the detection results of the improved enhanced SSD model in Embodiment 1 of the present invention on the actual application dataset.
[0023] Figure 4 Schematic diagram of the detection results of the improved enhanced SSD model and the traditional Transformer in Embodiment 1 of the present invention on the open-source dataset.
[0024] Figure 5 Schematic diagram of the detection results of the fusion of the improved enhanced SSD model and Swin Transformer in Embodiment 1 of the present invention on the open-source dataset.
[0025] Figure 6 Schematic diagram of the comparison of the prediction results of the adaptive target tracking algorithm, extended Kalman filter, and recursive least squares for the center position in Embodiment 1 of the present invention.
[0026] Figure 7 Schematic diagram of the comparison of the center position errors in different scenarios in Embodiment 1 of the present invention.
[0027] Figure 8 Schematic diagram of the comparison results of the algorithms in a simple scenario in Embodiment 2 of the present invention. Detailed implementation manners
[0028] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0029] Embodiment 1
[0030] As Figures 1 to 7 shown, this embodiment provides an adaptive recognition and tracking method for floating objects on the water surface, including: By generating default boxes of different scales on different feature layers of SSD, and dividing the detection layer input by SSD into several scales for processing, and at the same time removing the 1×1 spatial resolution inspection unit, an improved enhanced SSD model is constructed for identifying small-size floating object targets on the water surface. It should be noted that: The SSD (Single Shot MultiBox Detector) neural network is mainly used to detect and identify multiple objects in an image, while predicting the categories and positions of these objects. It can complete object detection in a single forward pass, improving the detection speed. The core idea of SSD is to add multiple prediction layers to the last few layers of a convolutional neural network (CNN) to achieve multi-scale object detection, which can capture the sizes and shapes of various objects of different sizes. Using default bounding boxes as the starting point for prediction reduces the computational burden and improves the detection speed. It can achieve a relatively fast detection speed while maintaining a high-precision level. However, the semantic information of the underlying feature maps of SSD is insufficient, which easily leads to missed detections of small-sized objects. To overcome the influence of the flight altitude and trajectory changes of drones on the imaging resolution of floating objects on the water surface, the present invention proposes an improved enhanced SSD algorithm. Specifically, the structures of the deep low-resolution detection layer and the shallow high-resolution detection layer are adjusted, and efficient object detection is achieved through multi-scale default boxes and joint optimization of classification and localization tasks: SSD generates default boxes of different scales on different feature layers to cover objects of multiple sizes. The scale formula is: ; where, S k is the default box of different scales on the k-th feature layer, S min is the minimum scale, corresponding to the size of the default box of the bottommost feature map, S max is the maximum scale, corresponding to the size of the highest-layer feature box, m is the total number of feature layers, k is the layer number. In this case, SSD300 is adopted, with a total of 6 feature layers, and the input image size is 300×300; the aspect ratio of each default box is usually selected as {1, 2, 3, 1 / 2, 1 / 3}{1, 2, 3, 1 / 2, 1 / 3}, and an additional scale is used as the default box.
[0031] The formula for the aspect ratio of the default box is as follows: ; where, is the width of the default box on the k-th feature layer, is the height of the default box on the k-th feature layer. In this case, the high-layer feature map is 3×3, and the scale of the default box is relatively large = 0.85, the low-layer feature map is 76×76, and the scale of the default box is relatively small = 0.15, and an automatic adjustment of the aspect ratio according to the extreme aspect ratios in the training data is established.
[0032] The six detection layers of the traditional SSD input are divided into three scales for processing, namely large size (76×76, 38×38), medium size (19×19, 10×10), and small size (5×5, 3×3). Since the sliding window below 3×3 in the traditional SSD has difficulties in extracting small-scale floating objects, the original 1×1 spatial resolution inspection unit is abandoned in this invention; after the improvement of the SSD algorithm, the highly variable nature during the machine or manual imaging process can be solved, resulting in the problem that the size of floating objects is not fixed, and small-size floating object targets can be effectively detected.
[0033] The improved enhanced SSD model is used to train the model on the open-source "Water Surface Floating Object Dataset - 2400" dataset shown in Table 1. This dataset is designed specifically for deep learning and contains 2400 high-quality pictures that are carefully collected and processed. It focuses on the recognition of water surface floating objects. All the pictures in the dataset are sourced from actual environment shootings, featuring authenticity and practicality, and have undergone professional post-processing to ensure image quality, thus guaranteeing data accuracy, which is very suitable for the training and validation phases of the model.
[0034] Table 1: Data types of the Water Surface Floating Object Dataset - 2400.
[0035] ; Floating objects such as polyethylene terephthalate (PET bottles), polystyrene foam boxes, metals, plastic papers, ropes, plastic buoys, nets, glass pots, and plastic products are detected from the images, and the detection accuracy is as Figure 2 shown; the improved enhanced SSD algorithm is applied to the collected image data, and the detection accuracy of the above floating objects is as Figure 3 shown. The input parameters for training the detection network of the improved enhanced SSD model are shown in Table 2.
[0036] Table 2: Input parameters for training the improved SSD detection network.
[0037] ; The Swin Transformer network is adopted to extract features in the improved enhanced SSD model, improving the accuracy of the improved enhanced SSD model in recognizing water surface dynamic floating object targets. It should be noted that: In the case of foggy days and large fluctuations on the water surface, traditional SSDs and optimized neural networks are unable to complete feature extraction under complex conditions. The Swin Transformer network is used to extract features of multi-scale floating objects on the water surface under complex conditions. The image is processed layer by layer through the local window self-attention mechanism; the shifted window mechanism is used to capture the long-range dependencies and relative position biases between different windows, effectively solving the problem that the SSD network is difficult to effectively extract features under complex conditions; through hierarchical window design and relative position bias, the Swin Transformer network effectively models the global dependencies of the image while reducing the computational complexity. The image is divided into non-overlapping equal-sized windows (regular windows), and self-attention is calculated independently within each window, reducing the amount of computation (the complexity is reduced from O(N 2 ) to O(M 2 ·N / M 2 ) = O(N), where M is the window size). In adjacent layers, the window is shifted pixels to the lower right corner to form an overlapping detection area (shifted window) between adjacent windows. The SwinBlock consists of multiple Stages. Each Stage alternately uses regular windows and shifted windows to achieve local feature capture and pixel interaction between different windows. Its receptive field is gradually expanded. After multiple layers are stacked, the model gradually covers the entire image to achieve global dependency detection.
[0038] In the Swin Transformer network, relative position encoding is introduced in the attention calculation. The attention weights are adjusted by the relative positions of the query (Q) and the key (K), replacing the traditional absolute position encoding. A relative position bias term P is added to the self-attention formula, is the learnable relative position matrix, is the number of pixels within the window; the relative position range of the pixels within the window is [−M + 1, M − 1], and the total number of parameters is (2M − 1) × (2M − 1). The complexity of Swin (window attention) is O(HW·M2).
[0039] As Figure 4 shown, by experimentally comparing the detection results of the improved enhanced SSD and the traditional Transformer algorithm on the open-source dataset, Table 3 shows the fusion accuracy of the improved enhanced SSD and the traditional Transformer algorithm, while the fusion of Swin Transformer and the improved enhanced SSD in the open-source dataset detection results is as Figure 5 shown; the detection accuracy of the four water surface scenarios is shown in Table 4, ensuring the accuracy of recognition in complex environments and the feasibility of inference on the embedded platform.
[0040] Table 3: Detection accuracy of the improved enhanced SSD + traditional Transformer on the open-source dataset.
[0041]
[0042] Table 4: Detection accuracy of four water surface scenarios.
[0043]
[0044] Preferably, the Kalman filter parameters are dynamically adjusted by LSTM to construct an adaptive water surface floating object target tracking model for tracking the water surface floating object target. It should be noted that: The actual trajectory of the drone presents various periodic flight motions. It is difficult for a simple target detection algorithm to accurately perform fast inference. Therefore, the fusion of LSTM and Kalman filter is adopted to ensure the optimal state estimation of the Kalman filter for dynamic floating objects by dynamically adjusting the process noise covariance Q and the observation noise covariance R.
[0045] Specifically, the steps to construct an adaptive water surface floating object target tracking model are as follows: 1) Initialization: Set the initial process noise covariance Q0 and the initial observation noise covariance R0, and initialize the LSTM hidden state.
[0046] 2) Perform loop tracking at each time step k: Input historical data: Input the errors and residuals of the past T steps into LSTM.
[0047] Generate adjustment factors: LSTM outputs α k and β k , and ensure positive numbers through the activation function Softplus().
[0048] Adjust the noise covariance: Calculate the process noise covariance at the current time step Q k and the observation noise covariance at the current time step R k as follows: ; ; where is the process noise adjustment factor, is the observation noise adjustment factor.
[0049] Adaptive noise adjustment is as follows: , ; where is the adjusted process noise covariance of the current time step, is the adjusted process noise covariance of the current time step.
[0050] Kalman filter prediction and update: using dynamic and Perform state estimation.
[0051] Record errors and residuals: for input at the next time step.
[0052] The tracking error is minimized as follows: ; in, is the actual target state at time k, is the updated state of the Kalman filter.
[0053] Back propagation optimizes LSTM parameters through time BPTT, fixes Kalman filter parameters (state transfer matrix F, observation matrix H, initial process noise covariance Q0, and initial observation noise covariance R0), and dynamically adjusts Kalman filter parameters through LSTM, which significantly improves the robustness to nonlinear and time-varying noise, and is suitable for target tracking in complex scenarios.
[0054] In order to prove the adaptive surface floating object tracking model of the present invention, the present embodiment first performs target detection on a large number of video frames, and then inputs the historical coordinates of the floating object into the LSTM to predict the Kalman filter parameters. In the experimental process, simple surface videos and videos with large surface fluctuations are used for experiments, and the algorithm of the present invention is compared with the extended Kalman filter (EKF) and the recursive least squares (RLS) for algorithm verification. The results are as follows: Figure 6 As shown in the figure, it shows that the present invention can achieve a relatively ideal effect in dealing with the variable motion trajectory of the observation equipment. The comparison of the center position error in different scenarios is as follows: Figure 7 shown.
[0055] This embodiment also provides a system for adaptively identifying and tracking floating objects on a water surface, including: The recognition module is used to build an improved enhanced SSD model to identify floating objects on the water surface; The tracking module is used to dynamically adjust the Kalman filter parameters through LSTM, build an adaptive surface floating object tracking model, and track the surface floating object targets.
[0056] Example 2 This embodiment is the second embodiment of the present invention. This embodiment is different from the first embodiment in that it provides a verification test of a method and system for adaptively identifying and tracking surface floating objects, and verifies and illustrates the technical effects used in this method.
[0057] Collect UAV aerial images of floating objects on the water surface from different water surface scenes and under different weather conditions, construct a dataset of floating objects on the water surface under four scenarios, and perform feature enhancement such as Gaussian blur and histogram processing on the dataset.
[0058] Furthermore, conduct an experiment on the target tracking of dynamic floating objects on the water surface by combining the Swin Transformer and the improved enhanced SSD model with the adaptive filtering algorithm: use object detection and object tracking algorithms for comparative experiments to verify the performance of the present invention; select two-stage detection algorithms, and use the improved SSD algorithm, YOLOV4 algorithm, and Faster−RCNN algorithm for object detection algorithms; use the MOT algorithm for object tracking algorithms; divide 116 video sequences into a training set and a test set according to a ratio of 7:3. Among them, the training set contains 82 video sequences and 2,817 images, and the test set contains 34 video sequences and 1,207 images. Both contain data of four different scenarios, which are divided into a simple water surface scenario (calm water surface) and a complex water surface scenario (dynamic light and shadow, water surface reflection, strong light reflection water surface, and foggy water surface) to verify the detection performance and generalization ability of the algorithm.
[0059] Specifically, in the simple scenario, the detection accuracy and success rate of the floating object detection algorithms including the four algorithms of the present invention are as Figure 8 shown. All algorithms can achieve good detection results, and the detection success rate of the present invention reaches more than 90%. As can be seen from Table 5, the curve areas of the detection accuracy and success rate of the algorithms in this embodiment are better than those of other algorithms. From the perspective of detection speed, since the improved SSD algorithm enhances the detection layer, the accuracy is improved, but the number of algorithm parameters increases, reducing the processing efficiency of the detector; the Faster−RCNN and YOLOV4 algorithms are significantly faster than the object detection algorithm, and the algorithm in this embodiment combines the advantages of the improved enhanced SSD algorithm and the Swin Transformer algorithm to reach 28.35 fps, achieving a balance between algorithm accuracy and efficiency. From the perspective of algorithm computational complexity, the improved SSD algorithm uses VGG–16 as the basic network, with a large number of network parameters and a large model memory capacity. YOLOv4 uses end-to-end detection, with improved speed but decreased accuracy.
[0060] The improved enhanced SSD model and Swin Transformer network in this embodiment can effectively extract the features of floating objects in a complex environment under the condition of reducing the algorithm complexity, and effectively improve the detection accuracy of floating objects at different heights and in different environments.
[0061] Table 5: Detection performance of different algorithms in the simple scenario.
[0062] ; Among them, the FLOPs in Table 5 are the number of floating-point operations required for the model to complete one forward propagation.
[0063] In complex water surface scenarios, due to problems such as dynamic light and shadow, water wave disturbance, and strong light reflection in the detection environment, the error of the target detection algorithm increases to a certain extent. Compared with the other three algorithms, the present invention has strong adaptability and robustness in complex scenarios, can better cope with the interference caused by complex environmental factors, and ensure the detection and tracking effect.
[0064] From the perspective of the computational complexity of the detection and tracking speed, due to the increased complexity of the floating object background environment in the target tracking algorithm, the processing efficiency of complex images decreases, and both the speed and computational complexity are lower than the detection performance in a simple water surface environment. However, in the complex scenario where the accuracy of the present invention reaches 83.24%, it maintains an average detection and tracking speed of 22.32 fps and a floating-point operation volume of 7.83×10 9 operations, solving the problems of dynamic light and shadow, water wave disturbance, and strong light reflection in the detection and tracking of floating objects on the water surface based on machine vision, as shown in Table 6.
[0065] Table 6: Detection and tracking performance of different algorithms in complex scenarios.
[0066] ; The computational complexity of the present invention is lower than that of the improved SSD, YOLOv4, and MOT algorithms.
[0067] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. An adaptive recognition and tracking method for floating objects on the water surface, characterized in that, Including: By adjusting the structures of the deep low-resolution detection layer and the shallow high-resolution detection layer of the SSD model, dividing the detection layers input to the SSD model into several scales for processing, and removing the set spatial resolution check unit, an improved enhanced SSD model is constructed to detect floating objects on the water surface; The open-source floating object dataset on the water surface is divided into a training set and a validation set for training and validating the improved enhanced SSD model, and the improved enhanced SSD model is applied to actual image data for testing to detect the performance and generalization ability of the improved enhanced SSD model; Among them, the improved enhanced SSD model also processes the image layer by layer by integrating the Swin Transformer local window self-attention mechanism; The historical coordinates of the detected floating objects on the water surface are input into the LSTM for predicting the Kalman filter parameters, ensuring that the Kalman filter dynamically performs the optimal state estimation on the floating objects on the water surface, and realizing the target tracking of the dynamic floating objects on the water surface.
2. The method for adaptively identifying and tracking water surface floating objects according to claim 1, wherein Generate default boxes of different scales on different feature layers of the SSD. In SSD300, the high-level feature map is 3×3, and the scale of the default box is relatively large, with the maximum scale = 0.85; the low-level feature map is 76×76, and the scale of the default box is relatively small, with the minimum scale = 0.
15. At the same time, automatically adjust the aspect ratio of the default box according to the extreme aspect ratios in the training set.
3. The method for adaptively identifying and tracking water surface floating objects according to claim 1, wherein The 6 detection layers input to the SSD300 are divided into three scales for processing.
4. The adaptive recognition and tracking method for floating objects on the water surface according to claim 3, characterized in that The spatial resolutions of 76×76 and 38×38 are regarded as large sizes, 19×19 and 10×10 are regarded as medium sizes, and the spatial resolutions of 5×5 and 3×3 are regarded as small sizes.
5. The method for adaptively identifying and tracking water surface floating objects according to claim 1, wherein Remove the 1×1 spatial resolution check unit.
6. The method for adaptively identifying and tracking floating objects on the water surface according to claim 1, characterized in that, Adopt LSTM to dynamically adjust the process noise covariance Q and the observation noise covariance R of the Kalman filter.
7. The method for adaptively identifying and tracking floating objects on the water surface according to claim 1, wherein, The steps to construct an adaptive target tracking model for floating objects on the water surface are as follows: 1) Initialization: Set the initial process noise covariance Q0 and the initial observation noise covariance R0, and initialize the LSTM hidden state; 2) Perform loop tracking at each time step k: Input historical data: Input the errors and residuals of the past T steps into the LSTM; Generate adjustment factors: LSTM outputs α k and β k , and ensure positive numbers through the activation function Softplus(); Adjust the noise covariance: Calculate the process noise covariance for the current time step Q k and the observation noise covariance for the current time step R k as follows: ; ; Among them, is the process noise adjustment factor, is the observation noise adjustment factor; Adaptive noise adjustment is as follows: , ; wherein, is the process noise covariance at the adjusted current time step, is the process noise covariance at the adjusted current time step; Kalman Filter Prediction and Update: Using dynamics and to perform state estimation; Record the errors and residuals: for input at the next time step; Minimize the tracking error as follows: ; Among them, is the true target state at time k, is the state after the update of the Kalman filter; Backpropagation through time (BPTT) is used to optimize the LSTM parameters, and the Kalman filter parameters are fixed.
8. An adaptive recognition and tracking system for floating objects on the water surface, characterized in that, Including: An identification module for constructing an improved enhanced SSD model to identify the target of floating objects on the water surface; A tracking module for dynamically adjusting the Kalman filter parameters through LSTM, constructing an adaptive target tracking model for floating objects on the water surface, and tracking the target of floating objects on the water surface.
Citation Information
Patent Citations
DeepSort water surface floating object multi-target tracking method based on lightweight SSD
CN114022812A
Method for detecting and identifying floating objects on water based on improved SSD (Solid State Disk) algorithm
CN114782772A
Water surface floating object target detection and tracking method based on space-time information fusion
CN116385915A
Method and apparatus for tracking moving object by using kalman filter
KR1020160110773A
Drone detection using modified one-stage object detectors
US20250117690A1