Water surface floating object adaptive identification and tracking method and system

By combining the improved SSD model with the Swin Transformer network and LSTM and Kalman filtering, an adaptive surface floating object recognition and tracking system is constructed, which solves the problem of floating object detection and tracking in complex scenarios and achieves high-precision and fast multi-scenario adaptability.

CN120355753BActive Publication Date: 2025-10-14CHUZHOU UNIV +1
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510820155.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-10-14
Estimated Expiration
2045-06-19

AI Technical Summary

Technical Problem

Existing technologies are not accurate enough in detecting and tracking floating objects on the water surface in scenarios such as dynamic floating objects on the water surface, small-sized target floating objects, water surface fluctuations, and strong light reflections.

Method used

An improved enhanced SSD model is fused with the Swin Transformer network, combined with LSTM and Kalman filtering, to construct an adaptive surface floating object target recognition and tracking system. By adjusting the detection layer structure and integrating the local window self-attention mechanism, high-precision detection and dynamic tracking are achieved.

Benefits of technology

It achieves high-precision detection and tracking of floating objects in complex scenarios, has low computational complexity, fast identification and dynamic monitoring capabilities, and is suitable for multi-scenario applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120355753B_ABST
    Figure CN120355753B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of image recognition, and discloses a water surface floating object adaptive recognition and tracking method and system, which comprises the following steps: constructing an improved enhanced SSD model and fusing a Swin Transformer local window self-attention mechanism to process images layer by layer; detecting water surface floating objects by using the improved enhanced SSD model; inputting the historical coordinates of the water surface floating objects in the detection result into an LSTM to predict Kalman filtering parameters, ensuring that Kalman filtering dynamically optimally estimates the state of the water surface floating objects, and realizing water surface dynamic floating object target tracking. The application has the advantages of low operation complexity, fast floating object recognition, high precision, accurate dynamic tracking and monitoring, strong consistency of application precision in multiple scenes, and is suitable for dynamic water surface floating object monitoring in complex scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image recognition, and in particular to a method and system for adaptively identifying and tracking water surface floating objects. Background Art

[0002] Water environment monitoring and protection are crucial components of global environmental monitoring, with monitoring of floating debris a key area of ​​this research. Floating debris refers to solid waste floating on the surfaces of rivers, lakes, and oceans. It impacts water quality and ecosystems, and has serious negative impacts on human life, particularly leading to river blockages and urban waterlogging. Therefore, monitoring and managing floating debris has become a crucial component of urban environmental monitoring and protection.

[0003] Traditional monitoring of floating objects on the water surface mainly relies on manual inspections or image processing technology. The monitoring data of floating objects on the water surface is collected through fixed or mobile image acquisition equipment. Then, image processing technology, such as image target recognition algorithm based on deep learning, is used to improve the recognition accuracy and efficiency of floating objects on the water surface. However, most surface floating recognition based on image processing technology is basically based on calm water scenes to obtain high recognition accuracy. Under natural conditions, complex scenes such as water surface fluctuations, water flow, strong light reflection on the water surface, foggy water surface, etc. are common. The use of traditional image processing technology can no longer meet the requirements of high-precision recognition and tracking of floating objects on the water surface in complex scenes.

[0004] The patent with publication number CN118918486B discloses a method and system for identifying targets in water environments to solve the technical problem that water surface reflections and floating object reflections significantly reduce detection accuracy. The patent proposes a yolov8-OMCIL network model, which is based on the yolov8 network model, the hybrid network inverse residual attention mechanism and the cross-modal attention mechanism. It enhances the ability to extract target features in water environments under different lighting conditions, but does not involve the detection of floating objects on the water surface in scenarios such as dynamic floating objects on the water surface, small-sized target floating objects and water surface fluctuations.

[0005] The patent with publication number CN118887572B discloses a method for identifying the density of floating objects on the water surface based on images. It focuses on solving the spatial area and water level calibration points of floating objects on the water surface, performs abnormality verification on the monitoring position of the image, and realizes an effective quantitative assessment of the density of floating objects. It does not involve issues such as the identification of floating objects on the water surface in multiple scenarios and accuracy assurance.

[0006] The patent with publication number CN216899018U discloses an integrated design of a real-time online monitoring and early warning system for floating objects on the water surface. It focuses on the construction of an integrated real-time data acquisition platform for floating objects on the water surface and a remote data analysis and early warning platform, and uses fixed observation point data acquisition and remote data analysis to achieve floating object monitoring. However, it does not involve the detection of floating objects on the water surface in scenarios such as dynamic floating objects on the water surface, small-sized target floating objects and water surface fluctuations.

[0007] Patent publication number CN114220044B discloses a method for detecting floating objects in rivers based on AI algorithms, which solves the problem of high-precision identification and tracking of floating objects on the water surface in existing complex scenarios.

[0008] Therefore, existing methods or systems fail to solve the problem of detecting and tracking floating objects on the water surface in scenarios such as dynamic floating objects on the water surface, small-sized target floating objects, water surface fluctuations and strong light reflections. Summary of the Invention

[0009] In response to the above-mentioned technical deficiencies, the technical problem to be solved by the present invention is to provide a method and system for adaptive identification and tracking of floating objects on the water surface, aiming to solve the problem that image processing technology is not accurate enough in detecting and tracking floating objects on the water surface in scenarios such as dynamic floating objects on the water surface, small-sized target floating objects, water surface fluctuations and strong light reflections.

[0010] To solve the above technical problems, the present invention adopts the following technical solutions: The present invention provides a method for adaptive identification and tracking of floating objects on the water surface, comprising: adjusting the structure of the deep low-resolution detection layer and the shallow high-resolution detection layer of the SSD model, dividing the detection layer input by the SSD model into several scales for processing, and removing the set spatial resolution inspection unit, thereby constructing an improved enhanced SSD model for detecting floating objects on the water surface;

[0011] The open-source surface floating object dataset was divided into a training set and a validation set for training and validating the improved enhanced SSD model. The improved enhanced SSD model was then applied to actual image data for testing to examine the performance and generalization ability of the improved enhanced SSD model.

[0012] The improved enhanced SSD model also integrates the Swin Transformer local window self-attention mechanism to process images layer by layer;

[0013] The historical coordinates of the detected floating objects on the water surface are input into the LSTM to predict the Kalman filter parameters, ensuring that the Kalman filter dynamically estimates the state of the floating objects on the water surface and realizes the tracking of dynamic floating objects on the water surface.

[0014] Furthermore, SSD generates default boxes of different scales on different feature layers. In SSD300, the high-level feature map is 3×3, and the default box scale is larger than the maximum scale. =0.85; the low-level feature map is 76×76, and the default box scale is smaller than the minimum scale =0.15, and automatically adjust the default box aspect ratio according to the extreme aspect ratios in the training set.

[0015] Furthermore, the six detection layers of SSD300 input are divided into three scales for processing.

[0016] Furthermore, 76×76 and 38×38 spatial resolutions are used as large sizes, 19×19 and 10×10 as medium sizes, and 5×5 and 3×3 spatial resolutions as small sizes.

[0017] Furthermore, the 1×1 spatial resolution inspection unit is removed.

[0018] Furthermore, LSTM is used to dynamically adjust the process noise covariance Q and observation noise covariance R of the Kalman filter.

[0019] Furthermore, the steps to construct an adaptive surface floating object tracking model are as follows:

[0020] 1) Initialization: Set the initial process noise covariance Q0 and the initial observation noise covariance R0, and initialize the LSTM hidden state;

[0021] 2) Perform cyclic tracking at each time step k:

[0022] Input historical data: Input the error and residual of the past T steps into LSTM;

[0023] Generate adjustment factor: LSTM output α k and β k , ensure positive numbers by activating the function Softplus();

[0024] Adjust noise covariance: Calculate the process noise covariance at the current time step Q k and the observation noise covariance at the current time step R k as follows:

[0025] ;

[0026] ;

[0027] in, is the process noise adjustment factor, is the observation noise adjustment factor;

[0028] The adaptive noise adjustment is as follows:

[0029] , ;

[0030] wherein, is the adjusted process noise covariance of the current time step, is the adjusted process noise covariance of the current time step;

[0031] Kalman filter prediction and update: using dynamic and for state estimation;

[0032] Record error and residual: for next time step input;

[0033] The tracking error minimization is as follows:

[0034] ;

[0035] wherein, is the real target state at k time, is the updated state of the Kalman filter;

[0036] Backpropagation through time (BPTT) optimizes LSTM parameters, and Kalman filter parameters are fixed.

[0037] The water surface floating object adaptive recognition and tracking system comprises:

[0038] A recognition module is used to construct an improved enhanced SSD model to recognize the water surface floating object target.

[0039] A tracking module is used to adjust the Kalman filter parameters through an LSTM to construct an adaptive water surface floating object target tracking model to track the water surface floating object target.

[0040] The present application has the following advantages:

[0041] 1. The improved enhanced SSD and Swin Transforme network fusion mode is adopted, so that the high-precision detection problem of water surface floating objects in foggy days is solved; and the LSTM and Kalman filter water surface floating object target tracking algorithm is adopted to ensure that the Kalman filter dynamically estimates the state of the floating object to the best, and the tracking and monitoring of the water surface moving floating object are ensured.

[0042] 2. The present application has the advantages of low computational complexity, fast floating object recognition, high precision, accurate dynamic tracking and monitoring, strong consistency of application precision in multiple scenes, and is suitable for dynamic water surface floating object monitoring in multiple scenes. BRIEF DESCRIPTION OF DRAWINGS

[0043] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0044] Figure 1 This is a flow chart of the method for adaptively identifying and tracking floating objects on the water surface provided in Example 1 of the present invention.

[0045] Figure 2 This is a schematic diagram of the detection results of the improved enhanced SSD model in Example 1 of the present invention on an open source dataset.

[0046] Figure 3 This is a schematic diagram of the detection results of the improved enhanced SSD model in Example 1 of the present invention in an actual application dataset.

[0047] Figure 4 Schematic diagram of the detection results of the improved enhanced SSD model in Example 1 of the present invention and the traditional Transformer on open source datasets.

[0048] Figure 5 This is a schematic diagram of the detection results of the open source dataset by fusing the improved enhanced SSD model and Swin Transformer in Example 1 of the present invention.

[0049] Figure 6 Schematic diagram comparing the adaptive target tracking algorithm, extended Kalman filtering, and recursive least squares center position prediction results in Example 1 of the present invention.

[0050] Figure 7 Schematic diagram of center position error comparison in different scenarios in Example 1 of the present invention.

[0051] Figure 8 This is a schematic diagram of algorithm comparison results in a simple scenario in Example 2 of the present invention. DETAILED DESCRIPTION

[0052] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0053] Example 1

[0054] like Figures 1-7As shown, this embodiment provides a method for adaptively identifying and tracking floating objects on a water surface, including:

[0055] By generating default boxes of different scales on different feature layers using SSD, dividing the detection layer of SSD input into several scales for processing, and removing the 1×1 spatial resolution inspection unit, an improved enhanced SSD model is constructed for identifying small-sized floating objects on the water surface. It should be noted that:

[0056] The SSD (Single Shot MultiBox Detector) neural network is mainly used to detect and identify multiple objects in an image and predict the categories and locations of these objects at the same time. It can complete target detection in a single forward pass, which improves detection speed. The core idea of ​​SSD is to achieve multi-scale target detection by adding multiple prediction layers to the last few layers of the convolutional neural network (CNN). It can capture the size and shape of various objects of different sizes and use the default bounding box as the starting point for prediction, which reduces the computational burden and improves detection speed. It can achieve faster detection speed while maintaining a high level of accuracy. However, the SSD underlying feature map lacks semantic information, which can easily lead to missed detection of small-sized targets. To overcome the impact of drone flight altitude and trajectory changes on the imaging resolution of floating objects on the water surface, this paper proposes an improved enhanced SSD algorithm. Specifically, the structure of the deep low-resolution detection layer and the shallow high-resolution detection layer is adjusted, and efficient target detection is achieved through multi-scale default boxes and joint optimization of classification and localization tasks:

[0057] SSD generates default boxes of different scales on different feature layers to cover multi-scale targets. The scale formula is:

[0058] ;

[0059] Among them, S k is the default box of different scales on the k feature layer, S min is the minimum scale, corresponding to the default box size of the lowest feature map, S max is the maximum scale, corresponding to the size of the highest-level feature frame, m is the total number of feature layers, k is the number of layers, SSD300 is used this time, with a total of 6 feature layers, and its input image size is 300×300; the aspect ratio of each default frame is Usually chosen as {1,2,3,1 / 2,1 / 3}{1,2,3,1 / 2,1 / 3}, and an additional scale is added is the default box.

[0060] The default box aspect ratio formula is as follows:

[0061] ;

[0062] in, is the width of the default box of the k feature layer, The height of the default box of the k-th feature layer. The high-level feature map is 3×3, and the default box size is larger. =0.85, the low-level feature map is 76×76, and the default box size is small =0.15, which automatically adjusts the aspect ratio based on the extreme aspect ratios in the training data.

[0063] The traditional SSD input six detection layers are divided into three scales for processing: large size (76×76, 38×38), medium size (19×19, 10×10), and small size (5×5, 3×3). Since sliding windows below 3×3 in traditional SSD have difficulty extracting small-scale floating objects, the present invention abandons the original 1×1 spatial resolution inspection unit. The improved SSD algorithm can solve the problem of high variability in the imaging process using machines or manual methods, which leads to the non-fixed size of floating objects, and can effectively detect small-sized floating objects.

[0064] The improved enhanced SSD model is used to train the model on the open source "Water Surface Floating Object Dataset-2400" dataset shown in Table 1. This dataset is designed for deep learning and contains 2,400 carefully collected and processed high-quality images, focusing on the identification of floating objects on the water surface. All images in the dataset are taken in real environments, which are authentic and practical. They have undergone professional post-processing to ensure image quality, thereby ensuring data accuracy, making them very suitable for the training and verification stages of the model.

[0065] Table 1: Water surface floating object dataset - 2400 data types.

[0066] ;

[0067] Detect floating objects such as polyethylene terephthalate (PET bottles), polystyrene foam boxes, metals, plastic sheets, ropes, plastic buoys, nets, glass pots, plastic products, etc. from images with detection accuracy as high as Figure 2 As shown in the figure, the improved enhanced SSD algorithm is applied to the collected image data, and the detection accuracy of the above floating objects is as follows: Figure 3 As shown in Table 2, the input parameters of the improved enhanced SSD model detection network training are shown in Table 2.

[0068] Table 2: Improved SSD detection network training input parameters.

[0069] ;

[0070] The Swin Transformer network is used for feature extraction in the improved enhanced SSD model to improve the accuracy of the improved enhanced SSD model in identifying dynamic floating objects on the water surface. It should be noted that:

[0071] In foggy and water surface with large fluctuations, the traditional SSD and the optimized neural network cannot complete feature extraction under complex conditions. The Swin Transformer network is used for feature extraction of multi-scale water surface floating objects under complex conditions. The local window self-attention mechanism is used to process the image layer by layer; the shift window mechanism is used to capture the long-range dependency and relative position bias between different windows, effectively solving the problem that the SSD network cannot effectively extract features under complex conditions; the Swin Transformer network uses hierarchical window design and relative position bias to effectively model the global dependency of the image while reducing the computational complexity, divides the image into non-overlapping equal-size windows (regular windows), and calculates the self-attention in each window independently to reduce the computational amount (the complexity is reduced from O(N 2 ) to O(M 2 ·N / M 2 )=O(N), M is the window size), and the window is moved pixels to the lower right corner in adjacent layers to form an overlapping detection area (moving window) between adjacent windows. The Swin Block is composed of multiple stages, each stage alternately uses regular windows and moving windows to capture local features and interact with different window pixels, and its receptive field is gradually expanded. After multiple layers of stacking, the model gradually covers the entire image to realize global dependency detection.

[0072] The Swin Transformer network introduces the position encoding in the attention calculation, adjusts the attention weight through the relative position of the query (Q) and the key (K), replaces the traditional absolute position encoding, and adds a relative position bias term P, to the self-attention formula, which is a learnable relative position matrix, is the number of pixels in the window; the relative position range of the pixels in the window is [−M+1,M−1], and the total number of parameters is (2M−1)×(2M−1). The complexity of Swin (window attention) is O(HW·M2).

[0073] As shown in Figure 4 , the improved enhanced SSD and the traditional Transformer algorithm are compared in the open source dataset detection results, and the fusion accuracy of the improved enhanced SSD and the traditional Transformer algorithm is shown in Table 3, and the fusion of SwinTransformer and the improved enhanced SSD in the open source dataset detection results is shown in Figure 5The detection accuracy of the four water surface scenes is shown in Table 4, which ensures the accuracy of recognition in complex environments and the feasibility of reasoning on embedded platforms.

[0074] Table 3: Detection accuracy of the improved enhanced SSD+traditional Transformer on open source datasets.

[0075]

[0076] Table 4: Detection accuracy of four water surface scenes.

[0077]

[0078] Preferably, the Kalman filter parameters are dynamically adjusted by LSTM to construct an adaptive surface floating object tracking model for tracking surface floating objects.

[0079] Actual drone trajectories exhibit a variety of periodic flight motions. Simple target detection algorithms struggle to accurately and rapidly infer the state. Therefore, we employ a fusion of LSTM and Kalman filtering to dynamically adjust the process noise covariance Q and the observation noise covariance R, ensuring that the Kalman filter optimally estimates the state of dynamic floating objects.

[0080] Specifically, the steps to build an adaptive surface floating object tracking model are as follows:

[0081] 1) Initialization: Set the initial process noise covariance Q0 and the initial observation noise covariance R0, and initialize the LSTM hidden state.

[0082] 2) Perform cyclic tracking at each time step k:

[0083] Input historical data: Input the errors and residuals of the past T steps into LSTM.

[0084] Generate adjustment factor: LSTM output α k and β k , ensured to be positive by activating the function Softplus().

[0085] Adjust noise covariance: Calculate the process noise covariance at the current time step Q k and the observation noise covariance at the current time step R k as follows:

[0086] ;

[0087] ;

[0088] in, is the process noise adjustment factor, is the observation noise adjustment factor.

[0089] Adaptive noise adjustment is as follows:

[0090] , ;

[0091] in, is the adjusted process noise covariance of the current time step, is the adjusted process noise covariance of the current time step.

[0092] Kalman filter prediction and update: using dynamic and Perform state estimation.

[0093] Record errors and residuals: used for input in the next time step.

[0094] The tracking error is minimized as follows:

[0095] ;

[0096] in, is the actual target state at time k, is the updated state of the Kalman filter.

[0097] Backpropagation optimizes LSTM parameters through time BPTT, fixes Kalman filter parameters (state transfer matrix F, observation matrix H, initial process noise covariance Q0, and initial observation noise covariance R0), and dynamically adjusts Kalman filter parameters through LSTM, significantly improving robustness to nonlinear and time-varying noise. It is suitable for target tracking in complex scenarios.

[0098] To demonstrate the adaptive surface floating object tracking model of the present invention, this embodiment first performs target detection on a large number of video frames, and then inputs the historical coordinates of the floating object into the LSTM to predict the Kalman filter parameters. During the experiment, simple surface videos and videos with large surface fluctuations were used for experiments. The algorithm of the present invention was compared with the extended Kalman filter (EKF) and recursive least squares (RLS) for algorithm verification. The results are shown in the figure. Figure 6 As shown in the figure, it shows that the present invention can achieve a relatively ideal effect in dealing with the variable motion trajectory of the observation equipment. The comparison of the center position error in different scenarios is as follows: Figure 7 shown.

[0099] This embodiment also provides an adaptive recognition and tracking system for floating objects on a water surface, including:

[0100] The recognition module is used to build an improved enhanced SSD model to identify floating objects on the water surface;

[0101] The tracking module is used to dynamically adjust the Kalman filter parameters through LSTM, build an adaptive surface floating object tracking model, and track surface floating objects.

[0102] Example 2

[0103] This embodiment is the second embodiment of the present invention. Different from the first embodiment, this embodiment provides a verification test of the method and system for adaptive identification and tracking of floating objects on the water surface, and verifies and explains the technical effects adopted in this method.

[0104] We collected aerial images of floating objects on the water surface from drones in different water scenes and under different weather conditions, constructed a dataset of floating objects on the water surface under four scenarios, and performed feature enhancement such as Gaussian blurring and histogram processing on the dataset.

[0105] Further experiments were conducted on tracking dynamic floating objects on the water surface by fusing the Swin Transformer with the improved enhanced SSD model and combining it with an adaptive filtering algorithm: comparative experiments were conducted using target detection and target tracking algorithms to verify the performance of the present invention; the detection algorithm was selected in two stages, and the target detection algorithm adopted the improved SSD algorithm, YOLOV4 algorithm, and Faster-RCNN algorithm; the target tracking algorithm adopted the MOT algorithm; the 116 video sequences were divided into training and test sets in a 7:3 ratio, of which the training set contained 82 video sequences and 2817 images, and the test set contained 34 video sequences and 1207 images, both of which contained data from four different scenes, divided into simple water surface scenes (calm water surface) and complex water surface scenes (dynamic light and shadow, water surface reflection, strong light reflection water surface, and foggy water surface), to verify the detection performance and generalization ability of the algorithm.

[0106] Specifically, in a simple scenario, the floating object detection accuracy and success rate of the four algorithms of the present invention are as follows: Figure 8As shown in the table, all algorithms achieve good detection results, with the detection success rate exceeding 90%. Table 5 shows that the algorithms in this embodiment outperform the other algorithms in both detection accuracy and success rate curve area. In terms of detection speed, the improved SSD algorithm enhances the detection layer, improving accuracy, but the increase in algorithm parameters reduces the detector's processing efficiency. The Faster-RCNN and YOLOv4 algorithms achieve significantly faster speeds than the target detection algorithms. The algorithm in this embodiment achieves a speed of 28.35 fps by combining the advantages of the improved enhanced SSD algorithm and the Swin Transformer algorithm, achieving a balance between accuracy and efficiency. In terms of computational complexity, the improved SSD algorithm uses VGG-16 as its base network, resulting in a large number of network parameters and a large model memory capacity. YOLOv4 uses end-to-end detection, which improves speed but reduces accuracy.

[0107] The improved enhanced SSD model and Swin Transformer network of this embodiment, and the fusion strategy can effectively extract floating object features in complex environments, effectively improving the accuracy of floating object detection at different heights and in different environments while reducing the complexity of the algorithm.

[0108] Table 5: Detection performance of different algorithms in simple scenarios.

[0109] ;

[0110] Among them, FLOPs in Table 5 is the number of floating-point operations required for the model to complete one forward propagation.

[0111] In complex water scenes, the target detection algorithm's error increases to a certain extent due to the presence of dynamic lighting, water wave disturbances, and strong light reflections. Compared to the other three algorithms, this proposed method has greater adaptability and robustness in complex scenes, effectively handling interference caused by complex environmental factors and ensuring detection and tracking effectiveness.

[0112] From the perspective of detection and tracking speed and computational complexity, the target tracking algorithm has lower processing efficiency of complex images due to the increased complexity of the floating object background environment. The speed and computational complexity are lower than the detection performance in a simple water surface environment. However, the accuracy of the proposed method reaches 83.24% in complex scenes, while maintaining an average detection and tracking speed of 22.32 fps and a floating point number of 7.83×10 9 The operation volume solves the problems of dynamic light and shadow, water wave disturbance and strong light reflection in the detection and tracking of floating objects on the water surface based on machine images, as shown in Table 6.

[0113] Table 6: Detection and tracking performance of different algorithms in complex scenarios.

[0114] ;

[0115] The computational complexity of the present invention is lower than that of the improved SSD, YOLOv4 and MOT algorithms.

[0116] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art will be able to modify the technical solutions described in the aforementioned embodiments or replace some of the technical features therein with equivalents. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.

Claims

1. The method for adaptively identifying and tracking floating objects on the water surface is characterized by: include: By adjusting the structure of the deep low-resolution detection layer and the shallow high-resolution detection layer of the SSD model, dividing the detection layer input to the SSD model into three scales, and removing the set spatial resolution inspection unit, an improved enhanced SSD model is constructed to detect floating objects on the water surface. The open-source surface floating object dataset was divided into a training set and a validation set for training and validating the improved enhanced SSD model. The improved enhanced SSD model was then applied to actual image data for testing to examine the performance and generalization ability of the improved enhanced SSD model. The improved enhanced SSD model also integrates the Swin Transformer local window self-attention mechanism to process images layer by layer; The historical coordinates of the detected floating objects on the water surface are input into the LSTM to predict the Kalman filter parameters, ensuring that the Kalman filter dynamically estimates the state of the floating objects on the water surface and realizes the tracking of dynamic floating objects on the water surface.

2. The method for adaptively identifying and tracking floating objects on a water surface according to claim 1, wherein: SSD generates default boxes of different scales on different feature layers. In SSD300, the high-level feature map is 3×3, and the maximum scale of the default box is S max =0.85; the low-level feature map is 76×76, and the default minimum box size is S min =0.15, and automatically adjust the default box aspect ratio according to the extreme aspect ratios in the training set.

3. The method for adaptively identifying and tracking floating objects on a water surface according to claim 1, wherein: The 6 detection layers of SSD300 input are divided into three scales for processing.

4. The method for adaptively identifying and tracking floating objects on a water surface according to claim 3, wherein: 76×76 and 38×38 spatial resolutions are used as large sizes, 19×19 and 10×10 as medium sizes, and 5×5 and 3×3 spatial resolutions as small sizes.

5. The method for adaptively identifying and tracking floating objects on a water surface according to claim 1, wherein: Remove the 1×1 spatial resolution inspection unit.

6. The method for adaptively identifying and tracking floating objects on a water surface according to claim 1, wherein: LSTM is used to dynamically adjust the process noise covariance Q and observation noise covariance R of the Kalman filter.

7. The method for adaptively identifying and tracking floating objects on a water surface according to claim 1, wherein: The steps to build an adaptive surface floating object tracking model are as follows: 1) Initialization: Set the initial process noise covariance Q0 and the initial observation noise covariance R0, and initialize the LSTM hidden state; 2) Perform cyclic tracking at each time step k: Input historical data: Input the error and residual of the past T steps into LSTM; Generate adjustment factor: LSTM output α k and β k , ensure positive numbers by activating the function Softplus(); Adjust noise covariance: Calculate the process noise covariance at the current time step Q k and the observation noise covariance at the current time step R k as follows: ; ; Among them, α k is the process noise adjustment factor, β k is the observation noise adjustment factor; Adaptive noise adjustment is as follows: , 8. Among them, is the adjusted process noise covariance of the current time step, is the adjusted process noise covariance of the current time step; Kalman filter prediction and update: using dynamic and Perform state estimation; Record errors and residuals: used as input for the next time step; The tracking error is minimized as follows: ; in, is the actual target state at time k, is the updated state of the Kalman filter; Backpropagation through time BPTT optimizes LSTM parameters and fixes Kalman filter parameters.

9. The adaptive identification and tracking system for floating objects on the water surface is characterized by: include: The recognition module is used to divide the detection layer into multi-scale processing through SSD, remove the spatial resolution inspection unit, combine with the Swin transformer for feature extraction, and build an improved enhanced SSD model to identify floating objects on the water surface; The tracking module is used to dynamically adjust the Kalman filter parameters through LSTM, build an adaptive surface floating object tracking model, and track surface floating objects.

Citation Information

Patent Citations

  • A method for detecting floating debris in river channels based on AI algorithms

    CN114220044B

  • Method for identifying density of floating objects on water surface based on image

    CN118887572B

  • A method and system for identifying water environment targets

    CN118918486B

  • Integrally-designed real-time online monitoring and early warning system for floating objects on water surface

    CN216899018U

  • DeepSort water surface floating object multi-target tracking method based on lightweight SSD

    CN114022812A