A method and system for monitoring river surface floating objects based on a drone

By constructing a riverbank model and combining it with multi-source fusion technology of lidar, visible light and infrared thermal imaging data, the problem of inaccurate monitoring of floating objects in rivers has been solved, and efficient and accurate identification and monitoring of floating objects on the river surface has been achieved.

CN121746970BActive Publication Date: 2026-05-29AEROSPACE INFORMATION RES INST CAS

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
AEROSPACE INFORMATION RES INST CAS
Filing Date
2025-12-12
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Current technologies for monitoring floating objects in rivers are not accurate enough. Traditional manual inspections are costly, have limited coverage, and poor real-time performance. Video analysis-based methods are susceptible to interference from water surface reflections and wave disturbances. Single-sensor drone solutions have reduced recognition accuracy in strong light, low light, or turbid water scenarios. Infrared thermal imaging technology is not sensitive enough to targets with low temperature differences. Multi-sensor fusion has not been deeply integrated with multi-source feature collaboration, resulting in low detection accuracy in complex scenarios.

Method used

Riverbank model is constructed by acquiring riverbank boundary data using drones. Multi-source data fusion is performed by combining lidar point cloud data, visible light image data, and infrared thermal imaging data. Lidar point cloud data is optimized using elevation filtering, and spatiotemporal alignment is applied to the data. Color, texture, and temperature difference features are extracted by combining the prediction probabilities and feature correlation constraints of different classification models. Weights are set to adjust the probability set, and the floating object confidence score is obtained through fusion.

Benefits of technology

It improves the accuracy and reliability of river surface floating object monitoring, can identify hidden floating objects that are easily overlooked under visible light, reduces false positives and false negatives, enhances the real-time performance and coverage of monitoring, and improves detection accuracy in complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121746970B_ABST
    Figure CN121746970B_ABST
Patent Text Reader

Abstract

The application provides a kind of based on unmanned aerial vehicle's river surface floating object monitoring method and system, belong to water environment protection monitoring field, including: obtaining the river bank boundary line data of target area and utilize the river bank boundary line data to build river bank model;Based on the river bank model, utilize unmanned aerial vehicle to cruise and obtain laser radar point cloud data, visible light image data and infrared thermal imaging data;The visible light image data and the infrared thermal imaging data are fused to obtain fused image data;Utilize the fused image data and the laser radar point cloud data to determine river surface floating object.The application can not only rely on image to determine the appearance characteristics of floating object, but also can accurately obtain its position, volume and other key information by means of point cloud data, greatly improving the accuracy of river surface floating object monitoring.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of water environment protection monitoring, and in particular to a method and system for monitoring floating objects on river surfaces based on unmanned aerial vehicles (UAVs). Background Technology

[0002] With global water pollution becoming increasingly severe, efficient monitoring of floating debris on river surfaces has become a crucial aspect of water environment protection. Traditional manual inspection methods are limited by high labor costs, small coverage areas, and poor real-time performance, making them unsuitable for handling sudden pollution incidents and monitoring large areas of water. While existing video analysis-based methods (such as Gaussian mixture model background difference) can achieve dynamic detection, they are susceptible to interference from water surface reflections and wave disturbances, and rely solely on two-dimensional visual information, failing to capture the three-dimensional morphological characteristics of floating debris, resulting in a persistently high rate of missed detections for small targets.

[0003] In recent years, monitoring solutions using drones equipped with visible light cameras have become increasingly popular (such as the YOLO series algorithms). However, these single sensors have significant limitations. In scenes with strong light, low light, or turbid water, the color and texture features of visible light images become less recognizable, resulting in a high rate of misjudgment of similar objects. While existing infrared thermal imaging technology can capture temperature differences, it is not sensitive enough to targets with low temperature differences. Furthermore, traditional multi-sensor fusion methods often remain at the level of simple data overlay without delving into the collaborative relationships of multi-source features, making it difficult to improve the overall detection accuracy in complex scenarios.

[0004] Traditional manual inspections are costly, have limited coverage, and poor real-time performance, making them unsuitable for large water areas and sudden pollution events. Video-based methods are susceptible to interference from water surface reflections and wave disturbances, relying solely on two-dimensional visual information and failing to capture the three-dimensional morphological features of floating objects, resulting in a high rate of missed detections for small targets. While drones equipped with visible light cameras offer advantages, single sensors suffer from reduced color and texture feature recognition in strong light, low light, or turbid water conditions, leading to high false positive rates. Infrared thermal imaging technology lacks sensitivity to targets with low temperature differences, and traditional multi-sensor fusion methods often involve simple data layer overlay without delving into the collaborative relationships between multi-source features, resulting in low overall detection accuracy in complex scenarios. This invention aims to address the core problem of inaccurate river floating object monitoring in existing technologies, improving the accuracy of floating object identification and its practical application capabilities. Summary of the Invention

[0005] To address the above technical problems, this invention provides a method and system for monitoring floating objects on river surfaces based on unmanned aerial vehicles (UAVs), which solves the problem of insufficient accuracy in monitoring floating objects in river channels in the prior art.

[0006] To achieve the above objectives, one aspect of the present invention provides a method for monitoring floating objects on a river surface based on an unmanned aerial vehicle (UAV), the method comprising: acquiring riverbank boundary line data of a target area and constructing a riverbank model using the riverbank boundary line data; acquiring lidar point cloud data, visible light image data, and infrared thermal imaging data using UAV cruising based on the riverbank model; fusing the visible light image data and the infrared thermal imaging data to obtain fused image data; and determining floating objects on the river surface using the fused image data and the lidar point cloud data.

[0007] Another aspect of the present invention provides a drone-based river surface floating object monitoring system, comprising: a processor, an input device, an output device, and a memory, wherein the processor, the input device, the output device, and the memory are interconnected, wherein the memory is used to store a computer program, the computer program including program instructions, and the processor is configured to invoke the program instructions to execute the drone-based river surface floating object monitoring method.

[0008] The present invention has the following beneficial effects:

[0009] 1. This invention constructs a riverbank model by acquiring riverbank boundary line data, thus defining a precise spatial range for subsequent monitoring and avoiding monitoring area deviation. Based on this riverbank model, it utilizes drone patrols to simultaneously acquire lidar point cloud data, visible light image data, and infrared thermal imaging data, achieving efficient multi-source data acquisition. This approach considers both the spatial dimension information of the data and the target characteristics under different spectra. By fusing visible light and infrared thermal imaging data, the advantages of both types of data can be integrated, improving the ability to identify floating objects on the river surface, especially those easily overlooked under visible light. Finally, by combining the fused image data and lidar point cloud data, the floating objects are identified. This approach not only clearly defines the appearance characteristics of the floating objects based on the images but also accurately obtains key information such as their location and volume using the point cloud data, significantly improving the accuracy of floating object monitoring on the river surface.

[0010] 2. This invention optimizes the original lidar point cloud data through elevation filtering, removes redundant interference, and then performs spatiotemporal alignment processing to ensure accurate matching of lidar point cloud, visible light and infrared thermal imaging data, thereby improving the scientific nature of multi-source data collaborative processing.

[0011] 3. This invention fits the river surface reference plane to the original LiDAR point cloud data to clarify and unify the elevation reference. Then, it extracts the vertical elevation component and calculates the vertical relative elevation to accurately quantify the elevation relationship between the point cloud and the river surface. Finally, it combines the vertical elevation range filtering of floating objects to efficiently remove non-floating object point clouds that are outside the range, thereby reducing the amount of calculation and interference in the later stage and improving the scientific nature of the original LiDAR point cloud filtering data.

[0012] 4. This invention extracts the average water surface temperature and calculates the temperature difference by using infrared thermal imaging data. This allows for the accurate capture of the temperature difference between floating objects and the water surface. By combining this temperature difference with visible light and infrared thermal imaging data, the invention can preserve the appearance details of floating objects in the visible light image and highlight the target identification brought about by the temperature difference in the infrared data. This effectively distinguishes between water surface reflections and floating objects, improving the scientific nature and accuracy of the fused image data.

[0013] 5. By integrating the prediction probabilities of different classification models, this invention leverages the advantages of image data in appearance and temperature identification, while also utilizing the spatial features of point cloud data to improve judgment accuracy. This effectively reduces misjudgments or omissions based on single data points, lowers the probability of erroneous observations, provides dual protection for the accurate identification of floating objects on the river surface, and further enhances the reliability of monitoring.

[0014] 6. This invention avoids feature redundancy or insufficient correlation by setting feature correlation constraints. It incorporates the cross-entropy loss function to form an optimized loss function, which can guide efficient feature collaboration while ensuring the accuracy of model classification, thereby improving the prediction accuracy of the first classification model.

[0015] 7. By normalizing the color, texture, and temperature difference features, this invention can eliminate the interference caused by the difference in the dimensions of different features, thereby improving the efficiency of the first classification model in utilizing features and the accuracy of prediction.

[0016] 8. This invention extracts mutually exclusive category pairs from two probability sets and calculates a first conflict coefficient. Combined with a conflict determination threshold, it assigns influence weights to both sets, accurately identifying data conflicts and assigning differentiated weights, avoiding reliance on a single data type. By adjusting the probability sets through weight adjustments and calculating a second conflict coefficient, the final fusion yields the floating object confidence score, effectively resolving data conflicts and making the fusion result more closely aligned with actual monitoring scenarios, further improving the accuracy of floating object monitoring and identification.

[0017] 9. This invention obtains a historical probability set through a binary classification model, calculates the confidence level and fusion accuracy of historical floating objects by combining multiple candidate thresholds, and finally selects the threshold corresponding to the optimal accuracy, thereby avoiding subjective setting bias and improving the rationality of probability set conflict determination and fusion. Attached Figure Description

[0018] Figure 1 This is a flowchart of a method for monitoring floating objects on a river surface based on an unmanned aerial vehicle (UAV) according to an embodiment of the present invention.

[0019] Figure 2 This is a schematic diagram of a river surface floating object monitoring system based on an unmanned aerial vehicle (UAV) according to an embodiment of the present invention. Detailed Implementation

[0020] Specific embodiments of the present invention will now be described in detail. It should be noted that the embodiments described herein are for illustrative purposes only and are not intended to limit the invention. In the following description, numerous specific details are set forth in order to provide a thorough understanding of the invention. However, it will be apparent to those skilled in the art that these specific details are not necessary to practice the invention. In other instances, well-known circuits, software, or methods have not been specifically described to avoid obscuring the invention.

[0021] Throughout this specification, references to "an embodiment," "an embodiment," "an example," or "an example" mean that a particular feature, structure, or characteristic described in connection with that embodiment or example is included in at least one embodiment of the invention. Therefore, the phrases "in an embodiment," "in an embodiment," "an example," or "an example" appearing in various places throughout the specification do not necessarily refer to the same embodiment or example. Furthermore, specific features, structures, or characteristics can be combined in one or more embodiments or examples in any suitable combination and / or sub-combination. Moreover, those skilled in the art will understand that the illustrations provided herein are for illustrative purposes and are not necessarily drawn to scale.

[0022] Please see Figure 1 The following is a flowchart of a method for monitoring floating objects on a river surface based on an unmanned aerial vehicle (UAV) according to the present invention, which includes the following steps:

[0023] Step S1: Obtain the riverbank boundary data of the target area and construct a riverbank model using the riverbank boundary data.

[0024] In this embodiment, the target area of ​​the river surface corresponding to the monitoring task is first defined. The UAV is responsible for river patrol within a range of 5-10 kilometers. Priority is given to retrieving the existing basic vector data of the riverbank boundary line in this area (including the coordinates of the shoreline inflection points, shoreline direction, etc.) through an authoritative geographic information database. If the basic vector data is missing or the accuracy is insufficient, the UAV can first conduct a low-altitude pre-cruise of the target area to collect high-resolution visible light initial images covering the riverbank and surrounding areas. The boundary lines between the riverbank and the land, and between the riverbank and the river surface, are identified and extracted from the initial images using an image segmentation algorithm to obtain the initial riverbank boundary line data. Subsequently, the obtained riverbank boundary line data is preprocessed to remove boundary anomalies caused by interference factors such as temporary hydraulic facilities and vegetation obstruction on the bank. At the same time, the boundary line data is uniformly converted into a coordinate system consistent with the subsequent UAV patrol data (such as the WGS84 coordinate system) to ensure the consistency of the data spatial reference. In the riverbank model construction stage, based on the preprocessed riverbank boundary data and combined with the topographic elevation information of the target area (which can be obtained from public topographic databases or extracted from the initial LiDAR data of the UAV pre-cruise), geospatial modeling technology is used to integrate the two-dimensional planar information of the riverbank boundary with the three-dimensional topographic information such as the bank slope elevation and the elevation of the bank top / bottom. This results in the construction of a three-dimensional riverbank model that includes the scope of the riverbank planar boundary, the geometry of the bank slope, and the vertical intersection relationship between the riverbank and the river surface. This model needs to clearly define the effective boundary of the UAV patrol area to avoid the patrol range exceeding the river surface or missing key water areas near the riverbank. At the same time, it provides a spatial constraint basis for the subsequent UAV patrol path planning, ensuring that the UAV can accurately collect monitoring data along the river surface area defined by the riverbank model.

[0025] Step S2, based on the riverbank model, uses UAV patrol to obtain lidar point cloud data, visible light image data, and infrared thermal imaging data, including the following sub-steps:

[0026] Step S201: Based on the riverbank model, use UAV cruise to obtain raw lidar point cloud data, raw visible light image data, and raw infrared thermal imaging data.

[0027] In this embodiment, a drone is used to acquire the current water flow velocity, and based on the current water flow velocity, a dynamic angle is determined for the drone to travel back and forth according to the riverbank model. The drone then cruises based on this dynamic angle. Multimodal data is acquired during the cruise, including lidar point cloud data, visible light image data, and infrared thermal imaging data. The dynamic angle satisfies the following formula:

[0028] ;

[0029] in, It is a dynamic included angle. The river's water flow velocity, This refers to the drone's cruising speed.

[0030] The UAV is equipped with an integrated data acquisition payload, which carries a lidar sensor, a high-resolution visible light camera, and an infrared thermal imager. The acquisition trigger signals of the three types of devices are synchronized with the GPS positioning module and the timestamp module through the UAV flight control system, ensuring that each set of acquired data is associated with precise WGS84 coordinate system geographic coordinates and millisecond-level timestamps. During the cruise execution phase, the UAV's flight altitude is set based on the river surface elevation information in the riverbank model. During flight, the flight control system calls the riverbank model in real time for boundary verification to prevent the UAV from deviating from the river surface area. In this process, the LiDAR continuously scans the river surface and near-shore three-dimensional space to generate raw LiDAR point cloud data. The visible light camera captures color images of the river surface at a set frame rate, recording the appearance and color information of floating objects to form raw visible light image data. The infrared thermal imager simultaneously collects images of the river surface temperature distribution, capturing the temperature difference between floating objects and the water body to generate raw infrared thermal imaging data. The three types of raw data are stored in real time in the UAV's local storage unit or transmitted to the ground terminal via 4G / 5G links. Each data segment is labeled with the corresponding cruise path segment and riverbank model position index, providing data traceability for subsequent elevation filtering and spatiotemporal alignment preprocessing.

[0031] Step S202 involves performing elevation filtering on the original lidar point cloud data to obtain the original filtered lidar point cloud data, including the following sub-steps:

[0032] Step S20201: Fit the river surface reference surface using the original point cloud data from the lidar.

[0033] In this embodiment, firstly, points whose x / y coordinates fall outside the river surface boundary defined by the riverbank model are selected from the original lidar point cloud data. Points from non-river surface areas such as nearshore land and hydraulic structures are initially eliminated. Then, a random sampling consensus algorithm is used to iteratively fit the preprocessed river surface point set. A small number of sample points (usually 3-5, the minimum number of points required to solve the plane equation) are randomly selected from the point set. Based on these sample points, an initial plane model (i.e., the candidate river surface reference surface) is calculated. Then, all other points in the river surface point set are traversed, and a preset elevation error threshold (such as ±) is used to determine the model. 0.1m) Determine whether each point conforms to the initial plane model (if it conforms, it is recorded as an interior point; if it does not conform, it is recorded as an exterior point). Count the number of interior points and the fitting error of the interior points relative to the initial plane model. Repeat the above sampling, modeling, and interior point counting steps multiple times (the number of iterations is set according to the size of the point set, usually hundreds to thousands of times). Select the plane model with the highest proportion of interior points and the smallest fitting error of interior points as the optimal river surface digital surface model. Finally, calculate the average elevation value of all interior points based on the optimal river surface digital surface model, and use the plane corresponding to the average elevation value as the final river surface reference plane.

[0034] Step S20202: Extract the vertical elevation component from the original lidar point cloud data.

[0035] In this embodiment, the z-value representing the vertical height in the three-dimensional coordinates of each point in the original LiDAR point cloud data is used as the core extraction object. Then, outlier removal is performed on the initially screened valid point set. A statistical filtering algorithm is used to calculate the mean and standard deviation of all z-values ​​in the point set, and points corresponding to z-values ​​exceeding the mean ± 3 times the standard deviation are removed to eliminate interference from LiDAR measurement errors on the extraction of vertical elevation components. Finally, according to the original index of the point cloud data, the z-value of each valid point is extracted separately and stored as a vertical elevation component dataset, while preserving the correlation between each vertical elevation component and the corresponding three-dimensional coordinates of the point.

[0036] Step S20203: Calculate the vertical relative elevation based on the river surface reference plane and the vertical elevation components.

[0037] In this embodiment, the average elevation benchmark of the river surface is obtained using the river surface reference surface, and the elevation value corresponding to the benchmark is defined as a unified benchmark elevation value. Then, for the previously extracted vertical elevation component dataset, a one-to-one correspondence calculation is performed according to the original index of the point cloud data, and the vertical relative elevation of each valid point is obtained by subtracting the benchmark elevation value from the vertical elevation component of each valid point.

[0038] Step S20204: Obtain the vertical elevation range of the floating object, and use the vertical elevation range of the floating object and the vertical relative elevation to perform elevation filtering on the original lidar point cloud data to obtain the original filtered lidar point cloud data.

[0039] In this embodiment, based on common floating object categories on the river surface (such as plastic waste, aquatic plants, dead branches, oil spills, large plastic foam, etc.), historical sample data of lidar point clouds for each category of floating object are collected. For the vertical relative elevation of the historical point cloud corresponding to each category of floating object, the minimum and maximum values ​​of the historical vertical relative elevation of each category of floating object are calculated to determine the elevation range of a single category of floating object. All single category intervals are then integrated to form an overall vertical elevation range of floating objects covering common categories (e.g., 0.1 m - 2 m). When using this range and vertical relative elevation for elevation filtering, the vertical relative elevation of each valid point is verified one by one according to the original index of the point cloud using the vertical relative elevation dataset. Points whose vertical relative elevation falls within the overall vertical elevation range of floating objects (i.e., potential floating object points) are retained, and points that exceed this range are removed. At the same time, the correlation between each point and the three-dimensional coordinates after filtering is retained to ensure that the spatial position of the point cloud is not lost. Finally, the original filtered lidar point cloud data containing only potential floating object points is obtained.

[0040] Step S203: Spatiotemporally align the raw visible light image data, the raw infrared thermal imaging data, and the raw filtered lidar point cloud data to obtain lidar point cloud data, visible light image data, and infrared thermal imaging data.

[0041] In this embodiment, when performing spatiotemporal alignment on the above three types of raw data, time dimension calibration must be completed first: retrieve the millisecond-level timestamps of the three types of data synchronously recorded by the flight control system during UAV cruise, and use the timestamps of the raw filtered data of lidar point cloud as a reference. Correct the time deviation caused by device triggering delay in the raw visible light image data and raw infrared thermal imaging data through linear interpolation, so as to ensure that the timestamp error of the three types of data is controlled within ±10ms at the same monitoring time, thus forming a time-synchronized dataset. Subsequently, spatial dimension alignment was performed. For the raw filtered data of the LiDAR point cloud, its built-in WGS84 coordinate system three-dimensional coordinates (x, y, z) were directly used. For the raw visible light image data and the raw infrared thermal imaging data, the POS attitude data (longitude, latitude, altitude, roll angle, pitch angle, yaw angle) during UAV cruise, as well as the intrinsic parameters (focal length, principal point coordinates, distortion coefficient) and extrinsic parameters (installation position and angle relative to the UAV body) of the two types of imaging devices, were combined. Through the photogrammetric spatial forward intersection algorithm, the coordinates of each pixel in the image were converted into ground three-dimensional coordinates in the WGS84 coordinate system. Then, the datasets were grouped according to the time synchronization, and the ground three-dimensional coordinates of the LiDAR point cloud and the image in each group were spatially matched. Isolated points in the point cloud that were outside the field of view of the corresponding image, as well as pixel areas in the image that were outside the river surface boundary of the riverbank model, were removed to ensure that the spatial coverage of the three types of data completely corresponded to the monitoring area defined by the riverbank model. Finally, data consistency verification is performed. By calculating the spatial overlap rate (≥95%) and time synchronization accuracy of each data set, datasets with qualified spatiotemporal matching are selected, ultimately obtaining lidar point cloud data, visible light image data, and infrared thermal imaging data with unified spatiotemporal reference and consistent spatial coverage.

[0042] Step S3 involves fusing the visible light image data and the infrared thermal imaging data to obtain fused image data, specifically including the following sub-steps:

[0043] Step S301: Extract the average water surface temperature using the infrared thermal imaging data.

[0044] In this embodiment, the temperature value of each pixel in the infrared thermal imaging data is extracted, and the average water surface temperature is obtained by dividing the sum of the temperature values ​​of each pixel by the total number of pixels.

[0045] Step S302: Calculate the temperature difference using the infrared thermal imaging data and the average water surface temperature.

[0046] In this embodiment, the coordinates of each effective pixel and its corresponding temperature value are determined. Then, using the previously extracted average water surface temperature as a unified benchmark, the difference is calculated pixel by pixel. That is, the temperature difference of each pixel is the pixel temperature value minus the average water surface temperature. Finally, the temperature difference values ​​of all pixels are associated with the corresponding pixel coordinates and stored to form a temperature difference dataset.

[0047] Step S303: Based on the temperature difference, the visible light image data and the infrared thermal imaging data are fused to obtain fused image data.

[0048] In this embodiment, a weighted fusion method is used to calculate the pixel grayscale values ​​of the visible light image data and the infrared thermal imaging data at the coordinates. The weight can be set according to the reliability of the data or according to the data acquisition environment. For example, if the lighting is sufficient and there is no obstruction, the weight of the visible light image data is generally set to 0.6 and the weight of the infrared thermal imaging data is set to 0.4.

[0049] ;

[0050] To fuse image data in coordinates The pixel grayscale value at that location, For visible light image data in coordinates The pixel grayscale value at that location, For infrared thermal imaging data in coordinates The pixel grayscale value at that location, For infrared thermal imaging data in coordinates Temperature value at that location, The average water surface temperature , As weight.

[0051] In the above formula A weighted fusion method is used to balance morphological recognizability and temperature saliency. It directly uses infrared grayscale to avoid the loss of features caused by water surface reflection and transparency interference in visible light, and relies on the sensitive characteristics of infrared heat distribution to capture subtle differences.

[0052] Step S4, using the fused image data and the lidar point cloud data to determine floating objects on the river surface, includes the following sub-steps:

[0053] Step S401: Extract color features, texture features, and temperature difference features from the fused image data.

[0054] In this embodiment, by fusing all pixels within the region represented by the image data, the Hue, Saturation, and Value channels of the HSV color space are extracted. Hue is divided into 16 intervals, and Saturation and Value are each divided into 8 intervals. The pixel percentage of each interval is calculated, resulting in 16+8+8=32-dimensional color features. The Local Binary Pattern (LBP) (8-neighborhood, radius 1) is calculated for all pixels within the region represented by the fusing image data. The LBP LBP is divided into 16 intervals (10 uniform patterns + 6 non-uniform patterns), and the pixel percentage of each pattern is calculated, resulting in 16-dimensional texture features. Infrared temperature values ​​are extracted from all pixels within the region represented by the fusing image data, and the mean and standard deviation of the temperature of the fusing image data are calculated using these infrared temperature values. Finally, the mean and standard deviation are concatenated to obtain 2-dimensional temperature difference features.

[0055] Step S402: Construct a first classification model and predict a first floating object probability set based on the color features, texture features, and temperature difference features using the first classification model.

[0056] The construction of the first classification model specifically includes the following sub-steps:

[0057] Step S40201: Set feature correlation constraints, which satisfy the following formula:

[0058] ;

[0059] in, For feature correlation constraint terms, To constrain the weights, For the sample size, The mutual information calculation function, For the first Color features of each sample For the first Temperature difference characteristics of individual samples For the first Texture features of each sample.

[0060] Minimize the above formula This is equivalent to maximizing the combined value within the parentheses. To maximize the effect, the temperature difference features (such as the high temperature of plastic and the medium temperature of aquatic plants) should be coordinated with the color and texture features to compensate for the blurring of appearance features caused by the reflection of the river surface. To minimize the overall size, we reduce the size of the feature set, avoiding redundant learning of color and texture (e.g., the high correlation between bright white and smooth texture in plastic). This forces them to differentiate their functions (color determines color gradation, texture determines surface morphology), improving feature utilization. The final formula optimizes feature relationships through a reward and penalty mechanism. For floating objects with significant temperature differences (plastic, aquatic plants), it rewards feature distributions with strong temperature-appearance correlation and weak color-texture redundancy. For floating objects with slight temperature differences (oil slicks), it reduces the constraint strength to avoid interference, allowing the first classification model to utilize the three types of features more efficiently and improving classification accuracy in complex river surface scenes (reflection, thin oil slicks).

[0061] Step S40202: Using the cross-entropy loss function as the base loss function, and adding the correlation constraint term to the base loss function to obtain the optimized loss function.

[0062] In this embodiment, the cross-entropy loss function is used. (Based on characterizing the difference in probability distribution between predicted results and true labels to ensure classification accuracy) the feature correlation constraint term (which regulates the association between color, texture, and temperature features through mutual information to guide feature collaboration) is weighted by coefficients. Weighted integration to construct an optimized loss function , This function drives model optimization learning, constrains the formation of temperature-anchored visual and complementary visual feature association patterns, and balances prediction accuracy with the rationality of feature representation.

[0063] Step S40203: Construct a first classification model based on the optimized loss function.

[0064] In this embodiment, based on historically collected fused image data, color, texture, and temperature difference features are extracted and labeled with floating object categories (such as floating objects formed from plastic and wood, aquatic plants, oil stains, and foam clusters). The extracted features are then normalized to construct a sample set with multi-class labels. Multinomial logistic regression is selected as the model framework, with the training objective being to optimize the loss function (cross-entropy loss + feature correlation constraint term, balancing classification accuracy and feature association rationality through weights). Gradient descent algorithm is used to optimize model parameters (learning the mapping relationship from features to category labels). During the training phase, parameters are iteratively updated using the training set, and overfitting is monitored in real-time using the validation set. After training, the classification performance is evaluated using metrics such as accuracy and F1-score on the test set. Simultaneously, the recognition gain of feature correlation constraints for complex scenes with weak temperature differences (such as thin oil stains) and weak visual differences (such as transparent plastic) is verified. Finally, a first classification model that combines classification accuracy and feature synergy is constructed.

[0065] It should be noted that, since the first classification model needs to output the probabilities of four categories (i.e., the first floating object probability set), and needs to satisfy the probability axiom that the sum of the probabilities of all categories is 1, the Softmax activation function is set in the output layer of the multinomial logistic regression framework.

[0066] The process of predicting the first floating object probability set using the first classification model based on the color features, texture features, and temperature difference features specifically includes the following sub-steps:

[0067] Step S40211: Normalize the color features, texture features, and temperature difference features to obtain normalized features.

[0068] In this embodiment, the color features and texture features are normalized by maximum and minimum values ​​to eliminate differences in the value range of different channels and to ensure that the statistics of uniform and non-uniform modes are on the same scale. The temperature difference features are normalized by mean and variance. Finally, the three normalized features are concatenated into a 50-dimensional feature vector in the order of 32-dimensional color → 16-dimensional texture → 2-dimensional temperature, which is the normalized feature.

[0069] Step S40212: Input the normalized features into the first classification model to obtain the first floating object probability set.

[0070] In this embodiment, normalized features are input into a multinomial logistic regression model trained with an optimized loss function (cross-entropy loss fused with feature correlation constraints). The model performs multi-class linear regression on the input vector using a pre-learned feature weight matrix, and then maps the regression output to a probability distribution using a Softmax function. Finally, it outputs the probability of the corresponding floating object subclass (floating objects, aquatic plants, oil slicks, and similar floating objects), forming the first floating object probability set.

[0071] Step S403: Extract the floating object volume features, floating object projected area features, and floating object shape factor features from the lidar point cloud data.

[0072] In this embodiment, the volume features of the floating objects are first extracted. A voxel accumulation algorithm is then used to divide the point cloud space into a 5cm×5cm×5cm voxel grid (the size of the voxel grid can be determined based on the size range of common floating objects on the river surface). All voxels are traversed, and the number of effective voxels containing at least 3 LiDAR points is counted (denoted as voxels). ), through formula The volume of the floating object is calculated, and the interference of isolated noise points on the volume calculation is eliminated. Then, the projected area feature of the floating object is extracted. The three-dimensional coordinates of the lidar point cloud are projected onto the XY plane where the river surface reference plane is located to obtain a two-dimensional projection point set. The Graham scanning method is used to generate the convex hull of this point set, and the planar area of ​​the convex hull is calculated, which is the projected area of ​​the floating object. Finally, the shape factor feature of the floating object is extracted. First, the area of ​​the minimum bounding rectangle of the projection point set on the XY plane is calculated by the rotating caliper method. Then, the ratio of the projected area to the area of ​​the minimum bounding rectangle is used as the shape factor. The closer the ratio is to 1, the more regular the shape of the floating object (such as a plastic box). The closer the ratio is to 0, the more irregular the shape (such as a clump of aquatic plants).

[0073] Step S404: Construct a second classification model. Based on the color features, texture features, floating object volume features, floating object projected area features, and floating object shape factor features, use the second classification model to predict the second floating object probability set.

[0074] In this embodiment, the construction of the second classification model is consistent with that of the first classification model, except for the training data. The training data consists of historical samples labeled with four categories: floating objects, aquatic plants, oil slicks, and floating similar objects. It includes color and texture features extracted from historical fused images, as well as volume, projected area, and shape factor features of floating objects extracted from historical LiDAR point clouds. Cross-entropy loss is used as the loss function, and feature correlation constraints can also be set. The second classification model is constructed based on multinomial logistic regression. Both the training data and the input data during prediction undergo the same normalization process to eliminate dimensional differences and improve the accuracy of prediction, ultimately yielding the second floating object probability set.

[0075] Step S405: Determine the floating objects on the river surface using the first floating object probability set and the second floating object probability set.

[0076] The process of determining floating objects on the river surface using the first floating object probability set and the second floating object probability set specifically includes the following sub-steps:

[0077] Step S40501: Extract mutually exclusive category pairs from the first floating object probability set and the second floating object probability set.

[0078] In this embodiment, based on the mutually exclusive rule that if the same monitoring sample belongs to one category, it must not belong to another category, all categories in the two probability sets are traversed, and pairs of the same category are excluded (such as floating object-floating object, aquatic plant-aquatic plant, etc., which do not satisfy the mutually exclusive property). All combinations that are not of the same category are extracted as mutually exclusive category pairs, such as floating object-aquatic plant, floating object-oil, floating object-floating similar object, aquatic plant-oil, aquatic plant-floating similar object, oil-floating similar object, etc., and finally a set of mutually exclusive category pairs is obtained.

[0079] Step S40502: Calculate the first conflict coefficient using the mutually exclusive category pair.

[0080] The first conflict coefficient satisfies the following formula:

[0081] ;

[0082] in, The first conflict coefficient, The number of mutually exclusive class pairs, For the first The probability of the first floating object in a pair of mutually exclusive class pairs , No. The probability of the second floating object in a pair of mutually exclusive class pairs .

[0083] The above formula measures the conflict strength by quantifying the joint confidence of mutually exclusive predictions across models. For each pair of mutually exclusive categories, it extracts the prediction probability of the first classification model for category A. The second classification model predicts the probability of mutually exclusive class B. The product of the two values ​​characterizes the conflict confidence level between the first classification model's belief that the value is A and the second classification model's belief that the value is B (A≠B). The first conflict coefficient is obtained by summing the conflict confidence levels of all mutually exclusive pairs. The larger the value, the more significant the discrepancy between the two models in determining the category of floating objects. This indicator provides a quantitative basis for subsequent multi-model fusion decisions and reveals the degree of prediction discrepancy caused by differences in feature dimensions (visual + temperature vs. visual + three-dimensional morphology).

[0084] Step S40503: Set a conflict determination threshold and compare the conflict coefficient with the conflict determination threshold.

[0085] Setting the conflict determination threshold specifically includes the following sub-steps:

[0086] Step S4050301: Based on the previously obtained historical sample data, the first historical probability set and the second historical probability set are predicted using the first classification model and the second classification model, respectively.

[0087] In this embodiment, for pre-collected and labeled historical sample data (categorized as floating objects, aquatic plants, oil slicks, and similar floating objects), color, texture, and temperature difference features are first extracted from the historical fused image data to construct a normalized feature set adapted to the input of the first classification model. For historical LiDAR point cloud data, the volume, projected area, and shape factor of the floating objects are calculated and concatenated with the color and texture features to construct a normalized feature set adapted to the input of the second classification model. Subsequently, the first classification model, which has been trained using an optimized loss function, is invoked, and the first model's input feature set is input sample by sample. The predicted probability for each historical sample corresponding to the four categories of floating objects, aquatic plants, oil slicks, and similar floating objects is output, forming the first historical probability set. Simultaneously, the second classification model, trained within the same framework, is invoked, and the second model's input feature set is input sample by sample. The predicted probability for each category is output, forming the second historical probability set. Both historical probability sets must be associated one-to-one with the true category labeling of the historical samples.

[0088] Step S4050302: Obtain multiple conflict determination candidate thresholds, and calculate the historical floating object confidence scores of the first historical probability set and the second historical probability set based on the conflict determination candidate thresholds.

[0089] In this embodiment, within the theoretical range of the conflict coefficient [0,1], a candidate threshold set is generated by equidistant traversal with a predetermined step size (e.g., an interval of 0.05, which can be adjusted according to accuracy requirements). When calculating the historical floating object confidence based on these candidate thresholds, the conflict coefficient is calculated for each sample in the first historical probability set and the second historical probability set, and conflict determination is performed. If the conflict coefficient is less than the current conflict determination candidate threshold, the weight coefficient is set to 1, that is, the two probability sets remain unchanged after adjustment, and the conflict coefficient multiplied by 1 remains unchanged. The fused comprehensive probability is further calculated as the historical floating object confidence of the sample. If the conflict coefficient is greater than the current conflict determination candidate threshold, the weight of the first historical probability set is set to 0.4, and the weight of the second historical probability set is set to 0.6 to increase the credibility of the visible light + radar point cloud prediction results. The two weights are multiplied by the corresponding elements in the historical probability set, and the conflict coefficient is adjusted (multiplied by the two weights). The fused historical floating object confidence is then calculated in the same way.

[0090] Step S4050303: Obtain the fusion accuracy based on the historical floating object confidence level.

[0091] In this embodiment, based on the historical floating object confidence scores, the category corresponding to the highest probability for each sample is selected as the predicted category. Simultaneously, the actual labels (manually labeled actual categories) of the historical samples are retrieved. The predicted category is then compared with the actual label. If they match (e.g., predicted oil slick and the actual label is oil slick), the prediction is considered correct; if they do not match (e.g., predicted floating object but the actual label is aquatic plants), the prediction is considered incorrect. After traversing all historical samples, the fusion accuracy under the current conflict determination candidate threshold is obtained by dividing the number of correctly predicted samples by the total number of historical samples.

[0092] Step S4050304: Select a conflict determination threshold from the conflict determination candidate thresholds based on the fusion accuracy.

[0093] In this embodiment, each conflict determination candidate threshold corresponds to a fusion accuracy, and the conflict determination candidate threshold corresponding to the highest fusion accuracy is selected as the conflict determination threshold.

[0094] Step S40504: Based on the comparison results, set influence weights for the first floating object probability set and the second floating object probability set respectively.

[0095] In this embodiment, the comparison results are divided into two categories: the first conflict coefficient is greater than the conflict determination threshold and the first conflict coefficient is less than the conflict determination threshold.

[0096] When the first conflict coefficient is greater than the conflict determination threshold, it is considered a serious conflict. The importance of the first floating object probability set is reduced, while the importance of the second floating object probability set is increased. Therefore, the influence weight is used to measure it. The influence weight of the first floating object probability set can be set to 0.4, and the influence weight of the second floating object probability set can be set to 0.6 (it can be adjusted according to the training effect of the model, but the sum of the two influence weights must be 1).

[0097] When the first conflict coefficient is less than the conflict determination threshold, it is considered a mild conflict. Trusting the synergy between the two models, the influence weights are all set to 1, which essentially means no adjustment is made.

[0098] The significance of this move is that, in the event of a severe conflict, relying on the physical objectivity of three-dimensional morphological features (such as the inability of floating objects to be disguised by color / temperature), visual + three-dimensional judgment is given priority to avoid misjudgment based on a single visual-temperature feature (such as oil stains reflecting light and misleading color classification). In the event of a mild conflict, the synergy between the two models is trusted to maintain the original weights and improve the robustness of floating object category fusion judgment by using the reliability differences of feature dimensions.

[0099] Step S40505: Adjust the first floating object probability set and the second floating object probability set based on the influence weight to obtain the first floating object probability adjustment set and the second floating object probability adjustment set.

[0100] In this embodiment, the predicted probability of each category in the first floating object probability set and the second floating object probability set is multiplied by the corresponding influence weight to obtain the first floating object probability adjustment set and the second floating object probability adjustment set.

[0101] Step S40506: Calculate the second conflict coefficient using the influence weight and the first conflict coefficient.

[0102] The second conflict coefficient satisfies the following formula:

[0103] ;

[0104] in, The second conflict coefficient, and These are the influence weights, This is the first conflict coefficient.

[0105] The above formula is the product of the model's influence weights. The original conflict coefficient K is co-scaled (derived by substituting the adjusted mutual exclusion category pair into the formula of the first conflict coefficient) to generate the second conflict coefficient. Its core logic is to incorporate the predictive credibility of the first and second classification models into the determination of the degree of conflict. If a certain model has a low weight, even if it disagrees with another model, after scaling by weight product, the apparent intensity of the conflict will decrease with the weight ratio of the low credibility model, which is more in line with the impact of model reliability differences on conflict decision-making, and makes the conflict coefficient more accurate in serving subsequent fusion judgment.

[0106] Step S40507: Based on the second conflict coefficient, the first floating object probability adjustment set and the second floating object probability adjustment set are fused to obtain the floating object confidence level.

[0107] The confidence level for floating objects satisfies the following formula:

[0108] ;

[0109] in, For category Confidence level of floating objects Adjust the set category for the probability of the first floating object. Confidence level, Adjust the set category for the probability of the second floating object. Confidence level, This is the second conflict coefficient.

[0110] In the above formula It is a two-model approach to categories The joint probability expresses that both models simultaneously support the category. The strength of consensus, This means that after deducting the conflict loss between the two models in mutually exclusive categories, normalization ensures that the confidence score focuses on effective consensus. Therefore, the confidence score for floating objects... Quantified multimodal features (visual-temperature + visual-3D morphology) to collaboratively support categories Furthermore, it overcomes the confidence strength of cross-category conflict interference.

[0111] Step S40508: Determine the floating objects on the river surface based on the confidence level of the floating objects.

[0112] In this embodiment, for the calculated confidence scores of floating objects (including probability values ​​for four categories: "floating objects, aquatic plants, oil pollution, and floating similar objects"), the category with the highest confidence score is selected for each sample. If the category is floating objects, it is determined that there are floating objects on the river surface in the area corresponding to the current lidar and image fusion data. If the maximum value corresponds to aquatic plants, oil pollution, or floating similar objects, it is classified into the corresponding category. Only when the highest confidence score category is clearly floating objects is the monitoring target confirmed as floating objects on the river surface. This process transforms the confidence scores of multi-model fusion into category determination results through probability ranking decision logic. It covers multi-category recognition and focuses on the accurate screening of floating objects, ensuring that the category with the highest confidence score (highest probability) is used as the final determination basis, thereby improving the reliability of the detection results.

[0113] like Figure 2 As shown, in another aspect, the present invention also provides a river surface floating object monitoring system based on unmanned aerial vehicles (UAVs), including: a processor, an input device, an output device, and a memory. The processor, the input device, the output device, and the memory are interconnected. The memory is used to store a computer program, the computer program including program instructions, and the processor is configured to call the program instructions to execute the relevant steps of the relevant embodiments of the UAV-based river surface floating object monitoring method of the present invention.

[0114] This invention provides a drone-based system for monitoring floating objects on a river surface. The functional components can be integrated into a single processing unit, or each component can exist independently, or two or more components can be integrated into one unit. The integrated components can be implemented in hardware or software.

[0115] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention, and they should all be covered within the scope of the claims and specification of the present invention.

Claims

1. A method for monitoring floating objects on a river surface based on unmanned aerial vehicles (UAVs), characterized in that, The method includes: Acquire riverbank boundary data for the target area and construct a riverbank model using the riverbank boundary data; Based on the riverbank model, UAV cruise was used to obtain lidar point cloud data, visible light image data and infrared thermal imaging data; The visible light image data and the infrared thermal imaging data are fused to obtain fused image data; The fused image data and the lidar point cloud data are used to determine floating objects on the river surface; The process of determining floating objects on the river surface using the fused image data and the lidar point cloud data includes: Color features, texture features, and temperature difference features are extracted from the fused image data; A first classification model is constructed, and a first floating object probability set is predicted based on the color features, texture features, and temperature difference features using the first classification model; The volume features, projected area features, and shape factor features of floating objects are extracted from the lidar point cloud data. A second classification model is constructed, and a second probability set of floating objects is predicted based on the color features, texture features, volume features, projected area features, and shape factor features of the floating objects. Floating objects on the river surface are determined using the first floating object probability set and the second floating object probability set; The step of determining floating objects on the river surface using the first floating object probability set and the second floating object probability set includes: Extract mutually exclusive category pairs from the first floating object probability set and the second floating object probability set; The first conflict coefficient is calculated using the mutually exclusive category pair; Set a conflict determination threshold, and compare the conflict coefficient with the conflict determination threshold; Based on the comparison results, influence weights are set for the first floating object probability set and the second floating object probability set respectively; Based on the influence weights, the first floating object probability set and the second floating object probability set are adjusted to obtain the first floating object probability adjustment set and the second floating object probability adjustment set; The second conflict coefficient is calculated using the influence weight and the first conflict coefficient; The first floating object probability adjustment set and the second floating object probability adjustment set are fused based on the second conflict coefficient to obtain the floating object confidence score; The floating objects on the river surface are determined based on the confidence level of the floating objects. The setting of the conflict determination threshold includes: Based on the pre-obtained historical sample data, the first historical probability set and the second historical probability set are predicted using the first classification model and the second classification model, respectively. Multiple conflict determination candidate thresholds are obtained, and the historical floating object confidence scores of the first historical probability set and the second historical probability set are calculated based on the conflict determination candidate thresholds. The fusion accuracy is obtained based on the confidence level of the historical floating objects. Based on the fusion accuracy, a conflict determination threshold is selected from the conflict determination candidate thresholds.

2. The method for monitoring floating objects on a river surface based on an unmanned aerial vehicle (UAV) according to claim 1, characterized in that, The acquisition of lidar point cloud data, visible light image data, and infrared thermal imaging data based on the riverbank model using UAV cruise includes: Based on the riverbank model, UAV cruise was used to obtain raw data of lidar point cloud, raw data of visible light image, and raw data of infrared thermal imaging. The original lidar point cloud data is obtained by performing elevation filtering on the original lidar point cloud data. Spatiotemporal alignment is performed on the raw visible light image data, the raw infrared thermal imaging data, and the raw filtered lidar point cloud data to obtain lidar point cloud data, visible light image data, and infrared thermal imaging data.

3. The method for monitoring floating objects on a river surface based on an unmanned aerial vehicle (UAV) according to claim 2, characterized in that, The process of performing elevation filtering on the raw lidar point cloud data to obtain the raw filtered lidar point cloud data includes: The river surface reference surface was fitted using the raw point cloud data from the lidar. Extract the vertical elevation component from the original lidar point cloud data; Calculate the vertical relative elevation based on the river surface reference plane and the vertical elevation components; Obtain the vertical elevation range of the floating object, and use the vertical elevation range of the floating object and the vertical relative elevation to perform elevation filtering on the original lidar point cloud data to obtain the original filtered lidar point cloud data.

4. The method for monitoring floating objects on a river surface based on an unmanned aerial vehicle (UAV) according to claim 1, characterized in that, The process of fusing the visible light image data and the infrared thermal imaging data to obtain fused image data includes: The average water surface temperature was extracted using the infrared thermal imaging data. The temperature difference is calculated using the infrared thermal imaging data and the average water surface temperature. Based on the temperature difference, the visible light image data and the infrared thermal imaging data are fused to obtain fused image data.

5. The method for monitoring floating objects on a river surface based on an unmanned aerial vehicle (UAV) according to claim 4, characterized in that, The construction of the first classification model includes: Set feature correlation constraints; The cross-entropy loss function is used as the base loss function, and the correlation constraint term is added to the base loss function to obtain the optimized loss function; The first classification model is constructed based on the optimized loss function.

6. The method for monitoring floating objects on a river surface based on an unmanned aerial vehicle (UAV) according to claim 5, characterized in that, The first floating object probability set predicted using the first classification model based on the color features, texture features, and temperature difference features includes: The color features, texture features, and temperature difference features are normalized to obtain normalized features; The normalized features are input into the first classification model to obtain the first floating object probability set.

7. A river surface floating object monitoring system based on unmanned aerial vehicles (UAVs), characterized in that, include: The system includes a processor, an input device, an output device, and a memory, all interconnected. The memory stores a computer program, which includes program instructions. The processor is configured to invoke the program instructions to execute a method for monitoring floating objects on a river surface based on an unmanned aerial vehicle (UAV) as described in any one of claims 1 to 6.