An air-ground collaborative road garbage precise cleaning method based on cross-view feature consistency verification
By using drones and sweepers for collaborative inspections, and employing cross-view feature consistency verification and IMU motion compensation methods, the problems of visual ambiguity and dynamic interference in road garbage sweeping were solved, achieving high-precision garbage identification and sweeping results.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- GUANGDONG CHENGJIN TECH CO LTD
- Filing Date
- 2026-05-08
- Publication Date
- 2026-07-31
AI Technical Summary
In existing ground-air collaborative road cleaning solutions, single-view recognition methods cannot distinguish visually ambiguous targets, dynamic environmental interference leads to misjudgment, and general cross-view matching methods do not consider the non-planarity of the road surface and the interference of UAV self-movement, resulting in low verification accuracy and high dynamic false rejection rate.
A cross-view feature consistency verification method is adopted. Through collaborative inspection by UAV and sweeper, dual-view images are collected and cross-view feature consistency verification is performed. Combined with local affine transformation and IMU motion compensation, reprojection error and UAV jitter interference are eliminated. A contrastive learning pre-trained feature network is used for accurate garbage identification, and adaptive sweeping and closed-loop feedback are used to ensure the cleaning effect.
It significantly reduced the false recognition rate (by 92%), improved the accuracy of removing dynamic interference (up to 96%), ensured the completion and accuracy of cleaning tasks, and adapted to different seasons and road conditions.
Smart Images

Figure CN122493330A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent sanitation and unmanned system collaborative control technology, specifically to a precise garbage identification method that utilizes drones and unmanned sweepers to perform cross-view geometric consistency verification of the same target from different perspectives, and combines dynamic compensation and adaptive sweeping. Background Technology
[0002] Current ground-air cooperative road cleaning solutions generally adopt an open-loop architecture of "UAV identification → sweeper execution," resulting in a high rate of false cleaning. Some improved solutions introduce secondary identification by the sweeper, but this involves repeated detection from the same perspective (ground level), failing to address the following fundamental problems: 1. Inherent ambiguity of a single perspective: From an aerial viewpoint, fallen leaves and their shadows, scraps of paper and road markings, plastic bags and puddles reflect light are highly similar in two-dimensional images, making it impossible for any single-frame image classifier to fundamentally distinguish them. 2. Dynamic environmental interference: Dynamic targets such as wind-blown leaves, falling willow catkins, and moving pedestrian shadows are misjudged as litter during UAV inspections, and by the time the sweeper arrives, the targets have disappeared or changed position, leading to ineffective scheduling. 3. Lack of cross-view geometric constraints: UAVs and sweepers perceive independently, making it impossible to verify the three-dimensional structural stability of targets using parallax information.
[0003] Existing cross-view matching methods (such as multi-view vehicle re-identification) typically assume that targets are located on the same plane or that the scene structure is known. However, in road cleaning scenarios, the road surface may have non-ideal planar factors such as slopes, potholes, and speed bumps, leading to reprojection errors when directly using homography transformation. Furthermore, the motion of the drone during hovering can cause optical flow methods to misidentify dynamic targets. This invention proposes a systematic solution to the aforementioned unresolved technical problems. Summary of the Invention
[0004] (a) Technical problems to be solved The technical problem to be solved by this invention is that in the joint ground and air cleaning of road debris, the existing single-view recognition method cannot distinguish visually ambiguous targets, while the general cross-view matching method does not take into account the non-planarity of the road surface and the interference of the self-movement of UAVs, resulting in low verification accuracy and high dynamic false rejection rate.
[0005] (II) Technical Solution To address the aforementioned technical problems, this invention provides a ground-air collaborative method for precise road debris cleaning based on cross-view feature consistency verification, applicable to cloud computing, drones, and unmanned sweeping vehicles, comprising the following steps: Step S1: Collaborative Inspection and Dual-View Image Acquisition The cloud sends the inspection path to the drone, which flies along the path. At the same time, it sends a "cooperative positioning" command to the idle unmanned sweepers in the area, and the sweepers drive to the designated cooperative observation point to stand by.
[0006] When the drone detects a suspected garbage target during flight, it adjusts the gimbal angle to acquire a top-down view image I_uav and records the drone's pose (3D coordinates P_uav and Euler angles), while simultaneously activating hover mode.
[0007] The drone sends the target coordinates (acquired via RTK-GPS, with an accuracy of ≤5cm) to the nearest unmanned sweeper. The unmanned sweeper then travels to a distance of 2-5 meters from the target, collects a ground-view image I_car, and records its own pose (P_car and Euler angles).
[0008] Step S2: Cross-view feature consistency verification (including non-planar road surface compensation) Based on the relative poses of the drone and the sweeper, the cloud-based system calculates the homography matrix H between the two views. Simultaneously, it calculates the disparity gradient of the target region in I_uav. If the disparity gradient exceeds a threshold (indicating significant non-planar undulations in the road surface), a local affine transformation is used: the target region is divided into multiple sub-blocks (e.g., a 4×4 grid), each sub-block independently calculates its affine transformation matrix, and then these sub-blocks are weighted and fused to obtain the transformed image I_warp. This step eliminates reprojection errors caused by uneven road surfaces.
[0009] Deep features of the target region in I_warp and I_car are extracted in the cloud, and a feature extraction network pre-trained by contrastive learning is used (see step S5 for specific training method) to obtain feature vectors f_warp and f_car.
[0010] The feature consistency score is calculated as C = α·cos(f_warp,f_car) + β·SSIM(I_warp,I_car), where α = 0.7 and β = 0.3 (obtained through statistical optimization of a large number of road waste samples, and α has a fixed value, which does not constitute a purely mathematical constraint).
[0011] If C ≥ threshold θ (θ is dynamically determined by offline ROC curve, with a typical value of 0.72~0.85), it is initially determined to be real garbage; otherwise, it is determined to be a false target and the process is terminated.
[0012] Step S3: Dynamic target removal based on motion compensation During the hovering period, the drone continuously acquires N frames of images (N≥10, frame interval 0.2 seconds), while simultaneously recording the drone's own IMU angular velocity and acceleration data.
[0013] For each frame, motion-compensated optical flow is calculated: first, the rotation matrix R_self and translation vector t_self of the UAV's self-motion are obtained by IMU integration, and the image is inversely compensated to obtain a corrected image sequence with "static background"; then, a lightweight optical flow network (such as RAFT-Lite) is run to calculate the optical flow field of the target region in the compensated image.
[0014] If the average amplitude of the compensated optical flow vector is greater than the dynamic threshold V_th (V_th = 3 pixels / frame, which is much smaller than the 5 pixels / frame before compensation, thus avoiding misjudgment caused by drone shaking), and the variance of the optical flow direction angle is less than 30°, then it is determined to be a dynamic interference object (such as wind-blown leaves or floating objects), and the process is terminated directly.
[0015] If the average amplitude is less than or equal to V_th, it is determined to be a static target, and proceed to step S4.
[0016] Step S4: Adaptive Cleaning and Closed-Loop Feedback For verified real-world waste, the cloud analyzes the outline area and shape characteristics of the waste region in I_car and selects one of the following cleaning modes: Scattered granular particles (area <0.01m²): Low-speed strong suction mode; Flake-shaped waste (area 0.01~0.1m² and perimeter-to-area ratio >0.8): roller brush winding mode; Large areas of debris (>0.1m²): Side brushes gather debris and multiple reciprocating cleaning modes.
[0017] After the sweeper performs its cleaning, it takes a picture of the cleaned area with the onboard camera and compares it with the picture before cleaning. If the residual rate is greater than 5%, it will automatically perform a second cleaning (up to 2 times).
[0018] The results of successful / failed cleaning, along with the corresponding images, are uploaded back to the cloud as positive / negative samples for model updates.
[0019] Step S5: Continuous training of the cross-view feature network The feature extraction network uses ResNet50 as its backbone and is pre-trained within a contrastive learning framework. The training dataset contains 100,000 pairs of real trash (positive samples) and 100,000 pairs of fake targets (negative samples). Positive samples are image pairs of the same trash from different viewpoints, while negative samples are image pairs of different trash or trash and distractors from different viewpoints. The loss function used is InfoNCE loss with a temperature coefficient of 0.07. After training, the feature vectors, based on cosine distance, can bring positive sample pairs closer together and push negative sample pairs further apart.
[0020] After each actual cleaning verification, the cloud stores the (I_uav, I_car) image pairs and manually confirmed labels (real trash or fake targets) into the database. Incremental training is performed weekly: 2000 pairs of samples (positive to negative ratio 1:1) are randomly selected from the database, and the fully connected layers and the last two residual blocks of the network are fine-tuned with a learning rate of 0.0001 for 5 epochs. The updated model is then deployed to the cloud inference node.
[0021] This training process ensures that the network continuously adapts to different seasons, road surfaces, and lighting conditions.
[0022] (III) The beneficial effects of the present invention are as follows: First, by introducing local affine transformation to handle non-planar road surfaces, the failure of traditional homography transformation in real sanitation scenarios is overcome, reducing the false recognition rate by more than 92% compared to a single-viewpoint approach (experimental data: in 1000 test samples, the false recognition rate of the traditional method was 18.7%, while that of this method was 1.5%). Second, optical flow analysis with IMU motion compensation eliminates false triggers caused by the drone's own hovering jitter, achieving a 96% accuracy rate in removing dynamic interference. Third, the comparative learning pre-training + incremental update framework makes the system increasingly accurate with each use. Fourth, the post-cleaning differential comparison and secondary cleaning mechanism ensures task completion and avoids the problem of "verification passed but cleaning failed". Attached Figure Description
[0023] Figure 1 This is a flowchart illustrating the intelligent road cleaning method that integrates ground and air transport, as provided in an embodiment of the present invention.
[0024] Figure 2 A spatiotemporal diagram illustrating the collaborative operation between drones and cleaning vehicles. Detailed Implementation
[0025] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to specific embodiments and the accompanying drawings. It should be understood that these descriptions are merely exemplary and not intended to limit the scope of the invention. Furthermore, descriptions of well-known structures and techniques are omitted in the following description to avoid unnecessarily obscuring the concept of the invention.
[0026] Example like Figure 1-2 As shown, this embodiment provides an intelligent road cleaning method that integrates ground and air transportation, and is applied to cloud 100, drone 200 and unmanned sweeper 300.
[0027] Equipment parameters: Drone 200: A six-rotor industrial drone equipped with RTK-GPS (horizontal accuracy 2cm), high-precision IMU (sampling rate 200Hz), three-axis gimbal (pitch -90°~+30°), 20-megapixel global shutter camera, onboard computing unit (NVIDIA Orin NX 16GB), and 5G communication module.
[0028] Sweeper 300: A wire-controlled chassis sweeper equipped with RTK-GPS, high-precision IMU, forward-facing gimbal camera (12 megapixels, tilt-adjustable), LiDAR (16 lines), onboard computing unit (NVIDIA Orin), and 5G communication module.
[0029] The Cloud 100 uses a GPU server cluster (4×NVIDIA A100) for model training and high-load inference, and communicates with the Drone 200 and the Sweeper 300.
[0030] The method includes the following steps: Step S1: Collaborative Inspection and Dual-View Data Acquisition A patrol route for a 2.5km long street is planned via cloud computing. The drone flies at an altitude of 12m and a speed of 4m / s. Simultaneously, the cloud sends coordination commands to two idle street sweepers in the area, each moving to a pre-set collaborative observation point (every 200m apart). After detecting a suspected target (confidence > 0.6), the drone hovers directly above the target (horizontal deviation < 0.5m), with its three-axis gimbal tilted at -90°, acquiring a top-view image I_uav (1920×1080). The drone then transmits the target's GPS coordinates via 5G to the nearest street sweeper (distance < 150m). The street sweeper automatically moves to a distance of 3m from the target, facing it, with its upper and lower gimbal tilted at -15°, acquiring a ground image I_car (1920×1080). Both images are timestamped (synchronization accuracy < 10ms) and have their pose information added before being sent to the cloud.
[0031] Step S2: Cross-view verification with non-planar compensation After receiving the image pairs and pose data from the cloud, perform the following operations: S21. Based on the UAV pose (R_uav, t_uav) and the sweeper pose (R_car, t_car), calculate the relative rotation R_rel = R_car^T·R_uav and the relative translation t_rel = R_car^T·(t_uav - t_car). Calculate the homography matrix H using R_rel, t_rel, and the camera intrinsic parameters.
[0032] S22. Uniformly select 16 feature points within the target bounding box I_uav and calculate their disparity gradient in I_uav. If the root mean square of the disparity gradient is >0.02 (in normalized coordinates), the road surface is determined to have significant non-planar undulations. Then, divide the target region into 4×4 sub-blocks. For each sub-block, fit an affine transformation matrix using the feature points within that sub-block. Then, perform bilinear interpolation to fuse the transformation results of the sub-blocks to obtain I_warp. If the disparity gradient is ≤0.02, directly use the homography matrix H to perform perspective transformation to obtain I_warp.
[0033] S23. Input the corresponding regions in I_warp and I_car (scaled to 224×224) into the feature extraction network. This network uses ResNet50, and is finally connected to a fully connected layer to output a 128-dimensional normalized feature vector. The network pre-training method is described in step S5.
[0034] S24. Calculate the cosine similarity S_cos=f_warp·f_car, calculate SSIM (window 11×11, Gaussian kernel σ=1.5), and the overall score C=0.7·S_cos+0.3·SSIM.
[0035] S25. In the offline experiment, the threshold θ=0.78 is determined by maximizing the Youden exponent on 500 pairs of positive and negative samples. If C≥0.78, it is judged as real garbage; otherwise, it is a false target, and the experiment is terminated.
[0036] Step S3: Dynamic target elimination with motion compensation The drone continuously acquires 15 frames of images during hovering (frame interval 0.2 seconds) and records IMU data (angular velocity, acceleration).
[0037] S31. For the i-th frame (i=2..15), use IMU integration to calculate the rotation matrix R_self_i and translation vector t_self_i relative to the 1st frame.
[0038] S32. Apply the inverse compensation transformation to the i-th frame image: I_comp_i=warp(I_i,R_self_i^T,-t_self_i), to make the background approximately static.
[0039] S33. Run the RAFT-Lite optical flow network (input resolution 640×480) on the compensated image sequence to extract the optical flow vector within the target region.
[0040] S34. Calculate the average value v_mean of the optical flow amplitude of all pixels in the target area, and the variance θ_var of the optical flow direction angle.
[0041] S35. If v_mean > 3 pixels / frame and θ_var < 30°, it is determined to be a dynamic interference object, the process is terminated and recorded; otherwise, it is determined to be a static target and the cleaning stage is entered.
[0042] Experiments show that static garbage has a v_mean of <1.5 pixels / frame after compensation, while windblown leaves have a v_mean of ≈6-10 pixels / frame, effectively compensating for the drone's own jitter.
[0043] Step S4: Adaptive Cleaning and Closed-Loop Feedback For real garbage, OpenCV is used to extract the garbage region contour in I_car and the convex hull area A is calculated.
[0044] If A < 0.01m² → Cleaning mode = strong suction (fan speed 100%, travel speed 0.3m / s).
[0045] If 0.01≤A<0.1m² and the perimeter / area ratio>0.8 → cleaning mode = roller brush winding (roller brush speed 120rpm).
[0046] If A≥0.1m² → Cleaning mode = side brush retraction + multiple reciprocating motions (first retraction 3 times, then low-speed strong suction 2 times).
[0047] After cleaning, a new image is captured and compared with the image before cleaning (absolute difference after grayscale conversion, threshold 20). The residual rate is calculated as: residual pixels / total number of pixels in the garbage area. If the residual rate is >5%, a second cleaning is automatically performed (maximum 2 times). If the residual rate is still >5% after the second cleaning, it is marked as "requiring manual intervention" and reported to the cloud.
[0048] Step S5: Feature Network Training Pre-training phase: Collect cross-view image pairs of road debris scenarios: under different seasons, weather, and road materials (asphalt, cement, brick), manually label 100,000 pairs of positive samples (top view + ground view of the same debris) and 100,000 pairs of negative samples (different types of debris or debris and interfering objects).
[0049] Network structure: ResNet50, outputting 128-dimensional features, followed by L2 normalization.
[0050] Contrastive learning loss function: InfoNCE, batch size 512, temperature τ=0.07, optimizer AdamW, learning rate 0.001, training for 50 epochs.
[0051] Evaluation: On the validation set, the mean cosine similarity of positive sample pairs was 0.89, and the mean cosine similarity of negative sample pairs was 0.31, indicating good discrimination.
[0052] Incremental training phase: Every week, 2000 pairs of samples (positive and negative 1:1) are randomly selected from the cloud database. Positive samples are image pairs with a residual rate of ≤5% after cleaning and which have passed manual review. Negative samples are image pairs with false targets or a residual rate of >20% after cleaning and which have been manually identified as misjudged.
[0053] Fine-tuning parameters: Only update the fully connected layers and the last two residual blocks of ResNet50, with a learning rate of 0.0001, a batch size of 32, and training for 5 epochs.
[0054] The updated model is deployed to cloud inference nodes, and can also be optionally deployed to onboard computing units of drones and cleaning vehicles.
[0055] In other embodiments, in scenarios with high real-time requirements and good road conditions, the non-planar determination and local affine transformation in step S2 can be omitted, and homography transformation can be used directly. Dynamic optical flow calculation can use Farneback dense optical flow (implemented in OpenCV) to replace the deep learning network, but the motion compensation step is still retained.
[0056] The above descriptions are merely some embodiments of the present invention. For those skilled in the art, various modifications and improvements can be made without departing from the inventive concept of the present invention, and these all fall within the protection scope of the present invention.
Claims
1. A ground-air cooperative method for precise road waste cleaning based on cross-view feature consistency verification, applicable to cloud computing, drones, and unmanned sweeping vehicles, characterized in that... Includes the following steps: Step S1: After the UAV detects a suspected garbage target in the air, it acquires a top-down view image I_uav of the target and records the UAV's pose, while the UAV hovers. The cloud-based system dispatches unmanned sweepers to the vicinity of the target, collects ground-view images (I_car) of the target, and records the position and pose of the unmanned sweeper. Step S2: The cloud calculates the homography matrix H between the two views based on the relative poses of the drone and the unmanned sweeper, and determines whether the disparity gradient of the target area in I_uav exceeds the preset threshold. If the target area exceeds the limit, the target area is divided into multiple sub-blocks. The affine transformation matrix is calculated for each sub-block and then weighted and fused to obtain the transformed image I_warp. If the target area does not exceed the limit, perspective transformation is directly performed using H to obtain I_warp. Step S3: Extract the depth features of the target region in I_warp and I_car from the cloud, calculate the feature consistency score C. If C ≥ preset threshold, proceed to step S4; otherwise, determine it as a false target and terminate. Step S4: During the hovering period of the drone, multiple frames of images are continuously acquired, and the drone's own IMU data is recorded at the same time. Motion compensation is performed on each frame of image using the IMU data to obtain a sequence of compensated images with a static background. The optical flow field of the target area in the compensated image sequence is calculated. If the average optical flow amplitude exceeds the dynamic threshold and the consistency of the optical flow direction meets the conditions, it is determined to be a dynamic interference object and the process is terminated; otherwise, it is determined to be real garbage. Step S5: The cloud sends a cleaning command to the unmanned sweeper, and the unmanned sweeper performs the cleaning.
2. A method for precise road surface waste cleaning based on cross-view feature consistency verification according to claim 1 or 2, characterized in that, In step S2, the disparity gradient threshold is 0.02 in normalized coordinates, the number of sub-blocks is 4×4, and bilinear interpolation is used for weighted fusion.
3. The method for precise cleaning of road debris based on cross-view feature consistency verification according to claim 1, characterized in that, The deep features are extracted by a convolutional neural network pre-trained by contrastive learning. This network uses ResNet50 as its backbone and outputs a 128-dimensional normalized feature vector. The feature consistency score is C = α·cos(f_warp,f_car) + β·SSIM(I_warp,I_car), where α = 0.7 and β = 0.
3.
4. The method for precise cleaning of road debris based on cross-view feature consistency verification according to claim 1, characterized in that, In step S4, the dynamic threshold is 3 pixels / frame, the optical flow direction consistency condition is that the variance of the optical flow vector angle is less than 30°, and the specific method of motion compensation is: using IMU integration to obtain the rotation matrix R_self and translation vector t_self of the UAV's self-motion, and applying the inverse compensation transformation I_comp=warp(I,R_self^T,-t_self) to each frame of image.
5. The method for precise cleaning of road debris based on cross-view feature consistency verification according to claim 1, characterized in that, Before the unmanned sweeper performs its cleaning, the cloud further selects the cleaning mode based on the outline area and shape characteristics of the target area: for an area <0.01m², it is a low-speed strong suction mode; for an area between 0.01 and 0.1m² and a perimeter-to-area ratio >0.8, it is a roller brush winding mode; and for an area ≥0.1m², it is a side brush gathering and multiple reciprocating cleaning mode.
6. The method for precise cleaning of road debris based on cross-view feature consistency verification according to claim 1, characterized in that, Both the drone and the unmanned sweeper are equipped with RTK-GPS, with a horizontal positioning accuracy of ≤5cm, and are equipped with a high-precision IMU; when calculating the homography matrix H, GPS and IMU attitude data are used simultaneously.
7. The method for precise cleaning of road debris based on cross-view feature consistency verification according to claim 1, characterized in that, In step S3, the feature extraction network undergoes incremental training weekly: 2000 pairs of samples (positive to negative ratio 1:1) are randomly selected from the cloud database, and the fully connected layer and the last two residual blocks of the network are fine-tuned with a learning rate of 0.0001 for 5 epochs; positive samples are image pairs with a residual rate of ≤5% after cleaning and verification, and negative samples are image pairs with false targets confirmed by manual verification or a residual rate of >20%.
8. The method for precise cleaning of road debris based on cross-view feature consistency verification according to claim 1, characterized in that, After the unmanned sweeper performs a cleaning operation, it takes a picture of the cleaned area with its onboard camera and compares it with the image before cleaning. If the residual rate exceeds 5%, it will automatically perform a second cleaning operation, which can be repeated up to 2 times.