An unmanned aerial vehicle intelligent detection and capture system based on cloud collaboration

CN121616994BActive Publication Date: 2026-07-03JIANGXI XINGHENG CHANGTIAN TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
JIANGXI XINGHENG CHANGTIAN TECHNOLOGY CO LTD
Filing Date
2025-12-01
Publication Date
2026-07-03

Smart Images

  • Figure CN121616994B_ABST
    Figure CN121616994B_ABST
Patent Text Reader

Abstract

This invention discloses a cloud-based collaborative intelligent UAV detection and capture system, comprising: a capture UAV configuration and data acquisition module, which acquires environmental images and attitude navigation data of the target area and uploads them to a cloud-based collaborative platform; an improved RepPoints point set detection module, which outputs candidate target 2D point sets; a 3D point set reconstruction module, which generates an initial 3D point set based on the 2D point set and attitude navigation data; a topology constraint correction module, which obtains a topologically consistent 3D point set; a capture prototype sparse coding module, which generates candidate capture 3D point sets; a capture feasible region screening module, which screens capture feasible 3D point sets based on dynamic constraints; and a spatial risk and capture trajectory planning module, which constructs a spatial risk field and outputs capture trajectory control commands. This invention achieves integrated multi-UAV collaborative point set reconstruction and trajectory planning, improving target capture stability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent sensing and spatial information processing technology for unmanned aerial vehicles (UAVs), and in particular to an intelligent detection and capture system for UAVs based on cloud-based collaboration. Background Technology

[0002] In existing technologies, UAVs in target detection and acquisition scenarios often employ a single or a small number of UAVs equipped with cameras for patrol, performing two-dimensional image detection and tracking of targets. Based on traditional target detection networks or regression bounding box methods, they output the target's position and category, and then combine this with simple navigation data to complete tracking and approach. These solutions mainly focus on target recognition and path following, and their characterization of target geometry, attitude distribution, and spatial structure is relatively coarse. They usually abstract targets as point targets or simplified bounding boxes, making it difficult to support fine-grained acquisition of complex target objects. At the same time, although some systems introduce multi-UAV collaboration, it is mostly limited to task allocation or view coverage. The image and attitude information of each UAV is often processed in a decentralized manner, lacking a fusion modeling mechanism under a unified coordinate system, and the target spatial position estimation still has errors.

[0003] In terms of 3D information acquisition, existing technologies have attempted to obtain 3D point cloud or contour information of targets through stereo vision, structured light, or multi-view reconstruction methods. However, these methods are mostly based on fixed camera arrays or close-range shooting scenarios, making them difficult to directly apply to airborne shooting environments in large target areas. For multi-view environmental images and attitude navigation data collected by multiple capture drones in the air, existing solutions often fail to fully utilize the spatial ray construction and geometric intersection mechanism under a unified coordinate system. They do not adequately consider the correspondence, topological connectivity, and structural integrity of representative points between different viewpoints, resulting in the reconstructed 3D point sets having problems such as high noise, numerous breaks, and shape distortion. Post-processing of 3D point sets often stops at filtering, smoothing, or simple pose fitting, lacking means to combine preset topological structure parameters for systematic topological correction and structural constraints. It is impossible to stably obtain topologically consistent 3D point sets that are consistent with the target object in terms of connectivity, branch structure, and local geometric features. There is also a lack of mechanisms for point set representation and capture-like pose generation using capture prototype priors.

[0004] In terms of capture decision-making and trajectory planning, existing UAV capture systems mostly adopt empirical rules or simple geometric reachability constraints. The modeling of the dynamic characteristics of the capturing UAV and the kinematic features of the capturing mechanism is relatively rough. Usually, the capture conditions are judged only by conditions such as distance threshold and angle threshold. A clear capture feasible domain model is not constructed, making it difficult to screen out the target representative points that can actually perform capture actions in three-dimensional space. In addition, most existing path planning methods are based on static obstacles or simplified cost functions for trajectory search, which do not fully consider environmental risk factors such as wind field disturbance and no-fly zones. They lack the ability to construct a spatial risk field in a unified coordinate system and jointly optimize the capture trajectory with the set of capture feasible points. As a result, the planned capture trajectory is insufficient in satisfying dynamic and capture constraints, avoiding environmental risks, and ensuring capture success rate. It also has limited adaptability to complex target areas and multi-UAV cooperative capture scenarios.

[0005] Therefore, how to provide a cloud-based collaborative intelligent drone detection and capture system is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0006] One objective of this invention is to propose a cloud-based collaborative intelligent detection and capture system for unmanned aerial vehicles (UAVs). This invention utilizes methods such as an airborne improved point set detection network, cloud-based multi-view 3D point set reconstruction, topology correction, and prototype sparse coding to obtain the topologically consistent 3D structure of the target object. By jointly planning the capture trajectory and generating executable control commands through a capture feasible domain model and a spatial risk field, it achieves refined target detection and stable capture in a multi-UAV collaborative environment. It has the advantages of high spatial positioning accuracy, reliable structure recovery, strong feasibility of capture decisions, and good adaptability to complex environments.

[0007] According to an embodiment of the present invention, a cloud-based collaborative intelligent detection and capture system for unmanned aerial vehicles includes:

[0008] The capture drone configuration and data acquisition module is used to deploy capture drones in the target area, acquire environmental images and attitude navigation data containing the target object in a unified coordinate system, and upload them to the cloud collaborative platform through the wireless communication unit.

[0009] The improved RepPoints point set detection module is used to input environmental images into the improved RepPoints point set detection network, output two-dimensional point sets of candidate target locations and target class probabilities, and obtain an initial two-dimensional point set;

[0010] The 3D point set reconstruction module is used to construct spatial rays and perform geometric intersections in a unified coordinate system based on the initial 2D point set and attitude navigation data, thereby generating an initial 3D point set corresponding to the target object.

[0011] The topology constraint correction module is used to perform topology correction on the initial 3D point set based on the preset proximity radius and topology parameters to obtain a topologically consistent 3D point set.

[0012] The capture prototype sparse coding module is used to perform sparse linear representation of topologically consistent 3D point sets using the capture prototype point set dictionary, and generate candidate capture 3D point sets.

[0013] The capture feasible region filtering module is used to establish a capture feasible region model based on the dynamic constraints of the capture UAV and the motion constraints of the capture mechanism, and to filter the capture feasible 3D point set from the candidate capture 3D point set.

[0014] The spatial risk and capture trajectory planning module is used to construct a spatial risk field by fusing environmental sensor data, set a capture trajectory cost function based on the feasible 3D point set and the spatial risk field, search for a capture trajectory that satisfies the dynamic constraints of the capture UAV and the motion constraints of the capture mechanism, generate control commands and send them to the target capture UAV to execute the capture operation.

[0015] Optionally, modules can be integrated using the following methods:

[0016] Deploy capture drones in the target area to collect environmental images containing the target object and corresponding attitude and navigation data;

[0017] On each capture drone, the environmental image is input into the improved RepPoints point set detection network, which outputs a two-dimensional point set consisting of a set of representative points and the target category probability for the candidate target location, thus obtaining the initial two-dimensional point set;

[0018] The initial two-dimensional point set and attitude navigation data of each captured UAV are uploaded to the cloud collaborative platform. The two-dimensional point set is back-projected into rays in a unified coordinate system and geometrically intersected to obtain the initial three-dimensional point set corresponding to the target object.

[0019] A topological constraint model is constructed in a cloud-based collaborative platform, and the initial 3D point set is topologically corrected to generate a topologically consistent 3D point set that satisfies the preset topological structure.

[0020] A capture prototype point set dictionary is used to sparsely encode the topologically consistent 3D point set, and candidate capture 3D point sets are generated based on the combination of sparse coefficients.

[0021] A feasible capture domain model is established based on the dynamic constraints of the capture UAV and the motion constraints of the capture mechanism. The feasible capture 3D point set is obtained by filtering the candidate capture 3D point set as input.

[0022] By integrating environmental sensor data to construct a spatial risk field, the capture trajectory and corresponding control commands are optimized based on the capture feasible 3D point set and spatial risk field, and the control commands are sent to the target capture UAV to execute the capture operation.

[0023] Optionally, the acquisition of the environmental image and the corresponding attitude navigation data includes:

[0024] Establish a unified coordinate system in the target area, and register the boundary information of the target area and capture the coordinates of the UAV takeoff point at the control terminal;

[0025] An image acquisition unit, attitude measurement unit, navigation and positioning unit, and wireless communication unit are installed on the airframe of each capture drone, and the image acquisition unit, attitude measurement unit, and navigation and positioning unit are connected to the airborne processing unit;

[0026] The control terminal assigns takeoff point, cruising altitude, cruising speed and cruising route parameters to each capture drone, sends takeoff command and cruising parameters, and controls the capture drone to fly in the target area according to the cruising route;

[0027] During the capture of the UAV flying along the cruise route, the image acquisition unit is triggered to acquire environmental images at preset time intervals. At the same time, the attitude data output by the attitude measurement unit and the position and velocity information output by the navigation and positioning unit are read. The environmental images are associated with the corresponding attitude data, position information and velocity information to form attitude navigation data corresponding to the environmental images.

[0028] Optionally, the establishment of the improved RepPoints point set detection network and the generation of the initial two-dimensional point set include:

[0029] An improved RepPoints point set detection network is deployed on each capture drone. The improved RepPoints point set detection network includes a backbone feature extraction subnetwork, a feature pyramid fusion subnetwork, and a point set detection subnetwork. The backbone feature extraction subnetwork, the feature pyramid fusion subnetwork, and the point set detection subnetwork are connected sequentially through feature map transmission. The backbone feature extraction subnetwork consists of convolutional layers, nonlinear activation layers, and downsampling layers. It performs convolution operations, feature channel transformations, and spatial scale reduction on the environmental image in a preset order to generate feature maps of different resolutions. The feature pyramid fusion subnetwork consists of lateral connection layers and upsampling layers. It performs channel alignment, point-by-point addition, and spatial interpolation operations between feature maps of different resolutions to generate pyramid feature maps with top-down fusion paths and bottom-up fusion paths. The point set detection subnetwork consists of shared convolutional layers, classification branch convolutional layers, and regression branch convolutional layers. It outputs the target class probability and the representative point position offset parameter at each position of the pyramid feature map.

[0030] On the drone, the environmental image is resized, pixel value normalized and color channel normalized to obtain a preprocessed environmental image. The preprocessed environmental image is input into the backbone feature extraction subnetwork to generate at least two scale feature maps. The scale feature maps are input into the feature pyramid fusion subnetwork to generate a fused feature map.

[0031] Discrete sampling locations are selected on the fused feature map as candidate target location centers. The point set regression branch outputs a set of representative point coordinate offsets in the horizontal and vertical directions based on each candidate target location center. The coordinates of the candidate target location center are added to the corresponding coordinate offsets to determine a set of two-dimensional coordinates of representative points corresponding to each candidate target location in the original image coordinate system, forming a two-dimensional point set candidate set.

[0032] The classification branch calculates the target category probability and detection confidence for each candidate target location in the two-dimensional point set candidate set. Based on the preset detection confidence threshold and the overlap suppression rule, the two-dimensional point set is filtered from the two-dimensional point set candidate set to obtain the initial two-dimensional point set. The preset detection confidence threshold is the minimum detection confidence value used to limit the retention of candidate target locations. The overlap suppression rule is to retain only the candidate target locations with higher detection confidence for candidate target locations in the two-dimensional point set candidate set whose overlap is greater than the overlap threshold.

[0033] Optionally, the generation of the initial three-dimensional point set includes:

[0034] Each capture drone packages the initial 2D point set, the attitude navigation data of the corresponding environmental image, and the timestamp into a data frame through the wireless communication unit and sends it to the cloud collaboration platform;

[0035] The cloud-based collaborative platform parses the data frames and, based on the position and attitude information in the attitude navigation data, determines the camera position and camera attitude corresponding to the environmental image in a unified coordinate system, which are recorded as camera imaging parameters.

[0036] For each representative point in the initial two-dimensional point set, the cloud-based collaborative platform constructs a spatial ray set in a unified coordinate system based on the pixel coordinates of the representative point in the environmental image and the camera imaging parameters, with the camera position as the starting point and the imaging direction of the representative point as the direction.

[0037] The cloud-based collaborative platform divides the spatial rays from different capture drones into ray groups according to the index of the representative points in the initial two-dimensional point set. Geometric intersection calculation is performed on each ray group. Specifically, the geometric intersection calculation is as follows: taking any two spatial rays in the ray group as a pair, the nearest point pair of each pair of spatial rays is calculated. The midpoint coordinates are calculated in a unified coordinate system based on the position of the nearest point pair. The sum of squared distances of all midpoint coordinates is used as the objective function to minimize the solution, and the spatial position with the minimum sum of squared distances is obtained. The spatial position is used as the three-dimensional coordinates of the corresponding representative point in the unified coordinate system. The three-dimensional coordinates of all representative points are arranged in index order to form an initial three-dimensional point set corresponding to the target object.

[0038] Optionally, the generation of the topologically consistent 3D point set includes:

[0039] The cloud-based collaborative platform reads the initial 3D point set corresponding to each target object, constructs adjacency relationships in the initial 3D point set according to the preset proximity radius, and generates a 3D point set topology map based on the adjacency relationship. The preset proximity radius is a distance threshold parameter used to determine whether any two points have established an adjacency relationship. It is statistically obtained from experimental data based on the target object size range and the sampling interval of the 3D point set before system deployment, and is registered as a fixed value in the cloud-based collaborative platform.

[0040] The cloud-based collaborative platform generates a topology template based on a preset set of topology parameters, including the number of connected components, branch nodes, loops, and endpoints. It then calculates the topology parameter set corresponding to the 3D point set topology graph, compares this set with the topology template, and determines the index of topological differences. The preset topology parameter set is obtained by performing connected component segmentation, node degree counting, and closed path detection on the labeled 3D point set collected during the training phase. The topology parameter set corresponding to the 3D point set topology graph is calculated by traversing the topology graph to calculate the number of connected components, counting the number of branch nodes and endpoints through node degree counting, and enumerating closed paths to calculate the number of loops.

[0041] The cloud-based collaborative platform performs topology correction operations on the topology difference location index. The topology correction operations include deleting isolated points from the initial 3D point set, inserting interpolation points between adjacent point pairs with a point spacing exceeding a first threshold, merging point coordinates on adjacent point pairs with a point spacing less than a second threshold, and translating the coordinates of relevant points near the branch node along the local principal direction to obtain an updated 3D point set.

[0042] The cloud-based collaborative platform reconstructs the adjacency relationships and topology parameter set based on the updated 3D point set. When the topology parameter set is consistent with the topology template, the updated 3D point set is recorded as the topology-consistent 3D point set.

[0043] Optionally, the generation of the candidate captured 3D point set includes:

[0044] In the cloud collaboration platform, a dictionary of captured prototype point sets is established. A set of captured prototype 3D point sets is registered as dictionary elements. Each captured prototype 3D point set consists of a set of 3D representative points under a unified coordinate system. Each 3D representative point contains a coordinate index and 3D coordinate values.

[0045] In the cloud-based collaborative platform, the topologically consistent 3D point set is expanded into a target feature vector according to the coordinate index, and the captured prototype 3D point set in the captured prototype point set dictionary is expanded into a set of base feature vectors according to the coordinate index.

[0046] In the cloud-based collaborative platform, a sparse linear representation model is constructed. In the sparse linear representation model, the target feature vector is represented as a linear combination of the base feature vector set. By solving the coefficient vector that satisfies that the number of non-zero coefficients does not exceed the sparsity threshold, the sparse coefficients corresponding one-to-one with the dictionary of captured prototype point sets are obtained.

[0047] In the cloud-based collaborative platform, based on the sparsity coefficient, capture prototype 3D point sets with sparsity coefficients greater than zero are selected from the capture prototype point set dictionary, and weighted superposition is performed using the sparsity coefficient as the weight to generate candidate capture 3D point sets.

[0048] Optionally, the generation of the capture feasible 3D point set includes:

[0049] The cloud-based collaborative platform reads the dynamic parameters of the captured drone and the motion parameters of the captured mechanism. The dynamic parameters include the maximum flight speed, maximum acceleration, maximum angular velocity, maximum pitch angle, and maximum roll angle. The motion parameters of the captured mechanism include the upper limit of the captured mechanism's working distance, the lower limit of the captured mechanism's working distance, the range of the captured mechanism's pitch angle, and the range of the captured mechanism's yaw angle.

[0050] The cloud-based collaborative platform establishes a capture feasible domain model in a unified coordinate system based on dynamic parameters and capture mechanism motion parameters. The capture feasible domain model is divided into position constraint subdomain, velocity constraint subdomain, attitude constraint subdomain, and capture distance angle constraint subdomain. The position constraint subdomain is used to limit the spatial distance range between the center of mass of the capture UAV and the target 3D representative point. The velocity constraint subdomain is used to limit the velocity range of the capture UAV at the moment of capture. The attitude constraint subdomain is used to limit the pitch angle and roll angle range at the moment of capture. The capture distance angle constraint subdomain is used to limit the angle range between the line connecting the center of mass of the capture UAV and the 3D representative point and the axis of the capture mechanism.

[0051] For each 3D representative point in the candidate capture 3D point set, the cloud-based collaborative platform calculates the spatial distance, pitch angle, and yaw angle between the centroid position and the 3D representative point based on the predicted centroid position and attitude parameters of the capture UAV. It also calculates the instantaneous velocity based on the predicted velocity vector of the capture UAV. The platform compares the spatial distance, pitch angle, yaw angle, and velocity magnitude with the value ranges of the position constraint subdomain, attitude constraint subdomain, and velocity constraint subdomain in the capture feasible domain model to obtain the constraint satisfaction mark of the 3D representative point in the capture feasible domain model.

[0052] The cloud-based collaborative platform calculates the capture feasibility value for each 3D representative point in the candidate capture 3D point set based on the constraint satisfaction. The capture feasibility value is a real number between zero and one. The capture feasibility value is obtained by weighted summation of the satisfaction of the position constraint subdomain, velocity constraint subdomain, attitude constraint subdomain, and capture distance angle constraint subdomain.

[0053] The cloud-based collaborative platform filters representative 3D points from the candidate capture 3D point set based on the capture feasibility value. It then combines the representative 3D points with capture feasibility values ​​greater than the capture feasibility threshold according to their original coordinate index order to form a capture feasible 3D point set.

[0054] Optionally, the capture operation includes:

[0055] The cloud-based collaborative platform receives environmental sensor data, converts obstacle detection data, atmospheric wind field data, terrain height data, and no-fly zone boundary data into a unified coordinate system, and synchronously processes the environmental sensor data according to time stamps to generate a set of three-dimensional environmental sampling points.

[0056] The cloud-based collaborative platform divides the target area into three-dimensional grid cells under a unified coordinate system. For each three-dimensional grid cell, the risk value is calculated based on the set of three-dimensional environmental sampling points. The risk components related to obstacle distance, atmospheric wind speed, and no-fly zone coverage are weighted and summed according to preset weights to obtain the risk value of the three-dimensional grid cell. The risk value of the three-dimensional grid cell is then registered as the grid weight of the spatial risk field.

[0057] The cloud-based collaborative platform establishes a capture trajectory cost function based on the capture feasible 3D point set and spatial risk field. The capture trajectory cost function is expressed as a weighted sum of time cost and risk cost. The time cost is calculated based on the flight time of the target capture UAV along the capture trajectory, and the risk cost is calculated by accumulating the grid weights corresponding to the sampling point positions of the spatial risk field. The time cost weight parameters and risk cost weight parameters are pre-registered as fixed coefficients in the cloud-based collaborative platform.

[0058] The cloud-based collaborative platform generates several candidate capture trajectories within the spanned space of the feasible 3D point set. Based on the capture trajectory cost function and the dynamic constraints of the capture UAV, the candidate capture trajectories are screened, and the capture trajectory with the smallest cost function value and satisfying the dynamic constraints is selected. The capture trajectory is discretized into a time-stamped track point sequence and an attitude command sequence, and control commands are generated. The control commands are then sent to the target capture UAV through the wireless communication unit to execute the capture operation.

[0059] The beneficial effects of this invention are:

[0060] This invention deploys an improved point set detection network at the capture drone end and integrates environmental images and attitude navigation data from multiple drones in a cloud-based collaborative platform. This enables multi-view point set reconstruction and topological correction of target objects within a unified coordinate system. Compared to existing schemes based solely on 2D bounding boxes or simple 3D fitting, this invention more accurately recovers the 3D structure of the target in terms of connectivity, branching structure, and local geometry. By leveraging the sparse coding of topologically consistent 3D point sets and the capture prototype point set dictionary, this invention can extract candidate capture 3D point sets that conform to typical capture postures from complex reconstruction results, maintaining high geometric stability and recognition accuracy even in scenarios with occlusion, significant viewpoint changes, and complex backgrounds.

[0061] In the capture decision and trajectory planning stages, this invention constructs a capture feasible region model by introducing the dynamic parameters of the capture UAV and the motion parameters of the capture mechanism. It then constrains and filters the spatial position, relative attitude, and instantaneous velocity of each representative point in the candidate capture 3D point set, retaining only the set of feasible 3D points that are actually reachable and capable of executing capture actions. Compared to methods relying on empirical thresholds and simplified geometric conditions, this significantly improves the executability and success rate of capture actions. Simultaneously, this invention integrates obstacle information, wind field information, and no-fly zone information in a unified coordinate system to construct a spatial risk field, incorporating both flight time and spatial risk into the capture trajectory cost function. Under the premise of satisfying dynamic and capture mechanism constraints, it searches for the capture trajectory with the minimum cost, enabling the capture process to balance safety, efficiency, and stability in complex environments, reducing collision risk and the probability of abnormal loss of control. Attached Figure Description

[0062] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:

[0063] Figure 1 This is a flowchart of a cloud-based collaborative intelligent detection and capture system for unmanned aerial vehicles (UAVs) proposed in this invention.

[0064] Figure 2This is a schematic diagram of three-dimensional point set reconstruction of a cloud-based collaborative UAV intelligent detection and capture system proposed in this invention. Detailed Implementation

[0065] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.

[0066] refer to Figure 1-2 A cloud-based collaborative intelligent drone detection and capture system includes:

[0067] The capture drone configuration and data acquisition module is used to deploy capture drones in the target area, acquire environmental images and attitude navigation data containing the target object in a unified coordinate system, and upload them to the cloud collaborative platform through the wireless communication unit.

[0068] The improved RepPoints point set detection module is used to input environmental images into the improved RepPoints point set detection network, output two-dimensional point sets of candidate target locations and target class probabilities, and obtain an initial two-dimensional point set;

[0069] The 3D point set reconstruction module is used to construct spatial rays and perform geometric intersections in a unified coordinate system based on the initial 2D point set and attitude navigation data, thereby generating an initial 3D point set corresponding to the target object.

[0070] The topology constraint correction module is used to perform topology correction on the initial 3D point set based on the preset proximity radius and topology parameters to obtain a topologically consistent 3D point set.

[0071] The capture prototype sparse coding module is used to perform sparse linear representation of topologically consistent 3D point sets using the capture prototype point set dictionary, and generate candidate capture 3D point sets.

[0072] The capture feasible region filtering module is used to establish a capture feasible region model based on the dynamic constraints of the capture UAV and the motion constraints of the capture mechanism, and to filter the capture feasible 3D point set from the candidate capture 3D point set.

[0073] The spatial risk and capture trajectory planning module is used to construct a spatial risk field by fusing environmental sensor data, set a capture trajectory cost function based on the feasible 3D point set and the spatial risk field, search for a capture trajectory that satisfies the dynamic constraints of the capture UAV and the motion constraints of the capture mechanism, generate control commands and send them to the target capture UAV to execute the capture operation.

[0074] In this embodiment, the modules are interconnected using the following method:

[0075] Deploy capture drones in the target area to collect environmental images containing the target object and corresponding attitude and navigation data;

[0076] On each capture drone, the environmental image is input into the improved RepPoints point set detection network, which outputs a two-dimensional point set consisting of a set of representative points and the target category probability for the candidate target location, thus obtaining the initial two-dimensional point set;

[0077] The initial two-dimensional point set and attitude navigation data of each captured UAV are uploaded to the cloud collaborative platform. The two-dimensional point set is back-projected into rays in a unified coordinate system and geometrically intersected to obtain the initial three-dimensional point set corresponding to the target object.

[0078] A topological constraint model is constructed in a cloud-based collaborative platform, and the initial 3D point set is topologically corrected to generate a topologically consistent 3D point set that satisfies the preset topological structure.

[0079] A capture prototype point set dictionary is used to sparsely encode the topologically consistent 3D point set, and candidate capture 3D point sets are generated based on the combination of sparse coefficients.

[0080] A feasible capture domain model is established based on the dynamic constraints of the capture UAV and the motion constraints of the capture mechanism. The feasible capture 3D point set is obtained by filtering the candidate capture 3D point set as input.

[0081] By integrating environmental sensor data to construct a spatial risk field, the capture trajectory and corresponding control commands are optimized based on the capture feasible 3D point set and spatial risk field, and the control commands are sent to the target capture UAV to execute the capture operation.

[0082] In this embodiment, the acquisition of the environmental image and the corresponding attitude navigation data includes:

[0083] Establish a unified coordinate system in the target area, and register the boundary information of the target area and capture the coordinates of the UAV takeoff point at the control terminal;

[0084] An image acquisition unit, attitude measurement unit, navigation and positioning unit, and wireless communication unit are installed on the airframe of each capture drone, and the image acquisition unit, attitude measurement unit, and navigation and positioning unit are connected to the airborne processing unit;

[0085] The control terminal assigns takeoff point, cruising altitude, cruising speed and cruising route parameters to each capture drone, sends takeoff command and cruising parameters, and controls the capture drone to fly in the target area according to the cruising route;

[0086] During the capture of the UAV flying along the cruise route, the image acquisition unit is triggered to acquire environmental images at preset time intervals. At the same time, the attitude data output by the attitude measurement unit and the position and velocity information output by the navigation and positioning unit are read. The environmental images are associated with the corresponding attitude data, position information and velocity information to form attitude navigation data corresponding to the environmental images.

[0087] In this embodiment, the establishment of the improved RepPoints point set detection network and the generation of the initial two-dimensional point set include:

[0088] An improved RepPoints point set detection network is deployed on each capture drone. The improved RepPoints point set detection network includes a backbone feature extraction subnetwork, a feature pyramid fusion subnetwork, and a point set detection subnetwork. The backbone feature extraction subnetwork, the feature pyramid fusion subnetwork, and the point set detection subnetwork are connected sequentially through feature map transmission. The backbone feature extraction subnetwork consists of convolutional layers, nonlinear activation layers, and downsampling layers. It performs convolution operations, feature channel transformations, and spatial scale reduction on the environmental image in a preset order to generate feature maps of different resolutions. The feature pyramid fusion subnetwork consists of lateral connection layers and upsampling layers. It performs channel alignment, point-by-point addition, and spatial interpolation operations between feature maps of different resolutions to generate pyramid feature maps with top-down fusion paths and bottom-up fusion paths. The point set detection subnetwork consists of shared convolutional layers, classification branch convolutional layers, and regression branch convolutional layers. It outputs the target class probability and the representative point position offset parameter at each position of the pyramid feature map.

[0089] On the drone, the environmental image is resized, pixel value normalized and color channel normalized to obtain a preprocessed environmental image. The preprocessed environmental image is input into the backbone feature extraction subnetwork to generate at least two scale feature maps. The scale feature maps are input into the feature pyramid fusion subnetwork to generate a fused feature map.

[0090] Discrete sampling locations are selected on the fused feature map as candidate target location centers. The point set regression branch outputs a set of representative point coordinate offsets in the horizontal and vertical directions based on each candidate target location center. The coordinates of the candidate target location center are added to the corresponding coordinate offsets to determine a set of two-dimensional coordinates of representative points corresponding to each candidate target location in the original image coordinate system, forming a two-dimensional point set candidate set.

[0091] The classification branch calculates the target category probability and detection confidence for each candidate target location in the two-dimensional point set candidate set. Based on the preset detection confidence threshold and the overlap suppression rule, the two-dimensional point set is filtered from the two-dimensional point set candidate set to obtain the initial two-dimensional point set. The preset detection confidence threshold is the minimum detection confidence value used to limit the retention of candidate target locations. The overlap suppression rule is to retain only the candidate target locations with higher detection confidence for candidate target locations in the two-dimensional point set candidate set whose overlap is greater than the overlap threshold.

[0092] This invention deploys an improved RepPoints point set detection network on a capture drone, consisting of a backbone feature extraction subnetwork, a feature pyramid fusion subnetwork, and a point set detection subnetwork. This network performs multi-scale feature extraction and bidirectional pyramid fusion on the preprocessed environmental image, and represents the candidate target location in the form of a representative point set on the fused feature map. The initial two-dimensional point set is selected by combining the detection confidence threshold and overlap suppression rules. This enables the drone to stably obtain a target point set representation with fine contours and accurate positioning even in complex scenes with large changes in viewpoint, significant differences in target size, and occlusion interference. This provides a unified and high-quality visual input foundation for subsequent three-dimensional point set reconstruction and capture trajectory planning.

[0093] In this embodiment, the generation of the initial three-dimensional point set includes:

[0094] Each capture drone packages the initial 2D point set, the attitude navigation data of the corresponding environmental image, and the timestamp into a data frame through the wireless communication unit and sends it to the cloud collaboration platform;

[0095] The cloud-based collaborative platform parses the data frames and, based on the position and attitude information in the attitude navigation data, determines the camera position and camera attitude corresponding to the environmental image in a unified coordinate system, which are recorded as camera imaging parameters.

[0096] For each representative point in the initial two-dimensional point set, the cloud-based collaborative platform constructs a spatial ray set in a unified coordinate system based on the pixel coordinates of the representative point in the environmental image and the camera imaging parameters, with the camera position as the starting point and the imaging direction of the representative point as the direction.

[0097] The cloud-based collaborative platform divides the spatial rays from different capture drones into ray groups according to the index of the representative points in the initial two-dimensional point set. Geometric intersection calculation is performed on each ray group. Specifically, the geometric intersection calculation is as follows: taking any two spatial rays in the ray group as a pair, the nearest point pair of each pair of spatial rays is calculated. The midpoint coordinates are calculated in a unified coordinate system based on the position of the nearest point pair. The sum of squared distances of all midpoint coordinates is used as the objective function to minimize the solution, and the spatial position with the minimum sum of squared distances is obtained. The spatial position is used as the three-dimensional coordinates of the corresponding representative point in the unified coordinate system. The three-dimensional coordinates of all representative points are arranged in index order to form an initial three-dimensional point set corresponding to the target object.

[0098] In this embodiment, the generation of the topologically consistent 3D point set includes:

[0099] The cloud-based collaborative platform reads the initial 3D point set corresponding to each target object, constructs adjacency relationships in the initial 3D point set according to the preset proximity radius, and generates a 3D point set topology map based on the adjacency relationship. The preset proximity radius is a distance threshold parameter used to determine whether any two points have established an adjacency relationship. It is statistically obtained from experimental data based on the target object size range and the sampling interval of the 3D point set before system deployment, and is registered as a fixed value in the cloud-based collaborative platform.

[0100] The cloud-based collaborative platform generates a topology template based on a preset set of topology parameters, including the number of connected components, branch nodes, loops, and endpoints. It then calculates the topology parameter set corresponding to the 3D point set topology graph, compares this set with the topology template, and determines the index of topological differences. The preset topology parameter set is obtained by performing connected component segmentation, node degree counting, and closed path detection on the labeled 3D point set collected during the training phase. The topology parameter set corresponding to the 3D point set topology graph is calculated by traversing the topology graph to calculate the number of connected components, counting the number of branch nodes and endpoints through node degree counting, and enumerating closed paths to calculate the number of loops.

[0101] The cloud-based collaborative platform performs topology correction operations on the topology difference location index. The topology correction operations include deleting isolated points from the initial 3D point set, inserting interpolation points between adjacent point pairs with a point spacing exceeding a first threshold, merging point coordinates on adjacent point pairs with a point spacing less than a second threshold, and translating the coordinates of relevant points near the branch node along the local principal direction to obtain an updated 3D point set.

[0102] The cloud-based collaborative platform reconstructs the adjacency relationships and topology parameter set based on the updated 3D point set. When the topology parameter set is consistent with the topology template, the updated 3D point set is recorded as the topology-consistent 3D point set.

[0103] In this embodiment, the generation of the candidate capture 3D point set includes:

[0104] In the cloud collaboration platform, a dictionary of captured prototype point sets is established. A set of captured prototype 3D point sets is registered as dictionary elements. Each captured prototype 3D point set consists of a set of 3D representative points under a unified coordinate system. Each 3D representative point contains a coordinate index and 3D coordinate values.

[0105] In the cloud-based collaborative platform, the topologically consistent 3D point set is expanded into a target feature vector according to the coordinate index, and the captured prototype 3D point set in the captured prototype point set dictionary is expanded into a set of base feature vectors according to the coordinate index.

[0106] In the cloud-based collaborative platform, a sparse linear representation model is constructed. In the sparse linear representation model, the target feature vector is represented as a linear combination of the base feature vector set. By solving the coefficient vector that satisfies that the number of non-zero coefficients does not exceed the sparsity threshold, the sparse coefficients corresponding one-to-one with the dictionary of captured prototype point sets are obtained.

[0107] In the cloud-based collaborative platform, based on the sparsity coefficient, capture prototype 3D point sets with sparsity coefficients greater than zero are selected from the capture prototype point set dictionary, and weighted superposition is performed using the sparsity coefficient as the weight to generate candidate capture 3D point sets.

[0108] In this embodiment, the generation of the capture feasible 3D point set includes:

[0109] The cloud-based collaborative platform reads the dynamic parameters of the captured drone and the motion parameters of the captured mechanism. The dynamic parameters include the maximum flight speed, maximum acceleration, maximum angular velocity, maximum pitch angle, and maximum roll angle. The motion parameters of the captured mechanism include the upper limit of the captured mechanism's working distance, the lower limit of the captured mechanism's working distance, the range of the captured mechanism's pitch angle, and the range of the captured mechanism's yaw angle.

[0110] The cloud-based collaborative platform establishes a capture feasible domain model in a unified coordinate system based on dynamic parameters and capture mechanism motion parameters. The capture feasible domain model is divided into position constraint subdomain, velocity constraint subdomain, attitude constraint subdomain, and capture distance angle constraint subdomain. The position constraint subdomain is used to limit the spatial distance range between the center of mass of the capture UAV and the target 3D representative point. The velocity constraint subdomain is used to limit the velocity range of the capture UAV at the moment of capture. The attitude constraint subdomain is used to limit the pitch angle and roll angle range at the moment of capture. The capture distance angle constraint subdomain is used to limit the angle range between the line connecting the center of mass of the capture UAV and the 3D representative point and the axis of the capture mechanism.

[0111] For each 3D representative point in the candidate capture 3D point set, the cloud-based collaborative platform calculates the spatial distance, pitch angle, and yaw angle between the centroid position and the 3D representative point based on the predicted centroid position and attitude parameters of the capture UAV. It also calculates the instantaneous velocity based on the predicted velocity vector of the capture UAV. The platform compares the spatial distance, pitch angle, yaw angle, and velocity magnitude with the value ranges of the position constraint subdomain, attitude constraint subdomain, and velocity constraint subdomain in the capture feasible domain model to obtain the constraint satisfaction mark of the 3D representative point in the capture feasible domain model.

[0112] The cloud-based collaborative platform calculates the capture feasibility value for each 3D representative point in the candidate capture 3D point set based on the constraint satisfaction. The capture feasibility value is a real number between zero and one. The capture feasibility value is obtained by weighted summation of the satisfaction of the position constraint subdomain, velocity constraint subdomain, attitude constraint subdomain, and capture distance angle constraint subdomain.

[0113] The cloud-based collaborative platform filters representative 3D points from the candidate capture 3D point set based on the capture feasibility value. It then combines the representative 3D points with capture feasibility values ​​greater than the capture feasibility threshold according to their original coordinate index order to form a capture feasible 3D point set.

[0114] In this embodiment, the capture operation includes:

[0115] The cloud-based collaborative platform receives environmental sensor data, converts obstacle detection data, atmospheric wind field data, terrain height data, and no-fly zone boundary data into a unified coordinate system, and synchronously processes the environmental sensor data according to time stamps to generate a set of three-dimensional environmental sampling points.

[0116] The cloud-based collaborative platform divides the target area into three-dimensional grid cells under a unified coordinate system. For each three-dimensional grid cell, the risk value is calculated based on the set of three-dimensional environmental sampling points. The risk components related to obstacle distance, atmospheric wind speed, and no-fly zone coverage are weighted and summed according to preset weights to obtain the risk value of the three-dimensional grid cell. The risk value of the three-dimensional grid cell is then registered as the grid weight of the spatial risk field.

[0117] The cloud-based collaborative platform establishes a capture trajectory cost function based on the capture feasible 3D point set and spatial risk field. The capture trajectory cost function is expressed as a weighted sum of time cost and risk cost. The time cost is calculated based on the flight time of the target capture UAV along the capture trajectory, and the risk cost is calculated by accumulating the grid weights corresponding to the sampling point positions of the spatial risk field. The time cost weight parameters and risk cost weight parameters are pre-registered as fixed coefficients in the cloud-based collaborative platform.

[0118] The cloud-based collaborative platform generates several candidate capture trajectories within the spanned space of the feasible 3D point set. Based on the capture trajectory cost function and the dynamic constraints of the capture UAV, the candidate capture trajectories are screened, and the capture trajectory with the smallest cost function value and satisfying the dynamic constraints is selected. The capture trajectory is discretized into a time-stamped track point sequence and an attitude command sequence, and control commands are generated. The control commands are then sent to the target capture UAV through the wireless communication unit to execute the capture operation.

[0119] Example 1:

[0120] To verify the feasibility of this invention in practice, it was applied to a low-altitude maneuvering target acquisition scenario within a closed test airspace. This test airspace was equipped with multi-faceted simulated building facades, steel frame structures, and flexible obstructions, divided into multi-layered three-dimensional mesh units in the height, lateral, and longitudinal directions to simulate a complex urban canyon environment. The target was a small aircraft, approximately 0.4m long, 0.4m wide, and 0.25m high, maneuvering along a random trajectory within the test airspace, continuously changing altitude and heading to create multi-view, multi-attitude, and partially obstructed observation conditions. In comparative testing, the traditional approach, using a "single-machine two-dimensional detection straight-line interception" method, relied solely on a single acquisition drone equipped with a conventional two-dimensional target detection network. The target's planar position was calculated based on the image bounding box center, approaching the target via a simple straight line or polygonal line. Emergency avoidance was only triggered by a distance threshold when approaching obstacles. This method is prone to problems such as large deviations in three-dimensional position estimation, mismatch between the target's attitude at the acquisition moment, and excessively close proximity to obstacles when the target is rapidly maneuvering or obstructed, significantly limiting both the acquisition success rate and safety.

[0121] In the same test scenario, this invention deploys multiple capture drones, which form a cross-view coverage of the test airspace from different altitudes and orientations. A unified coordinate system is pre-established in the control terminal, and the boundaries of the test airspace and the positions of each takeoff point are registered. Each capture drone cruises in the airspace according to a predetermined cruise altitude and route. During the cruise, the onboard image acquisition unit is triggered at fixed time intervals to acquire environmental images containing the target object. At the same time, the roll angle, pitch angle, and yaw angle output by the attitude measurement unit, as well as the position and velocity information output by the navigation and positioning unit, are recorded. The environmental images are correlated with the attitude and navigation data, and the onboard processing unit completes the size adjustment, pixel normalization, and color channel labeling. Standardization processing: The preprocessed image is input into the improved RepPoints point set detection network. On the multi-scale feature map, it outputs a two-dimensional point set consisting of several representative points and the target category probability for the candidate target location. Based on the detection confidence threshold and overlap suppression rules, the initial two-dimensional point set is obtained. Each capture UAV uploads the initial two-dimensional point set and corresponding attitude navigation data to the cloud collaborative platform through the wireless communication unit. The cloud collaborative platform constructs a set of spatial rays according to the camera position and attitude in a unified coordinate system. The rays from different UAVs are grouped according to the representative point index. The spatial position with the minimum sum of squared distances is obtained through geometric intersection, forming the initial three-dimensional point set corresponding to the target object. The cloud-based collaborative platform further performs topology correction on the initial 3D point set based on a preset topology template. This involves deleting isolated points, filling broken links, merging overly dense points, and fine-tuning point coordinates along the local principal direction in branch regions. This ensures that topology parameters such as the number of connected components, branch nodes, and endpoints are consistent with the template, resulting in a topology-consistent 3D point set. Subsequently, the topology-consistent 3D point set is expanded into a target feature vector. A capture prototype point set dictionary containing various capture posture arrangements is introduced. The sparse coefficient vector is solved using sparse linear representation. Candidate capture 3D point sets are generated based on combinations of non-zero coefficients. Finally, based on the maximum acceleration and maximum... A feasible capture domain model is established using parameters such as angular velocity, the deployment angle of the capture mechanism, the net bag diameter, and the jet velocity. The candidate capture 3D point set is screened for position, attitude, and velocity constraints to obtain a feasible capture 3D point set. A spatial risk field is constructed by combining obstacle distance sensor and wind speed sensor data. A capture trajectory cost function is constructed with flight time and spatial risk as weights. Under the premise of satisfying dynamics and capture mechanism constraints, the capture trajectory with minimum cost is searched. The trajectory is discretized into a time-stamped track point sequence and an attitude command sequence, which are sent to the target capture UAV to achieve collaborative capture of maneuvering targets.

[0122] To verify the effectiveness of this invention in complex environments, the same test airspace, similar target maneuvering intensity, and wind disturbance conditions were selected. Multiple batches of capture missions were executed using both the traditional single-aircraft two-dimensional detection straight-line interception scheme and the scheme of this invention. The capture success rate, average target three-dimensional positioning error, relative attitude error at the moment of capture, average completion time per mission, and the number of dangerous approaches and forced mission terminations within each fixed number of missions were statistically analyzed. The results are shown in Table 1.

[0123] Table 1. Results of Capture Performance Comparison Test

[0124] index Traditional single-machine two-dimensional detection line-of-sight interception solution This invention is a cloud-based collaborative intelligent detection and capture system for unmanned aerial vehicles (UAVs). Capture success rate 65% 93% Target 3D positioning average error / m 0.80 0.26 Mean relative attitude error at the moment of capture / (°) 18.5 6.3 Average completion time per task / seconds 110 74 Number of dangerous proximity events (per 50 missions) 8 1 Number of times a mission is forced to be aborted (every 50 missions) 6 0

[0125] The above comparison shows that, under the same target and experimental conditions, this invention, relying on multi-UAV multi-view point set detection and cloud-based 3D point set reconstruction, combined with topological constraints, sparse coding of the capture prototype point set dictionary, capture feasible domain model, and capture trajectory jointly optimized by spatial risk field, significantly reduces the target's 3D positioning error and the instantaneous attitude error during capture. This significantly improves the capture success rate and mission efficiency, and greatly reduces the number of dangerous approaches and forced mission terminations. Thus, it effectively solves the problems of inaccurate capture position, mismatched capture attitude, and insufficient consideration of environmental risks in traditional single-UAV 2D detection straight-line interception schemes under complex obstacle environments and target maneuvering conditions.

[0126] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A cloud-based collaborative intelligent detection and capture system for unmanned aerial vehicles (UAVs), characterized in that, include: The capture drone configuration and data acquisition module is used to deploy capture drones in the target area, acquire environmental images and attitude navigation data containing the target object in a unified coordinate system, and upload them to the cloud collaborative platform through the wireless communication unit. The improved RepPoints point set detection module is used to input environmental images into the improved RepPoints point set detection network, output two-dimensional point sets of candidate target locations and target class probabilities, and obtain an initial two-dimensional point set; The 3D point set reconstruction module is used to construct spatial rays and perform geometric intersections in a unified coordinate system based on the initial 2D point set and attitude navigation data, thereby generating an initial 3D point set corresponding to the target object. The topology constraint correction module is used to perform topology correction on the initial 3D point set based on the preset proximity radius and topology parameters to obtain a topologically consistent 3D point set. The capture prototype sparse coding module is used to perform sparse linear representation of topologically consistent 3D point sets using the capture prototype point set dictionary, and generate candidate capture 3D point sets. The capture feasible region filtering module is used to establish a capture feasible region model based on the dynamic constraints of the capture UAV and the motion constraints of the capture mechanism, and to filter the capture feasible 3D point set from the candidate capture 3D point set. The spatial risk and capture trajectory planning module is used to construct a spatial risk field by fusing environmental sensor data, set a capture trajectory cost function based on the feasible 3D point set and the spatial risk field, search for a capture trajectory that satisfies the dynamic constraints of the capture UAV and the motion constraints of the capture mechanism, generate control commands and send them to the target capture UAV to execute the capture operation.

2. The cloud-based collaborative UAV intelligent detection and capture system according to claim 1, characterized in that, The modules are connected in the following way: Deploy capture drones in the target area to collect environmental images containing the target object and corresponding attitude and navigation data; On each capture drone, the environmental image is input into the improved RepPoints point set detection network, which outputs a two-dimensional point set consisting of a set of representative points and the target category probability for the candidate target location, thus obtaining the initial two-dimensional point set; The initial two-dimensional point set and attitude navigation data of each captured UAV are uploaded to the cloud collaborative platform. The two-dimensional point set is back-projected into rays in a unified coordinate system and geometrically intersected to obtain the initial three-dimensional point set corresponding to the target object. A topological constraint model is constructed in a cloud-based collaborative platform, and the initial 3D point set is topologically corrected to generate a topologically consistent 3D point set that satisfies the preset topological structure. A capture prototype point set dictionary is used to sparsely encode the topologically consistent 3D point set, and candidate capture 3D point sets are generated based on the combination of sparse coefficients. A feasible capture domain model is established based on the dynamic constraints of the capture UAV and the motion constraints of the capture mechanism. The feasible capture 3D point set is obtained by filtering the candidate capture 3D point set as input. By integrating environmental sensor data to construct a spatial risk field, the capture trajectory and corresponding control commands are optimized based on the capture feasible 3D point set and spatial risk field, and the control commands are sent to the target capture UAV to execute the capture operation.

3. The cloud-based collaborative intelligent detection and capture system for unmanned aerial vehicles (UAVs) according to claim 2, characterized in that, The acquisition of the environmental images and corresponding attitude navigation data includes: Establish a unified coordinate system in the target area, and register the boundary information of the target area and capture the coordinates of the UAV takeoff point at the control terminal; An image acquisition unit, attitude measurement unit, navigation and positioning unit, and wireless communication unit are installed on the airframe of each capture drone, and the image acquisition unit, attitude measurement unit, and navigation and positioning unit are connected to the airborne processing unit; The control terminal assigns takeoff point, cruising altitude, cruising speed and cruising route parameters to each capture drone, sends takeoff command and cruising parameters, and controls the capture drone to fly in the target area according to the cruising route; During the capture of the UAV flying along the cruise route, the image acquisition unit is triggered to acquire environmental images at preset time intervals. At the same time, the attitude data output by the attitude measurement unit and the position and velocity information output by the navigation and positioning unit are read. The environmental images are associated with the corresponding attitude data, position information and velocity information to form attitude navigation data corresponding to the environmental images.

4. The cloud-based collaborative intelligent detection and capture system for unmanned aerial vehicles (UAVs) according to claim 2, characterized in that, The establishment of the improved RepPoints point set detection network and the generation of the initial two-dimensional point set include: An improved RepPoints point set detection network is deployed on each capture drone. The improved RepPoints point set detection network includes a backbone feature extraction subnetwork, a feature pyramid fusion subnetwork, and a point set detection subnetwork. The backbone feature extraction subnetwork, the feature pyramid fusion subnetwork, and the point set detection subnetwork are connected sequentially through feature map transmission. The backbone feature extraction subnetwork consists of convolutional layers, nonlinear activation layers, and downsampling layers. It performs convolution operations, feature channel transformations, and spatial scale reduction on the environmental image in a preset order to generate feature maps of different resolutions. The feature pyramid fusion subnetwork consists of lateral connection layers and upsampling layers. It performs channel alignment, point-by-point addition, and spatial interpolation operations between feature maps of different resolutions to generate pyramid feature maps with top-down fusion paths and bottom-up fusion paths. The point set detection subnetwork consists of shared convolutional layers, classification branch convolutional layers, and regression branch convolutional layers. It outputs the target class probability and the representative point position offset parameter at each position of the pyramid feature map. On the drone, the environmental image is resized, pixel value normalized and color channel normalized to obtain a preprocessed environmental image. The preprocessed environmental image is input into the backbone feature extraction subnetwork to generate at least two scale feature maps. The scale feature maps are input into the feature pyramid fusion subnetwork to generate a fused feature map. Discrete sampling locations are selected on the fused feature map as candidate target location centers. The point set regression branch outputs a set of representative point coordinate offsets in the horizontal and vertical directions based on each candidate target location center. The coordinates of the candidate target location center are added to the corresponding coordinate offsets to determine a set of two-dimensional coordinates of representative points corresponding to each candidate target location in the original image coordinate system, forming a two-dimensional point set candidate set. The classification branch calculates the target category probability and detection confidence for each candidate target location in the candidate set of two-dimensional point sets. Based on the preset detection confidence threshold and overlap suppression rules, the candidate set of two-dimensional point sets is filtered to obtain the initial two-dimensional point set.

5. The cloud-based collaborative intelligent detection and capture system for unmanned aerial vehicles (UAVs) according to claim 2, characterized in that, The generation of the initial three-dimensional point set includes: Each capture drone packages the initial 2D point set, the attitude navigation data of the corresponding environmental image, and the timestamp into a data frame through the wireless communication unit and sends it to the cloud collaboration platform; The cloud-based collaborative platform parses the data frames and, based on the position and attitude information in the attitude navigation data, determines the camera position and camera attitude corresponding to the environmental image in a unified coordinate system, which are recorded as camera imaging parameters. For each representative point in the initial two-dimensional point set, the cloud-based collaborative platform constructs a spatial ray set in a unified coordinate system based on the pixel coordinates of the representative point in the environmental image and the camera imaging parameters, with the camera position as the starting point and the imaging direction of the representative point as the direction. The cloud-based collaborative platform divides the spatial rays from different capture drones into ray groups according to the index of the representative points in the initial two-dimensional point set. It performs geometric intersection calculations on each ray group to obtain the spatial position with the minimum sum of squared distances. The spatial position is then used as the three-dimensional coordinates of the corresponding representative point in a unified coordinate system. All the three-dimensional coordinates of the representative points are arranged in index order to form an initial three-dimensional point set corresponding to the target object.

6. The cloud-based collaborative intelligent detection and capture system for unmanned aerial vehicles (UAVs) according to claim 2, characterized in that, The generation of the topologically consistent 3D point set includes: The cloud-based collaborative platform reads the initial 3D point set corresponding to each target object, constructs adjacency relationships in the initial 3D point set according to the preset proximity radius, and generates a 3D point set topology graph based on the adjacency relationships; The cloud-based collaborative platform generates a topology template based on a preset set of topology parameters, which includes the number of connected components, the number of branch nodes, the number of loops, and the number of endpoints. It calculates the topology parameter set corresponding to the 3D point set topology graph, compares the topology parameter set with the topology template, and determines the index of the topology difference location. The cloud-based collaborative platform performs topology correction operations on the topology difference location index. The topology correction operations include deleting isolated points from the initial 3D point set, inserting interpolation points between adjacent point pairs with a point spacing exceeding a first threshold, merging point coordinates on adjacent point pairs with a point spacing less than a second threshold, and translating the coordinates of relevant points near the branch node along the local principal direction to obtain an updated 3D point set. The cloud-based collaborative platform reconstructs the adjacency relationships and topology parameter set based on the updated 3D point set. When the topology parameter set is consistent with the topology template, the updated 3D point set is recorded as the topology-consistent 3D point set.

7. The cloud-based collaborative intelligent detection and capture system for unmanned aerial vehicles (UAVs) according to claim 2, characterized in that, The generation of the candidate captured 3D point set includes: In the cloud collaboration platform, a dictionary of captured prototype point sets is established. A set of captured prototype 3D point sets is registered as dictionary elements. Each captured prototype 3D point set consists of a set of 3D representative points under a unified coordinate system. Each 3D representative point contains a coordinate index and 3D coordinate values. In the cloud-based collaborative platform, the topologically consistent 3D point set is expanded into a target feature vector according to the coordinate index, and the captured prototype 3D point set in the captured prototype point set dictionary is expanded into a set of base feature vectors according to the coordinate index. In the cloud-based collaborative platform, a sparse linear representation model is constructed. In the sparse linear representation model, the target feature vector is represented as a linear combination of the base feature vector set. By solving the coefficient vector that satisfies that the number of non-zero coefficients does not exceed the sparsity threshold, the sparse coefficients corresponding one-to-one with the dictionary of captured prototype point sets are obtained. In the cloud-based collaborative platform, based on the sparsity coefficient, capture prototype 3D point sets with sparsity coefficients greater than zero are selected from the capture prototype point set dictionary, and weighted superposition is performed using the sparsity coefficient as the weight to generate candidate capture 3D point sets.

8. The cloud-based collaborative intelligent detection and capture system for unmanned aerial vehicles (UAVs) according to claim 2, characterized in that, The generation of the capture feasible 3D point set includes: The cloud-based collaborative platform reads the dynamic parameters of the capture drone and the motion parameters of the capture mechanism, and establishes a capture feasible domain model based on the dynamic constraints and capture mechanism constraints. For each 3D representative point in the candidate capture 3D point set, the cloud-based collaborative platform calculates the spatial distance between the centroid position and the 3D representative point, the relative attitude angle, and the instantaneous velocity at capture time based on the predicted centroid position, attitude parameters, and velocity parameters of the capture UAV, and compares them with the position constraints, attitude constraints, and velocity constraints in the capture feasible domain model. The cloud-based collaborative platform filters candidate capture 3D point sets according to the comparison results, and combines the 3D representative points that meet all the constraints of the capture feasible domain model in the order of coordinate index to generate a capture feasible 3D point set.

9. A cloud-based collaborative intelligent detection and capture system for unmanned aerial vehicles (UAVs) according to claim 2, characterized in that, The capture operation includes: The cloud-based collaborative platform receives environmental sensor data and divides the target area into spatial regions based on obstacle information, wind field information, and no-fly zone information in a unified coordinate system. It assigns a risk value to each spatial unit and generates a spatial risk field. The cloud-based collaborative platform sets a capture trajectory cost function based on capturing feasible 3D point sets and spatial risk fields, and obtains the cost value by linearly combining flight time and spatial risks through weight parameters; The cloud-based collaborative platform searches for capture trajectories that satisfy the dynamic constraints and capture mechanism constraints of the capture UAV within the spanned space of the feasible 3D point set. According to the capture trajectory cost function, it selects the capture trajectory with the lowest replacement value from the search results, discretizes the capture trajectory into a time-stamped track point sequence and an attitude command sequence, generates control commands, and sends them to the target capture UAV through the wireless communication unit to execute the capture operation.

Citation Information

Patent Citations

  • Multi-unmanned aerial vehicle cooperative hunting trajectory planning method and system based on three-dimensional Voronoi diagram

    CN117724531A

  • News scene three-dimensional reconstruction and visualization method based on multi-source remote sensing data

    CN119904592A