A multi-unmanned aerial vehicle cooperative interception method based on a gimbal visual guide

By constructing a multi-UAV collaborative network and distributed perception technology, and utilizing multi-view stereo vision and data fusion algorithms, the problems of target positioning accuracy and collaborative perception in UAV interception technology were solved, achieving efficient and reliable multi-UAV collaborative interception.

CN122431399APending Publication Date: 2026-07-21CHANGZHOU YUNSHAO LOW ALTITUDE INTELLIGENT TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHANGZHOU YUNSHAO LOW ALTITUDE INTELLIGENT TECHNOLOGY CO LTD
Filing Date
2026-04-24
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Existing drone interception technologies suffer from problems such as insufficient target 3D positioning accuracy, low utilization of multi-view information, and lack of collaborative perception and decision-making capabilities among multiple drones in complex dynamic environments.

Method used

A multi-UAV collaborative network is constructed. A distributed perception network is established through adjustable gimbal cameras. Multi-view stereo vision matching is used to calculate the target's three-dimensional spatial position and velocity. The observation data is fused by combining an attention mechanism. A distributed predictive control algorithm is applied to optimize the interception path and dynamically adjust the gimbal parameters to achieve continuous tracking.

Benefits of technology

It improves the high-precision spatial positioning and stable continuous tracking capabilities of targets, enhances the robustness and efficiency of multi-UAV cooperative interception, and ensures continuous and reliable interception of highly maneuverable targets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122431399A_ABST
    Figure CN122431399A_ABST
Patent Text Reader

Abstract

The application discloses a multi-unmanned aerial vehicle cooperative interception method based on a gimbal visual guide, relates to the technical field of unmanned aerial vehicle cooperative control and intelligent visual perception, and comprises the following steps: a distributed cooperative network of multiple unmanned aerial vehicles is constructed to realize sharing of identification information, position information and resource states; a target multi-view image is acquired by using a gimbal camera of each unmanned aerial vehicle; a three-dimensional position and a velocity vector of the target are calculated through stereo visual matching and multi-frame difference calculation; multi-source observation data are weighted and fused based on an attention mechanism to obtain a target motion state; an optimal interception path is planned by using a distributed predictive control algorithm in combination with resource states and relative positions of the unmanned aerial vehicles, and an elastic formation is formed; gimbal parameters are dynamically adjusted, and a main tracking unmanned aerial vehicle is adaptively switched to realize continuous and stable tracking. The application improves target positioning accuracy and cooperative interception efficiency, and enhances system robustness and environmental adaptability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of collaborative control and intelligent visual perception technology for unmanned aerial vehicles (UAVs), and in particular to a multi-UAV collaborative interception method based on gimbal-guided vision. Background Technology

[0002] With the opening of low-altitude airspace and the rapid expansion of UAV applications, UAVs are increasingly used in logistics, inspection and monitoring, and security countermeasures. This has led to the growing problem of monitoring and intercepting unauthorized UAV flights, which has become a key technological challenge in intelligent airspace management. Against this backdrop, vision-based autonomous UAV interception methods have become a research hotspot due to their advantages, including independence from external navigation signals, strong anti-interference capabilities, and high environmental adaptability. Existing technologies mostly employ a single UAV equipped with a visual sensor to track and intercept target UAVs through target detection, state estimation, and trajectory planning. However, in complex dynamic environments, a single observation perspective is susceptible to occlusion, changes in lighting, and target maneuverability, leading to decreased target positioning accuracy and insufficient tracking stability. Furthermore, traditional vision-guided methods are mostly based on monocular or weakly coupled multi-sensor fusion mechanisms, lacking system modeling for collaborative perception and decision-making among multiple UAVs, making it difficult to achieve robust spatial target state reconstruction and collaborative path optimization. In addition, existing methods typically employ centralized control or fixed-weight fusion strategies, failing to fully consider the differences in resource status among multiple UAVs and the dynamic changes in observation reliability. This results in significant room for improvement in interception efficiency and success rate in multi-target or highly dynamic scenarios.

[0003] CN121209574B discloses a method for intercepting UAV collisions based on monocular vision. This method acquires the azimuth and angle information of the target UAV through monocular vision, and uses orthogonal projection matrices and algebraic transformations to convert nonlinear measurements into pseudolinear forms. It further combines an adaptive Kalman filter algorithm to estimate the target position, and finally achieves collision interception through maneuver control. This method improves the system's robustness and target state estimation accuracy in noisy environments to some extent. However, it relies heavily on monocular vision information, has limited spatial depth information acquisition capabilities, and struggles to achieve high-precision 3D positioning in complex environments. Furthermore, this method lacks a collaborative perception and information sharing mechanism among multiple UAVs and lacks a target state fusion process under multi-view geometric constraints, resulting in insufficient tracking continuity and interception reliability in scenarios with high-speed target maneuvering or occlusion.

[0004] CN110132060A discloses a method for intercepting unmanned aerial vehicles (UAVs) based on visual navigation. This method acquires target image information using an onboard image sensor and achieves target recognition and autonomous tracking based on a visual navigation algorithm. Upon approaching the target, a capture net is deployed to intercept it. This method improves the automation level of the UAV interception process, reduces reliance on ground control, and shows good application results in low-speed, small target interception scenarios. However, this method still relies on a single UAV to perform the interception task, lacking a multi-UAV collaborative mechanism, making it difficult to cope with highly maneuverable targets or complex airspace environments. Furthermore, its visual perception mainly relies on a single perspective, without establishing a multi-view spatial geometric constraint model, and does not consider dynamic weight allocation and fusion optimization of observation data. This results in limitations in system stability and interception accuracy under conditions of target occlusion, discontinuous observation, or sensor performance fluctuations. Summary of the Invention

[0005] The purpose of this section is to outline some aspects of the embodiments of the present invention and to briefly introduce some preferred embodiments. Some simplifications or omissions may be made in this section, as well as in the abstract and title of the present application, to avoid obscuring the purpose of this section, the abstract and title of the invention. Such simplifications or omissions shall not be used to limit the scope of the present invention.

[0006] In view of the problems of insufficient target 3D positioning accuracy, low utilization of multi-view information, and lack of collaborative perception and decision-making capabilities among multiple UAVs in existing vision-based UAV interception technologies, this invention is proposed.

[0007] Therefore, the problem to be solved by this invention is how to achieve high-precision spatial positioning, stable and continuous tracking, and efficient collaborative interception among multiple UAVs in complex dynamic environments.

[0008] To solve the above-mentioned technical problems, the present invention provides the following technical solution: This invention provides a multi-UAV collaborative interception method based on gimbal vision guidance, which includes: constructing a multi-UAV collaborative network, wherein each UAV is equipped with an adjustable gimbal camera, establishing a communication link through a preset protocol to form a distributed sensing network, and each UAV sharing its own identifier, initial position information and local resource status; Image sequences of the target are acquired using the gimbal cameras of each UAV. Based on the multi-view geometric constraints of multiple UAVs on the same target, the three-dimensional spatial position of the target relative to each UAV is calculated through stereo vision matching, and the velocity vector of the target is calculated through multi-frame difference. Selective fusion of the three-dimensional spatial position and velocity vectors of each UAV target is performed, and dynamic weight coefficients are assigned to the observation data of different UAVs through an attention mechanism to generate the fused target motion state; Based on the fused target motion state, each UAV determines its optimal interception path using a distributed predictive control algorithm according to its own resource status and relative position, and works together to form a flexible formation structure that surrounds the target. The orientation and imaging parameters of the gimbal camera are dynamically adjusted according to the real-time motion state of the target. Each UAV maintains continuous tracking of the target. At the same time, based on the relative geometric relationship between each UAV and the target in the flexible formation structure and the local resource status, the main tracking UAV is adaptively switched to maintain continuous visual guidance of the target.

[0009] Compared with existing technologies, the beneficial effects of this invention are as follows: By constructing a multi-UAV collaborative network and sharing resource status, a distributed perception foundation is laid, improving the robustness and efficiency of task collaboration; by utilizing multi-view stereo vision and multi-frame differential calculation to calculate the target's three-dimensional position and velocity, high-precision motion perception is achieved, enhancing the observation dimensionality and data reliability; based on the attention mechanism, selective fusion of observation data from each UAV effectively suppresses low-quality input, improving the noise resistance and consistency of the fused state; by applying distributed predictive control and combining resource status and formation safety constraints, the interception path and flexible formation structure are jointly optimized, balancing timeliness, energy consumption, and flight safety, achieving adaptive trajectory planning in dynamic environments; by dynamically adjusting gimbal parameters and achieving seamless switching of the main tracking UAV based on adaptation score and switching gain, the continuity of visual guidance and system fault tolerance are ensured, thereby completing the continuous and reliable interception of highly maneuverable targets. Attached Figure Description

[0010] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. Wherein: Figure 1 This is a flowchart of a multi-UAV collaborative interception method based on gimbal vision guidance. Detailed Implementation

[0011] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments.

[0012] Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without inventive effort should fall within the scope of protection of this invention.

[0013] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.

[0014] As mentioned in the background section, existing technologies generally rely on monocular vision or a single UAV to perform interception tasks, lacking multi-view geometric constraint modeling and adaptive fusion mechanisms for observation data. Furthermore, they fail to incorporate dynamic collaborative control based on UAV resource status, resulting in low interception success rates in scenarios involving high-speed target maneuvering or occlusion. To address these issues, this invention provides a multi-UAV collaborative interception method based on gimbal-guided vision.

[0015] Reference Figure 1 , Figure 1 This is a flowchart illustrating a multi-UAV cooperative interception method based on gimbal vision guidance according to an embodiment of the present invention. Figure 1 As shown, a multi-UAV cooperative interception method based on gimbal vision guidance includes: S1: Construct a multi-drone collaborative network. Each drone is equipped with an adjustable gimbal camera and establishes a communication link through a preset protocol to form a distributed sensing network. Each drone shares its own identifier, initial position information, and local resource status.

[0016] S1.1: Assign a unique local identifier to each UAV. After powering on, each UAV broadcasts a network access request frame carrying its local identifier. After receiving the network access request frame, the ground control station returns a network confirmation command to each UAV to complete the node registration.

[0017] It should be noted that the local identifier includes the drone number and the network access timestamp, which is used to distinguish the identity of each node in subsequent communication. The local identifier adopts a three-part structure of formation code + sequence number + timestamp, where the timestamp is accurate to the millisecond level to avoid identifier conflicts when multiple drones in the same formation are powered on at the same time.

[0018] S1.2: Collect the initial position information of the local machine, which consists of longitude, latitude and altitude output by the airborne GNSS module; each UAV encapsulates the initial position information and its own identifier into a position broadcast frame and broadcasts it to the other UAVs in the formation through the communication link. Each UAV receives and stores the initial position information of all nodes in the formation to form an initial topology map.

[0019] Preferably, when the GNSS signal is blocked, causing the accuracy of the initial position information to be lower than a preset threshold, each UAV switches to the position estimate output by the inertial navigation module as a supplement to the initial position information in order to maintain the integrity of the initial topology map.

[0020] S1.3: Each UAV collects its own resource status, encapsulates the resource status into resource broadcast frames at fixed intervals, broadcasts them to the other UAVs in the formation through the communication link, and summarizes the resource status of all nodes to form a resource status table.

[0021] It should be noted that the local resource status includes the current remaining battery percentage, the gimbal camera working status flag, and the communication link signal strength; the broadcast period of the resource broadcast frame is set to 500 milliseconds to achieve a balance between bandwidth usage and status timeliness; when the formation size exceeds eight drones, the broadcast period can be extended to 1000 milliseconds to reduce the probability of channel congestion.

[0022] S1.4: Based on the initial topology map and resource status table, the ground control station assigns master node and slave node roles to the formation according to the preset communication protocol.

[0023] Specifically, when the remaining power percentage of the master node is lower than the preset switching threshold, the ground control station re-searches the resource status table, upgrades the currently qualified suboptimal node to the new master node, and demotes the original master node to a slave node. The communication link structure is updated accordingly to maintain the continuous and stable operation of the distributed sensing network.

[0024] It should be noted that the drone with the highest remaining battery percentage and the best communication link signal strength in its local resource status is designated as the master node; the master node is responsible for aggregating the data from each slave node and forwarding it to the ground control station; the slave nodes receive the task instructions issued by the master node, thus forming a hierarchical distributed sensing network; the switching threshold is set when the remaining battery percentage is less than 30%, and at least two overlapping transmission windows of resource broadcast frames are retained during the switching process between the old and new master nodes to avoid data frame breaks during the switching moment.

[0025] Preferably, when the remaining power percentage of the master node falls below a preset switching threshold, the ground control station re-searches the resource status table and upgrades the currently qualified suboptimal node to a new master node, while the original master node is downgraded to a slave node. The communication link structure is updated accordingly to maintain the continuous and stable operation of the distributed sensing network. Here, the suboptimal node refers to the node with the highest remaining power percentage and the best communication link signal strength, excluding the master node, in the resource status table.

[0026] S1.5: After the distributed sensing network is established, the gimbal camera on each UAV performs an initial self-test. The self-test includes verifying the rotation range of the gimbal's pitch and yaw axes, confirming the camera image acquisition frame rate, and measuring the delay in image transmission to the communication link. The self-test results are written into the gimbal camera working status flag in the form of a gimbal status report and synchronized to the resource status table through resource broadcast frames.

[0027] Furthermore, if the gimbal camera's working status flag of a certain UAV shows an abnormality, the ground control station will mark the UAV as a visual degradation node in the resource status table. In subsequent steps, the visual degradation node will only participate in communication relay and will not participate in image acquisition tasks.

[0028] S2: Use the gimbal cameras of each UAV to collect image sequences of the target. Based on the multi-view geometric constraints of multiple UAVs on the same target, calculate the three-dimensional spatial position of the target relative to each UAV through stereo vision matching, and calculate the velocity vector of the target through multi-frame difference.

[0029] S2.1: Select drones whose gimbal camera working status flag is normal in the resource status table as valid acquisition nodes. The gimbal camera of the valid acquisition node continuously shoots the target area at a preset image acquisition frame rate and outputs the target image sequence.

[0030] In an optional embodiment, when the number of effective acquisition nodes in the formation is not less than three, the ground control station selects the three nodes with the best spatial distribution as the main acquisition nodes and the remaining nodes as backup acquisition nodes to reduce communication bandwidth usage.

[0031] Preferably, each frame in the target image sequence is accompanied by a timestamp of acquisition and the local identifier of the corresponding UAV, so as to align the acquisition time sequence of each node during subsequent multi-view matching; the image acquisition frame rate is set to thirty frames per second, and the target image sequence of each effective acquisition node is uploaded to the ground control station in real time through the communication link; the criteria for determining the optimal spatial distribution are that the baseline distance between each pair of main acquisition nodes is greater than a preset baseline threshold, wherein the preset baseline threshold is dynamically adjusted according to the estimated target distance. If the target distance from the formation center is greater than 100 meters, the preset baseline threshold is set to 10 meters.

[0032] S2.2: The ground control station performs target detection on each frame of the target image sequence to obtain the pixel coordinates and target confidence score of the target in each frame.

[0033] Furthermore, when the target confidence score is lower than the preset target detection confidence threshold of 0.75, the ground control station marks the image frame as a low-quality frame. Low-quality frames will not participate in subsequent stereo vision matching calculations to avoid introducing position calculation errors due to low-quality detection results. In the case where the target confidence score is continuously lower than the preset target detection confidence threshold due to partial occlusion of the target, the ground control station calls the target image sequence of the backup acquisition node to replace the low-quality frames in the corresponding time period to maintain the continuity of target detection.

[0034] Specifically, pixel coordinates are represented with the top-left corner of the image as the origin, the horizontal axis as the x-axis, and the vertical axis as the y-axis. The ground control station uses a pre-trained detection model based on a lightweight convolutional neural network to output the pixel coordinates of the target in each frame and the target confidence score. The main features of the detection model are as follows: It adopts a three-segment structure of backbone network, feature fusion network and detection head; The backbone network is composed of stacked depthwise separable convolutional modules. The depthwise separable convolutional modules split the standard convolution into two steps: channel-wise convolution and pointwise convolution, which reduces the amount of computation while retaining the feature extraction capability. The backbone network has four downsampling stages, and the feature map size output by each stage is one-half, one-quarter, one-eighth, and one-sixteenth of the input image size, respectively, with corresponding channel numbers of 32, 64, 128, and 256. The detection head of the detection model is constructed, which takes the fused feature map as input and outputs a target bounding box regression branch and a target confidence branch. The target bounding box regression branch outputs four regression values: the x-coordinate of the center point, the y-coordinate of the center point, the width of the bounding box, and the height of the bounding box. The target confidence branch outputs the predicted confidence value that the target exists in the current anchor box region. The detection head presets three anchor boxes with different aspect ratios at each spatial location in the fused feature map. The size of the anchor boxes is determined by performing cluster analysis on the target bounding box sizes in the training dataset. The anchor box aspect ratios are set to 1:1, 1:2, and 2:1, corresponding to the target appearing as a square, vertical rectangle, and horizontal rectangle in the image. The activation function of the detector output layer uses the sigmoid function for the center point x-coordinate and center point y-coordinate branches to constrain the center point coordinates within the corresponding grid cells. No activation function is applied to the bounding box width and bounding box height branches, and the logarithmic spatial offset relative to the anchor box size is directly output. The target confidence branch uses the sigmoid function to constrain the confidence prediction value to the range of zero to one. The anchor box size clustering analysis uses the width and height of all labeled bounding boxes in the training dataset as input, and uses the K-means clustering algorithm to divide the bounding box sizes into three categories. The width and height corresponding to the center of each cluster are taken as the three anchor box sizes to ensure the matching degree between the preset anchor boxes and the actual size distribution of the targets in the training dataset. The backbone network is composed of stacked depthwise separable convolutional modules, with the total number of parameters controlled within 1.56 million. In the depthwise separable convolution module, the kernel size for channel-wise convolution is set to 3×3, and the kernel size for pointwise convolution is set to 1×1. A batch normalization layer and a linear rectified activation function are inserted between two convolution steps to accelerate training convergence and suppress gradient vanishing. The feature fusion network employs a feature pyramid structure to fuse feature maps at different scales; The feature fusion network adopts a feature pyramid structure, which merges the feature maps output from the third and fourth downsampling stages of the backbone network through upsampling and channel concatenation operations to generate a fused feature map. The fused feature map retains the detailed information of the large-scale feature map and the semantic information of the small-scale feature map, so as to improve the adaptability of the lightweight convolutional neural network detection model to the target size changes at different distances. 4) The detection head outputs the target bounding box regression branch and the target confidence branch; 5) Post-training static quantization is used to convert the model weights from 32-bit floating-point numbers to 8-bit integers, thereby improving inference speed; A training dataset for the detection model was collected and constructed. The training dataset consists of image frames of multi-rotor UAVs, fixed-wing UAVs, and other small flying targets under different lighting conditions, background environments, and flight altitudes. The target bounding box annotations for each image frame were generated using a combination of manual and semi-automatic annotation. The target bounding box annotations were based on the top left corner of the image and recorded the x-coordinate of the center point, y-coordinate of the center point, the width of the bounding box, and the height of the bounding box. The training dataset was divided into a training subset, a validation subset, and a test subset in an 8:1:1 ratio. The total number of image frames in the training dataset is no less than 20,000. The ratio of image frames containing targets to background image frames without targets is set to 7:3 to balance the distribution of positive and negative samples and reduce the false detection rate of the lightweight convolutional neural network detection model. To address the characteristics of small target size and complex background in interception scenarios, data augmentation processing was performed on the training subset. This data augmentation included four operations: random horizontal flipping, random brightness and contrast adjustment, random cropping and scaling, and mosaic stitching enhancement. Mosaic stitching enhancement stitched four different image frames into a single training image to expand the target scale and background diversity in a single training iteration. In the data augmentation, the random brightness adjustment range was set to 0.6 to 1.4 times the original brightness value, and the random contrast adjustment range was set to 0.7 to 1.3 times the original contrast value to simulate the imaging differences of the UAV under different lighting conditions such as strong light, backlight, and cloudy skies. Training the detection model specifically includes: The training loss function consists of a weighted sum of three parts: bounding box regression loss, target confidence loss, and classification loss. The bounding box regression loss uses the complete intersection-union (CIU) loss function, which introduces a center point distance penalty term and an aspect ratio consistency penalty term on the basis of the standard CIU loss to accelerate the convergence of bounding box regression. The target confidence loss uses the binary cross-entropy loss function, which assigns positive sample weights and negative sample weights to anchor boxes containing targets and background anchor boxes, respectively. The positive sample weight is set to 1, and the negative sample weight is set to 0.5 to alleviate the impact of the imbalance between the number of positive and negative samples on training. The weighting coefficients for the three loss components are 0.5 for bounding box regression loss, 0.3 for target confidence loss, and 0.2 for classification loss. The training optimizer uses the adaptive moment estimation optimization algorithm, with an initial learning rate of 0.001 and a total training epoch of 300. The learning rate is reduced to one-tenth of its original value at the 200th and 250th epochs, respectively. During the training process, the mean average precision index is calculated on the validation subset after each round. The model weight file corresponding to the highest mean average precision index is used as the final deployment weight. The final deployment weight is loaded into the lightweight convolutional neural network detection model at the ground control station inference end. The final deployment weights are quantized into integers before being loaded onto the ground control station. This converts the model weights from 32-bit floating-point numbers to 8-bit integers. After quantization, the model size is compressed to one-quarter of the original size, and the inference speed is increased to more than twice the original inference speed to meet the latency requirements of real-time inference at the ground control station. The quantization process uses a static quantization method after training, and the quantization calibration dataset is randomly selected from the validation subset of 500 frames.

[0035] Furthermore, the ground control station inputs each frame of the target image sequence into the detection model. After forward inference, it obtains the bounding box information and confidence prediction value corresponding to each anchor box. The ground control station performs non-maximum suppression processing on the confidence prediction values ​​of all anchor boxes, filters overlapping bounding boxes, and retains the bounding box with the highest confidence prediction value as the final detection result. The confidence filtering threshold for non-maximum suppression is set to 0.5, meaning that anchor boxes with a confidence prediction value lower than 0.5 are directly discarded before non-maximum suppression. It should be noted that this threshold is used to filter low-confidence bounding boxes within the detection model, which is different from the aforementioned preset target detection confidence threshold of 0.75 used to judge low-quality frames. The intersection-union ratio (IUGR) suppression threshold for non-maximum suppression is set to 0.45, meaning that candidate bounding boxes with an IUGR greater than 0.45 with the current highest confidence bounding box are suppressed and deleted. After non-maximum suppression processing, the ground control station directly outputs the x-coordinate and y-coordinate of the center point of the bounding box in the final detection result as the pixel coordinates of the target in the current frame image. The pixel coordinates are represented with the upper left corner of the image as the origin, with the horizontal direction as the positive x-axis and the vertical direction as the positive y-axis, and the unit is pixels. The confidence prediction value is directly output as the target confidence score after being retained by non-maximum suppression. The target confidence score ranges from zero to one. The higher the value, the higher the confidence level of the target's existence in the current frame image. When multiple detection results are retained after non-maximum suppression processing in the same frame, the ground control station uses the pixel coordinates and target confidence score corresponding to the detection result with the highest target confidence score as the valid output of this frame, and discards the other detection results to avoid multiple target false detections interfering with subsequent stereo vision matching calculations.

[0036] S2.3: Based on the intrinsic and extrinsic parameter matrices of the gimbal camera of each master acquisition node, perform a projection transformation from the image coordinate system to the world coordinate system on the pixel coordinates of each effective frame in the target image sequence to obtain the observation rays of each master acquisition node on the target.

[0037] It should be noted that the intrinsic parameter matrix of the gimbal camera is obtained during factory calibration, while the extrinsic parameter matrix of the gimbal camera is calculated in real time based on the initial position information of each UAV and the current pitch and yaw angles of the gimbal. The current pitch and yaw angles of the gimbal are output by the built-in encoder of the gimbal, and the encoder data is synchronized to the ground control station through the communication link. The calculation frequency is consistent with the image acquisition frame rate to avoid timing misalignment between the extrinsic parameter matrix and the image frame.

[0038] S2.4: Based on the multi-view geometric constraints of multiple UAVs on the same target, the ground control station performs stereo vision matching on the observation rays of each main acquisition node. By minimizing the spatial distance error between each observation ray and the common intersection point, the least squares method is used to solve the three-dimensional spatial position of the target relative to each UAV. The three-dimensional spatial position is represented by three-axis coordinates in the world coordinate system, with the origin of the coordinates set to the geographical location of the ground control station.

[0039] Specifically, during the stereo vision matching process, each observation ray involved in the calculation is assigned a weighting coefficient according to the target confidence score of the corresponding master acquisition node. The weighting coefficient is positively correlated with the target confidence score to reduce the interference of observation rays with poor detection quality on the three-dimensional spatial position calculation results.

[0040] Preferably, when the number of effective observation rays participating in stereo vision matching is not less than three, the solution result of the least squares method includes an estimate of the position uncertainty, which is used as one of the input parameters for the dynamic weight coefficient allocation in S3.

[0041] S2.5: The ground control station performs multi-frame differential calculation on the three-dimensional spatial position of multiple consecutive frames. The difference between the three-dimensional spatial positions of two adjacent frames is divided by the corresponding inter-frame time interval to obtain the component velocities of the target in each coordinate axis direction. The three-axis component velocities together constitute the target's velocity vector.

[0042] Preferably, the inter-frame time interval is obtained by subtracting the acquisition timestamps of two adjacent frames to eliminate the inter-frame time non-uniformity error caused by communication link transmission jitter; in order to suppress the influence of single-frame position calculation error on velocity vector, the ground control station uses sliding window mean filtering to smooth the velocity vector. The sliding window length is set to five frames, that is, the mean of the difference results of the current frame and the previous four frames is taken as the smooth velocity vector output of the current frame.

[0043] Furthermore, when the target undergoes a sudden maneuver that causes the abrupt change in the three-dimensional spatial position of adjacent frames to exceed the preset abrupt change threshold, the ground control station skips the sliding window mean filtering and directly uses the difference result of the current frame as the velocity vector to maintain the timeliness of the response to the target maneuver; the preset abrupt change threshold is set according to the maximum acceleration estimated based on the target type.

[0044] S3: Perform selective fusion of the three-dimensional spatial position and velocity vectors of each UAV target, and assign dynamic weight coefficients to the observation data of different UAVs through an attention mechanism to generate the fused target motion state.

[0045] S3.1: The ground control station receives the three-dimensional spatial position and velocity vectors output by each valid acquisition node, performs a quality pre-assessment on the input data of each node, and obtains the observation data buffer queue.

[0046] Furthermore, the quality pre-assessment uses the target confidence score and position uncertainty estimate of each node in the corresponding frame as the initial quality indicators. After aligning the node number with the acquisition timestamp, the data is stored in the observation data buffer queue. The observation data buffer queue maintains the three-dimensional spatial position and velocity vectors of each node in the last five frames in a sliding window manner for subsequent dynamic weight coefficient calculation.

[0047] It should be noted that each record in the observation data buffer queue contains six fields: the node's local identifier, the acquisition timestamp, the three-dimensional spatial location, the velocity vector, the target confidence score, and the estimated value of the location uncertainty. The local identifier and the acquisition timestamp are used as a joint index to avoid misalignment when multiple nodes' data are stored together.

[0048] Preferably, if a node fails to write new data to the observation data buffer queue within three consecutive frames, the ground control station marks the node as a missing data node. The missing data node will not participate in the dynamic weight coefficient allocation in the current fusion cycle, so as to avoid expired data from polluting the fusion results.

[0049] S3.2: Perform image sharpness assessment on the image data of the current frame of each node in the observation data buffer queue. Use the Laplacian operator to calculate the gradient energy value of the target region of the current frame image of each node, and use the gradient energy value as the image sharpness score.

[0050] It should be noted that a higher image sharpness score indicates richer target edge details in the current frame image of the corresponding node, and a higher reliability of the corresponding observation data. The range of the target region is taken as the bounding box of the target detection output extended outward by 10% of the pixel area, so as to include the target edge texture in the gradient energy value calculation range and improve the sensitivity of the image sharpness score to the sharpness of the target contour.

[0051] Specifically, the ground control station performs normalization processing on the image sharpness score, dividing the image sharpness score of each node by the sum of the image sharpness scores of all valid acquisition nodes in the current frame to obtain the normalized image sharpness score, where the normalized image sharpness score ranges from zero to one.

[0052] S3.3: Evaluate the integrity of the target in the current frame image of each node, and use the ratio of the target bounding box area output by the detection model to the preset standard target area after distance correction as the target integrity score.

[0053] Furthermore, the target bounding box area is obtained by multiplying the bounding box width by the bounding box height. The ground control station first performs an upper limit truncation on the original target integrity score. When the original target integrity score is greater than one, one is taken as the target integrity score. When the original target integrity score is less than one, the original value is retained, and the truncated result is taken as the target integrity score.

[0054] It should be noted that when the original target integrity score is greater than one, the target imaging size exceeds the preset standard target area. This usually occurs when the target is close to the node. In this case, the target is fully presented in the image in its complete form. After being truncated to one, the target integrity score takes the maximum value, indicating that this node has the most complete visual coverage of the target.

[0055] Furthermore, the target integrity score is compared with a preset integrity threshold: when the target integrity score is greater than the preset integrity threshold, it is determined that the target is well presented in the current frame image, and the target integrity score of the corresponding node participates in the dynamic weight coefficient calculation of S3.5 according to the actual value; when the target integrity score is less than or equal to the preset integrity threshold, it is determined that the target has a large area of ​​occlusion, and the target integrity score of the corresponding node is given a penalty coefficient of 0.3 in the dynamic weight coefficient calculation.

[0056] Preferably, the preset integrity threshold is set to 0.5, meaning that occlusion detection is triggered when the area of ​​the target bounding box is less than 50% of the preset standard target area after distance correction. The preset standard target area after distance correction is obtained by calculating the current distance between the target and the corresponding node based on the three-dimensional spatial position of each node, and dynamically correcting the preset standard target area according to the inverse square law of distance. The specific calculation formula is: Preset standard target area after distance correction = Preset standard target area × (Reference standard distance / Current actual distance) 2 The reference standard distance is 50 meters, and the preset standard target area is 2500 square pixels (based on the typical imaging size of the target at the reference standard distance). This correction formula is used to eliminate the influence of the change in the distance between the target and the node on the target integrity score, ensuring that the target integrity score only reflects the degree to which the target is occluded.

[0057] Specifically, when the target integrity score of a node is lower than the preset integrity threshold for five consecutive frames, the ground control station marks the node as a continuously occluded node and compresses the upper limit of the dynamic weight coefficient of the continuously occluded node in the subsequent fusion cycle to 0.1 until the target integrity score of the node recovers to above the preset integrity threshold for three consecutive frames. Only then is the marking of the continuously occluded node removed and the normal dynamic weight coefficient calculation process resumed.

[0058] S3.4: Perform environmental interference assessment on the current frame image of each node, use the pixel grayscale variance of the background region of the image as the environmental interference score, and perform reverse normalization on the environmental interference score to obtain the normalized environmental confidence score, where the background region is defined as the remaining region in the current frame image after removing the target bounding box.

[0059] Furthermore, a higher environmental interference score indicates a more complex background texture, stronger environmental interference on the corresponding node, and lower reliability of the observation data; reverse normalization processing: normalize the environmental interference score of each node by taking the reciprocal, where the normalized environmental reliability score is negatively correlated with the environmental interference score, and the value range is from zero to one.

[0060] In an optional embodiment, when there is a known strong interference background (such as a densely built-up urban area or a water surface with strong light reflection) in the area where the formation is located, the ground control station can add a preset background interference compensation value to the environmental interference score during system initialization in order to perform prior correction on the interference level of the known scene; in the main embodiment, the preset background interference compensation value is zero by default, that is, no prior correction is performed.

[0061] S3.5: Using normalized image sharpness score, target integrity score and normalized environment credibility score as input, calculate the dynamic weight coefficients of each node through an attention mechanism.

[0062] Furthermore, the attention mechanism weights and sums the above three scores according to a preset ratio to obtain the comprehensive quality score of each node. Then, it performs softmax normalization on the comprehensive quality scores of all valid acquisition nodes and outputs the dynamic weight coefficients of each node, where the sum of the dynamic weight coefficients is 1.

[0063] Preferably, the weight ratios of image sharpness score, target integrity score, and normalized environment confidence score in the preset ratio are set to 0.4, 0.4, and 0.2, respectively. Among them, image sharpness and target integrity have a more significant impact on fusion quality, so they are given higher weights. To avoid the fusion result from over-reliance on a single observation source due to excessively high dynamic weight coefficients of a single node, the ground control station sets an upper limit constraint on the dynamic weight coefficients. Under normal circumstances, the dynamic weight coefficient of a single node does not exceed 0.6. For example, in S3.3, for a continuously occluded node, its dynamic weight coefficient upper limit will be compressed to 0.1 until normal visual conditions are restored. When the softmax output exceeds the corresponding upper limit, the excess part is redistributed to the remaining nodes according to the comprehensive quality score ratio of the remaining nodes.

[0064] S3.6: Based on the dynamic weight coefficients of each node, perform weighted fusion on the three-dimensional spatial position and velocity vector of each node in the current frame of the observation data buffer queue to generate the fused target motion state.

[0065] Specifically, weighted fusion is performed separately: the dynamic weight coefficient of each node is multiplied by the three-axis coordinate components of the corresponding node's three-dimensional spatial position and then summed to obtain the fused three-dimensional spatial position; the dynamic weight coefficient of each node is multiplied by the three-axis components of the corresponding node's velocity vector and then summed to obtain the fused velocity vector; the fused three-dimensional spatial position and the fused velocity vector together constitute the fused target motion state.

[0066] Preferably, the fused target motion state is accompanied by a fusion confidence label, which is determined by the number of effective acquisition nodes participating in this fusion and the uniformity of the distribution of the dynamic weight coefficients of each node. When the number of effective acquisition nodes participating in the fusion is less than two, or the dynamic weight coefficient of a single node reaches the upper limit constraint, the fusion confidence label is set to low confidence, and the ground control station issues a perception degradation warning to the formation. The fused target motion state and the fusion confidence label together serve as input data for the distributed predictive control algorithm in S4, which is used for subsequent optimal interception path planning by each UAV.

[0067] S4: Based on the fused target motion state, each UAV determines its optimal interception path according to its own resource status and relative position using a distributed predictive control algorithm, and works together to form a flexible formation structure that surrounds the target. S4.1: Send the fused target motion state and fused confidence label to each UAV in the formation, and calculate the relative position vector between the UAV and the target based on the fused three-dimensional spatial position and the current position of the UAV.

[0068] It should be noted that the relative position vector is represented by three-axis components in the world coordinate system; each UAV synchronously reads its current local resource status and uses the relative position vector and local resource status together as local input parameters for the distributed predictive control algorithm.

[0069] Specifically, the current position of the UAV is output by the fusion of the airborne GNSS module and the inertial navigation module, with an update frequency of no less than ten times per second to ensure the timeliness of the relative position vector. When the fusion confidence level is marked as low, each UAV will add an uncertainty mark to the calculation result of the relative position vector and expand the corresponding constraint margin in the subsequent solution of the distributed predictive control algorithm.

[0070] It should be noted that when calculating the relative position vector, each UAV uses the current frame data of the fused three-dimensional spatial position as the benchmark, and does not use the historical frame data cached locally, so as to avoid systematic deviations in the relative position vectors due to inconsistent data timing among multiple UAVs. After completing the above relative position vector calculation, each UAV reports the results to the ground control station to provide a decision-making basis for the subsequent allocation of interception responsibility sectors.

[0071] S4.2: Based on the fused velocity vector of the fused target motion state, perform target motion prediction on the local machine. Starting from the current fused three-dimensional spatial position, extrapolate according to the fused velocity vector to generate the predicted position sequence of the target in the prediction time domain.

[0072] Furthermore, to improve the adaptability of the predicted position sequence to target maneuvering, each UAV introduces a target acceleration estimate during the extrapolation process. The target acceleration estimate is obtained by dividing the difference between the velocity vectors of two adjacent frames after fusion in the observation data buffer queue by the inter-frame time interval, and an upper limit constraint on the magnitude of the target acceleration estimate is imposed. If the fusion confidence is marked as low confidence, each UAV shortens the prediction time domain from three seconds to one second to reduce the cumulative error impact of low-quality fusion data on the far nodes of the predicted position sequence.

[0073] Preferably, the prediction time domain length is set to 3 seconds, the prediction step size is consistent with the inter-frame time interval corresponding to the image acquisition frame rate, and a total of 90 prediction position nodes are generated to form a prediction position sequence; the upper limit of the amplitude is determined based on the maximum maneuvering acceleration preset according to the target type.

[0074] S4.3: Based on the remaining power percentage in the local resource status and the local current location, combined with the predicted location sequence, the ground control station assigns interception responsibility sectors to each UAV.

[0075] Furthermore, the interception responsibility sector is divided equally along the target's predicted trajectory, based on the azimuth angle between each UAV and the target. Each UAV is responsible for the interception task within the corresponding azimuth angle range. The division result of the interception responsibility sector is sent to each UAV through the communication link, and each UAV stores the boundary angle of the interception responsibility sector in its own task parameter table.

[0076] Specifically, when dividing the interception responsibility sector, priority is given to assigning drones with higher remaining battery percentages to the sector in front of the target's predicted flight path to ensure the first-mover advantage in interception; drones with less than 40% remaining battery percentages are assigned to the sector to the side or rear of the target's predicted flight path to shorten their flight distance and extend their endurance margin.

[0077] In an optional embodiment, the number of interception responsibility sectors is consistent with the number of effective acquisition nodes participating in the interception. If the number of effective acquisition nodes is 3, the azimuth range corresponding to each interception responsibility sector is 120 degrees. If the number of effective acquisition nodes exceeds 4, the ground control station will merge two adjacent interception responsibility sectors into one, which will be handled by a single unit, in order to avoid the risk of collision caused by overly dense formation. In the main embodiment, the number of effective acquisition nodes is 3 by default.

[0078] It should be noted that the intercept responsibility sector division results are sent to each UAV in real time. Each UAV then determines the desired interception point and calculates the optimal interception path according to the S4.4 process. When the target changes direction, causing a change in the predicted trajectory, the ground control station will re-execute the intercept responsibility sector division and trigger each UAV to recalculate the optimal interception path. The two steps are closely linked in time, forming a closed-loop feedback mechanism.

[0079] S4.4: Each UAV executes a distributed predictive control algorithm on its own machine, taking the target prediction node in the corresponding interception responsibility sector in the predicted position sequence as the desired interception point, and taking the current position of the UAV as the starting point, to solve the optimal interception path in the prediction time domain.

[0080] Furthermore, the boundary angle of the interception responsibility sector assigned to each UAV is determined by selecting target prediction nodes that fall within the azimuth angle range of the local UAV's interception responsibility sector from the predicted position sequence. The target prediction node that is the closest in time sequence to the predicted position sequence and located within the interception responsibility sector is determined as the local UAV's expected interception point. The expected interception point is represented by three-axis coordinates in the world coordinate system and is updated synchronously with each update of the fused target motion state to ensure the consistency between the expected interception point and the actual motion trajectory of the target.

[0081] Specifically, when no target prediction node in the predicted position sequence falls within the azimuth range of the local interception responsibility sector, each UAV reports a sector gap signal to the ground control station. The ground control station then re-executes the interception responsibility sector division process in S4.3, updates the interception responsibility sector boundary angles of each UAV, and reissues the data. Each UAV then redetermines its desired interception point based on the updated interception responsibility sector boundary angles.

[0082] Furthermore, each UAV uses its current location as the optimization starting point and the desired interception point as the optimization endpoint to construct an optimization problem of a distributed predictive control algorithm within the prediction time domain. The objective function of the optimization problem is composed of a weighted sum of flight time cost and flight energy consumption cost. The flight time cost is represented by the estimated flight time required for the UAV to reach the desired interception point within the prediction time domain. The flight energy consumption cost is represented by the sum of the thrust command magnitudes of each control step within the prediction time domain. The weighting coefficients of the flight time cost and flight energy consumption cost are dynamically determined based on the percentage of remaining battery power in the UAV's resource status. When the remaining battery power is sufficient, the weighting coefficient of the flight time cost is increased to prioritize rapid arrival. When the remaining battery power is low, the weighting coefficient of the flight energy consumption cost is increased to prioritize reducing energy consumption.

[0083] Furthermore, the dynamic adjustment of the weighting coefficients takes the remaining power percentage in the local resource status as input and outputs the corresponding weighting coefficient combination according to the piecewise linear mapping relationship, which is configured during system initialization.

[0084] In an optional embodiment, if the remaining battery percentage is higher than the first battery threshold of 60%, the weighting coefficient for the flight time cost is set to 0.7 and the weighting coefficient for the flight energy consumption cost is set to 0.3; if the remaining battery percentage is lower than the second battery threshold of 40%, the weighting coefficient for the flight time cost is adjusted to 0.4 and the weighting coefficient for the flight energy consumption cost is adjusted to 0.6; when the remaining battery percentage is between the two thresholds, the weighting coefficients are determined by linear interpolation. This is consistent with the classification standard of remaining battery percentage below 40% mentioned in S4.3.

[0085] Specifically, when constructing the optimization problem, local kinematic constraints and formation safety constraints are applied to the optimization problem. Local kinematic constraints include the local maximum flight speed constraint, maximum turning angular velocity constraint, and maximum thrust constraint. Each constraint parameter is configured during system initialization based on the parameters of each UAV model. Formation safety constraints use the predicted motion trajectory of other UAVs in the formation as the dynamic obstacle boundary. It requires that the spatial distance between the predicted position of the local UAV at each control step in the prediction time domain and the predicted position of the corresponding control step of other UAVs is greater than the minimum safe distance between UAVs. The minimum safe distance between UAVs is uniformly configured during system initialization.

[0086] Furthermore, the predicted motion trajectories of the remaining UAVs are extrapolated from the current position of each UAV broadcast through the communication link and the fused velocity vector. At the beginning of each solution cycle, each UAV receives the latest broadcast data from the other UAVs in the formation and updates the dynamic obstacle boundaries to ensure the consistency between the formation safety constraints and the real-time state of the formation.

[0087] Furthermore, the boundary angle of the interception responsibility sector is incorporated into the optimization problem as a soft constraint on the flight path. This requires that the angle by which the flight path of the aircraft deviates from the central axis of the interception responsibility sector in the prediction time domain does not exceed the preset upper limit of the flight path deviation angle, so as to avoid the flight path of the aircraft from encroaching on the interception responsibility sector of the adjacent UAV. The preset upper limit of the flight path deviation angle is set to half of the azimuth angle range corresponding to the interception responsibility sector. When the formation size is 3 aircraft, the preset upper limit of the flight path deviation angle corresponds to 60 degrees.

[0088] In its specific implementation, the key steps of the quadratic programming algorithm for the predicted location sequence include: 1) Construct a state variable vector and a control variable vector. The state variables include the UAV's three-dimensional position and three-dimensional velocity, and the control variable is the three-dimensional thrust. 2) Establish a discrete-time state transition model, which is obtained by discretizing Newton's equations of motion; 3) Construct a quadratic objective function, which includes the time cost to the desired interception point and the control energy consumption cost; 4) Set linear constraints, including the maximum speed constraint, maximum acceleration constraint, and safe distance constraint from other UAVs; 5) Within each solution cycle, the problem is initialized based on the current state, and the optimal control sequence is obtained through iterative solution.

[0089] It should be noted that the distributed predictive control algorithm adopts a rolling time-domain execution mode, that is, each UAV executes only the first control step of the optimal thrust command sequence in each solution cycle. In the next solution cycle, the latest fused target motion state and the current position of the UAV are used as inputs to continuously correct the following error of the optimal interception path to the target maneuver. The maximum number of iterations of the sequential quadratic programming algorithm is set according to the computing power constraints of the ground control station, under the premise of meeting the solution accuracy, so as to ensure that the solution time of the distributed predictive control algorithm does not exceed the duration of a single solution cycle.

[0090] In an optional embodiment, the solution cycle duration is set to 0.1, and the maximum number of iterations is set to 20. If the number of iterations reaches 20, the current optimal thrust command sequence is directly output, and the iteration continues.

[0091] Furthermore, after each solution cycle, each UAV broadcasts the coordinates of the key nodes of the optimal interception path output by itself in the current solution cycle, along with the execution control commands, to the other UAVs in the formation via a communication link. Upon receiving the broadcasts, the other UAVs update the dynamic obstacle boundaries in their respective optimization problems, forming a distributed collaborative solution update mechanism. The ground control station aggregates the coordinates of the key nodes of the optimal interception path broadcast by each UAV and monitors whether the overall path planning results of the formation meet the encirclement integrity requirements of the flexible formation structure. When the optimal interception paths of any two UAVs intersect in space, the ground control station issues a conflict resolution command to the corresponding UAV, triggering the corresponding UAV to re-execute the sequential quadratic programming algorithm and add the repulsive potential field term of the conflict path segment to the formation safety constraints until the spatial intersection of the optimal interception paths of the two UAVs is eliminated.

[0092] Preferably, the repulsive potential field term uses the reciprocal of the distance between the predicted positions of the two drones in the conflict path segment as the potential field strength. The closer the distance, the greater the potential field strength, which drives the solution results of the corresponding UAV to shift away from the conflict area. The influence range of the repulsive potential field term is limited to one control step before and after the conflict path segment, so as to avoid the repulsive potential field term interfering with the solution results of the non-conflict path segment.

[0093] Furthermore, when the optimal interception paths of all UAVs in the formation meet the requirements of the integrity of the flexible formation structure and there is no spatial intersection, the ground control station issues a path confirmation command to the formation. Each UAV then drives its flight controller to perform flight actions according to the execution control command, and enters the formation flight monitoring process.

[0094] S4.5: While each UAV flies toward the desired interception point according to the optimal interception path, the ground control station monitors the overall spatial distribution of the formation based on the real-time reports of the UAV's current position and resource status, and determines whether each UAV maintains the preset formation spacing within the boundary of the interception responsibility sector.

[0095] Specifically, when the lateral deviation of any UAV's actual position from the optimal interception path exceeds a preset deviation threshold, the ground control station issues a path correction command to the UAV. The path correction command triggers the UAV to re-execute the distributed predictive control algorithm to update the optimal interception path.

[0096] It should be noted that the preset deviation threshold is set to 2 meters, and the lateral deviation is calculated from the vertical distance from the current position of the UAV to the nearest node on the optimal interception path; the path correction command is issued at a frequency of no more than 5 times per second to avoid frequent correction commands occupying too much communication link bandwidth and affecting the real-time transmission of the fused target motion status.

[0097] It should be further explained that the interception responsibility sector is the flight responsibility area divided in S4.3 based on the azimuth angle relationship between each UAV and the target, while the minimum encirclement angle mentioned in S4.6 is the minimum azimuth angle interval requirement between each UAV in the flexible formation structure formed by the interception responsibility sector. Both together ensure the effective encirclement of the target by the formation.

[0098] S4.6: During the flight along the optimal interception path, the azimuth and pitch angles of each UAV relative to the target are adjusted in coordination, with the interception responsibility sector as a constraint, so that each UAV forms a flexible formation structure surrounding the target.

[0099] Furthermore, the flexible formation structure is defined as follows: the azimuth angle interval between each UAV and the target is not less than the preset minimum encirclement angle, and the spatial distance between each UAV is not less than the minimum safe distance between UAVs. In the flexible formation structure, the azimuth angle position of each UAV is dynamically adjusted as the motion state of the fused target is updated. When the target turns and causes the predicted position sequence to deviate, the ground control station re-divides the interception responsibility sector and issues the updated expected interception point. Each UAV then re-executes the distributed predictive control algorithm to maintain the continuous encirclement of the target by the flexible formation structure.

[0100] Furthermore, the preset minimum encirclement angle is set to 60 degrees. When the number of effective acquisition nodes in the formation is 3, the azimuth angle interval of the three drones is 120 degrees, which satisfies the preset minimum encirclement angle constraint and forms a uniform triangular encirclement. When a drone leaves the formation due to the depletion of its remaining power percentage, the ground control station redistributes the interception responsibility sector of the remaining drones to maintain the integrity of the encirclement of the flexible formation structure.

[0101] S5: Dynamically adjusts the orientation and imaging parameters of the gimbal camera according to the real-time motion state of the target, and each UAV maintains continuous tracking of the target. At the same time, based on the relative geometric relationship between each UAV and the target in the flexible formation structure and the local resource status, it adaptively switches the main tracking UAV to maintain continuous visual guidance of the target.

[0102] S5.1: Based on the flexible formation structure and the fused target motion state, calculate in real time the desired pitch angle and desired yaw angle of the local gimbal camera required to point at the target.

[0103] Specifically, the desired pitch angle of the gimbal is obtained by solving the arctangent function using the height difference and horizontal distance between the current position of the drone and the fused 3D spatial position. The desired yaw angle of the gimbal is obtained by solving the arctangent function using the azimuth angle of the current position of the drone and the fused 3D spatial position in the horizontal plane. Each drone outputs the desired pitch angle and desired yaw angle of the gimbal to the gimbal servo controller, driving the gimbal to rotate to the corresponding angle.

[0104] Furthermore, the gimbal servo controller uses a proportional-integral-derivative control law to drive the gimbal rotation. The proportional, integral, and derivative coefficients are calibrated at the factory, and the gimbal rotation response time does not exceed 0.1 seconds to ensure that the gimbal orientation and target position are updated in a timely manner.

[0105] It should be noted that the calculation of the gimbal's expected pitch angle and expected yaw angle is based on the current frame data of the fused three-dimensional spatial position. When the fusion confidence is marked as low confidence, each UAV uses the nearest predicted node in the predicted position sequence to replace the fused three-dimensional spatial position in the above angle calculation, so as to maintain the continuity of the gimbal's orientation.

[0106] S5.2: After the gimbal rotates to the desired pitch angle and yaw angle, each UAV dynamically adjusts the imaging parameters of the gimbal camera based on the magnitude of the relative position vector between the UAV and the target.

[0107] Furthermore, the magnitude of the relative position vector is the current distance between the device and the target, and the imaging parameters include the camera focal length and the image acquisition frame rate. When the current distance between the device and the target is greater than the preset far-distance threshold, the gimbal camera focal length is adjusted to the maximum focal length setting, and the image acquisition frame rate is maintained at the standard frame rate. When the current distance between the device and the target is less than the preset near-distance threshold, the gimbal camera focal length is adjusted to the minimum focal length setting, and the image acquisition frame rate is increased to a high-speed frame rate to obtain a high temporal resolution image sequence of the target in the near field.

[0108] Specifically, when the current distance between the camera and the target is between the preset near-distance threshold and the preset far-distance threshold, the focal length of the gimbal camera is determined by linear interpolation between the minimum focal length setting and the maximum focal length setting based on the current distance between the camera and the target. The image acquisition frame rate is maintained at the standard frame rate to achieve a balance between the imaging field of view and the target resolution.

[0109] Preferably, the preset long-distance threshold is set to 80 meters, the preset short-distance threshold is set to 30 meters, the maximum focal length corresponds to a focal length of 30 mm, the minimum focal length corresponds to a focal length of 8 mm, the standard frame rate is 30 frames per second, and the high-speed frame rate is 60 frames per second. The imaging parameter adjustment command is generated by each UAV and sent directly to the gimbal camera control interface. The adjustment delay does not exceed 50 milliseconds to avoid blurring of the target image or the target exceeding the field of view due to the lag of imaging parameters caused by distance changes.

[0110] S5.3: After completing the imaging parameter adjustment, each UAV performs target lock verification on the target image sequence of the current frame output by its own gimbal camera.

[0111] Furthermore, the deviation between the target's pixel coordinates in the current frame image and the image center point is used as the target locking deviation. When the target locking deviation exceeds the preset locking deviation threshold, each UAV outputs a compensation rotation command to the gimbal servo controller. The compensation angle is superimposed on the gimbal's desired pitch angle and desired yaw angle to drive the gimbal to correct its orientation toward the target pixel coordinates until the target locking deviation converges to within the preset locking deviation threshold.

[0112] It should be noted that the preset lock-on deviation threshold is set to 5% of the image width, that is, for an image with a resolution of 1920 x 1080 pixels, the preset lock-on deviation threshold corresponds to 96.0 pixels; the compensation angle is calculated by multiplying the target lock-on deviation by the camera angular resolution coefficient, which is determined based on the ratio of the field of view to the image resolution corresponding to the current focal length setting; the execution priority of the compensation rotation command is higher than the regular update commands for the gimbal's desired pitch angle and desired yaw angle, so as to ensure that the target lock-on correction action is not interrupted by the periodic angle update and to avoid oscillation of the gimbal orientation during the correction process.

[0113] S5.4: The ground control station periodically calculates the main tracking adaptation score of each UAV based on the relative geometric relationship between each UAV and the target in the flexible formation structure and the status of its own resources.

[0114] Furthermore, the main tracking adaptation score is composed of a weighted sum of three components: the line-of-sight occlusion assessment score, the current distance score between the drone and the target, and the remaining battery percentage score. The line-of-sight occlusion assessment score is determined based on the target integrity score of each drone in the current frame. The current distance score between the drone and the target is assessed based on the degree to which the current distance between the drone and the target falls within the preset optimal tracking distance range. The remaining battery percentage score is positively correlated with the remaining battery percentage in the drone's resource status.

[0115] Preferably, the preset optimal tracking distance range corresponds to the preset near distance threshold and preset far distance threshold in S5.2, which are uniformly set to 30 meters to 80 meters. When the current distance between the machine and the target falls within the preset optimal tracking distance range, the current distance score between the machine and the target is full one. When it exceeds the preset optimal tracking distance range, the current distance score between the machine and the target is linearly reduced according to the deviation. S5.5: Sort the main tracking adaptation scores of each UAV in the formation and designate the UAV with the highest main tracking adaptation score as the current main tracking UAV.

[0116] Specifically, the main tracking UAV is responsible for uploading the target image sequence output by its own gimbal camera to the ground control station via a priority transmission channel. The ground control station prioritizes the target image sequence of the main tracking UAV as the main input source for target detection and pixel coordinate extraction in S2. Other UAVs, as auxiliary tracking UAVs, upload their target image sequences via ordinary transmission channels as a supplementary source of data for the main tracking UAV.

[0117] It should be noted that the priority transmission channel and the normal transmission channel share the physical channel of the communication link and are distinguished by the quality of service priority flag. The quality of service priority flag of the data packets corresponding to the main tracking UAV is set to the highest level, and they are transmitted first when the channel is congested, so as to ensure the low-latency arrival of the target image sequence of the main tracking UAV.

[0118] S5.6: After the main tracking adaptation score calculation cycle ends, extract the local identifier of the node with the highest main tracking adaptation score in the current cycle, record it as the current optimal tracking node identifier, and compare it with the main tracking UAV local identifier recorded in the previous cycle.

[0119] Furthermore, when the current optimal tracking node identifier is inconsistent with the native identifier of the main tracking UAV, that is, the node with the highest current main tracking adaptation score is not the main tracking UAV specified in the previous cycle, and the main tracking adaptation score of the node corresponding to the current optimal tracking node identifier exceeds the preset switching gain multiple of the main tracking adaptation score of the current main tracking UAV, the main tracking UAV switching is triggered; the ground control station issues a main tracking upgrade command to the new main tracking UAV and an auxiliary tracking downgrade command to the original main tracking UAV. After receiving the main tracking upgrade command, the new main tracking UAV switches to the priority transmission channel, and after receiving the auxiliary tracking downgrade command, the original main tracking UAV switches to the normal transmission channel.

[0120] Preferably, the preset switching gain factor is set to 1.2, meaning that the main tracking adaptation score of the new candidate node must exceed 20% of the main tracking adaptation score of the current main tracking drone before switching can be triggered, in order to avoid frequent switching when the scores of the two nodes are similar, which would cause the visual guidance data source to jitter.

[0121] Furthermore, during the switching process of the main tracking UAV, there is an overlapping transmission window of no less than two target image sequences. That is, the first two target image sequences after the upgrade of the new main tracking UAV and the last two target image sequences before the downgrade of the original main tracking UAV are uploaded to the ground control station at the same time. The ground control station completes the seamless switching of the data source based on the comparison of the target confidence scores of the two target image sequences within the overlapping transmission window, so as to avoid the occurrence of target tracking frame breaks at the moment of switching.

[0122] The timing process for switching the master tracking drone is as follows: 1) The ground control station calculates and sorts the main tracking adaptation scores of each UAV; 2) Identify new primary tracking drone candidates and compare them with the current primary tracking drone; 3) When the switching conditions are met, a switching command is sent to both the old and new master tracking drones simultaneously; 4) The old and new master tracking drones process their respective commands in parallel and switch transmission channels synchronously; 5) During the overlapping transmission window, the ground control station simultaneously receives two image data streams; 6) Based on the target confidence score, the ground control station smoothly switches data sources to complete a seamless transition.

[0123] Furthermore, when the fusion confidence level remains low for more than the preset number of frames, the ground control station does not wait for the main tracking adaptation score calculation cycle to end. Instead, it immediately performs an emergency main tracking UAV switch based solely on the target integrity score of the latest frame of each UAV. The preset number of frames is set to ten, so as to quickly restore continuous visual guidance to the target when the perception quality is severely degraded.

[0124] In an optional embodiment, when the main tracking adaptation score of all UAVs in the formation is lower than the preset minimum tracking threshold, the ground control station issues a diffusion command to the formation. Each UAV appropriately increases its current distance from the target according to the optimal interception path to improve the field of view coverage of each UAV. After the target integrity score of each UAV recovers to above the preset integrity threshold, the main tracking adaptation score sorting and main tracking UAV designation process is re-executed. In the main embodiment, the diffusion command is not triggered by default, and each UAV maintains a flexible formation structure to continue to perform the interception mission.

[0125] In summary, this invention lays the foundation for distributed perception by constructing a multi-UAV collaborative network and sharing resource status, thereby improving the robustness and efficiency of task collaboration. High-precision motion perception is achieved by utilizing multi-view stereo vision and multi-frame differential calculation of the target's 3D position and velocity, enhancing the observation dimensionality and data reliability. Selective fusion of observation data from various UAVs based on an attention mechanism effectively suppresses low-quality input, improving the noise resistance and consistency of the fused state. Distributed predictive control, combined with resource status and formation safety constraints, jointly optimizes the interception path and flexible formation structure, balancing timeliness, energy consumption, and flight safety, achieving adaptive trajectory planning in dynamic environments. Dynamic adjustment of gimbal parameters and seamless switching of the main tracking UAV based on adaptation scores and switching gains ensure the continuity of visual guidance and system fault tolerance, thus enabling continuous and reliable interception of highly maneuverable targets.

[0126] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A multi-UAV cooperative interception method based on gimbal vision guidance, characterized in that, include: Construct a multi-drone collaborative network, where each drone is equipped with an adjustable gimbal camera and establishes a communication link through a preset protocol to form a distributed sensing network. Each drone shares its own identifier, initial position information, and local resource status. Image sequences of the target are acquired using the gimbal cameras of each UAV. Based on the multi-view geometric constraints of multiple UAVs on the same target, the three-dimensional spatial position of the target relative to each UAV is calculated through stereo vision matching, and the velocity vector of the target is calculated through multi-frame difference. Selective fusion is performed on the three-dimensional spatial position and velocity vector of each UAV target, and dynamic weight coefficients are assigned to the observation data of different UAVs through an attention mechanism to generate the fused target motion state; Based on the fused target motion state, each UAV, according to its own resource status and relative position, uses a distributed predictive control algorithm to determine its own optimal interception path and collaboratively forms an elastic formation structure surrounding the target. The orientation and imaging parameters of the gimbal camera are dynamically adjusted according to the real-time motion state of the target. Each UAV maintains continuous tracking of the target. At the same time, based on the relative geometric relationship between each UAV and the target in the elastic formation structure and the local resource status, the main tracking UAV is adaptively switched to maintain continuous visual guidance of the target.

2. The multi-UAV cooperative interception method based on gimbal vision guidance as described in claim 1, characterized in that, Adaptive switching of the master tracking drone, including: After each UAV completes the imaging parameter adjustment, it performs target lock verification on the target image sequence of the current frame output by its own gimbal camera; Based on the relative geometric relationship between each UAV and the target in the flexible formation structure and the status of its own resources, the ground control station periodically calculates the main tracking adaptation score of each UAV. The main tracking adaptation scores of each UAV in the formation are sorted, and the UAV with the highest main tracking adaptation score is designated as the current main tracking UAV. After the main tracking adaptation score calculation cycle ends, the local identifier of the node with the highest main tracking adaptation score in the current cycle is extracted and recorded as the current optimal tracking node identifier, and compared with the main tracking UAV local identifier recorded in the previous cycle.

3. The multi-UAV cooperative interception method based on gimbal vision guidance as described in claim 2, characterized in that, The comparison was performed with the native identifier of the main tracking drone recorded in the previous period, including: When the current optimal tracking node identifier is inconsistent with the local identifier of the main tracking drone, that is, the node with the highest current main tracking adaptation score is not the main tracking drone specified in the previous cycle, and the main tracking adaptation score of the node corresponding to the current optimal tracking node identifier exceeds the preset switching gain multiple of the main tracking adaptation score of the current main tracking drone, the main tracking drone switching is triggered.

4. The multi-UAV cooperative interception method based on gimbal vision guidance as described in claim 3, characterized in that, The switching of the main tracking drone includes: The ground control station calculates and sorts the main tracking adaptation scores of each UAV. Identify new primary tracking drone candidates and compare them with the current primary tracking drone; When the switching conditions are met, a switching command is sent to both the old and new master tracking drones simultaneously; The old and new master tracking drones process their respective commands in parallel and switch transmission channels synchronously. During the overlapping transmission window, the ground control station simultaneously receives two image data streams; Based on the target confidence score, the ground control station smoothly switches data sources, completing a seamless transition.

5. The multi-UAV cooperative interception method based on gimbal vision guidance as described in claim 2, characterized in that, The flexible formation structure includes: The fused target motion state and fused confidence label are sent to each UAV in the formation. The relative position vector between the UAV and the target is calculated based on the fused three-dimensional spatial position and the current position of the UAV. Based on the fused velocity vector of the fused target motion state, target motion prediction is performed on the local machine. Starting from the current fused three-dimensional spatial position, the target's predicted position sequence in the prediction time domain is generated by extrapolating the fused velocity vector. Based on the remaining battery percentage in the local resource status and the local current location, combined with the predicted location sequence, the ground control station assigns interception responsibility sectors to each UAV. Each UAV executes a distributed predictive control algorithm on its own machine, taking the target prediction node corresponding to the interception responsibility sector in the predicted position sequence as the desired interception point, and taking the current position of the UAV as the starting point, to solve the optimal interception path in the prediction time domain; While each UAV flies toward the desired interception point according to the optimal interception path, the ground control station monitors the overall spatial distribution of the formation based on the real-time reported current position and resource status of each UAV, and determines whether each UAV maintains the preset formation spacing within the boundary of the interception responsibility sector. During flight along the optimal interception path, the azimuth and pitch angles of each UAV relative to the target are adjusted in coordination with the interception responsibility sector as a constraint, so that each UAV forms an elastic formation structure surrounding the target.

6. The multi-UAV cooperative interception method based on gimbal vision guidance as described in claim 5, characterized in that, An optimization problem for constructing the distributed predictive control algorithm within the prediction time domain is defined, wherein the objective function of the optimization problem is a weighted sum of a flight time cost term and a flight energy consumption cost term; the flight time cost term is characterized by the estimated flight time required for the aircraft to reach the desired interception point within the prediction time domain; the flight energy consumption cost term is characterized by the sum of the thrust command magnitudes of each control step of the aircraft within the prediction time domain; and the weighting coefficients of the flight time cost term and the flight energy consumption cost term are dynamically determined based on the percentage of remaining battery power in the aircraft's resource status.

7. The multi-UAV cooperative interception method based on gimbal vision guidance as described in claim 6, characterized in that, When constructing the optimization problem, local kinematic constraints and formation safety constraints are applied to the optimization problem. The local kinematic constraints include the local maximum flight speed constraint, the maximum turning angular velocity constraint, and the maximum thrust constraint. Each constraint parameter is configured during system initialization based on the parameters of each UAV model. The formation safety constraints use the predicted motion trajectory of other UAVs in the formation as the dynamic obstacle boundary, requiring that the spatial distance between the predicted position of the local UAV at each control step in the prediction time domain and the predicted position of the corresponding control step of other UAVs is greater than the minimum safe distance between UAVs.

8. The multi-UAV cooperative interception method based on gimbal vision guidance as described in claim 5, characterized in that, The fused target motion state includes: The ground control station receives the three-dimensional spatial position and velocity vectors output by each effective acquisition node, and performs a quality pre-assessment on the input data of each node to obtain the observation data buffer queue. Image sharpness assessment is performed on the image data of the current frame of each node in the observation data buffer queue. The Laplacian operator is used to calculate the gradient energy value of the target region of the current frame image of each node, and the gradient energy value is used as the image sharpness score. The integrity of the target in the current frame image of each node is evaluated, and the ratio of the target bounding box area output by the detection model to the preset standard target area after distance correction is used as the target integrity score. An environmental interference level assessment is performed on the current frame image of each node. The pixel grayscale variance of the image background region is used as the environmental interference score, and the environmental interference score is subjected to inverse normalization to obtain a normalized environmental confidence score. Using the normalized image sharpness score, the target integrity score, and the normalized environment credibility score as inputs, the dynamic weight coefficients of each node are calculated through an attention mechanism. Based on the dynamic weight coefficients of each node, the three-dimensional spatial position and velocity vector of each node in the current frame of the observation data buffer queue are weighted and fused to generate the fused target motion state.

9. The multi-UAV cooperative interception method based on gimbal vision guidance as described in claim 8, characterized in that, The ground control station performs normalization processing on the image sharpness score, dividing the image sharpness score of each node by the sum of the image sharpness scores of all valid acquisition nodes in the current frame to obtain the normalized image sharpness score.

10. The multi-UAV cooperative interception method based on gimbal vision guidance as described in claim 8, characterized in that, The detection model comprises a backbone network, a feature fusion network, and a detection head. The backbone network is composed of stacked depthwise separable convolutional modules. The detection head takes the fused feature map as input and outputs a target bounding box regression branch and a target confidence score. The feature fusion network adopts a feature pyramid structure, which fuses the feature maps output from the third and fourth downsampling stages of the backbone network through upsampling and channel concatenation operations to generate a fused feature map.