A multi-robot active vision positioning method based on spatial geometry and target background semantic cooperation

CN122544775APending Publication Date: 2026-08-11NANJING UNIV OF AERONAUTICS & ASTRONAUTICS
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-08
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

[0003]在多无人机协同视觉目标定位任务中,采用固定编队构型策略,通过预先设定观测队形并在任务执行过程中保持不变,忽略了环境遮挡与背景干扰对观测质量的动态影响,且系统未主动调整编队的位置以获取更高质量的量测信息,导致对非合作目标的定位误差大幅上升

Benefits of technology

[0035]Beneficial Effects: This invention presents a multi-machine active visual localization method based on spatial geometry and target background semantic coordination. It constructs a multi-machine coupled extended Kalman filter active localization framework based on the maximum correlation entropy criterion, achieving robust estimation of non-cooperative target states through adaptive weighting of abnormal measurements. A two-layer active perception triggering mechanism is constructed from two dimensions: the localization metric layer and the target visibility layer. After triggering, a two-stage active observation method is designed. The optimal observation direction is determined by optimizing the cooperative localization accuracy factor and ensuring a safe flight area. Background complexity is quantified using a multi-channel spatial gradient energy field and adaptive quadtree partitioning. Finally, a joint cost function is constructed by integrating the above multi-dimensional constraints to solve for the optimal three-dimensional observation waypoint that combines target observability and angle measurement accuracy. Compared with schemes that do not employ active configuration optimization, this invention improves the localization accuracy and tracking robustness of non-cooperative targets in complex scenarios with degraded initial configurations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122544775A_ABST
    Figure CN122544775A_ABST
Patent Text Reader

Abstract

This invention discloses a multi-machine active visual localization method based on spatial geometry and target background semantic coordination, belonging to the field of localization and navigation technology. A multi-machine coupled extended Kalman filter active localization framework based on the maximum correlation entropy criterion is constructed. Robust estimation of the state of non-cooperative targets is achieved through adaptive weighting of abnormal measurements. A two-layer active perception triggering mechanism is built from two dimensions: the localization metric layer and the target visibility layer. A two-stage active observation method is designed, optimizing the cooperative localization accuracy factor while ensuring a safe flight area to determine the optimal observation direction. Background complexity is quantified using a multi-channel spatial gradient energy field and adaptive quadtree partitioning. Finally, the optimal three-dimensional observation waypoint, which combines target observability and angular measurement accuracy, is solved by comprehensively considering multi-dimensional constraints. Compared with schemes that do not employ active configuration optimization, this invention can improve the localization accuracy of non-cooperative targets in complex scenarios with degraded initial configurations.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention discloses a multi-machine active visual localization method based on spatial geometry and target background semantic coordination, belonging to the field of localization and navigation technology. Background Technology

[0002] In recent years, the continuous development of unmanned aerial vehicle (UAV) technology has led to its widespread application in surveying, mapping, rescue, and other fields. Compared to single-UAV systems, multi-UAV systems can effectively enhance the system's collaborative capabilities through information sharing and dynamic task allocation, making them a core support for autonomous intelligent technologies such as swarm search and target tracking in complex environments.

[0003] In multi-UAV collaborative visual target localization tasks, a fixed formation strategy is adopted. By pre-setting the observation formation and keeping it unchanged during mission execution, the dynamic impact of environmental occlusion and background interference on observation quality is ignored. Furthermore, the system does not actively adjust the formation position to obtain higher-quality measurement information, leading to a significant increase in localization errors for non-cooperative targets. Secondly, existing active perception strategies typically prioritize maximizing information gain, forcing observations along the normal to the major axis of the target error ellipsoid. This ignores the direct impact of background complexity in the real physical environment, causing the observers in the formation to be guided to geometrically optimal but visually poor blind spots, resulting in missed target detection and further amplifying the target localization error. Simultaneously, visual angular resolution is strictly constrained by the observation distance. When the observation distance is too close, the target occupies too large a proportion in the image, and tiny pixel-level jitter at the edge of the detection box is drastically amplified into angular measurement noise. When the observation distance is too far, limited by the camera's finite angular resolution, the target needs to perform significant spatial maneuvers to generate an effective response on the pixel array, leading to a lag in target state estimation. Summary of the Invention

[0004] To address the aforementioned issues, this invention discloses a multi-machine active visual localization method based on spatial geometry and target background semantic collaboration. While adaptively adjusting the observation formation configuration, it also considers the target's visual observability and angular measurement accuracy, effectively improving the collaborative localization performance of non-cooperative targets.

[0005] Technical solution: A multi-machine active visual localization method based on spatial geometry and target-background semantic collaboration, comprising the following steps:

[0006] Step 1: Based on the visual observation information of multiple observation UAVs on non-cooperative targets, the target state is estimated using a filter, and a preset two-layer trigger decision mechanism is used to determine whether to trigger active observation adjustment. If triggered, Step 2 is executed; otherwise, the system maintains the current observation state and continues the target state estimation for the next cycle based on the new observation information.

[0007] Step 2: Based on the triggering reason, dynamically select one of the multiple observation drones as the dispatched drone;

[0008] Step 3: Under the premise of fixing the positions of the other observation UAVs besides the scheduled UAV, with the goal of optimizing the multi-view cooperative positioning accuracy factor and considering the safe flight area constraint, a coarse-adjusted three-dimensional waypoint is calculated for the scheduled UAV.

[0009] Step 4: Obtain the image from the current view of the scheduled machine, quantify the background complexity outside the area where the non-cooperative target is located in the image, and select a target background area from the areas with clean backgrounds; based on the target background area, perform pixel-level correction on the coarse-tuned 3D waypoints to obtain fine-tuned 3D waypoints.

[0010] Step 5: Construct a distance cost function that couples near-field angle measurement jitter with far-field pixel resolution, and solve for the optimal observation distance of the scheduled machine relative to the non-cooperative target; combine the optimal observation distance, the direction of the coarse-tuned 3D waypoint, and the fine-tuning correction amount to calculate and output the final 3D observation waypoint of the scheduled machine.

[0011] Furthermore, in step 1, the filter is an extended Kalman filter based on the maximum correlation entropy criterion, which adaptively reduces the weight of outlier measurements in the following manner:

[0012] According to the The observation residual of the observation drone at time k Calculate the weights:

[0013] ,

[0014] in, The relevant entropy kernel width parameter, Let be the Jacobian matrix of the measurement function with respect to the target state;

[0015] The measurement noise covariance matrix of the observation instrument is dynamically adjusted using the aforementioned weights. ,in This is the magnification factor. The original measurement noise covariance matrix; using the adjusted... Perform Kalman filtering update.

[0016] Furthermore, in step 1, the dual-layer triggering decision mechanism includes: a metric layer triggering based on positioning uncertainty, and a visibility layer triggering based on whether the target has left the field of view.

[0017] Furthermore, the condition for triggering the metric layer is: the posterior covariance matrix of the target state estimate. Trace of position submatrix Greater than the preset positioning error tolerance threshold The condition for triggering the visibility layer is: there exists at least one observation drone whose normalized information squared... It is zero.

[0018] Furthermore, in step 2, if the triggering reason is the visibility layer triggering, then the observation UAV with a normalized innovation square of zero is directly designated as the scheduled machine; if the triggering reason is the metric layer triggering, then the normalized innovation square of each observation UAV is calculated. and will The largest observation drone was selected as the scheduled drone.

[0019] Furthermore, in step 3, optimizing the multi-view cooperative positioning accuracy factor specifically involves minimizing the cooperative positioning accuracy factor. ,in It is a matrix consisting of the unit observation vectors from each observation device to the target.

[0020] Furthermore, in step 4, the quantization of background complexity includes:

[0021] Step 4-1: Construct the position of each pixel in the image. Spatial gradient energy field :

[0022] ,

[0023] in, , These represent the horizontal and vertical gradients of the image in the RGB channels, respectively.

[0024] Step 4-2: Using the detection bounding box size of the non-cooperative target as a sliding window, calculate the position of each candidate pixel within the effective background search domain. Total gradient energy within the center window ;

[0025] Step 4-3, calculate the total gradient energy. With threshold Compare and generate a binarized background semantic matrix. ,in For a clean background, It is complex.

[0026] Furthermore, in step 4, selecting a target background region from a region with a clean background includes:

[0027] Step 4-4, for the binarized background semantic matrix The row-adaptive quadtree is decomposed, and all leaf nodes with pure semantic states are extracted to form a candidate node set: ,in The center pixel coordinates of the node. Node-scale features;

[0028] Steps 4-5: Project the prior position of the non-cooperative target onto the current image coordinate system to obtain the prior projected pixel coordinates. ;

[0029] Steps 4-6, based on the comprehensive cost function Select the optimal node from the set of candidate nodes. ,in Project the target prior pixel to the candidate node center Euclidean pixel distance, Penalty weights are applied to complex backgrounds.

[0030] Furthermore, in step 5, the distance cost function for: ,in, The physical distance between the observation vehicle and the target. To suppress the exponential repulsion term of near-field angle measurement jitter, To suppress the cost of far-field secondary attraction due to sluggish changes in target pixels, The furthest reference distance; by solving... The optimal observation distance is obtained by finding the minimum point.

[0031] Furthermore, in step 5,

[0032] Based on the optimal node Coarse adjustment of 3D waypoints Pixel-level fine-tuning is performed to obtain the finely tuned 3D waypoints. :

[0033]

[0034] in, For the rotation matrix from the camera to the world frame, , For camera focal length, The physical distance from the target to the optical center of the camera.

[0035] Beneficial Effects: This invention presents a multi-machine active visual localization method based on spatial geometry and target background semantic coordination. It constructs a multi-machine coupled extended Kalman filter active localization framework based on the maximum correlation entropy criterion, achieving robust estimation of non-cooperative target states through adaptive weighting of abnormal measurements. A two-layer active perception triggering mechanism is constructed from two dimensions: the localization metric layer and the target visibility layer. After triggering, a two-stage active observation method is designed. The optimal observation direction is determined by optimizing the cooperative localization accuracy factor and ensuring a safe flight area. Background complexity is quantified using a multi-channel spatial gradient energy field and adaptive quadtree partitioning. Finally, a joint cost function is constructed by integrating the above multi-dimensional constraints to solve for the optimal three-dimensional observation waypoint that combines target observability and angle measurement accuracy. Compared with schemes that do not employ active configuration optimization, this invention improves the localization accuracy and tracking robustness of non-cooperative targets in complex scenarios with degraded initial configurations. Attached Figure Description

[0036] Figure 1 This is a schematic diagram illustrating the principle and flow of the method of the present invention;

[0037] Figure 2 A comparison chart showing the uncertainty of non-cooperative target localization after optimization using the method of this invention and after fixed formation configuration;

[0038] Figure 3 This is a comparison chart of the axis errors for locating non-cooperative targets using the method of this invention and a fixed formation configuration. Detailed Implementation

[0039] The invention will now be further explained with reference to the accompanying drawings.

[0040] The method of this invention integrates multiple constraints such as information gain and target visualization, and adjusts the observation position of the selected active sensing observation machine from coarse to fine, which overcomes the problem of insufficient information acquisition capability of fixed formation configuration, so as to further improve the positioning accuracy of visual cooperative formation for non-cooperative targets, and at the same time enhance the adaptability of visual cooperative positioning algorithm in complex scenarios such as target occlusion.

[0041] like Figure 1 As shown, a multi-machine active visual localization method based on spatial geometry and target background semantic collaboration includes the following steps:

[0042] Step 1: Construct a multi-machine coupled extended Kalman filter framework based on the maximum correlation entropy criterion to obtain the posterior estimate of the target state in the cooperative localization system. Construct a two-layer trigger decision mechanism from two dimensions: the localization metric layer and the target visibility layer, to evaluate the global observation quality and activate the active adjustment event.

[0043] In step 1, filtering and localization are performed based on the observation information from multiple machines and the motion model of the non-cooperative target, and an event adjustment mechanism is triggered, including the following specific steps:

[0044] Step 1-1, define the state vector of the non-cooperative target cooperative localization system as the target's position and velocity in the navigation coordinate system:

[0045]

[0046] in, For the target location, Let the target velocity be denoted as . Assuming the non-cooperative target follows a constant velocity model, the discretized state equation and prediction equation are as follows:

[0047]

[0048] in, Here is the state transition matrix. Assuming the process noise follows a zero-mean Gaussian distribution, its covariance matrix is: , For prior state estimation, Let be the prior covariance matrix.

[0049] Steps 1-2: Define the cooperating unit The observation residual vector is obtained through the target detector. ,

[0050]

[0051] in, These are the measurement vectors for the visual azimuth and pitch angles. For the spatial location of the cooperating unit, Let be a nonlinear measurement function that maps from the target position and individual unit position to the line-of-sight angle. The original measurement noise covariance matrix of this vision camera is defined as follows: .

[0052] Steps 1-3 address the issue of filter contamination caused by anomalous observations in real-world environments by introducing the Maximum Correlation Entropy (MCC) criterion into measurement updates, dynamically adjusting its weight based on the confidence level of each observation. An exponential kernel function is constructed using residuals to assign lower weights to anomalous measurements.

[0053]

[0054] in, The relevant entropy kernel width parameter, Let be the Jacobian matrix of the measurement function with respect to the target state. The measurement noise variance matrix is ​​dynamically adjusted using weights.

[0055]

[0056] In the formula This is the amplification factor; the smaller the weight, the greater the amplification.

[0057] Kalman gain calculation In the posterior state and covariance update, the adjusted [value] is used. :

[0058]

[0059]

[0060]

[0061] For outlier observations, the filter automatically reduces its confidence level, minimizing their impact on the final estimation results.

[0062] Steps (1-4) establish a dual active perception triggering mechanism, comprising a visual observation layer and a state estimation metric layer, to cover degraded scenarios during collaborative perception. At the metric level, the posterior covariance matrix is ​​extracted. position submatrix Calculate the trace As a quantitative indicator of global positioning uncertainty. It reflects the degree of uncertainty in the positioning accuracy at the current moment; the larger the value, the less reliable the positioning result.

[0063] The system is valid if and only if the following conditions are met.

[0064]

[0065] in, A preset positioning error tolerance threshold is used to determine if the system's perception quality has degraded. Furthermore, at the visual observation level, when the system faces extreme conditions such as background interference, line-of-sight occlusion, or targets escaping the field of view, the front-end target detector fails, causing a break in the multi-machine collaborative observation topology, which in turn degrades the normalized innovation square to zero, i.e.:

[0066]

[0067] Combining the two-tiered judgment, the system proactively adjusts the triggering condition as defined as follows:

[0068]

[0069] Among them, observation topological fractures have a higher scheduling priority to avoid unbounded covariance divergence. Then proceed to step 2.

[0070] Step 2: Based on the different triggering reasons in Step 1, select the most suitable UAV as the scheduled machine to perform subsequent position adjustments. Calculate the normalized squared innovation of each observer and dynamically allocate the active adjustment role.

[0071] If the active adjustment is triggered by the target visibility layer, that is, at least one drone (individual) completely loses the target (its normalized innovation square degenerates to zero), then the unobservable drone (individual) contributes zero to the estimated measurement, and is directly designated as the active adjustment actuator without the need for normalized innovation square comparison.

[0072] If the proactive adjustment is triggered by global positioning uncertainty, then, based on the innovation residual in step 1... With new information covariance Calculate the normalized square of the innovation of each cooperating unit. ,

[0073]

[0074]

[0075] When the consistency filter assumption holds It follows a chi-square distribution with degrees of freedom equal to the measurement dimension. An excessively large value indicates a severe degradation in the contribution of the measurement from the current perspective to the estimated information. Therefore, choosing... The drone with the highest value is the one being scheduled. By changing its position, it is most likely that its observation quality can be improved, thus achieving the greatest overall positioning accuracy improvement with minimal adjustment cost (moving only one drone).

[0076] Step 3: With the positions of the remaining UAVs fixed, determine the optimal observation direction by optimizing the cooperative positioning accuracy factor and ensuring a safe flight area, thus determining the coarse-adjusted target waypoint. For the selected active adjustment cooperative unit... Step 3 involves selecting a coarse adjustment position based on the current observed formation topology, including the following specific steps:

[0077] Step 3-1, Define the multi-view cooperative positioning accuracy factor. This factor measures the degree of influence of the observation geometry on the target positioning accuracy.

[0078] Let the first The position of the drone in the horizontal plane is The estimated location of the target is , for the A drone is deployed to calculate the estimated position from which it points towards the target. unit observation vector :

[0079]

[0080] Stack the observation vectors of N UAVs to construct the observation direction matrix. Multi-view collaborative positioning accuracy factor Defined as The square root of the trace is expressed as follows:

[0081]

[0082] when Approaching the singularity At this point, the observation directions are approximately collinear, which mathematically explains why clustered observations have poor positioning accuracy. Therefore, the optimal geometric configuration requires minimizing... , that is, maximize .

[0083] Step 3-2, considering the safety constraints of no-fly zones for single aircraft Online search is used to find a safe observation location for the scheduled drones that minimizes C-GDOP. To avoid collision risks and increased communication complexity caused by simultaneous maneuvers of multiple drones, a single-drone sequential scheduling strategy is adopted, with the remaining drones fixed. The current location of the drone is used to optimize the scheduling of the drone. Target direction angle and Given a desired observation distance The corresponding three-dimensional candidate positions can be calculated. Let the candidate hovering point of the scheduled machine be:

[0084]

[0085] in The desired observation distance is determined by solving for the optimal distance in step 5. The position optimized based on the multi-view cooperative positioning accuracy factor may fall into unsafe areas such as no-fly zones. Therefore, it is necessary to ensure the optimal positioning configuration while keeping the observation aircraft operating within a safe area, obtaining the safe area angle in discrete angle space. ,

[0086]

[0087] in, Candidate positions The shortest Euclidean distance to the no-fly zone, This is the safety margin radius between the observation aircraft and the no-fly zone. The rest are fixed. The frame position, at an angle within a safe area. Inner traversal, for each group Construct a complete observation direction matrix containing candidate locations. Calculate the corresponding C-GDOP value. Find the safe angle that minimizes C-GDOP through discretized angle search, resulting in the optimized set of observation direction angles:

[0088]

[0089] Combined with the current use The final output is the coarsely adjusted target waypoint:

[0090]

[0091] Step 4: Quantize the image background complexity using a multi-channel spatial gradient energy field and an adaptive quadtree spatial partitioning, extract a sparse and pure semantic node set, and complete the precise selection of fine-tuned positions through a comprehensive cost function.

[0092] For the selected active adjustment cooperative unit Step 4, based on semantic background perception using gradient energy manifold and quadtree decomposition, utilizes image processing to use background complexity as a reference for observation location selection, and includes the following specific steps:

[0093] Step 4-1: Define the non-cooperative target pixel region and the background search region to clarify the calculation range and avoid interference from the target's own texture in the evaluation. Obtain the coordinates of the non-cooperative target pixels output by the visual detector at the current moment. and the width of the detection box of the target in the image. and height In the pixel domain Internally built target exclusive mask The target and its extended neighborhood are designated as invalid background regions.

[0094]

[0095] in, To avoid the target's own boundary dilation coefficient, a safety buffer is established around the detection box through dilation operations. Pixels within this area are marked and not included in the evaluation, preventing the target's own texture from contaminating the background purity assessment. The effective background search domain is... That is, in the image except outside areas This is a valid background search domain.

[0096] Step 4-2: Construct a multi-channel spatial gradient energy field to establish an index quantifying the texture complexity of each pixel location in the image. This involves processing the RGB continuous signals in the image. By applying spatial discrete differential operators and performing convolution filtering, the first-order spatial partial derivatives of the entire image in the horizontal and vertical directions are extracted, i.e., the horizontal gradient field. and vertical gradient field :

[0097]

[0098] in This represents a two-dimensional convolution operation. This represents the corresponding orthogonal direction gradient convolution kernel.

[0099] Define pixel coordinates Spatial gradient energy field A scalar superposition of gradient magnitudes for all color channels:

[0100]

[0101] in, It quantifies the degree of drastic change in the physical structure at any location in the image. A higher value indicates more dramatic changes in color and brightness near that point, resulting in richer textures and more numerous edges. For mask areas... For pixels within the target area, force their energy field to be set to infinity to ensure that the target area itself is not selected as a background candidate.

[0102] Step 4-3 considers the sliding window integral of the target projection size, upgrading from point evaluation to surface evaluation and simplifying the representation. Since a clean background needs to accommodate the projection of the entire non-cooperative target, the energy field of a single pixel cannot be relied upon alone. In the search area... Within, at the expected target detection box size Using the integration window, calculate the candidate pixels. Total cost of local window background centered Its expression is as follows:

[0103]

[0104] To reduce the communication bandwidth consumption of multi-machine collaboration, a background purity tolerance threshold is set, and the dense window cost matrix of the entire graph is mapped to a binary semantic matrix in the state space containing only pure and complex elements.

[0105]

[0106] For a clean background, It is complex. The threshold is among the factors considered. The bimodal boundary point of the full-map cost matrix can be determined by an adaptive method.

[0107] Step 4-4, Dense Binary Semantic Matrix of the Entire Image Direct communication transmission is costly, therefore, it is necessary to perform binary semantic matrix processing. Adaptive multi-scale spatial sparsity partitioning is performed to significantly compress the amount of data that needs to be processed and transmitted, and to obtain a structured set of candidate regions.

[0108] right Adaptive quadtree partitioning is performed: Initialize the quadtree with the entire graph as the root node, and recursively check the semantic state consistency within the area covered by the current node. If the semantic state of all pixels within the area is completely consistent (all 1s or all 0s), stop splitting, encapsulate it as a leaf node, and record its center pixel coordinates. and scale characteristics If semantic mixing exists within a region, it is recursively split into four sub-regions until the minimum resolution or semantic consistency is achieved. All leaf nodes with a pure state (i.e., a value of 1) are extracted as a sparse topological node set.

[0109]

[0110] in, The coordinates of the node center are This refers to the node scale (area size). Step 4 compresses the dense pixel map into a small number of large, homogeneous semantic blocks.

[0111] Steps 4-5 involve optimal node evaluation based on the comprehensive cost function. From all clean candidate nodes, the node with the best comprehensive value is selected as the target background region for fine-tuning. The target's prior 3D position is then determined using the camera projection model. Projecting onto the current image coordinate system yields the prior projected pixels. For any candidate pure node extracted from a quadtree Define the overall cost score:

[0112]

[0113] in, Project the target prior pixel to the candidate node center Euclidean pixel distance; This represents the quadtree-scale feature of the node; Penalty weights are applied to complex backgrounds.

[0114]

[0115] choose smallest node As the final preferred option.

[0116] Compared to a greedy selection strategy that only uses pixel distance, the above cost function can effectively guide the UAV to avoid the closest but less pure local areas in scenes with complex background structures, and select a larger range of pure background observation view that is more conducive to target detection.

[0117] Step 5: To address the issues of near-field jitter and far-field sluggishness in observation distance, and to integrate all the optimization results from the previous steps into a specific, executable 3D coordinate instruction, Step 5 constructs an asymmetric distance cost function that couples near-field angular jitter with far-field pixel resolution to solve for the optimal observation distance. Combining this with the coarse-tuned geometric optimal direction and the fine-tuned point pixel correction, Step 5 outputs the optimal 3D observation waypoint.

[0118] Step 5, based on the obtained semantic information, constructs computational physical distance constraints, including the following specific steps:

[0119] Step 5-1: Construct an asymmetric distance cost function and solve for the optimal observation distance. Visual angular measurement accuracy is subject to bidirectional constraints from the observation distance, exhibiting significant asymmetric degradation characteristics. When the distance is too close, the target pixel ratio is too large, and the jitter at the edge of the detection box is drastically amplified into angular noise; when the distance is too far, the target angular resolution is insufficient, and minute movements cannot be captured by the camera. Therefore, an asymmetric distance cost function coupling detection box jitter and pixel resolution is constructed. :

[0120]

[0121] in, The physical distance between the observation vehicle and the target is obtained from the prior position and the state of the observation vehicle itself. To mitigate the near-field exponential repulsion cost caused by a large target proportion leading to detection box jitter, To suppress the cost of far-field secondary attraction due to sluggish changes in target pixels, Let this be the furthest reference distance. Set the first derivative of the cost function to zero, and use Newton's method to solve for the optimal observation distance. :

[0122]

[0123] It is the point where repulsion and attraction are balanced, and it is the theoretically optimal compromise distance.

[0124] Step 5-2: Combine the output of the three-dimensional optimal waypoint of the fine adjustment point with the optimal observation angle determined in step 3. The pixel-level correction of the fine-tuning points in steps 4-5 and the optimal observation distance obtained in step 5-1. The final three-dimensional ideal observation waypoint of the scheduled machine is output. First, the base waypoint is determined by the coarse-adjusted target, and then the UAV flies to the coarse-adjusted waypoint. Then, perform pixel-level fine-tuning corrections in steps 4-5, directing the target prior projected pixels toward the optimal semantic node. Center alignment yields the finely tuned local reference waypoints (fine-tuned waypoints) from the current viewpoint:

[0125]

[0126] in, This is the output of step 3, the geometric optimization stage. It represents the optimal observation position calculated from the perspective of multi-aircraft collaborative geometric accuracy (minimizing C-GDOP) under the premise of considering flight safety. After a fine adjustment, it is updated to the current position of the UAV. The coordinates of the center pixel of the optimal semantic node selected by the adaptive quadtree and the comprehensive cost function in step 4 semantic optimization stage represent a region with a clean background and suitable for observation. It is the prior estimated position of the target in the pixel coordinate system, which is projected onto the current image plane by the camera model; The optical axis depth from the target to the camera's optical center. For the rotation matrix from the camera to the world frame, , , where is the camera focal length. Since changes in the observer's position cause background parallax, this paper does not consider the waypoint output in step 5-2 as a one-time globally optimal waypoint, but rather defines it as a local reference waypoint for fine-tuning. The observer continuously acquires images as it moves towards this waypoint, repeatedly performing background complexity assessment, clean node extraction, and pixel-level correction until visual error convergence and positioning uncertainty are reduced, ultimately achieving a balance between geometric information gain and target visual observability.

[0127] To verify the effectiveness of the proposed multi-machine active visual localization method based on spatial geometry and target background semantic collaboration, a high-fidelity robot simulation analysis was conducted. In the simulation, each observation UAV was equipped with a monocular camera with a resolution of 1280*720, and its real-time pose information could be acquired through the flight controller. A non-cooperative target UAV was set up, and the multiple observation UAVs collaborated to observe and locate the target. Figure 2 The curve represents the uncertainty convergence curve for positioning before and after using the method of this invention. Figure 3 The figures show the positioning error curves before and after using the method of this invention.

[0128] like Figure 2 As shown, after entering the active adjustment phase, although the uncertainty of target positioning fluctuates briefly, it exhibits a significant overall downward trend. Compared to the traditional fixed formation configuration, this invention significantly reduces the uncertainty of multi-machine cooperative positioning, further demonstrating that this invention can effectively improve the system's positioning reliability. Furthermore, from Figure 3As can be seen, compared with no active adjustment, the target positioning error is significantly reduced to approximately 0.22 meters after adopting the active observation and adjustment mechanism of this invention. The above experimental results fully demonstrate that this invention can effectively improve the cooperative positioning performance of multiple UAVs for non-cooperative targets, and has good engineering practice and application value.

[0129] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A multi-machine active visual localization method based on spatial geometry and target-background semantic collaboration, characterized in that, Includes the following steps: Step 1: Based on the visual observation information of multiple observation UAVs on non-cooperative targets, the target state is estimated using a filter, and a preset two-layer trigger decision mechanism is used to determine whether to trigger active observation adjustment. If triggered, Step 2 is executed; otherwise, the system maintains the current observation state and continues the target state estimation for the next cycle based on the new observation information. Step 2: Based on the triggering reason, dynamically select one of the multiple observation drones as the dispatched drone; Step 3: Under the premise of fixing the positions of the other observation UAVs besides the scheduled UAV, with the goal of optimizing the multi-view cooperative positioning accuracy factor and considering the safe flight area constraint, a coarse-adjusted three-dimensional waypoint is calculated for the scheduled UAV. Step 4: Obtain the image from the current perspective of the scheduled machine, quantify the background complexity outside the area where the non-cooperative target is located in the image, and select a target background area from the areas with clean backgrounds; Based on the target background region, the coarse-tuned 3D waypoints are corrected at the pixel level to obtain fine-tuned 3D waypoints. Step 5: Construct a distance cost function that couples near-field angle measurement jitter with far-field pixel resolution, and solve for the optimal observation distance of the scheduled machine relative to the non-cooperative target; combine the optimal observation distance, the direction of the coarse-tuned 3D waypoint, and the fine-tuning correction amount to calculate and output the final 3D observation waypoint of the scheduled machine.

2. The multi-machine active visual positioning method according to claim 1, characterized in that, In step 1, the filter is an extended Kalman filter based on the maximum correlation entropy criterion, which adaptively reduces the weight of outlier measurements in the following way: According to the The observation residual of the observation drone at time k Calculate the weights: , in, The relevant entropy kernel width parameter, Let be the Jacobian matrix of the measurement function with respect to the target state; The measurement noise covariance matrix of the observation instrument is dynamically adjusted using the aforementioned weights. ,in This is the magnification factor. The original measurement noise covariance matrix; using the adjusted... Perform Kalman filtering update.

3. The multi-machine active visual positioning method according to claim 1, characterized in that, In step 1, the dual-layer triggering decision mechanism includes: a metric layer triggering based on positioning uncertainty, and a visibility layer triggering based on whether the target has left the field of view.

4. The multi-machine active visual positioning method according to claim 3, characterized in that, The condition for triggering the metric layer is: the posterior covariance matrix of the target state estimate. Trace of position submatrix Greater than the preset positioning error tolerance threshold The condition for triggering the visibility layer is: there exists at least one observation drone whose normalized information squared... It is zero.

5. The multi-machine active visual positioning method according to claim 3 or 4, characterized in that, In step 2, if the triggering reason is the visibility layer triggering, then the observation UAV with a normalized innovation square of zero is directly designated as the scheduled machine; if the triggering reason is the metric layer triggering, then the normalized innovation square of each observation UAV is calculated. and will The largest observation drone was selected as the scheduled drone.

6. The multi-machine active visual positioning method according to claim 1, characterized in that, In step 3, optimizing the multi-view cooperative positioning accuracy factor specifically involves minimizing the cooperative positioning accuracy factor. ,in It is a matrix consisting of the unit observation vectors from each observation device to the target.

7. The multi-machine active visual positioning method according to claim 1, characterized in that, In step 4, the quantization of background complexity includes: Step 4-1: Construct the position of each pixel in the image. Spatial gradient energy field : , in, , These represent the horizontal and vertical gradients of the image in the RGB channels, respectively. Step 4-2: Using the detection bounding box size of the non-cooperative target as a sliding window, calculate the position of each candidate pixel within the effective background search domain. Total gradient energy within the center window ; Step 4-3, calculate the total gradient energy. With threshold Compare and generate a binarized background semantic matrix. ,in For a clean background, It is complex.

8. The multi-machine active vision localization method according to claim 7, characterized in that, In step 4, selecting a target background region from a region with a clean background includes: Step 4-4, for the binarized background semantic matrix The row-adaptive quadtree is decomposed, and all leaf nodes with pure semantic states are extracted to form a candidate node set: ,in The center pixel coordinates of the node. Node-scale features; Steps 4-5: Project the prior position of the non-cooperative target onto the current image coordinate system to obtain the prior projected pixel coordinates. ; Steps 4-6, based on the comprehensive cost function Select the optimal node from the set of candidate nodes. ,in Project the target prior pixel to the candidate node center Euclidean pixel distance, Penalty weights are applied to complex backgrounds.

9. The multi-machine active visual positioning method according to claim 1, characterized in that, In step 5, the distance cost function for: ,in, The physical distance between the observation vehicle and the target. To suppress the exponential repulsion term of near-field angle measurement jitter, To suppress the cost of far-field secondary attraction due to sluggish changes in target pixels, The furthest reference distance; by solving... The optimal observation distance is obtained by finding the minimum point.

10. The multi-machine active vision localization method according to claim 7 or 9, characterized in that, In step 5, Based on the optimal node Coarse adjustment of 3D waypoints Pixel-level fine-tuning is performed to obtain the finely tuned 3D waypoints. : , in, For the rotation matrix from the camera to the world frame, , For camera focal length, The physical distance from the target to the optical center of the camera.