Target aircraft formation flight relative position correction method and system based on visual detection

CN122813873APending Publication Date: 2026-09-25NANJING AOKONG EQUIPMENT TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611300460.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-26
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

把单条视觉结果直接写入编队控制状态,容易使局部误匹配或短时低质量观测传入飞行控制

Benefits of technology

1、本发明通过曝光时刻记录、观察机状态插值和被观察机状态外推的组合,使视觉图像、相机外参和两机导航状态落在同一时间基准;视觉相对位姿在转换到编队坐标系时直接使用该同步状态,减少把通信到达时刻或相邻采样周期的运动差误当成编队几何偏差,提高高速飞行和成员相对运动条件下相对位姿比较的时间一致性,降低图像观测与导航状态错时造成的校正方向偏移问题。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122813873A_ABST
    Figure CN122813873A_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of unmanned aerial vehicle, and particularly relates to a target aircraft formation flight relative pose correction method and system based on visual detection. The present application combines exposure time recording, observed aircraft state interpolation and observed aircraft state extrapolation, so that the visual image, camera external parameter and navigation state of the two aircrafts fall on the same time reference. The synchronous state is directly used when the visual relative pose is converted to the formation coordinate system, the time consistency of the relative pose comparison under the conditions of high-speed flight and member relative motion is improved, and the correction direction deviation caused by the time difference between the image observation and the navigation state is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of unmanned aerial vehicle technology, specifically relating to a method and system for relative attitude correction of target drone formation flight based on vision detection. Background Technology

[0002] Target drones are unmanned aerial platforms used for flight training, range testing, and target simulation. When multiple target drones fly in formation with predetermined relative positions and attitudes, the formation controller needs to continuously acquire the relative positions and attitudes of the members and form track, speed, or attitude control variables accordingly.

[0003] Formation members typically output their own navigation status and then exchange positions and attitudes via a communication link. The sampling times, communication arrival times, and airborne image exposure times of different members may differ. Directly subtracting the navigation status at arrival will not yield a completely accurate relative pose to the actual image recording time. Furthermore, navigation errors from each member may accumulate at different rates, causing the relative pose to gradually deviate from the actual formation relationship.

[0004] Airborne cameras can directly observe adjacent target drones, but the visual results of a single frame vary with distance, observation angle, occlusion range, and the number of visible feature points. Directly writing a single visual result into the formation control state can easily lead to local mismatches or short-term low-quality observations being transmitted to flight control. If each visual observation is processed separately between the corresponding two aircraft, it is difficult to utilize the interconnected observation relationships between multiple crew members to consistently correct the relative state of the entire formation.

[0005] To address the aforementioned problems, this invention proposes a method and system for relative pose correction of target drone formation flight based on visual detection. Summary of the Invention

[0006] To address the aforementioned problems in the prior art, the present invention aims to provide a method and system for relative attitude correction of target drone formation flight based on visual detection.

[0007] To achieve the above-mentioned objectives, the technical solution adopted by this invention is as follows: A visual detection-based method for relative pose correction in target drone formation flight includes: Receive image frames and exposure times, and synchronize the navigation states of the observer and the observed aircraft based on the exposure times to form a synchronous observation record; The image search area is determined based on the predicted position of the observed machine in the synchronous observation record. The observed target machine is detected in the search area and the correspondence between two-dimensional feature points and three-dimensional reference points is established to form a feature observation record. The visual relative pose is solved based on the feature observation records, and the observation weights are calculated by combining the synchronization time deviation, detection confidence and reprojection error to form a visual pose edge that includes visual relative pose, residual and observation weights. Collect visual pose edges within a sliding time window, construct a formation pose graph with formation members as nodes and visual pose edges as edges, and jointly solve the correction increment of each non-reference node after fixing the reference node. The correction state is determined based on the cumulative weight of nodes, graph connectivity, and correction increment boundary. The correction increment in the effective state is limited by amplitude and rate of change, and the limited correction amount is output. The limited correction amount is superimposed on the relative state used for formation control to generate the corrected formation state and output it to the formation controller.

[0008] Preferably, the navigation states of the observer and the observed are synchronized in time according to the exposure time, including: For images whose exposure time falls between two adjacent navigation state sampling times of the observer, linear interpolation is performed on position and velocity according to the time interval, and spherical linear interpolation is performed on attitude to obtain the observer state at the exposure time; for the most recent broadcast state of the observed machine, its velocity and angular velocity are extrapolated to the exposure time; the time deviation between the adopted member state and the exposure time is calculated, and written into the synchronous observation record along with the image identifier, observer identifier, observed machine identifier, exposure time, interpolated observer state, predicted observed machine state, and camera calibration parameters.

[0009] Preferably, detecting the observed target machine within the search area and establishing the correspondence between two-dimensional feature points and three-dimensional reference points includes: The target aircraft bounding box is detected within the search area, and two-dimensional feature points are extracted within the bounding box. The correspondence between the two-dimensional feature points and the three-dimensional reference points is established based on the reference point index, local descriptors, and geometric arrangement. The three-dimensional reference points are set at at least one position among the nose, wingtip, and tail tip of the target aircraft. When the number of feature points reaches a set lower limit and the spatial distribution meets the non-collinearity condition, a feature observation record is formed, which includes image identifier, observer identifier, observed aircraft identifier, two-dimensional feature point coordinates, corresponding three-dimensional reference point index, detection confidence, and number of visible points.

[0010] Preferably, the observation weights are calculated by combining the synchronization time deviation, detection confidence, and reprojection error, including: The observation weight is determined by the detection confidence, inlier sufficiency coefficient, reprojection error penalty term, and time deviation penalty term. The inlier sufficiency coefficient is determined by the ratio of the number of reprojected inliers to a set lower limit. The reprojection error penalty term decreases exponentially as the reprojection error increases, and the time deviation penalty term decreases exponentially as the time deviation increases. The higher the detection confidence and the higher the inlier sufficiency coefficient, the greater the observation weight.

[0011] Preferably, after fixing the reference node, the joint solution of the correction increments for each non-reference node includes: The translation correction increment and attitude correction increment of the reference node are fixed at zero, and the correction increments of the other nodes are the variables to be solved. The difference between the visual relative pose and the navigation relative pose of each edge is used as the edge residual. The objective function is to solve the problem by using the weighted robust sum of the edge residuals and the sum of the adjacent periodic smoothing terms. The correction records of each node and the weighted residuals after the solution are output.

[0012] Preferably, the correction state is determined based on the cumulative weight of nodes, graph connectivity, and correction increment boundary, including: When a reference node is connected to a non-reference node through an edge that reaches the lower limit of the weight, the cumulative weight of a node reaches a set value, the weighted residual after solving is lower than the weighted residual before solving, and the correction increment is within the allowable boundary, it is determined to be in a valid state. When the number of valid edges decreases but the most recent valid record is still within the retention period, it is determined to be in a retention state, and the correction increment decays periodically starting from the most recent valid correction amount. When the valid connectivity of the graph changes or the residual is continuously outside the reconstruction threshold, it is determined to be in a reconstruction state.

[0013] Preferably, the correction increment in the effective state is limited in amplitude and rate of change, including: The translation correction increment and attitude correction increment are respectively component-limited, and the change in adjacent correction cycles is limited to a preset upper limit; in the holding state, the correction is periodically decayed starting from the most recent effective correction, forming a continuous limited correction.

[0014] Preferably, it further includes: transmitting back the actual correction amount used in this cycle as the previous cycle correction increment for the smoothing term in the pose graph optimization of the next cycle.

[0015] This invention also discloses a visual detection-based target drone formation flight relative attitude correction system, comprising: The image access and timing alignment unit is used to receive image frames and exposure times, and to synchronize the navigation states of the observer and the observed machine according to the exposure times to form a synchronous observation record. The target and feature detection unit is used to determine the image search area based on the predicted position of the observed machine in the synchronous observation record, detect the observed target machine in the search area and establish the correspondence between two-dimensional feature points and three-dimensional reference points to form a feature observation record. The visual pose calculation unit is used to solve the visual relative pose based on the feature observation record, and calculate the observation weight by combining the synchronization time deviation, detection confidence and reprojection error to form a visual pose edge containing the visual relative pose, residual and observation weight. The formation pose graph correction unit is used to collect visual pose edges within a sliding time window, construct a formation pose graph with formation members as nodes and visual pose edges as edges, and jointly solve the correction increment of each non-reference node after fixing the reference node. The correction state determination unit is used to determine the correction state based on the cumulative weight of nodes, graph connectivity, and correction increment boundary. It limits the amplitude and rate of change of the correction increment in the effective state and outputs the limited correction amount. The correction application unit is used to superimpose the limited correction amount onto the relative state used for formation control, generate the corrected formation state, and output it to the formation controller.

[0016] Preferably, the formation pose graph correction unit fixes the translation correction increment and attitude correction increment of the reference node to zero, and uses the correction increment of the remaining nodes as variables to be solved; the difference between the visual relative pose and the navigation relative pose of each side is used as the side residual, and the weighted robust sum of the side residuals and the sum of the smoothing terms of the adjacent period are used as the objective function for solving; the correction application unit also feeds back the actual correction amount used in this period to the formation pose graph correction unit as the previous period correction increment of the smoothing term of the next period.

[0017] Beneficial effects 1. This invention combines exposure time recording, observer state interpolation, and observed state extrapolation to ensure that visual images, camera extrinsic parameters, and navigation states of both aircraft fall on the same time reference. When the visual relative pose is converted to the formation coordinate system, the synchronized state is used directly, reducing the misinterpretation of communication arrival time or motion errors between adjacent sampling periods as formation geometric deviations. This improves the time consistency of relative pose comparison under high-speed flight and relative motion conditions of crew members, and reduces the correction direction offset problem caused by time mismatch between image observation and navigation state.

[0018] 2. This invention obtains visual pose edges based on visible feature points and combines detection confidence, reprojection error, time deviation, and visible point ratio to form edge weights. Then, it uses a formation pose graph with fixed reference nodes to jointly solve the correction increment of each member, so that the same member is constrained by multiple adjacent observation edges. The effect of abnormal edges decreases with quality weight and robust processing, maintaining the consistency of the correction relationship of each pair of members and reducing the propagation of local errors caused by single-machine single-frame results directly covering the global state. This reduces the problem of contradictions when multiple target machines are corrected individually.

[0019] 3. This invention uses graph connectivity, cumulative node weights, residual changes, and correction boundaries to jointly determine the correction state. The effective correction increment is then superimposed on the relative state used for formation control after being limited by translation and attitude component amplitude and rate of change. When the quality deteriorates for a short time, the most recent effective quantity is maintained or decayed according to a set rule, so that the visual correction quantity enters the flight control link in a continuous and constrained manner, reducing the sudden change in control state caused by the switching of observation quality, thereby improving the continuity of formation control input and solving the problem of correction quantity jump under the condition of visual observation interval. Attached Figure Description

[0020] Figure 1 This is a flowchart of the method of the present invention; Figure 2 This is a system module diagram of the present invention. Detailed Implementation

[0021] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to specific embodiments. It should be understood that the specific embodiments described herein are merely for explaining the invention and are not intended to limit the scope of protection of the invention.

[0022] Example 1 Please see Figure 1 As shown, this embodiment provides a visual detection-based method for relative pose correction in target drone formation flight. The method includes the following steps: S1. Access images and member status and generate synchronous observation records; The system receives image frames, image exposure times, and camera calibration parameters output by the observer's camera, and simultaneously reads the navigation state sequences of the observer and candidate observed machines from member communication data. Each navigation state includes at least the member identifier, sampling time, position, velocity, attitude, and angular velocity; the camera calibration parameters include intrinsic parameters and fixed extrinsic parameters from the camera coordinate system to the observer's body coordinate system.

[0023] For images whose exposure time falls between two adjacent navigation state sampling times of the observer, linear interpolation is performed on position and velocity over time intervals, and spherical linear interpolation is performed on attitude to obtain the observer state at the exposure time. The most recent broadcast state of the observed machine is extrapolated to the exposure time based on its velocity and angular velocity; when communication data simultaneously covers both sides of the exposure time, interpolation is also used to obtain the observed machine state. The time deviation between the adopted member state and the exposure time is calculated (if there is a time deviation at both the observer and observed machine ends, the larger value is taken), and written into the synchronous observation record along with the image identifier, observer identifier, candidate observed machine identifier, exposure time, interpolated observer state, predicted observed machine state, and camera calibration parameters.

[0024] The synchronous observation record uses the exposure time as a unified index and includes image data, the state of the observer at the same time, the state of the observed device at the same time, camera calibration parameters, a set of 3D reference points matching the candidate target, and time deviation. This allows step S2 to directly obtain the image, the search range of the observed device, and the 3D reference points when reading the record. The time deviation is used in step S3 to calculate the visual pose edge weights. If the time deviation is outside the allowable range, the record enters a waiting-to-be-completed state and is re-formed after obtaining the member states covering the exposure time. This step outputs the synchronous observation record for S2 to perform target and feature point detection, and together with the coordinate transformation in S3, it reduces the relative pose deviation caused by the time difference.

[0025] S2. Detect the observed target and establish a feature observation record; Read the synchronous observation record generated in step S1. Based on the predicted state of the observed aircraft, camera extrinsic parameters, and camera projection model, project the predicted positions of candidate observed aircraft onto the image plane to obtain the search area. Detect the target aircraft outline or target box within the search area, and determine the two-dimensional feature points corresponding to the target aircraft's three-dimensional reference points within the detection area. The three-dimensional reference points are set at geometrically clear locations on the nose, wingtip, tail tip, or other parts of the fuselage that are identifiable from multiple viewpoints; each reference point has three-dimensional coordinates and a unique index in the target aircraft's coordinate system.

[0026] During detection, the target bounding box and detection confidence score are first obtained based on the target's appearance. Then, corner points, contour turning points, or feature points output by the key point detection model are extracted within the target bounding box. A correspondence between 2D and 3D points is established based on reference point indices, local descriptors, and geometric arrangement relationships. Candidate points falling outside the target area are eliminated based on the relationship between the target contour and the target area. To maintain pose solvability, feature points advancing to the next step must cover at least two non-parallel directions of the target body. When the number of points reaches a set lower limit and the spatial distribution satisfies the non-collinearity condition, a feature observation record is formed.

[0027] The detection confidence score is defined as the probability of target presence or the classification score of the bounding box output by the target detection network, used to characterize the reliability that the observed target machine is contained within the current bounding box. If the output of the detection network is an unnormalized raw score, it is mapped to [0,1] using the sigmoid function; if the output is already in probabilistic form, it is directly taken as the detection confidence score. For multi-class detection networks, the probability value of the corresponding class of the observed target machine is selected as the detection confidence score.

[0028] Each feature observation record includes image identifier, observer identifier, observed device identifier, 2D feature point coordinates, corresponding 3D reference point index, detection confidence, number of visible points, target bounding box scale, and occlusion ratio. Step S3 solves for the visual relative pose using 2D and 3D point pairs and calculates the observation weights using the detection confidence, number of visible points, and occlusion ratio. Based on this, step S2 converts the temporal consistency record from step S1 into a feature observation record that can be geometrically solved, and together with S3, converts the image target detection result into a physically meaningful relative pose.

[0029] S3. Solve for the visual relative pose and form the visual pose edge; The feature observation records output by S2 and the camera intrinsic parameters recorded by S1 are read, and the rotation and translation of the observed machine relative to the camera coordinate system are obtained by perspective multi-point pose calculation. During the solution process, the minimum point set is repeatedly extracted from the 2D and 3D point pairs to calculate candidate poses, and the 3D reference points are reprojected onto the image according to the candidate poses; points whose distance between the projected point and the measured 2D point is less than the inlier threshold are recorded as inliers. Candidate poses with a large number of inliers and small reprojection errors are selected, and then nonlinear refinement is performed using all inliers. The number of feature points with reprojection errors less than the inlier threshold under this final pose is recorded as follows: This is used for subsequent observation weight calculation; the number of reprojected points here is different from the number of visible points in S2, which respectively reflect the geometric optimization quality and the image detection quality.

[0030] The relative pose in the camera coordinate system is sequentially transformed to the formation reference coordinate system using the camera extrinsic parameters and the exposure time in S1. For the observation machine... and the observed machine The obtained visual relative pose is recorded as This includes visual relative translation and visual relative rotation. Simultaneously, based on the navigation states of the two aircraft in S1, the navigation relative pose at the same exposure moment is calculated, and a six-dimensional edge residual is formed by the difference between the visual relative pose and the navigation relative pose. The first three dimensions are the translation residual, and the last three dimensions are the rotation vector resulting from the combination of the inverse visual relative rotation transformation and the navigation relative rotation.

[0031] The observation weights are calculated according to the following relationship:

[0032] The interior point sufficiency coefficient is calculated according to the following relationship:

[0033] in, For observation machine To the observed machine Dimensionless observation weights; The detection confidence level is a value between 0 and 1; This represents the number of points within the reprojection; The interior point sufficiency coefficient; The root mean square error of reprojection is expressed in pixels. The reprojection error attenuation scale is set to 3 pixels in the example; The time deviation is in milliseconds, the same as the time deviation recorded in the synchronous observation record in S1; The time deviation decay scale is set to 50 milliseconds in the example; This means truncating the calculated value to between 0 and 1.

[0034] Higher detection confidence and interior point sufficiency coefficients result in greater weights; conversely, greater reprojection errors and time deviations lead to exponentially decreasing weights. (Identify the observation machine.) Observed machine identification , Edge residuals, reprojection errors, temporal biases, interior point counts, and observation weights are written into the visual pose edge. This visual pose edge serves as the input for S4 to construct the formation pose graph; its edge weights also allow the detection quality of S2 and the temporal quality of S1 to directly participate in the joint solution of S4.

[0035] S4. Construct the formation pose graph and solve for the node correction record; Within a sliding time window, visual pose edges output by S3 are collected, with formation members as nodes and the visual pose edges from the observer to the observed machine as directed edges. The reference machine or a member with a high cumulative observation weight specified in the formation task is selected as the reference node, and the translation correction increment and attitude correction increment of the reference node in this round are fixed to zero. The correction increments of the other nodes are used as variables to be determined.

[0036] For each slave node Pointing to node The difference between the visual relative pose and the navigation relative pose of an edge is denoted as . Let the undetermined correction increment of the two nodes be denoted as and The corrected node increment difference Approaching The translational residual is first divided by 1 meter, and the attitude angle residual is first divided by 1 degree, making the six components dimensionless. The objective function is the sum of the values ​​of each side. Multiply The sum of, plus Multiply adjacent periodic smoothing terms in the form of; This represents the actual node correction amount used in the previous cycle. Preferably, the smoothing coefficient... Take 0.1. The Huber cost is applied: when the residual norm is no greater than 1.5, the square value is taken; when it exceeds 1.5, it increases with a linear slope of 3 and is subtracted by 2.25, thus maintaining the continuity of the function.

[0037] A feasible iterative process is as follows: input the visual pose edges within the time window, the navigation status of each member at the exposure time, and the correction increment of the previous cycle; fix the correction increment of the reference node to zero; and use the Gauss-Newton method for iterative solution. In each iteration, the residuals of each edge are linearly approximated at the current estimated value. Based on the relationship between the residuals and the correction increments of each node, combined with the edge weights and Huber robustness coefficient, a corresponding system of linear equations is constructed. Solving this system of equations yields the incremental correction for each non-reference node. After updating the node pose, the edge residuals are recalculated. The process stops when the incremental correction is less than the convergence threshold or the maximum number of iterations is reached, and the correction records of each member node and the weighted residuals after the solution are output.

[0038] Each node correction record includes the node identifier, translation correction increment relative to the reference node, attitude correction increment, cumulative weight of the edges connected to that node, weighted residuals before and after the solution, and time window identifier. S5 reads these fields to determine the correction status. Since the same node can be connected to multiple visual pose edges simultaneously, S4 constrains the local measurements generated by different viewing directions onto the same set of node variables; the quality weights in S3 determine the role of each edge in the joint solution, and the two together reduce the propagation of local anomalies to the entire formation correction.

[0039] S5. Determine the correction status and generate the limited correction amount; Read the node correction records and connectivity information of the current pose graph output by S4. The system processes the data according to three states: The state is considered valid when the reference node can be connected to a non-reference node through a visual pose edge that reaches the lower weight limit, the cumulative weight of the node reaches a set value, the weighted residual after solving is lower than the weighted residual before solving, and the correction increment is within the allowable boundary; the state is considered maintained when the number of valid edges decreases in a short period of time but the most recent valid record is still within the retention period; the state is considered reconstructed when the valid connectivity of the graph changes or the residual remains outside the reconstruction threshold. The reconstruction threshold is preferably 0.5.

[0040] In the effective state, component limiting is applied to the translation and attitude correction increments, and the changes in adjacent correction cycles are restricted. In the hold state, the most recent effective correction is used as the starting point, and the correction is decayed periodically to form a continuous restricted correction. In the reconstruction state, the current correction state used by the flight control is retained, and S2 is notified to expand the search area, and S4 is notified to re-establish the reference node and visual pose edge set. Preferably, the upper limit of the translation change in a single correction cycle can be 0.20 meters, and the upper limit of the attitude angle change can be 0.20 degrees; these values ​​are tuned according to the control cycle and flight envelope.

[0041] This step outputs constrained correction values ​​with valid time, status identifier, translation component, and attitude component, which are then superimposed onto the formation control state by S6. S5 converts the graph structure quality and residual quality of S4 into a decisionable state and connects it to S6 with amplitude and rate of change constraints, so that changes in visual observation quality enter the control link as continuous correction values ​​rather than abrupt states.

[0042] S6. Generate and correct the formation state and drive the flight control; Read the constrained correction values ​​output from S5, as well as the relative positions and attitudes of the members currently used by the formation controller. For non-reference nodes, superimpose the constrained translation correction values ​​onto their position state relative to the reference node, and superimpose the constrained attitude correction values ​​onto their relative attitude state in rotational combination order. Record the correction state and effective time to form the corrected formation state. For control variables formed by the relationships between non-reference nodes, recalculate their relative positions and attitudes based on the correction states of each node relative to the reference node, maintaining consistency in the relative relationships throughout the formation pose diagram.

[0043] The formation controller calculates the formation error by comparing the corrected formation state with the predetermined formation state, and converts the formation error into track, speed, altitude, or attitude setpoints; the flight controller drives the actuators based on these setpoints. The correction application unit simultaneously sends back the actual correction amount used in the current cycle to S4 as the previous cycle's correction increment for the smoothing term in the next time window. Thus, the exposure state at S1 is converted into visual pose edges via S2 and S3; S4 converts multiple edges into nodal correction records; S5 converts the nodal records into constrained correction amounts; and S6 forms the final corrected formation state and returns the actual adopted amount to the next cycle, achieving continuous iteration.

[0044] Example 2 Please see Figure 2 As shown, this embodiment provides a target drone formation flight relative pose correction system based on vision detection, including an image access and timing alignment unit, a target and feature detection unit, a visual pose calculation unit, a formation pose map correction unit, a correction state determination unit, a correction application unit, and a member communication and recording unit.

[0045] Each unit can be deployed on a formation reference machine, an airborne processor of each member, or a ground processing device capable of receiving formation data; when deployed in a distributed manner, records are associated with member identifiers, exposure times, and time window identifiers.

[0046] The member communication and recording unit receives the navigation status of each target drone, the member relationships in the formation mission, and the predetermined formation, and saves the camera calibration parameters, the target drone's three-dimensional reference points, and the actual correction values ​​used in the previous cycle. The image acquisition and timing alignment unit acquires images with exposure times from the cameras and reads the navigation status of the corresponding time period from the member communication and recording unit, interpolating or extrapolating to form a synchronous observation record; this record, along with the images, is sent to the target and feature detection unit.

[0047] The target and feature detection unit reads image data and a set of 3D reference points from the synchronous observation record, determines the image search area based on the predicted position of the candidate target, outputs a feature observation record with a 2D-3D correspondence, and sends this record, along with the synchronous observation record, to the visual pose calculation unit. The visual pose calculation unit reads the feature observation record, the synchronous navigation state, and camera calibration parameters, generates visual pose edges with edge residuals and observation weights, and writes the visual pose edges into the current time window of the member communication and recording unit. The formation pose graph correction unit reads the visual pose edges within the current time window and the actual correction amount used in the previous cycle, fixes the reference nodes, and outputs the correction records for each member node and the pose graph connectivity information.

[0048] The correction state determination unit reads node correction records and pose graph connectivity information, and generates a state identifier and restricted correction amount based on cumulative weights, residual changes, and correction boundaries. The correction application unit receives the restricted correction amount, reads the current formation relative state and forms a corrected formation state, which is then output to the formation controller. The correction application unit also writes the actual correction amount used in this cycle back to the member communication and recording unit, so that the formation pose graph correction unit can establish a smoothing term in the next cycle. Each unit transmits data sequentially using synchronous observation records, feature observation records, visual pose edges, node correction records, and corrected formation states, and the system structure and method main chain are consistent.

[0049] The airborne implementation of the system can consist of a processor, memory, a camera interface, a navigation data interface, and a crew communication interface. The memory stores the programs used to execute S1 through S6, camera calibration parameters, 3D reference points, and correction records. When the processor executes the program, it performs timing alignment, visual inspection, pose solving, graph optimization, state determination, and correction application. The camera interface provides image and exposure trigger times, the navigation data interface provides the local state, and the crew communication interface exchanges the states and visual pose edges of other crew members.

[0050] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A method for relative attitude correction in target drone formation flight based on vision detection, characterized in that, include: Receive image frames and exposure times, and synchronize the navigation states of the observer and the observed aircraft based on the exposure times to form a synchronous observation record; The image search area is determined based on the predicted position of the observed machine in the synchronous observation record. The observed target machine is detected in the search area and the correspondence between two-dimensional feature points and three-dimensional reference points is established to form a feature observation record. The visual relative pose is solved based on the feature observation records, and the observation weights are calculated by combining the synchronization time deviation, detection confidence and reprojection error to form a visual pose edge that includes visual relative pose, residual and observation weights. Collect visual pose edges within a sliding time window, construct a formation pose graph with formation members as nodes and visual pose edges as edges, and jointly solve the correction increment of each non-reference node after fixing the reference node. The correction state is determined based on the cumulative weight of nodes, graph connectivity, and correction increment boundary. The correction increment in the effective state is limited by amplitude and rate of change, and the limited correction amount is output. The limited correction amount is superimposed on the relative state used for formation control to generate the corrected formation state and output it to the formation controller.

2. The method for relative pose correction of target drone formation flight based on vision detection according to claim 1, characterized in that, Synchronize the navigation states of the observer and the observed aircraft based on the exposure time, including: For images whose exposure time falls between two adjacent navigation state sampling times of the observer, linear interpolation is performed on position and velocity according to the time interval, and spherical linear interpolation is performed on attitude to obtain the observer state at the exposure time; for the most recent broadcast state of the observed machine, its velocity and angular velocity are extrapolated to the exposure time; the time deviation between the adopted member state and the exposure time is calculated, and written into the synchronous observation record along with the image identifier, observer identifier, observed machine identifier, exposure time, interpolated observer state, predicted observed machine state, and camera calibration parameters.

3. The method for relative pose correction of target drone formation flight based on vision detection according to claim 1, characterized in that, Detect the observed target within the search area and establish the correspondence between two-dimensional feature points and three-dimensional reference points, including: The target aircraft bounding box is detected within the search area, and two-dimensional feature points are extracted within the bounding box. The correspondence between the two-dimensional feature points and the three-dimensional reference points is established based on the reference point index, local descriptors, and geometric arrangement. The three-dimensional reference points are set at at least one position among the nose, wingtip, and tail tip of the target aircraft. When the number of feature points reaches a set lower limit and the spatial distribution meets the non-collinearity condition, a feature observation record is formed, which includes image identifier, observer identifier, observed aircraft identifier, two-dimensional feature point coordinates, corresponding three-dimensional reference point index, detection confidence, and number of visible points.

4. The method for relative pose correction of target drone formation flight based on vision detection according to claim 3, characterized in that, The observation weights are calculated by combining synchronization time deviation, detection confidence, and reprojection error, including: The observation weight is determined by the detection confidence, inlier sufficiency coefficient, reprojection error penalty term, and time deviation penalty term. The inlier sufficiency coefficient is determined by the ratio of the number of reprojected inliers to a set lower limit. The reprojection error penalty term decreases exponentially as the reprojection error increases, and the time deviation penalty term decreases exponentially as the time deviation increases. The higher the detection confidence and the higher the inlier sufficiency coefficient, the greater the observation weight.

5. The method for relative pose correction of target drone formation flight based on vision detection according to claim 1, characterized in that, After fixing the reference node, jointly solve for the correction increments of each non-reference node, including: The translation correction increment and attitude correction increment of the reference node are fixed at zero, and the correction increments of the other nodes are the variables to be solved. The difference between the visual relative pose and the navigation relative pose of each edge is used as the edge residual. The objective function is to solve the problem by using the weighted robust sum of the edge residuals and the sum of the adjacent periodic smoothing terms. The correction records of each node and the weighted residuals after the solution are output.

6. The method for relative pose correction of target drone formation flight based on vision detection according to claim 1, characterized in that, The correction state is determined based on the cumulative node weights, graph connectivity, and correction increment boundaries, including: When a reference node is connected to a non-reference node through an edge that reaches the lower limit of the weight, the cumulative weight of a node reaches a set value, the weighted residual after solving is lower than the weighted residual before solving, and the correction increment is within the allowable boundary, it is determined to be in a valid state. When the number of valid edges decreases but the most recent valid record is still within the retention period, it is determined to be in a retention state, and the correction increment decays periodically starting from the most recent valid correction amount. When the valid connectivity of the graph changes or the residual is continuously outside the reconstruction threshold, it is determined to be in a reconstruction state.

7. The method for relative pose correction of target drone formation flight based on vision detection according to claim 1, characterized in that, Limits the amplitude and rate of change of the correction increment in the effective state, including: The translation correction increment and attitude correction increment are respectively component-limited, and the change in adjacent correction cycles is limited to a preset upper limit; in the holding state, the correction is periodically decayed starting from the most recent effective correction, forming a continuous limited correction.

8. The method for relative pose correction of target drone formation flight based on vision detection according to claim 1, characterized in that, Also includes: The actual correction amount used in this cycle is fed back as the previous cycle correction increment for the smoothing term in the pose graph optimization of the next cycle.

9. A target drone formation flight relative attitude correction system based on vision detection, characterized in that, include: The image access and timing alignment unit is used to receive image frames and exposure times, and to synchronize the navigation states of the observer and the observed machine according to the exposure times to form a synchronous observation record. The target and feature detection unit is used to determine the image search area based on the predicted position of the observed machine in the synchronous observation record, detect the observed target machine in the search area and establish the correspondence between two-dimensional feature points and three-dimensional reference points to form a feature observation record. The visual pose calculation unit is used to solve the visual relative pose based on the feature observation record, and calculate the observation weight by combining the synchronization time deviation, detection confidence and reprojection error to form a visual pose edge containing the visual relative pose, residual and observation weight. The formation pose graph correction unit is used to collect visual pose edges within a sliding time window, construct a formation pose graph with formation members as nodes and visual pose edges as edges, and jointly solve the correction increment of each non-reference node after fixing the reference node. The correction state determination unit is used to determine the correction state based on the cumulative weight of nodes, graph connectivity, and correction increment boundary. It limits the amplitude and rate of change of the correction increment in the effective state and outputs the limited correction amount. The correction application unit is used to superimpose the limited correction amount onto the relative state used for formation control, generate the corrected formation state, and output it to the formation controller.

10. The target drone formation flight relative attitude correction system based on vision detection according to claim 9, characterized in that, The formation pose graph correction unit keeps the translation correction increment and attitude correction increment of the reference node fixed at zero, and uses the correction increment of the other nodes as variables to be solved. The difference between the visual relative pose and the navigation relative pose of each side is used as the side residual, and the weighted robust sum of the side residuals and the sum of the smoothing terms of the adjacent period are used as the objective function to solve the problem. The correction application unit also feeds back the actual correction amount used in this period to the formation pose graph correction unit as the previous period correction increment of the smoothing term of the next period.