A relative pose estimation method for UAV formation based only on monocular vision information
Through the monocular vision-based UAV formation relative pose estimation method, the relationship between UAV motion prediction and camera image features is utilized to solve the time correlation ambiguity and feature point correlation ambiguity problems in pose estimation in UAV formation, achieving stable pose tracking and precise formation control.
Patent Information
- Application Number
- CN202210685392.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-17
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2042-06-17
AI Technical Summary
In the existing UAV formation collaborative control, the pose estimation method based on monocular vision has problems of time correlation ambiguity and feature point correlation ambiguity, which leads to large errors in pose information and makes it difficult to achieve precise collaborative control of UAV clusters.
By determining the three-dimensional coordinates of the preset significant feature points on the drone in the body coordinate system, a camera projection model is established, and the relationship between the drone motion prediction and the camera image features is used to perform feature point detection and reprojection, calculate the three-dimensional position and posture, and update the posture in combination with the particle filter algorithm to solve the reconstruction and tracking problems.
It achieves stable UAV posture tracking results, simplifies information measurement and processing, reduces occlusion sensitivity, improves the UAV's ability to cope with target occlusion, and improves the accuracy of formation control.
Smart Images

Figure CN115170656B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of unmanned aerial vehicles (UAVs), and in particular to a method for estimating the relative pose of an UAV formation based solely on monocular vision information. Background Art
[0002] A drone generally refers to an unmanned aerial vehicle controlled by a radio remote control device and its own programmable control unit. Due to their low cost, high maneuverability, and flexible deployment, drones have been widely used in various military and civilian fields in recent years to perform surveillance, reconnaissance, and search and rescue missions. In particular, in the field of military confrontation, new combat concepts such as swarm warfare and saturation strikes, based on multi-drone collaborative control technology, are emerging.
[0003] Currently, drone swarms primarily use electromagnetic communication to transmit information such as position and pose for collaborative operations. However, in some communication-restricted conditions, vision-based drone formation coordination is often employed. In this vision-based drone formation coordination approach, a single drone platform uses visual sensors, such as optical cameras, to obtain the position and motion information of the coordinated drones for coordinated control. However, drones equipped with only a single optical camera cannot track the position and pose of other drones using stereo measurement. Therefore, the target's three-dimensional position and pose must be calculated using four or more feature points of known geometric configuration on the target drone. This leads to pose tracking facing both temporal correlation ambiguity and feature point-drone correlation ambiguity. The former is often referred to as the temporal correlation problem, and the process of resolving it is called the tracking process. The latter is often referred to as the spatial correlation problem, and the process of resolving it is called the reconstruction process.
[0004] Relative pose estimation methods can be divided into two categories based on the order of reconstruction and tracking: tracking first, then reconstruction, and reconstruction first, then tracking. Tracking first, then reconstruction, first correlates feature point trajectories in the adjacent image frame, then determines spatial correlations to resolve 3D information. Reconstruction first, then tracking, first calculates the 3D pose represented by each possible spatial correlation and then tracks the target using the temporal correlation of these 3D poses.
[0005] However, the method of tracking first and then reconstructing cannot effectively use the target's three-dimensional motion prediction to deal with occlusion problems, which can easily lead to jumps and interruptions between trajectories. The method of reconstructing first and then tracking is prone to generating ghost images, leading to false detection of trajectories. Therefore, when applied to UAV formation control, both of the above methods require filtering the measured pose information due to the large errors in the pose information. However, since the method of separating reconstruction and tracking cannot effectively track densely similar targets, even if the measured pose information is filtered, it still cannot effectively deal with occlusion problems, making it difficult for both methods to achieve precise coordinated control of UAV clusters. Summary of the Invention
[0006] In order to solve some or all of the technical problems existing in the above-mentioned prior art, the present invention provides a method for estimating the relative pose of a UAV formation based only on monocular vision information.
[0007] The technical solutions of the present invention are as follows:
[0008] A method for estimating relative pose of a UAV formation based solely on monocular vision information is provided, the method comprising:
[0009] S1, determining the three-dimensional coordinates of at least four preset significant feature points on the drone in its body coordinate system, determining the internal and external parameters of the camera carried by each drone, and establishing a camera projection model;
[0010] S2, acquire the first frame of image from the camera, perform feature point detection on the acquired image, and determine the detection cells and isolated feature point sets corresponding to the drones based on the geometric relationship between the detected feature points and each drone. The detection cells corresponding to the drones represent the set of feature points that are associated with the current drone and have no ambiguity in their association with the current drone, and the isolated feature point set represents the set of feature points that have ambiguity in their association with the drones but are not associated with any drones.
[0011] S3, determining a complete detection cell, and calculating the 3D position and posture of the drone corresponding to the complete detection cell based on the 3D coordinates of the preset significant feature points, where a complete detection cell refers to a detection cell that includes more than four feature points;
[0012] S4, acquiring the next frame of image from the camera, performing feature point detection on the acquired image, and determining the detection cells and isolated feature point sets corresponding to the drones based on the geometric relationships between the detected feature points and each drone;
[0013] For each target drone, while keeping its attitude unchanged, predict the current 3D position of the target drone based on its 3D position at the previous moment, where the target drone is a drone whose 3D position and attitude at the previous moment are known;
[0014] S5, for each target drone, according to the camera projection model, reproject the preset significant feature points of the target drone under the three-dimensional position prediction value to the camera pixel coordinate system, obtain the reprojected complete pixel cell corresponding to the three-dimensional position prediction value, calculate the minimum difference between the reprojected complete pixel cell corresponding to the three-dimensional position prediction value and all the detection cells and isolated feature point sets obtained in step S4, if the minimum difference is not greater than the preset threshold, update the posture of the target drone according to the three-dimensional position prediction value, output the three-dimensional position prediction value and posture as the tracking result of the corresponding target drone, and mark the detection cell or isolated feature point subset corresponding to the minimum difference as associated, if the minimum difference is greater than the preset threshold, the corresponding target drone is considered to be lost in tracking;
[0015] S6: Re-detect and screen the remaining unassociated detection cells and isolated feature point sets to determine whether there are any unassociated complete detection cells. If there are any unassociated complete detection cells, calculate the 3D position and attitude of the drone corresponding to all unassociated complete detection cells based on the 3D coordinates of the preset significant feature points, and output the 3D position and attitude as the tracking result of the corresponding drone.
[0016] S7, return to step S4 until the UAV formation mission is completed.
[0017] In some possible implementations, the preset significant feature points include: the wing tips, tail tips, and / or manually placed cooperation signs of the UAV.
[0018] In some possible implementations, the camera coordinate system O-XYZ is set as the reference coordinate system, and the camera coordinate system coincides with the drone body coordinate system, the camera image coordinate system is o-xy, the camera pixel coordinate system is o-uv, and the three-dimensional coordinates of point P in the j-th drone body coordinate system are
[0019] The camera projection model is:
[0020]
[0021] in, is the center of gravity O of the jth UAV j The three-dimensional coordinates in the camera coordinate system, [u P ,v P ] T is the coordinate of point P in the camera pixel coordinate system, R j is the attitude rotation matrix of the j-th UAV, represents a 3x3 matrix, f x and f y Indicates the equivalent focal length of the camera in the x and y directions, c x ,c yRepresents the coordinates of the camera principal point in the camera pixel coordinate system, and K represents the intrinsic parameter matrix composed of the camera intrinsic parameters.
[0022] In some possible implementations, calculating the three-dimensional position and posture of the drone corresponding to the complete detection cell based on the three-dimensional coordinates of the preset significant feature points includes:
[0023] Determine the three-dimensional coordinates of the drone's preset salient feature points corresponding to the complete detection cell in the aircraft coordinate system. Based on the Perspective-n-Point principle, calculate the distance from the preset salient feature points to the camera's optical center.
[0024] Calculate the three-dimensional coordinates of the preset significant feature points in the camera coordinate system according to the iterative closest point algorithm;
[0025] The three-dimensional position and attitude of the UAV are calculated based on the relative position relationship between the preset significant feature points and the center of gravity of the UAV.
[0026] In some possible implementations, predicting the three-dimensional position of the target drone at a current moment based on the three-dimensional position of the target drone at a previous moment includes:
[0027] According to the 3D position of the target UAV at the previous moment, the initial predicted value of the 3D position of the target UAV at the current moment is calculated and predicted by the dynamic model of the target UAV;
[0028] Based on the initial predicted value of the 3D position at the current moment, multiple particle prediction values are generated through particle sampling. According to the camera projection model, the preset significant feature points of the target drone under each particle prediction value are reprojected to the camera pixel coordinate system to obtain the reprojected complete pixel cell corresponding to each particle prediction value;
[0029] According to the difference between the reprojected complete pixel cell corresponding to the particle prediction value and all the detection cells and isolated feature point sets obtained in step S4, the weight corresponding to the particle prediction value is updated. According to each particle prediction value and its corresponding weight, the three-dimensional position prediction value of the target drone at the current moment is calculated.
[0030] In some possible implementations, the reprojected complete pixel cell corresponding to the predicted value of the i-th particle of the j-th target drone is set to Obtain it in the following ways:
[0031] The following formula is used to calculate the coordinates of the preset significant feature points of the j-th target drone under the i-th particle prediction value in the camera coordinate system;
[0032]
[0033] According to the coordinates of the preset salient feature points in the camera coordinate system, the coordinates of the preset salient feature points in the camera pixel coordinate system are reprojected using the following formula to obtain the reprojected pixel cells;
[0034]
[0035] in, represents the coordinates of the preset significant feature point Q of the j-th target drone under the predicted value of the i-th particle in the camera coordinate system, represents the predicted value of the i-th particle of the j-th target drone at time t-1, R j,t-1 represents the attitude rotation matrix of the j-th target UAV at time t-1, Q j represents the three-dimensional coordinates of the preset significant feature point Q of the j-th target UAV in its body coordinate system, ξ j represents the set of preset significant feature points of the j-th target drone, Represents the coordinates of the preset salient feature point Q in the camera pixel coordinate system.
[0036] In some possible implementations, the difference between the reprojected complete pixel cell and the detection cell is calculated using the following formula:
[0037]
[0038] in,
[0039]
[0040] Represents the reprojected complete pixel cell and detection of cell χ k The difference between M and the reprojected complete pixel cell The number of feature points in , Represents the reprojected complete pixel cell The mth feature point in the cell, p represents the image feature point in the detection cell, and ε represents the allowable pixel error.
[0041] In some possible implementations, the difference between the reprojected complete pixel cell and the isolated feature point set is calculated using the following formula:
[0042]
[0043] in,
[0044]
[0045] Represents the reprojected complete pixel cell The difference from the isolated feature point set ξ, Represents isolated feature point set ξ and feature point The nearest point.
[0046] In some possible implementations, if Then use the following formula to update the weight corresponding to the particle prediction value:
[0047]
[0048] like Then use the following formula to update the weight corresponding to the particle prediction value:
[0049]
[0050] Among them, χ * Represents the reprojected complete pixel cell The best matching detection cell, Represents the particle prediction value The corresponding weight, Represents the reprojected complete pixel cell The best matching subset of isolated feature points,
[0051] In some possible implementations, the following formula is used to calculate the predicted three-dimensional position of the target drone at the current moment:
[0052]
[0053] in, represents the predicted three-dimensional position of the j-th target UAV at time t, N p Represents the number of particle prediction values generated for the j-th target drone.
[0054] The main advantages of the technical solution of the present invention are as follows:
[0055] The relative pose estimation method of a UAV formation based solely on monocular vision information of the present invention utilizes the relationship between UAV motion prediction and camera image features. Only camera image information needs to be input to obtain stable UAV pose tracking results. This method can simultaneously solve reconstruction and tracking problems, simplify the information measurement and processing process of the UAV cluster, effectively eliminate phantoms, reduce the tracking sensitivity to occlusion, and significantly improve the UAV's ability to cope with target occlusion. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0057] Figure 1 This is a flow chart of a method for estimating relative pose of a UAV formation based only on monocular vision information according to an embodiment of the present invention;
[0058] Figure 2 is a schematic diagram of a coordinate system and a camera projection model according to an embodiment of the present invention;
[0059] Figure 3 A schematic diagram of a reconstruction tracking framework according to an embodiment of the present invention. DETAILED DESCRIPTION
[0060] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with specific embodiments of the present invention and corresponding drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0061] The technical solutions provided by the embodiments of the present invention are described in detail below with reference to the accompanying drawings.
[0062] See also Figure 1 An embodiment of the present invention provides a method for estimating the relative pose of a UAV formation based only on monocular vision information, the method comprising the following steps:
[0063] S1, determining the three-dimensional coordinates of at least four preset significant feature points on the drone in its body coordinate system, determining the internal and external parameters of the camera carried by each drone, and establishing a camera projection model;
[0064] S2, acquire the first frame of image from the camera, perform feature point detection on the acquired image, and determine the detection cells and isolated feature point sets corresponding to the drones based on the geometric relationship between the detected feature points and each drone. The detection cells corresponding to the drones represent the set of feature points that are associated with the current drone and have no ambiguity in their association with the current drone, and the isolated feature point set represents the set of feature points that have ambiguity in their association with the drones but are not associated with any drones.
[0065] S3, determining a complete detection cell, and calculating the 3D position and posture of the drone corresponding to the complete detection cell based on the 3D coordinates of the preset significant feature points, where a complete detection cell refers to a detection cell that includes more than four feature points;
[0066] S4, acquiring the next frame of image from the camera, performing feature point detection on the acquired image, and determining the detection cells and isolated feature point sets corresponding to the drones based on the geometric relationships between the detected feature points and each drone;
[0067] For each target drone, while keeping its attitude unchanged, predict the current 3D position of the target drone based on its 3D position at the previous moment, where the target drone is a drone whose 3D position and attitude at the previous moment are known;
[0068] S5, for each target drone, according to the camera projection model, reproject the preset significant feature points of the target drone under the three-dimensional position prediction value to the camera pixel coordinate system, obtain the reprojected complete pixel cell corresponding to the three-dimensional position prediction value, calculate the minimum difference between the reprojected complete pixel cell corresponding to the three-dimensional position prediction value and all the detection cells and isolated feature point sets obtained in step S4, if the minimum difference is not greater than the preset threshold, update the posture of the target drone according to the three-dimensional position prediction value, output the three-dimensional position prediction value and posture as the tracking result of the corresponding target drone, and mark the detection cell or isolated feature point subset corresponding to the minimum difference as associated, if the minimum difference is greater than the preset threshold, the corresponding target drone is considered to be lost in tracking;
[0069] S6: Re-detect and screen the remaining unassociated detection cells and isolated feature point sets to determine whether there are any unassociated complete detection cells. If there are any unassociated complete detection cells, calculate the 3D position and attitude of the drone corresponding to all unassociated complete detection cells based on the 3D coordinates of the preset significant feature points, and output the 3D position and attitude as the tracking result of the corresponding drone.
[0070] S7, return to step S4 until the UAV formation mission is completed.
[0071] An embodiment of the present invention provides a method for estimating the relative pose of a UAV formation based solely on monocular visual information. By utilizing the relationship between UAV motion prediction and camera image features, a stable UAV pose tracking result can be obtained by simply inputting camera image information. This method can simultaneously solve reconstruction and tracking problems, simplify the information measurement and processing process of the UAV cluster, effectively eliminate phantoms, reduce the sensitivity of tracking to occlusion, and significantly improve the UAV's ability to cope with target occlusion.
[0072] The following describes in detail the steps and principles of a method for estimating the relative pose of a UAV formation based solely on monocular vision information, provided by an embodiment of the present invention.
[0073] Step S1: Determine the three-dimensional coordinates of at least four preset significant feature points on the drone in its body coordinate system, determine the internal and external parameters of the camera carried by each drone, and establish a camera projection model.
[0074] See also Figure 2 In a vision-based drone swarm, drone tracking of aerial targets is solely about the relative position of the target and the drone. Therefore, in one embodiment of the present invention, the camera coordinate system O-XYZ is used as the reference coordinate system, and the camera coordinate system is assumed to coincide with the drone's body coordinate system. The drone's body coordinate system is rigidly connected to the drone, with its origin at the drone's center of gravity. The X-axis points in the direction of the drone's nose, the Y-axis points from the origin to the right of the drone, and the Z-axis is determined by the right-hand rule using the X- and Y-axes.
[0075] Furthermore, let us assume that the camera image coordinate system is o-xy, the camera pixel coordinate system is o-uv, and the three-dimensional coordinates of point P in the body coordinate system of the j-th UAV are Based on this setting, the camera projection model can be expressed as:
[0076]
[0077] in, is the center of gravity O of the jth UAV j The three-dimensional coordinates in the camera coordinate system, [u P ,v P ] T is the coordinate of point P in the camera pixel coordinate system, R j is the attitude rotation matrix of the j-th UAV, represents a 3x3 matrix, f x and f y Indicates the equivalent focal length of the camera in the x and y directions, c x ,c y represents the coordinates of the camera's principal point in the camera's pixel coordinate system, and K represents the intrinsic parameter matrix composed of the camera's intrinsic parameters. The intrinsic parameter matrix K is calibrated before the drone takes off.
[0078] Furthermore, in one embodiment of the present invention, the preset significant feature points include: the wing tips, tail tips, and / or artificially placed cooperation signs of the UAV.
[0079] Optionally, in order to facilitate data processing and improve data processing efficiency, four significant feature points can be preset, including two wing tip points and two tail tip points. Figure 2 , taking the jth drone as an example, Figure 2 Point A in j , point B j , point D j and point E j There are four preset significant feature points, point a j 、Point b j 、Point d j and e j Point A j , point B j , point D j and point E j Position on the image.
[0080] Step S2: Acquire the first frame of image from the camera, perform feature point detection on the acquired image, and determine the detection cells and isolated feature point sets corresponding to the drones based on the geometric relationship between the detected feature points and each drone.
[0081] Specifically, to achieve target reconstruction and tracking, in the first frame of the acquired image, multiple feature points corresponding to the drone can be obtained through image feature detection methods. Based on the geometric relationship between the detected feature points and each drone, the association relationship between the feature points and the drone, as well as whether the association relationship between the feature points and the drone has association ambiguity, can be determined. Then, based on the association relationship between the feature points and the drone, as well as whether the association relationship between the feature points and the drone has association ambiguity, the detection cells and isolated feature point sets corresponding to the drone can be determined, that is, the input image state. Among them, the detection cells corresponding to the drone represent the set of feature points associated with the current drone and without association ambiguity with the current drone, and the isolated feature point set represents the set of feature points that have association ambiguity with the drone but are not associated with any drone.
[0082] Before target tracking, due to the fuzzy correlation between feature points in different frames, it is impossible to determine whether two detection cells in different images belong to the same drone. It can only be determined that the feature points within the detection cells belong to the same drone.
[0083] Step S3: Determine a complete detection cell and calculate the three-dimensional position and posture of the drone corresponding to the complete detection cell based on the three-dimensional coordinates of the preset significant feature points. A complete detection cell refers to a detection cell including more than four feature points.
[0084] For a three-dimensional space target, at least four feature points are required to determine its position and posture. In one embodiment of the present invention, when a detection cell includes more than four feature points, the detection cell is a complete detection cell, otherwise it is an incomplete detection cell.
[0085] In one embodiment of the present invention, after determining a complete detection cell in the detection cells, calculating the three-dimensional position and attitude of the drone corresponding to the complete detection cell based on the three-dimensional coordinates of the preset significant feature points may include the following steps:
[0086] Step S31: Determine the three-dimensional coordinates of the preset salient feature point of the drone corresponding to the complete detection cell in the drone coordinate system, and calculate the distance from the preset salient feature point to the optical center of the camera based on the Perspective-n-Point principle (PnP principle);
[0087] Step S32, calculating the three-dimensional coordinates of the preset significant feature points in the camera coordinate system according to the iterative closest point algorithm (ICP algorithm);
[0088] Step S33 , calculating and determining the three-dimensional position and posture of the drone based on the relative position relationship between the preset significant feature points and the center of gravity of the drone.
[0089] Taking the j-th UAV as an example, assume that the complete detection cell of the j-th UAV is χ j , the corresponding three-dimensional coordinates of the preset significant feature point of the j-th UAV in its body coordinate system are {P j By complete detection of cell χ j and the three-dimensional coordinates of the preset significant feature points of the corresponding drone in its body coordinate system {P j}, according to the PnP principle, it can be uniquely determined that {P j}Distance from the camera optical center O to {|OP j |}, and then according to the ICP algorithm, {P j}The three-dimensional coordinates in the camera coordinate system, thereby determining the three-dimensional position X of the j-th UAV in the camera coordinate system j,0 and Posture R j,0 .
[0090] Through the above method, the three-dimensional position and posture of the drones corresponding to all complete detection cells in the camera coordinate system corresponding to the current image can be initialized. For the drones represented by incomplete detection cells and / or isolated feature point sets, their three-dimensional position and posture are still unknown at the current moment and need to be reconstructed based on subsequent processing.
[0091] Step S4: Acquire the next frame of image from the camera, perform feature point detection on the acquired image, and determine the detection cells and isolated feature point sets corresponding to the drones based on the geometric relationship between the detected feature points and each drone. For each target drone, predict the 3D position of the target drone at the current moment based on the 3D position of the target drone at the previous moment while keeping the posture unchanged. The target drone is a drone whose 3D position and posture at the previous moment are known.
[0092] Specifically, in the next frame image acquired, multiple feature points corresponding to the drone can be obtained through the image feature detection method. According to the geometric relationship between the detected feature points and each drone, the association relationship between the feature points and the drone, and whether there is association ambiguity in the association relationship between the feature points and the drone can be determined. According to the association relationship between the feature points and the drone, and whether there is association ambiguity in the association relationship between the feature points and the drone, the detection cells and isolated feature point sets corresponding to the drone, that is, the input image state, can be determined.
[0093] Furthermore, in one embodiment of the present invention, in order to improve the prediction accuracy of the three-dimensional position of the drone at the current moment and improve the accuracy of drone formation control, predicting the three-dimensional position of the target drone at the current moment based on the three-dimensional position of the target drone at the previous moment may include the following steps:
[0094] Step S41, based on the three-dimensional position of the target UAV at the previous moment, an initial predicted value of the three-dimensional position of the target UAV at the current moment is calculated and predicted by the dynamic model of the target UAV;
[0095] Step S42: Based on the initial predicted value of the three-dimensional position at the current moment, multiple particle predicted values are generated through particle sampling. According to the camera projection model, the preset significant feature points of the target drone under each particle predicted value are reprojected to the camera pixel coordinate system to obtain the reprojected complete pixel cell corresponding to each particle predicted value.
[0096] In step S43, the weight corresponding to the particle prediction value is updated according to the difference between the reprojected complete pixel cell corresponding to the particle prediction value and all the detection cells and isolated feature point sets obtained in step S4. The three-dimensional position prediction value of the target drone at the current moment is calculated based on each particle prediction value and its corresponding weight.
[0097] Specifically, taking the jth target UAV as an example, assume that the three-dimensional position of the jth target UAV at time t-1 is X j,t-1 , that is, the three-dimensional position at the previous moment is X j,t-1 , the posture of the j-th target UAV remains unchanged. According to the Bayesian inference particle filter solution, the dynamic model of the UAV can predict the initial prediction value of the three-dimensional position of the j-th target UAV at time t as X j,t|t-1 , that is, the initial predicted value of the three-dimensional position at the current moment is X j,t|t-1 ; Then, the importance resampling particle filter method is used to select a specific standard deviation σ j And X j,t|t-1 N is the mean sample p particles Set the initial weight of each particle to Get N pThe predicted values of particles and their corresponding weights. At this time, the state of the target UAV can be parallelized using N p Particle predictions can be established as Figure 3 The reconstruction tracking framework shown in Figure 3 midpoint point point and point is the Nth jth UAV p Particle predicted value state The following 4 preset significant feature points.
[0098] Furthermore, for the different particle prediction values of the target UAV, when the three-dimensional coordinates of the preset significant feature points of the target UAV in its body coordinate system are known, the reprojected complete pixel cells corresponding to the different particle prediction values of the target UAV, that is, the reprojected image state, can be calculated according to the camera projection model.
[0099] Specifically, taking the j-th target drone as an example, the reprojected complete pixel cell corresponding to the i-th particle prediction value of the j-th target drone is set to but You can obtain it in the following ways:
[0100] The following formula is used to calculate the coordinates of the preset significant feature points of the j-th target drone under the i-th particle prediction value in the camera coordinate system;
[0101]
[0102] According to the coordinates of the preset salient feature points in the camera coordinate system, the coordinates of the preset salient feature points in the camera pixel coordinate system are reprojected using the following formula to obtain the reprojected pixel cells;
[0103]
[0104] in, represents the coordinates of the preset significant feature point Q of the j-th target drone under the predicted value of the i-th particle in the camera coordinate system, represents the predicted value of the i-th particle of the j-th target drone at time t-1, R j,t-1 represents the attitude rotation matrix of the j-th target UAV at time t-1, Q j represents the three-dimensional coordinates of the preset significant feature point Q of the j-th target UAV in its body coordinate system, ξ j represents the set of preset significant feature points of the j-th target drone, Represents the coordinates of the preset salient feature point Q in the camera pixel coordinate system.
[0105] By using the above method, the reprojected complete pixel cell corresponding to each particle prediction value, i.e., the reprojected image state, can be determined. t target drones, if each target drone has N p particle prediction values, we can get N t ×N p Reprojected complete pixel cells.
[0106] Furthermore, in order to improve the tracking accuracy and the accuracy of UAV formation control, in one embodiment of the present invention, by calculating the difference between the reprojected complete pixel cell and all the detection cells and the isolated feature point set, that is, the difference between the input image state and the reprojected image state, the weight corresponding to the particle prediction value is updated according to the difference to perform the correction of the predicted state of the target UAV and the inter-frame matching of the image detection.
[0107] Specifically, the reprojected complete pixel cell corresponding to the predicted value of the i-th particle of the j-th target drone is calculated Taking the difference between all detection cells and isolated feature point sets as an example, the difference between the reprojected complete pixel cell and the detection cell can be defined as:
[0108]
[0109] in,
[0110]
[0111] Represents the reprojected complete pixel cell and detection of cell χ k The difference between M and the reprojected complete pixel cell The number of feature points in , Represents the reprojected complete pixel cell The mth feature point in the cell, p represents the image feature point in the detection cell, and ε represents the allowable pixel error.
[0112] The number of feature points in the reprojected complete pixel cell of the target drone is the same as the number of preset salient feature points of the target drone. Due to occlusion, missed detection, and other conditions, in actual applications, image feature points corresponding to feature points in the reprojected complete pixel cell may be missing.
[0113] According to the above definition and calculation formula of the difference between the reprojected complete pixel cell and the detection cell, the detection cell with the smallest difference from the reprojected complete pixel cell can be found. This detection cell is the detection cell that best matches the reprojected complete pixel cell.
[0114] Specifically, the reprojected complete pixel cell corresponding to the predicted value of the i-th particle of the j-th target drone is For example, reproject the complete pixel cell The best matching detection cell can be determined using the following formula:
[0115]
[0116] Among them, χ * Represents the reprojected complete pixel cell The best matching detection cell.
[0117] Furthermore, the reprojected complete pixel cell corresponding to the predicted value of the i-th particle of the j-th target drone is calculated Taking the difference between all detection cells and isolated feature point sets as an example, the difference between the reprojected complete pixel cell and the isolated feature point set can be defined as:
[0118]
[0119] in,
[0120]
[0121] Represents the reprojected complete pixel cell The difference from the isolated feature point set ξ, Represents isolated feature point set ξ and feature point The nearest point.
[0122] Among them, the isolated feature point set ξ and the feature point Nearest point The formula can be Please help.
[0123] Furthermore, in order to improve the pose estimation accuracy of the target UAV, so as to improve the tracking accuracy and the UAV formation control accuracy, in one embodiment of the present invention, different methods are used to update the weights corresponding to the particle prediction values according to the different differences between the reprojected complete pixel cells and all detection cells and isolated feature point sets.
[0124] Specifically, if That is, if the difference between the reprojected complete pixel cell and the detected cell is smaller, the weight corresponding to the particle prediction value is updated using the following formula:
[0125]
[0126] like That is, if the difference between the reprojected complete pixel cell and a subset of the isolated feature point set is smaller, the weight corresponding to the particle prediction value is updated using the following formula:
[0127]
[0128] in, Represents the particle prediction value The corresponding weight, Represents the reprojected complete pixel cell The best matching subset of isolated feature points,
[0129] Furthermore, after completing the weight update of each particle prediction value, the three-dimensional position prediction value of the target UAV at time t, that is, the three-dimensional position prediction value at the current moment, can be calculated using the following formula based on each particle prediction value and its corresponding weight:
[0130]
[0131] in, represents the predicted three-dimensional position of the j-th target UAV at time t, N p Represents the number of particle prediction values generated for the j-th target drone.
[0132] In step S5, for each target UAV, according to the camera projection model, the preset significant feature points of the target UAV under the three-dimensional position prediction value are reprojected to the camera pixel coordinate system to obtain the reprojected complete pixel cell corresponding to the three-dimensional position prediction value, and the minimum difference between the reprojected complete pixel cell corresponding to the three-dimensional position prediction value and all the detection cells and isolated feature point sets obtained in step S4 is calculated. If the minimum difference is not greater than the preset threshold, the posture of the target UAV is updated according to the three-dimensional position prediction value, and the three-dimensional position prediction value and posture are output as the tracking result of the corresponding target UAV, and the detection cell or isolated feature point subset corresponding to the minimum difference is marked as associated. If the minimum difference is greater than the preset threshold, the corresponding target UAV is deemed to be lost in tracking.
[0133] Specifically, taking the jth target UAV as an example, after obtaining the three-dimensional position prediction value of the target UAV Then, referring to the solution method of the reprojected complete pixel cell in step S4 above, the three-dimensional position prediction value of the target drone can be calculated. The corresponding reprojected complete pixel cell; then, referring to the calculation formula of the difference defined in step S4 above, the three-dimensional position prediction value can be calculated The minimum difference between the corresponding reprojected complete pixel cell and all detection cells and isolated feature point sets obtained in step S4, that is, the three-dimensional position prediction value The minimum difference between the corresponding reprojected complete pixel cell and all the detection cells and isolated feature point sets detected in the camera image acquired at the current moment. If the minimum difference is not greater than the preset threshold, the iterative closest point algorithm (ICP algorithm) is used according to the three-dimensional position prediction value to calculate and update the target drone's posture, output the three-dimensional position prediction value and posture as the corresponding target drone tracking result, and the detection cell corresponding to the minimum difference is set as the target drone's tracking result. Or isolated feature point subset ψ j Marked as associated, if the minimum difference is greater than the preset threshold, the corresponding target drone is considered lost.
[0134] In step S6, the remaining unassociated detection cells and isolated feature point sets are re-detected and screened to determine whether there are any unassociated complete detection cells. If there are any unassociated complete detection cells, the three-dimensional position and posture of the drone corresponding to all unassociated complete detection cells are calculated based on the three-dimensional coordinates of the preset significant feature points, and the three-dimensional position and posture are output as the tracking results of the corresponding drone.
[0135] After completing the tracking and confirmation of the corresponding target drone at the current moment based on the position and posture of the target drone at the previous moment, considering that there may be complete detection cells that are not associated in the remaining unassociated detection cells and isolated feature point sets of all detection cells and isolated feature point sets detected in the camera image acquired at the current moment, in one embodiment of the present invention, the remaining unassociated detection cells and isolated feature point sets are also re-detected and screened to determine whether there are unassociated complete detection cells. If there are unassociated complete detection cells, the three-dimensional position and posture of the drones corresponding to all unassociated complete detection cells are calculated based on the three-dimensional coordinates of the preset significant feature points, and the three-dimensional position and posture are output as the tracking results of the corresponding drones, thereby realizing target reconstruction.
[0136] Among them, according to the three-dimensional coordinates of the preset significant feature points, the three-dimensional position and posture of the drone corresponding to all unassociated complete detection cells are calculated, and the above steps S31-S33 can be referred to.
[0137] In one embodiment of the present invention, by re-detecting and screening the remaining unassociated detection cells and isolated feature point sets, it is possible to re-detect the target and timely discover and generate new target drones.
[0138] Step S7, return to step S4 until the drone formation mission is completed.
[0139] When a drone cluster performs a formation mission, the execution process of the formation mission is usually a relatively long process. Therefore, during the execution of the drone formation mission, after completing steps S1 to S6 once, the drone returns to step S4 to loop through steps S4 to S6 until the drone formation mission is completed.
[0140] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or apparatus that includes a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or apparatus.
[0141] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A method for estimating relative pose of a UAV formation based only on monocular vision information, characterized in that: The following steps are involved: S1, determining the three-dimensional coordinates of at least four preset significant feature points on the drone in its body coordinate system, determining the internal and external parameters of the camera carried by each drone, and establishing a camera projection model; S2, acquire the first frame of image from the camera, perform feature point detection on the acquired image, and determine the detection cells and isolated feature point sets corresponding to the drones based on the geometric relationship between the detected feature points and each drone. The detection cells corresponding to the drones represent the set of feature points that are associated with the current drone and have no ambiguity in their association with the current drone, and the isolated feature point set represents the set of feature points that have ambiguity in their association with the drones but are not associated with any drones. S3, determining a complete detection cell, and calculating the 3D position and posture of the drone corresponding to the complete detection cell based on the 3D coordinates of the preset significant feature points, where a complete detection cell refers to a detection cell that includes more than four feature points; S4, acquiring the next frame of image from the camera, performing feature point detection on the acquired image, and determining the detection cells and isolated feature point sets corresponding to the drones based on the geometric relationships between the detected feature points and each drone; For each target drone, while keeping its attitude unchanged, predict the current 3D position of the target drone based on its 3D position at the previous moment, where the target drone is a drone whose 3D position and attitude at the previous moment are known; S5, for each target drone, according to the camera projection model, reproject the preset significant feature points of the target drone under the three-dimensional position prediction value to the camera pixel coordinate system, obtain the reprojected complete pixel cell corresponding to the three-dimensional position prediction value, calculate the minimum difference between the reprojected complete pixel cell corresponding to the three-dimensional position prediction value and all the detection cells and isolated feature point sets obtained in step S4, if the minimum difference is not greater than the preset threshold, update the posture of the target drone according to the three-dimensional position prediction value, output the three-dimensional position prediction value and posture as the tracking result of the corresponding target drone, and mark the detection cell or isolated feature point subset corresponding to the minimum difference as associated, if the minimum difference is greater than the preset threshold, the corresponding target drone is considered to be lost in tracking; S6: Re-detect and screen the remaining unassociated detection cells and isolated feature point sets to determine whether there are any unassociated complete detection cells. If there are any unassociated complete detection cells, calculate the 3D position and attitude of the drone corresponding to all unassociated complete detection cells based on the 3D coordinates of the preset significant feature points, and output the 3D position and attitude as the tracking result of the corresponding drone. S7, return to step S4 until the UAV formation mission is completed.
2. The method for estimating relative pose of a UAV formation based only on monocular vision information according to claim 1, characterized in that: The preset significant feature points include: the wing tips, tail tips, and / or artificially placed cooperation signs of the UAV.
3. The method for estimating relative pose of a UAV formation based only on monocular vision information according to claim 1 or 2, characterized in that: Setting: The camera coordinate system O-XYZ is the reference coordinate system, and the camera coordinate system coincides with the drone body coordinate system. The camera image coordinate system is o-xy, the camera pixel coordinate system is o-uv, and the three-dimensional coordinates of point P in the j-th drone body coordinate system are The camera projection model is: in, is the center of gravity O of the jth UAV j The three-dimensional coordinates in the camera coordinate system, is the coordinate of point P in the camera pixel coordinate system, R j is the attitude rotation matrix of the j-th UAV, represents a 3x3 matrix, f x and f y Indicates the equivalent focal length of the camera in the x and y directions, c x ,c y Represents the coordinates of the camera principal point in the camera pixel coordinate system, and K represents the intrinsic parameter matrix composed of the camera intrinsic parameters.
4. The method for estimating relative pose of a UAV formation based only on monocular vision information according to claim 3, characterized in that: The method of calculating the three-dimensional position and posture of the drone corresponding to the complete detection cell according to the three-dimensional coordinates of the preset significant feature points includes: Determine the three-dimensional coordinates of the drone's preset salient feature points corresponding to the complete detection cell in the aircraft coordinate system. Based on the Perspective-n-Point principle, calculate the distance from the preset salient feature points to the camera's optical center. Calculate the three-dimensional coordinates of the preset significant feature points in the camera coordinate system according to the iterative closest point algorithm; The three-dimensional position and attitude of the UAV are calculated based on the relative position relationship between the preset significant feature points and the center of gravity of the UAV.
5. The method for estimating relative pose of a UAV formation based only on monocular vision information according to claim 4, characterized in that: The method of predicting the three-dimensional position of the target drone at the current moment based on the three-dimensional position of the target drone at the previous moment includes: According to the 3D position of the target UAV at the previous moment, the initial predicted value of the 3D position of the target UAV at the current moment is calculated and predicted by the dynamic model of the target UAV; Based on the initial predicted value of the 3D position at the current moment, multiple particle prediction values are generated through particle sampling. According to the camera projection model, the preset significant feature points of the target drone under each particle prediction value are reprojected to the camera pixel coordinate system to obtain the reprojected complete pixel cell corresponding to each particle prediction value; According to the difference between the reprojected complete pixel cell corresponding to the particle prediction value and all the detection cells and isolated feature point sets obtained in step S4, the weight corresponding to the particle prediction value is updated. According to each particle prediction value and its corresponding weight, the three-dimensional position prediction value of the target drone at the current moment is calculated.
6. The method for estimating relative pose of a UAV formation based only on monocular vision information according to claim 5, characterized in that: Assume that the reprojected complete pixel cell corresponding to the predicted value of the i-th particle of the j-th target drone is Obtain it in the following ways: The following formula is used to calculate the coordinates of the preset significant feature points of the j-th target drone under the i-th particle prediction value in the camera coordinate system; According to the coordinates of the preset salient feature points in the camera coordinate system, the coordinates of the preset salient feature points in the camera pixel coordinate system are reprojected using the following formula to obtain the reprojected pixel cells; in, represents the coordinates of the preset significant feature point Q of the j-th target drone under the predicted value of the i-th particle in the camera coordinate system, represents the predicted value of the i-th particle of the j-th target drone at time t-1, R j,t-1 represents the attitude rotation matrix of the j-th target UAV at time t-1, Q j represents the three-dimensional coordinates of the preset significant feature point Q of the j-th target UAV in its body coordinate system, ξ j represents the set of preset significant feature points of the j-th target drone, Represents the coordinates of the preset salient feature point Q in the camera pixel coordinate system.
7. The method for estimating relative pose of a UAV formation based solely on monocular vision information according to claim 6, characterized in that: The difference between the reprojected complete pixel cell and the detection cell is calculated using the following formula: in, Represents the reprojected complete pixel cell and detection of cell χ k The difference between M and the reprojected complete pixel cell ζ j (i) The number of feature points in , Represents the reprojected complete pixel cell The mth feature point in the cell, p represents the image feature point in the detection cell, and ε represents the allowable pixel error.
8. The method for estimating relative pose of a UAV formation based solely on monocular vision information according to claim 7, characterized in that: The difference between the reprojected complete pixel cell and the isolated feature point set is calculated using the following formula: in, Represents the reprojected complete pixel cell The difference from the isolated feature point set ξ, Represents isolated feature point set ξ and feature point The nearest point.
9. The method for estimating relative pose of a UAV formation based solely on monocular vision information according to claim 8, characterized in that: like Then use the following formula to update the weight corresponding to the particle prediction value: like Then use the following formula to update the weight corresponding to the particle prediction value: Among them, χ * Represents the reprojected complete pixel cell The best matching detection cell, Represents the particle prediction value The corresponding weight, Represents the reprojected complete pixel cell The best matching subset of isolated feature points, 10. The method for estimating relative pose of a UAV formation based solely on monocular vision information according to claim 9, characterized in that: The following formula is used to calculate the predicted three-dimensional position of the target drone at the current moment: in, represents the predicted three-dimensional position of the j-th target UAV at time t, N p Represents the number of particle prediction values generated for the j-th target drone.
Citation Information
Patent Citations
Monocular vision odometer positioning method and positioning system based on semi-direct method
CN108986037A
Unmanned aerial vehicle-based projection method and apparatus, device, and storage medium
WO2021227359A1