Refueling taper sleeve positioning method based on visible light and infrared fusion
By using the method of visible light and infrared fusion, combined with convolutional neural networks and visual algorithms, the two-dimensional and three-dimensional coordinates of the refueling drogue are solved, which solves the complexity of refueling drogue positioning during aerial refueling and achieves precise docking in complex flight scenarios.
Patent Information
- Application Number
- CN202510788446.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-13
- Publication Date
- 2025-09-12
AI Technical Summary
Existing visual relative positioning technology has difficulty coping with complex flight scenarios during aerial refueling. The single-source imaging mode cannot effectively identify the refueling drogue, especially during night flight, where it cannot be identified or is easily interfered by stray light.
A refueling drogue positioning method based on the fusion of visible light and infrared is adopted. A convolutional neural network is used to solve the appearance features of the visible light imaging image, and a visual algorithm is combined to solve the LED infrared lamp features of the infrared imaging image. The three-dimensional coordinates of the refueling drogue are obtained through pose estimation and Kalman filtering.
The accuracy and robustness of the refueling drogue positioning are improved, enabling precise docking in complex flight scenarios and reducing dependence on GPS absolute positioning.
Smart Images

Figure CN120635206A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of visual navigation, and in particular relates to a refueling drogue positioning method based on the fusion of visible light and infrared. Background Art
[0002] Aerial refueling is a refueling technique in which a tanker aircraft simultaneously refuels one or more receiving aircraft in flight, effectively increasing the range and flight time of the receiving aircraft. The complete aerial refueling process includes five phases: rendezvous, formation, docking, refueling, and exit. Aerial refueling can be performed using both soft and hard refueling.
[0003] During soft aerial refueling, the flexible hose is subject to nonlinear, fast-changing, and strongly coupled oscillating motions in the refueling drogue, often under complex wind disturbances such as the tanker's wake vortex, gusts, and atmospheric turbulence. This introduces uncertainty and risk to the precise docking of the refueling drogue. With the rapid development of visual relative positioning technology, the development of a visual navigation system for aerial refueling is expected to address these technical challenges. This system utilizes visual relative positioning technology for close-range perception, reducing reliance on GPS absolute positioning during docking and enabling direct relative positioning between the refueling drogue and the receiving plug.
[0004] However, existing visual relative positioning technologies often employ single-source imaging, using either visible light cameras or infrared cameras. These two imaging modes correspond to two distinct image processing methods: visible light camera imaging corresponds to a convolutional neural network-based refueling drogue localization method, while infrared camera imaging uses a visual algorithm-based refueling drogue feature point localization method. For visual relative positioning, single-source positioning methods struggle to cope with demanding and complex aerial scenarios. For example, a convolutional neural network-based refueling drogue localization method cannot identify the refueling drogue during nighttime flight operations, while a visual algorithm-based refueling drogue feature point localization method is susceptible to noise interference caused by stray light.
[0005] At present, there is an urgent need to develop a refueling drogue positioning method based on the fusion of visible light and infrared. Summary of the Invention
[0006] The technical problem to be solved by the present invention is to provide a refueling drogue positioning method based on the fusion of visible light and infrared, so as to overcome the limitations of the single-source refueling drogue positioning method.
[0007] The refueling drogue positioning method based on visible light and infrared fusion of the present invention comprises the following steps: S10. Establish a visual navigation system for aerial refueling; The tanker and receiver form a two-plane formation; both the visible light camera and the infrared camera are fixed to the receiver, and the refueling drogue is mounted on the tanker. The refueling drogue is equipped with eight infrared LEDs along its circumference. Conduct a two-aircraft formation flight. When the receiving aircraft reaches the visual navigation area, the tanker releases the refueling drogue and lights up the eight infrared LEDs in the refueling drogue. The receiving aircraft then turns on its visible light camera and infrared camera to capture images, obtaining visible light and infrared images respectively. S20. Using a convolutional neural network to calculate the appearance features of the refueling drogue in the visible light imaging image, obtain the two-dimensional coordinates of the refueling drogue in the pixel coordinate system; S30. Based on the infrared imaging image, use a visual algorithm to calculate the characteristics of the LED infrared lamp beads of the refueling drogue to obtain the two-dimensional coordinates of the refueling drogue in the pixel coordinate system; S40. For the two-dimensional coordinates of the refueling drogue in the pixel coordinate system obtained in S20 and S30, perform pose estimation, coordinate transformation, and fusion to obtain an estimated three-dimensional coordinate value of the refueling drogue; S50. Use Kalman filtering to predict and gradually update the estimated three-dimensional coordinate value of the refueling cone sleeve, continuously improve the prediction accuracy, and finally output the three-dimensional coordinate prediction value of the refueling cone sleeve that meets the accuracy requirements.
[0008] Furthermore, the convolutional neural network in S20 is a YOLOv4 object detection network. The trained YOLOv4 object detection network model is used to identify the refueling drogue, calculate the bounding box information of the refueling drogue, and then solve the two-dimensional coordinates of the refueling drogue in the pixel coordinate system. The specific process includes the following: The YOLOv4 network model consists of three parts: the backbone feature extraction layer BackBone, the enhanced feature extraction layer Neck, and the processing output layer Prediction. Among them, the backbone feature extraction layer BackBone is CSPDarknet-53, the enhanced feature extraction layer Neck is composed of the SPP module and the PAN module, and the processing output layer Prediction adopts the YOLO Head structure of YOLOv3. In the backbone feature extraction layer BackBone, the basic component of CSPDarknet-53 is CSPX, which consists of CBM and X Res_unit modules Concat; CBM is composed of convolutional layer Conv, normalization layer BN and activation function Mish; CBL is composed of convolutional layer Conv, normalization layer BN and activation function LeakyReLU; The Res_unit structure is derived from the residual network ResNet, which includes two CBM layers and a feature fusion layer add. The input of the first CBM layer and the output of the second CBM layer are added through the feature fusion layer add, making the YOLOv4 network model deeper. The SPP module in the enhanced feature extraction layer Neck is located in the last feature layer of CSPDarknet-53. After performing three convolution operations on the last feature layer, maximum pooling is performed using four pooling kernels of different scales; the sizes of the four pooling kernels are 13×13, 9×9, 5×5, and 1×1, respectively. After processing by the SPP module, the perception can be enhanced and more significant image features can be separated. In the YOLOv4 network model, the PAN module is applied to the three effective feature layers to complete feature fusion and improve feature extraction capabilities; The output layer, Prediction, predicts targets of three different sizes. It is composed of the YOLOHead structure in the YOLOv4 network model. When the input image size is 608×608×3, the output sizes of the YOLOv4 network model are 76×76, 38×38, and 19×19, respectively, for detecting small, medium, and large targets. In the YOLOv4 network model, the result of the anchor box prediction is the offset value relative to the prior box. It is necessary to solve the actual center pixel coordinates of the bounding box according to formula (1): : ; in, is the coordinate offset value predicted by the YOLOv4 network model, is the true prior box size, is the pixel coordinate of the center of the true prior frame; The YOLOv4 network model designs three output feature maps of different sizes at different locations in the network structure to detect the receiving aircraft and the refueling drogue respectively. The final output tensor results are: ; in, represents the size of the image; C is the number of categories to be detected; the number 3 means that each grid requires 3 anchor boxes; the number 4 represents the center pixel coordinate of the bounding box and target width and height Four parameters, the number 1 represents the confidence of the bounding box; Finally, the target width and height of the output The value is used to judge whether the pixel area currently occupied by the refueling drogue is greater than a given value, so as to select whether to output the two-dimensional center coordinates of the refueling drogue, and the two-dimensional center coordinates of the refueling drogue are used as the two-dimensional coordinates of the refueling drogue in the pixel coordinate system.
[0009] Furthermore, the visual algorithm in S30 is a combination of traditional visual methods. The image acquired by the infrared camera is a single-channel grayscale image. The algorithm steps for the two-dimensional coordinates in the pixel coordinate system include the following processes: Binarize the single-channel grayscale image to obtain a binary image. The binarization method uses global threshold binarization. Set the threshold size to , the pixel gray value is greater than The pixel value is set to white at 255, and the pixel grayscale value is less than The pixels are set to black with a pixel value of 0; the binarization formula is: ; in, Represents the grayscale value of each pixel of a single-channel grayscale image; Then, the connected domain labeling method based on region growing is used to search and label the connected domains of the binary image to obtain the connected domain information set, which includes the coordinates of the centroid, area and perimeter of the connected domain; and the connected domain information set with an area value less than and greater than The connected domains are eliminated to obtain the effective connected domain information set; Perform ellipse fitting on the centroid coordinates of the connected domain in the valid connected domain information set to find the coordinate points of 8 LED infrared lamp beads; the standard equation of the ellipse is: ; in, are the coordinates of the ellipse center, is the long axis, is the minor axis; the general equation of an ellipse is: ; definition ; The fitting method is as follows, input a set of point sets , is the serial number of the point to be fitted, and the six parameters of the ellipse AF are obtained; the polynomial is the coordinate point on the ellipse algebraic distance to a given conic section; ; in, are the six constant values of the ellipse, satisfying the constraints ; Defining a vector , Write the polynomial as a vector: ; Formula (7) is solved by the least square method, but the fitting result is a conic section instead of an ellipse. In order to ensure that the solution is an ellipse, the constraint condition is reconsidered. ; When a takes any non-zero value, that is, a≠0, a•a represents the same conic section; under a pre-set scale, the inequality constraint is transformed into an equality constraint; Define an N×6 matrix D: ; in, are the x-axis coordinates of the points to be fitted in the pixel coordinate system, are the y-axis coordinates of the points to be fitted in the pixel coordinate system; The ellipse fitting problem can be reformulated as: , when it meets Under the condition of , define the 6X6 constant matrix C: ; The constraints Written ;right Problem decomposition: exist Under the condition of ; construct the Lagrangian function: , solve the partial derivative and set it to 0, and we get ,Right now ,make , the equation becomes , is the Lagrange multiplier, which is then simplified to the form of obtaining the eigenvalue At this point, the ellipse fitting problem is transformed into finding The eigenvalue and eigenvector problems are solved. By selecting the eigenvectors that meet the conditions, the ellipse parameters are calculated. The ellipse parameters include the center coordinates, major and minor axis parameters of the ellipse. After obtaining the ellipse parameters, the coordinate points of each set of valid connected domain information sets are brought into the ellipse standard equation to calculate the reprojection error. The error is expressed as: ; The coordinate points of the effective connected domain information set with the smallest reprojection error are taken as the most suitable coordinate points of the 8 LED infrared lamp beads; the coordinates of the center of the ellipse corresponding to this effective connected domain information set are taken as the two-dimensional coordinates of the refueling cone sleeve in the pixel coordinate system.
[0010] Furthermore, the specific solution process for the estimated three-dimensional coordinates of the refueling drogue of S40 is as follows: Before fusing the visible light imaging and infrared imaging results, the camera pose and coordinate system transformation need to be calculated for the two 2D coordinate points of the refueling drogue. Once the coordinates of the 3D space point and the coordinates of the 2D image feature points corresponding to the 3D space point are known, the camera pose is solved using the PnP algorithm. PnP algorithms include P3P, EPnP, UPnP, direct linear transformation, and nonlinear optimization. When using the direct linear transformation method to solve, for a given three-dimensional point , in the image, the projected feature point coordinates are ; At this time, the camera pose is the unknown quantity to be solved, and the corresponding projection equation is as follows: ; in, is the camera intrinsic parameter matrix, obtained through camera calibration; is the rotation matrix, is the translation vector; are the parameters in the camera intrinsic parameter matrix respectively; Camera pose After the solution is completed, the three-dimensional coordinates of the refueling drogue in different coordinate systems are converted to a unified coordinate system through coordinate system conversion; the three-dimensional coordinates of the refueling drogue are based on the center of the cone as the origin and follow the right-hand coordinate system principle. Direction vertically upward, Horizontal and vertical to the right, Facing outward, 8 LED infrared lamp beads are evenly distributed on the refueling cone sleeve in a three-dimensional coordinate system; Four cameras are distributed on the receiving aircraft. Two of them are infrared cameras, located on the left and right sides of the receiving aircraft fuselage; the other two are visible light cameras, located on the left and right wings of the receiving aircraft. The two-dimensional coordinates of the refueling drogue in the pixel coordinate system obtained by the four cameras correspond to the left infrared coordinate system, the right infrared coordinate system, the left visible light coordinate system, and the right visible light coordinate system, respectively. The two-dimensional coordinates of the refueling drogue in the pixel coordinate system obtained by the four cameras are transformed relative to the body coordinate system of the nose of the refueling aircraft, and then the three-dimensional coordinate estimation value of the refueling drogue is obtained by summing and taking the average value.
[0011] Furthermore, the specific solution process for the three-dimensional coordinate prediction value of the refueling drogue of S50 is as follows: In order to ensure the smoothness of the output of the three-dimensional coordinate estimation value of the refueling drogue, the Kalman filter is used to filter the three-dimensional coordinate estimation value of the refueling drogue; the prediction and update equations of the Kalman filter are as follows: ; The order of Kalman filtering is as follows: S51. Initialization; Initialization is performed only once and provides two parameters: initial state , initial state variance , make predictions after initialization; S52.Measurement; The measurement process also provides two parameters: the measured value Measurement variance ; S53. Status update; The input to the state update process is: measurement value , measurement variance , state prediction and the state prediction variance Based on the input, the state update process calculates the Kalman gain and produces the output: state estimation and the state estimation variance ; S54. Forecast; The prediction process uses the system dynamics model to estimate the current state of the system and the state estimation variance Extrapolate to the next moment; in the first iteration, the initialization parameters are used as the value and variance of the state prediction, and the predicted output is used as the value and variance of the state prediction at the next moment; after the prediction, the three-dimensional coordinate prediction value of the refueling drogue that meets the preset accuracy requirements is obtained.
[0012] The present invention's refueling drogue positioning method, based on the fusion of visible light and infrared, uses a convolutional neural network to calculate the refueling drogue's appearance features in a visible light image to obtain its two-dimensional coordinates. It also uses a visual algorithm to calculate the characteristics of the refueling drogue's LED infrared beads in an infrared image to obtain its two-dimensional coordinates. The two-dimensional coordinates calculated from the two imagery are then subjected to pose estimation and Kalman filtering to obtain the final predicted three-dimensional coordinates. The present invention's refueling drogue positioning method, based on the fusion of visible light and infrared, has been experimentally verified to demonstrate accuracy and robustness, demonstrating practical engineering value. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Figure 1 Schematic diagram of the network structure of YOLOv4; Figure 2 Schematic diagram of the basic structure of CSPX Figure 3 This is a schematic diagram of the basic structure of CBM; Figure 4 This is a schematic diagram of the basic structure of Res_unit; Figure 5 Schematic diagram of the basic structure of SPP; Figure 6 This is a schematic diagram of the basic structure of PAN; Figure 7 It is a flowchart of traditional visual method; Figure 8 is a binary image; Figure 9 is a schematic diagram of the coordinate system; Figure 10 This is the overall flow chart of the refueling drogue positioning method based on the fusion of visible light and infrared. DETAILED DESCRIPTION
[0014] The present invention will be described in detail below with reference to the accompanying drawings and embodiments.
[0015] Example: Figure 10 As shown, the refueling drogue positioning method based on visible light and infrared fusion of this embodiment includes the following steps: S10. Establish a visual navigation system for aerial refueling; The tanker and receiver form a two-plane formation; both the visible light camera and the infrared camera are fixed to the receiver, and the refueling drogue is mounted on the tanker. The refueling drogue is equipped with eight infrared LEDs along its circumference. Conduct a two-aircraft formation flight. When the receiving aircraft reaches the visual navigation area, the tanker releases the refueling drogue and lights up the eight infrared LEDs in the refueling drogue. The receiving aircraft then turns on its visible light camera and infrared camera to capture images, obtaining visible light and infrared images respectively. S20. Using a convolutional neural network to calculate the appearance features of the refueling drogue in the visible light imaging image, obtain the two-dimensional coordinates of the refueling drogue in the pixel coordinate system; S30. Based on the infrared imaging image, use a visual algorithm to calculate the characteristics of the LED infrared lamp beads of the refueling drogue to obtain the two-dimensional coordinates of the refueling drogue in the pixel coordinate system; S40. For the two-dimensional coordinates of the refueling drogue in the pixel coordinate system obtained in S20 and S30, perform pose estimation, coordinate transformation, and fusion to obtain an estimated three-dimensional coordinate value of the refueling drogue; S50. Use Kalman filtering to predict and gradually update the estimated three-dimensional coordinate value of the refueling cone sleeve, continuously improve the prediction accuracy, and finally output the three-dimensional coordinate prediction value of the refueling cone sleeve that meets the accuracy requirements.
[0016] Furthermore, the convolutional neural network in S20 is a YOLOv4 object detection network. The trained YOLOv4 object detection network model is used to identify the refueling drogue, calculate the bounding box information of the refueling drogue, and then solve the two-dimensional coordinates of the refueling drogue in the pixel coordinate system. The specific process includes the following: The YOLOv4 network model (YOLOv4 Structure) is a single-stage deep learning target detection algorithm that offers higher accuracy and real-time performance in identifying small targets. It is suitable for deployment in receiving aircraft equipped with edge computing devices. like Figure 1 As shown in the figure, the YOLOv4 network model consists of three parts: the backbone feature extraction layer BackBone, the enhanced feature extraction layer Neck, and the processing output layer Prediction; among them, the backbone feature extraction layer BackBone is CSPDarknet-53, the enhanced feature extraction layer Neck is composed of the SPP module and the PAN module, and the processing output layer Prediction adopts the YOLO Head structure of YOLOv3; like Figure 2 As shown in the figure, in the backbone feature extraction layer BackBone, the basic component of CSPDarknet-53 is CSPX, which is composed of CBM and X Res_unit modules Concat; like Figure 3 As shown in the figure, CBM is composed of convolution layer Conv, normalization layer BN and activation function Mish; CBL is composed of convolution layer Conv, normalization layer BN and activation function LeakyReLU; like Figure 4 As shown in the figure, the Res_unit structure is derived from the residual network ResNet, which includes two CBM layers and a feature fusion layer add. The input of the first CBM layer and the output of the second CBM layer are added through the feature fusion layer add, making the YOLOv4 network model deeper. SPP module structure see Figure 5 The SPP module in the enhanced feature extraction layer Neck is located in the last feature layer of CSPDarknet-53. After performing three convolution operations on the last feature layer, maximum pooling is performed using four pooling kernels of different scales; the sizes of the four pooling kernels are 13×13, 9×9, 5×5, and 1×1, respectively. After processing by the SPP module, the perception can be enhanced and more significant image features can be separated. PAN module structure see Figure 6 In the YOLOv4 network model, the PAN module is applied to the three effective feature layers to complete feature fusion and improve feature extraction capabilities; The output layer, Prediction, predicts targets of three different sizes. It is composed of the YOLOHead structure in the YOLOv4 network model. When the input image size is 608×608×3, the output sizes of the YOLOv4 network model are 76×76, 38×38, and 19×19, respectively, for detecting small, medium, and large targets. In the YOLOv4 network model, the result of the anchor box prediction is the offset value relative to the prior box. It is necessary to solve the actual center pixel coordinates of the bounding box according to formula (1): : ; in, is the coordinate offset value predicted by the YOLOv4 network model, is the true prior box size, is the pixel coordinate of the center of the true prior frame; The YOLOv4 network model designs three output feature maps of different sizes at different locations in the network structure to detect the receiving aircraft and the refueling drogue respectively. The final output tensor results are: ; in, represents the size of the image; C is the number of categories to be detected; the number 3 means that each grid requires 3 anchor boxes; the number 4 represents the center pixel coordinate of the bounding box and target width and height Four parameters, the number 1 represents the confidence of the bounding box; Finally, the target width and height of the output The value is used to judge whether the pixel area currently occupied by the refueling drogue is greater than a given value, so as to select whether to output the two-dimensional center coordinates of the refueling drogue, and the two-dimensional center coordinates of the refueling drogue are used as the two-dimensional coordinates of the refueling drogue in the pixel coordinate system.
[0017] Furthermore, if Figure 7 As shown, the visual algorithm in S30 is a combination of traditional visual methods. The image acquired by the infrared camera is a single-channel grayscale image. The algorithm steps for the two-dimensional coordinates in the pixel coordinate system include the following processes: For example Figure 8 The single-channel grayscale image shown is binarized to obtain a binary image. The binarization method uses global threshold binarization, and the threshold size is set to , the pixel gray value is greater than The pixel value is set to white at 255, and the pixel grayscale value is less than The pixels are set to black with a pixel value of 0; the binarization formula is: ; in, Represents the grayscale value of each pixel of a single-channel grayscale image; Then, the connected domain labeling method based on region growing is used to search and label the connected domains of the binary image to obtain the connected domain information set, which includes the coordinates of the centroid, area and perimeter of the connected domain; and the connected domain information set with an area value less than and greater than The connected domains are eliminated to obtain the effective connected domain information set; Perform ellipse fitting on the centroid coordinates of the connected domain in the valid connected domain information set to find the coordinate points of 8 LED infrared lamp beads; the standard equation of the ellipse is: ; in, are the coordinates of the ellipse center, is the long axis, is the minor axis; the general equation of an ellipse is: ; definition ; The fitting method is as follows, input a set of point sets , is the serial number of the point to be fitted, and the six parameters of the ellipse AF are obtained; the polynomial is the coordinate point on the ellipse algebraic distance to a given conic section; ; in, are the six constant values of the ellipse, satisfying the constraints ; Defining a vector , Write the polynomial as a vector: ; Formula (7) is solved by the least square method, but the fitting result is a conic section instead of an ellipse. In order to ensure that the solution is an ellipse, the constraint condition is reconsidered. ; When a takes any non-zero value, that is, a≠0, a•a represents the same conic section; under a pre-set scale, the inequality constraint is transformed into an equality constraint; Define an N×6 matrix D: ; in, are the x-axis coordinates of the points to be fitted in the pixel coordinate system, are the y-axis coordinates of the points to be fitted in the pixel coordinate system; The ellipse fitting problem can be reformulated as: , when it meets Under the condition of , define the 6X6 constant matrix C: ; The constraints Written ;right Problem decomposition: exist Under the condition of ; construct the Lagrangian function: , solve the partial derivative and set it to 0, and we get ,Right now ,make , the equation becomes , is the Lagrange multiplier, which is then simplified to the form of obtaining the eigenvalue At this point, the ellipse fitting problem is transformed into finding The eigenvalue and eigenvector problems are solved. By selecting the eigenvectors that meet the conditions, the ellipse parameters are calculated. The ellipse parameters include the center coordinates, major and minor axis parameters of the ellipse. After obtaining the ellipse parameters, the coordinate points of each set of valid connected domain information sets are brought into the ellipse standard equation to calculate the reprojection error. The error is expressed as: ; The coordinate points of the effective connected domain information set with the smallest reprojection error are taken as the most suitable coordinate points of the 8 LED infrared lamp beads; the coordinates of the center of the ellipse corresponding to this effective connected domain information set are taken as the two-dimensional coordinates of the refueling cone sleeve in the pixel coordinate system.
[0018] Furthermore, the specific solution process for the estimated three-dimensional coordinates of the refueling drogue of S40 is as follows: Before fusing the visible light imaging and infrared imaging results, the camera pose and coordinate system transformation need to be calculated for the two 2D coordinate points of the refueling drogue. Once the coordinates of the 3D space point and the coordinates of the 2D image feature points corresponding to the 3D space point are known, the camera pose is solved using the PnP algorithm. PnP algorithms include P3P, EPnP, UPnP, direct linear transformation, and nonlinear optimization. When using the direct linear transformation method to solve, for a given three-dimensional point , in the image, the projected feature point coordinates are ; At this time, the camera pose is the unknown quantity to be solved, and the corresponding projection equation is as follows: ; in, is the camera intrinsic parameter matrix, obtained through camera calibration; is the rotation matrix, is the translation vector; are the parameters in the camera intrinsic parameter matrix respectively; Camera pose After the solution is completed, the three-dimensional coordinates of the refueling drogue in different coordinate systems are converted to a unified coordinate system through coordinate system conversion; the three-dimensional coordinates of the refueling drogue are based on the center of the cone as the origin and follow the right-hand coordinate system principle. Direction vertically upward, Horizontal and vertical to the right, Facing outward, 8 LED infrared lamp beads are evenly distributed on the refueling cone sleeve in a three-dimensional coordinate system; like Figure 9 As shown, there are four cameras distributed on the receiving aircraft. Two of them are infrared cameras, located on the left and right sides of the receiving aircraft fuselage respectively; the other two cameras are visible light cameras, located on the left and right wings of the receiving aircraft respectively. The two-dimensional coordinates of the refueling drogue in the pixel coordinate system obtained by the four cameras correspond to the left infrared coordinate system, the right infrared coordinate system, the left visible light coordinate system, and the right visible light coordinate system respectively. The two-dimensional coordinates of the refueling drogue in the pixel coordinate system obtained by the four cameras are transformed relative to the body coordinate system of the nose of the refueling aircraft, and then the three-dimensional coordinate estimation value of the refueling drogue is obtained by summing and taking the average value.
[0019] Furthermore, the specific solution process for the three-dimensional coordinate prediction value of the refueling drogue of S50 is as follows: In order to ensure the smoothness of the output of the three-dimensional coordinate estimation value of the refueling drogue, the Kalman filter is used to filter the three-dimensional coordinate estimation value of the refueling drogue; the prediction and update equations of the Kalman filter are as follows: ; The order of Kalman filtering is as follows: S51. Initialization; Initialization is performed only once and provides two parameters: initial state , initial state variance , make predictions after initialization; S52.Measurement; The measurement process also provides two parameters: the measured value Measurement variance ; S53. Status update; The input to the state update process is: measurement value , measurement variance , state prediction and the state prediction variance Based on the input, the state update process calculates the Kalman gain and produces the output: state estimation and the state estimation variance ; S54. Forecast; The prediction process uses the system dynamics model to estimate the current state of the system and the state estimation variance Extrapolate to the next moment; in the first iteration, the initialization parameters are used as the value and variance of the state prediction, and the predicted output is used as the value and variance of the state prediction at the next moment; after the prediction, the three-dimensional coordinate prediction value of the refueling drogue that meets the preset accuracy requirements is obtained.
[0020] Although the embodiments of the present invention have been disclosed as above, they are not limited to the applications listed in the description and implementation methods. For those familiar with the art, all features disclosed in the present invention, or all steps in the disclosed methods or processes, except for mutually exclusive features and / or steps, can be combined in any way without departing from the principles of the present invention. The present invention is not limited to the specific details and illustrations shown and described herein.
Claims
1. A refueling drogue positioning method based on visible light and infrared fusion, characterized in that: The following steps are involved: S10. Establish a visual navigation system for aerial refueling; The tanker and receiver form a two-plane formation; both the visible light camera and the infrared camera are fixed to the receiver, and the refueling drogue is mounted on the tanker. The refueling drogue is equipped with eight infrared LEDs along its circumference. Conduct a two-aircraft formation flight. When the receiving aircraft reaches the visual navigation area, the tanker releases the refueling drogue and lights up the eight infrared LEDs in the refueling drogue. The receiving aircraft then turns on its visible light camera and infrared camera to capture images, obtaining visible light and infrared images respectively. S20. Using a convolutional neural network to calculate the appearance features of the refueling drogue in the visible light imaging image, obtain the two-dimensional coordinates of the refueling drogue in the pixel coordinate system; S30. Based on the infrared imaging image, use a visual algorithm to calculate the characteristics of the LED infrared lamp beads of the refueling drogue to obtain the two-dimensional coordinates of the refueling drogue in the pixel coordinate system; S40. For the two-dimensional coordinates of the refueling drogue in the pixel coordinate system obtained in S20 and S30, perform pose estimation, coordinate transformation, and fusion to obtain an estimated three-dimensional coordinate value of the refueling drogue; S50. Use Kalman filtering to predict and gradually update the estimated three-dimensional coordinate value of the refueling cone sleeve, continuously improve the prediction accuracy, and finally output the three-dimensional coordinate prediction value of the refueling cone sleeve that meets the accuracy requirements.
2. The refueling drogue positioning method based on visible light and infrared fusion according to claim 1 is characterized in that: The convolutional neural network in S20 is a YOLOv4 object detection network. It uses the trained YOLOv4 object detection network model to identify the refueling drogue, calculate the bounding box information of the refueling drogue, and then solve the two-dimensional coordinates of the refueling drogue in the pixel coordinate system. The specific process includes the following: The YOLOv4 network model consists of three parts: the backbone feature extraction layer BackBone, the enhanced feature extraction layer Neck, and the processing output layer Prediction. Among them, the backbone feature extraction layer BackBone is CSPDarknet-53, the enhanced feature extraction layer Neck is composed of the SPP module and the PAN module, and the processing output layer Prediction adopts the YOLO Head structure of YOLOv3. In the backbone feature extraction layer BackBone, the basic component of CSPDarknet-53 is CSPX, which consists of CBM and X Res_unit modules Concat; CBM is composed of convolutional layer Conv, normalization layer BN and activation function Mish; CBL is composed of convolutional layer Conv, normalization layer BN and activation function LeakyReLU; The Res_unit structure is derived from the residual network ResNet, which includes two CBM layers and a feature fusion layer add. The input of the first CBM layer and the output of the second CBM layer are added through the feature fusion layer add, making the YOLOv4 network model deeper. The SPP module in the enhanced feature extraction layer Neck is located in the last feature layer of CSPDarknet-53. After performing three convolution operations on the last feature layer, maximum pooling is performed using four pooling kernels of different scales; the sizes of the four pooling kernels are 13×13, 9×9, 5×5, and 1×1, respectively. After processing by the SPP module, the perception can be enhanced and more significant image features can be separated. In the YOLOv4 network model, the PAN module is applied to the three effective feature layers to complete feature fusion and improve feature extraction capabilities; The output layer, Prediction, predicts targets of three different sizes. It is composed of the YOLO Head structure in the YOLOv4 network model. When the input image size is 608×608×3, the output sizes of the YOLOv4 network model are 76×76, 38×38, and 19×19, respectively, for detecting small, medium, and large targets. In the YOLOv4 network model, the result of the anchor box prediction is the offset value relative to the prior box. It is necessary to solve the actual center pixel coordinates of the bounding box according to formula (1): : ; in, is the coordinate offset value predicted by the YOLOv4 network model, is the true prior box size, is the pixel coordinate of the center of the true prior frame; The YOLOv4 network model designs three output feature maps of different sizes at different locations in the network structure to detect the receiving aircraft and the refueling drogue respectively. The final output tensor results are: ; in, represents the size of the image; C is the number of categories to be detected; the number 3 means that each grid requires 3 anchor boxes; the number 4 represents the center pixel coordinate of the bounding box and target width and height Four parameters, the number 1 represents the confidence of the bounding box; Finally, the target width and height of the output The value is used to judge whether the pixel area currently occupied by the refueling drogue is greater than a given value, so as to select whether to output the two-dimensional center coordinates of the refueling drogue, and the two-dimensional center coordinates of the refueling drogue are used as the two-dimensional coordinates of the refueling drogue in the pixel coordinate system.
3. The refueling drogue positioning method based on visible light and infrared fusion according to claim 1 is characterized in that: The visual algorithm in S30 is a combination of traditional visual methods. The image captured by the infrared camera is a single-channel grayscale image. The algorithm steps for the two-dimensional coordinates in the pixel coordinate system include the following process: Binarize the single-channel grayscale image to obtain a binary image. The binarization method uses global threshold binarization. Set the threshold size to , the pixel gray value is greater than The pixel value is set to white at 255, and the pixel grayscale value is less than The pixels are set to black with a pixel value of 0; the binarization formula is: ; in, Represents the grayscale value of each pixel of a single-channel grayscale image; Then, the connected domain labeling method based on region growing is used to search and label the connected domains of the binary image to obtain the connected domain information set, which includes the coordinates of the centroid, area and perimeter of the connected domain; and the connected domain information set with an area value less than and greater than The connected domains are eliminated to obtain the effective connected domain information set; Perform ellipse fitting on the centroid coordinates of the connected domain in the valid connected domain information set to find the coordinate points of 8 LED infrared lamp beads; the standard equation of the ellipse is: ; in, are the coordinates of the ellipse center, is the long axis, is the minor axis; the general equation of an ellipse is: ; definition ; The fitting method is as follows, input a set of point sets , is the serial number of the point to be fitted, and the six parameters of the ellipse AF are obtained; the polynomial is the coordinate point on the ellipse algebraic distance to a given conic section; ; in, are the six constant values of the ellipse, satisfying the constraints ; Defining a vector , Write the polynomial as a vector: ; Formula (7) is solved by the least square method, but the fitting result is a conic section instead of an ellipse. In order to ensure that the solution is an ellipse, the constraint condition is reconsidered. ; When a takes any non-zero value, that is, a≠0, a•a represents the same conic section; under a pre-set scale, the inequality constraint is transformed into an equality constraint; Define an N×6 matrix D: ; in, are the x-axis coordinates of the points to be fitted in the pixel coordinate system, are the y-axis coordinates of the points to be fitted in the pixel coordinate system; The ellipse fitting problem can be reformulated as: , when it meets Under the condition of , define the 6X6 constant matrix C: ; The constraints Written ;right Problem decomposition: exist Under the condition of ; construct the Lagrangian function: , solve the partial derivative and set it to 0, and we get ,Right now ,make , the equation becomes , is the Lagrange multiplier, which is then simplified to the form of obtaining the eigenvalue At this point, the ellipse fitting problem is transformed into finding The eigenvalue and eigenvector problems are solved. By selecting the eigenvectors that meet the conditions, the ellipse parameters are calculated. The ellipse parameters include the center coordinates, major and minor axis parameters of the ellipse. After obtaining the ellipse parameters, the coordinate points of each set of valid connected domain information sets are brought into the ellipse standard equation to calculate the reprojection error. The error is expressed as: ; The coordinate points of the effective connected domain information set with the smallest reprojection error are taken as the most suitable coordinate points of the 8 LED infrared lamp beads; the coordinates of the center of the ellipse corresponding to this effective connected domain information set are taken as the two-dimensional coordinates of the refueling cone sleeve in the pixel coordinate system.
4. The refueling drogue positioning method based on visible light and infrared fusion according to claim 1 is characterized in that: The specific solution process for the estimated three-dimensional coordinates of the S40 refueling drogue is as follows: Before fusing the visible light imaging and infrared imaging results, the camera pose and coordinate system transformation need to be calculated for the two 2D coordinate points of the refueling drogue. Once the coordinates of the 3D space point and the coordinates of the 2D image feature points corresponding to the 3D space point are known, the camera pose is solved using the PnP algorithm. PnP algorithms include P3P, EPnP, UPnP, direct linear transformation, and nonlinear optimization. When using the direct linear transformation method to solve, for a given three-dimensional point , in the image, the projected feature point coordinates are ; At this time, the camera pose is the unknown quantity to be solved, and the corresponding projection equation is as follows: ; in, is the camera intrinsic parameter matrix, obtained through camera calibration; is the rotation matrix, is the translation vector; are the parameters in the camera intrinsic parameter matrix respectively; Camera pose After the solution is completed, the three-dimensional coordinates of the refueling drogue in different coordinate systems are converted to a unified coordinate system through coordinate system conversion; the three-dimensional coordinates of the refueling drogue are based on the center of the cone as the origin and follow the right-hand coordinate system principle. Direction vertically upward, Horizontal and vertical to the right, Facing outward, 8 LED infrared lamp beads are evenly distributed on the refueling cone sleeve in a three-dimensional coordinate system; Four cameras are distributed on the receiving aircraft. Two of them are infrared cameras, located on the left and right sides of the receiving aircraft fuselage; the other two are visible light cameras, located on the left and right wings of the receiving aircraft. The two-dimensional coordinates of the refueling drogue in the pixel coordinate system obtained by the four cameras correspond to the left infrared coordinate system, the right infrared coordinate system, the left visible light coordinate system, and the right visible light coordinate system, respectively. The two-dimensional coordinates of the refueling drogue in the pixel coordinate system obtained by the four cameras are transformed relative to the body coordinate system of the nose of the refueling aircraft, and then the three-dimensional coordinate estimation value of the refueling drogue is obtained by summing and taking the average value.
5. The refueling drogue positioning method based on visible light and infrared fusion according to claim 1 is characterized in that: The specific solution process of the three-dimensional coordinate prediction value of the refueling drogue of S50 is as follows: In order to ensure the smoothness of the output of the three-dimensional coordinate estimation value of the refueling drogue, the Kalman filter is used to filter the three-dimensional coordinate estimation value of the refueling drogue; the prediction and update equations of the Kalman filter are as follows: ; The order of Kalman filtering is as follows: S51. Initialization; Initialization is performed only once and provides two parameters: initial state , initial state variance , make predictions after initialization; S52.Measurement; The measurement process also provides two parameters: the measured value Measurement variance ; S53. Status update; The input to the state update process is: measurement value , measurement variance , state prediction and the state prediction variance Based on the input, the state update process calculates the Kalman gain and produces the output: state estimation and the state estimation variance ; S54. Forecast; The prediction process uses the system dynamics model to estimate the current state of the system and the state estimation variance Extrapolate to the next moment; in the first iteration, the initialization parameters are used as the value and variance of the state prediction, and the predicted output is used as the value and variance of the state prediction at the next moment; after the prediction, the three-dimensional coordinate prediction value of the refueling drogue that meets the preset accuracy requirements is obtained.