Fusion Tracking Method for Visible and Infrared Images during UAV Dynamic Flight
Through simplified ground calibration and running registration algorithms on airborne computers, combined with multi-scale descriptors and PID algorithms, the rapid and accurate registration and fusion of visible light and infrared images in drones is achieved, solving the problems of complex calibration and real-time tracking in the prior art, and improving the reliability of drone target tracking.
Patent Information
- Application Number
- CN202510256457.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-05
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-03-05
AI Technical Summary
The prior art lacks suitable image registration methods when using visible light and infrared cameras with zoom lenses in drones, resulting in complex ground calibration processes and real-time tracking difficulties.
A fusion tracking method for visible light and infrared images in dynamic flight of UAV is proposed, including ground calibration, rough field angle matching, image descriptor extraction and registration, image fusion and target tracking, etc., and fast and accurate image registration is achieved by simplifying the calibration process and using multi-scale descriptors and PID algorithms.
It significantly reduces the workload of image registration, improves the reliability and real-time of target tracking, and is suitable for visible light and infrared camera conditions of zoom lenses in drones.
Smart Images

Figure CN119762541B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of unmanned aerial vehicle (UAV) target tracking, and particularly to a method for fusing and tracking visible light and infrared images during the dynamic flight of a UAV. Background Art
[0002] Traditional image registration methods usually require a complex ground calibration process, which includes precise measurement of the offset of the optical axis centers of visible light and infrared cameras, edge distortion, and pixel mapping relationships. These calibration processes are often time-consuming and laborious, and are only applicable to the case of fixed focal length lenses. However, in the UAV airborne optoelectronic turret system, visible light and infrared cameras are usually equipped with variable focal length lenses, which makes the traditional calibration methods no longer applicable because they need to be calibrated separately for each focal length, resulting in a huge workload.
[0003] In addition, due to the dynamic nature of the UAV platform and the real-time requirements of tasks, traditional image registration methods face challenges in practical applications. When the UAV is performing monitoring tasks, it is normal for the camera focal length to change, which requires the image registration algorithm to be able to quickly adapt to the focal length change and achieve real-time processing.
[0004] The main limitations of the existing technology lie in its complex ground calibration process and the lack of an image registration solution applicable to variable focal length lenses; the lack of an algorithm that can adapt to the focal length changes of the variable focal length visible camera and variable focal length infrared camera of the UAV airborne optoelectronic turret and achieve fast and accurate registration. Summary of the Invention
[0005] To solve the above problems, the present invention provides a method for fusing and tracking visible light and infrared images during the dynamic flight of a UAV.
[0006] The object of the present invention is to provide a method for fusing and tracking visible light and infrared images during the dynamic flight of a UAV, which specifically includes the following steps:
[0007] S1. Ground calibration: Measure and calibrate the relationship between the focal length of the visible light camera and the field of view angle in the Y direction of the visible light, and the relationship between the focal length of the infrared camera and the field of view angle in the Y direction of the infrared on the ground, record the relationship between the field of view angle and the focal length as a file, and store it on the camera control board;
[0008] S2. Coarse matching of field of view angles: The camera control board controls the movement of the visible light camera or the infrared camera through the PID algorithm to make the field of view angle of a single pixel of the visible light image equal to the field of view angle of a single pixel of the infrared image;
[0009] S3. Extraction and registration of image descriptors: Extract the heterogeneous image descriptors of the visible light image and the infrared image, match the descriptors, store the projection relationship of each pixel coordinate of the visible light image and the infrared image as matrix A; call matrix A to complete the registration;
[0010] S4. Image fusion: According to the registered images, convert the visible light image format from RGB to YUV, overlay the corresponding brightness values of the infrared image on the Y channel, calculate the local contrast and information entropy of the visible light image and the infrared image, dynamically adjust the fusion ratio of the visible light image and the infrared image, and generate a registered and fused image;
[0011] S5. Input the registered and fused image into a relevant tracking algorithm for target tracking.
[0012] Preferably, the measurement and calibration method in step S1 specifically includes: using a folding visible-infrared dual-optical co-field optical tube to gradually measure the Y-direction field of view angle of the visible light camera and the Y-direction field of view angle of the infrared camera; the folding visible-infrared dual-optical co-field optical tube contains a crosshair, and the stepping length is 0.1 m.
[0013] Preferably, the method for adjusting the field of view angle of a single pixel of the visible light image to be equal to that of a single pixel of the infrared image in step S2 is as follows:
[0014] When the UAV performs a mission, if the visible lens is the main lens and the infrared is fused, the following steps are executed: taking the visible lens as the main lens, when the ground station sends the large-field-of-view and small-field-of-view commands, only the visible lens responds; after the ground station control ends, the visible lens is in a static state. At this time, read out the focal length value g of the visible light CCD ; According to the focal length value g of the visible light CCD 、the visible pixel size and the infrared pixel size, calculate the infrared focal length under the condition of the same field of view for a single pixel. The camera control board controls the repeated movement of the infrared camera through the PID algorithm until the visible light focal length is 1 / 6 of the infrared focal length;
[0015] When the UAV performs a mission, if the infrared lens is the main lens and the visible lens is fused, the following steps are executed: taking the infrared lens as the main lens, when the ground station sends the large-field-of-view and small-field-of-view commands, only the infrared lens responds; after the ground station control ends, the infrared lens is in a static state. At this time, read out the focal length value g of the infrared lens IR ; According to the focal length value g of the infrared lens IR , according to the visible pixel size and the infrared pixel size, calculate the visible light focal length under the condition of the same field of view for a single pixel. The camera control board controls the repeated movement of the visible light camera through the PID algorithm until the visible light focal length is 1 / 6 of the infrared focal length.
[0016] Preferably, the heterogeneous image descriptors of the visible light image and the infrared image in step S3 include a phase consistency descriptor and a gradient descriptor; the method of weighted correlation distance is used for matching; the specific method is as follows:
[0017] S301. Calculate the phase consistency descriptor of the heterogeneous images to measure the phase of specific frequency components in the heterogeneous images; for each pixel, the phase can be calculated by the following formula:
[0018] ;
[0019] where, is the pixel coordinate, w is different scales, is the phase at scale w, and M represents the number of sizes;
[0020] S302. Calculate the gradient descriptor of the heterogeneous images;
[0021] The gradient direction descriptor of the heterogeneous images is as follows:
[0022] ; ;
[0023] where, and are the gradients in the x - direction and y - direction at scale m respectively, I is the brightness of the image; thus, the gradient descriptor of the heterogeneous images is expressed as:
[0024]
[0025] ;
[0026] In the formula, represents the direction of the gradient, that is, the direction of the brightness change of the image at the point (x, y);
[0027] S303. Set the statistical scale as m, and the gradient histogram of 10×10 pixels centered on the (x, y) pixel is as follows:
[0028] ;
[0029] S304. Based on the target size of the image, take M / 2 scales upward and M / 2 scales downward, for a total of M scales; ; the phase consistency descriptor , the gradient descriptor , and the gradient histogram at the m - th scale form the descriptor vector at the m - th scale:
[0030] ;
[0031] S305. Use the method of weighted correlation distance for matching; the expression is as follows:
[0032] ;
[0033] Wherein, WCD represents the weighted correlation distance, that is, the weighted sum of the differences between the visible light image and the infrared image descriptor vectors from scale 1 to scale M; m is the scale of the descriptor; M represents the number of sizes; and respectively represent the multi-scale descriptor vectors at the m-th scale in the visible light image and the infrared image; is the weight of the m-th scale, expressed as: ; is the descriptor vector of the m-th scale and variance;
[0034] When the WCD value is greater than 0.5, it is considered that and match successfully; store the projection relationship of each pixel coordinate in the visible light image and the infrared image as matrix A.
[0035] Preferably, step S4 specifically includes the following sub-steps:
[0036] S401. Calculate the local contrast of the visible light image and the infrared image; normalize the local contrast of the visible light image and the infrared image to obtain the weights of the local contrast of the visible light image and the infrared image;
[0037] S402. Calculate the information entropy of the visible light image and the infrared image, normalize the information entropy of the visible light image and the infrared image to obtain the weights of the information entropy of the visible light image and the infrared image;
[0038] S403. Convert the RGB format file of the visible light image to YUV, and make the Y-channel brightness value Y ccd of the visible light image Y ir of the infrared image Y fused , the expression is as follows:
[0039] ;
[0040] Wherein, Y fused represents the registration fusion image, Y ccd represents the Y-channel brightness value of the visible light image, Y ir is the Y-channel brightness value of the infrared image; represents the weight of the visible light image based on the local contrast, represents the weight of the infrared image based on the local contrast; represents the weight of the visible light image based on information entropy, represents the weight of the infrared image based on information entropy.
[0041] Preferably, step S401 specifically includes the following sub-steps:
[0042] S4011. Determine the local area centered on the feature, apply a mean filter to the local area to obtain the local average brightness, calculate the standard deviation of the deviation between the pixel values in the local area and the local average brightness, and take the maximum value of the standard deviation as the local contrast;
[0043] The local standard deviation calculation formulas for the visible light image and the infrared image are as follows:
[0044] ;
[0045] ;
[0046] In the formula, is the pixel value of the visible image at the coordinate ( x , y ), is the pixel value of the infrared image at the coordinate ( x , y ), is the average pixel value of the local area Ω f in the visible light image or the infrared image, and ∣Ω f ∣ represents the total number of pixels in the local area Ω f in the visible light image or the infrared image;
[0047] The local contrast expressions for the visible light image or the infrared image are as follows:
[0048] CCDLocalContrast = max( )
[0049] IRLocalContrast = max( )
[0050] S4012. Normalize the local contrasts of the visible light image and the infrared image so that the sum of CCDLocalContrast and IRLocalContrast is 1; obtain the weights of the local contrasts of the visible light image and the infrared image:
[0051] ;
[0052] In the formula, represents the weight of the visible light image based on local contrast, represents the weight of the infrared image based on local contrast.
[0053] Preferably, step S402 specifically includes the following sub-steps:
[0054] S4021. Use a histogram to represent the number of times the gray value i appears in the visible light image or the infrared image; for the gray value of each pixel point in the gray image or , perform the following operations:
[0055] ;
[0056] In the formula, and represent the visible histogram and the infrared histogram; i represents the gray value; represents the indicator function; when ; is equal to i, is 1, otherwise it is 0; when is equal to i, is 1, otherwise it is 0; represents the gray value of the pixel point with coordinates (x, y) in the visible image, represents the gray value of the pixel point with coordinates (x, y) in the infrared image;
[0057] S4022. Perform normalization processing to convert the histogram into a probability distribution by dividing the frequency of each gray value by the total number of pixels in the image:
[0058] ;
[0059] ;
[0060] In the formula, and are the probability distributions of each gray value in the visible light image and the infrared image respectively, and are the width and height of the visible light image respectively, and are the width and height of the infrared image respectively;
[0061] S4023. Use the Shannon information entropy formula to multiply the probability of each gray value by its logarithm to the base 2 and sum over all 256 possible gray values to calculate the visible light image information entropy H ccd and the infrared image information entropy H ir; The calculation formula is as follows:
[0062] ;
[0063] ;
[0064] S4024. Normalize the information entropy of the visible light image and the infrared image so that H ccd and H the sum of ir is 1, obtaining the weights of the information entropy of the visible light image and the infrared image:
[0065] ;
[0066] In the formula, represents the weight of the visible light image based on the information entropy, represents the weight of the infrared image based on the information entropy.
[0067] Preferably, the value range of the information entropy is 0 to 8 bits; when all pixels have the same gray value, H ccd = 0, H ir = 0; when the probability of each gray value appearing is equal, the information entropy H ccd = 8, H ir = 8.
[0068] Preferably, the correlation tracking algorithm in step S5 specifically includes the following steps:
[0069] S501. Target area initialization: In the fused tracking image, determine the initial area of the target through the target detection algorithm, and use this initial area as the starting point of the tracking algorithm;
[0070] S502. Template matching and correlation calculation: Use the initial area of the target as the template, and search for the target by calculating the correlation between the template and each possible position in the image; the correlation can be calculated by the following formula:
[0071] ;
[0072] Among them, R ( x , y ) is the correlation score at the registration and fusion image position ( x , y ), T ( i , j ) is the pixel value of the fused template image, I ( x + i , y + j ) is the pixel value of the target fusion image, μT and μI are the means of the template and the target area respectively;
[0073] S503. Peak Detection and Target Location: In the correlation graph, the new position of the target is determined by finding the local maximum; if the peak is higher than the preset threshold, it is considered that the target is successfully tracked at this position, and the model of the target is updated according to the tracking result.
[0074] Compared with the prior art, the present invention can achieve the following beneficial effects:
[0075] The present invention only needs to simply calibrate the visible and infrared field of view angles on the ground. After zooming, on the airborne computer, run the registration algorithm, and combine multi-scale descriptors and an optimized search strategy to improve the efficiency and accuracy of feature matching. It is applicable to visible-infrared image registration and fusion tracking under different resolutions and when the visible and infrared cameras are zoom lenses, especially applicable to fields such as image registration and tracking of variable-focus visible cameras and variable-focus infrared cameras on unmanned aerial vehicle (UAV) airborne optoelectronic turrets, with broad application prospects and practical application value, significantly reducing the workload of registration calibration and improving the reliability of target tracking. Description of the Drawings
[0076] Figure 1 is a flowchart of a method for fusing and tracking visible and infrared images during the dynamic flight of a UAV according to an embodiment of the present invention.
[0077] Figure 2 is a visible light image taken by a UAV during aerial photography according to an embodiment of the present invention.
[0078] Figure 3 is an infrared image taken by a UAV during aerial photography according to an embodiment of the present invention.
[0079] Figure 4 is a fused registration image generated after fusing the visible light image and the infrared image on a UAV according to an embodiment of the present invention. Detailed Embodiments
[0080] In the following, embodiments of the present invention will be described with reference to the drawings. In the following description, the same modules are denoted by the same reference numerals. In the case of the same reference numerals, their names and functions are also the same. Therefore, their detailed descriptions will not be repeated.
[0081] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the drawings and specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and do not constitute a limitation to the present invention.
[0082] The present invention aims to solve the limitations of the prior art and provide a new image registration method, which is applicable to the integrated payload system of an airborne optoelectronic radar of an unmanned aerial vehicle (UAV), can reduce the calibration workload, and improve the efficiency and accuracy of image registration. By performing simple field-of-view calibration on the ground and combining multi-scale descriptors and an optimized search strategy, the registration algorithm can run in real time on an airborne computer to meet the requirements of real-time image processing of the UAV.
[0083] See Figure 1 , the present invention provides a method for fusing and tracking visible light and infrared images during the dynamic flight of a UAV, specifically including the following steps:
[0084] S1. Ground calibration: Install the visible light camera and the infrared camera on the airborne optoelectronic turret of the UAV; Measure and calibrate the relationship between the focal length of the visible light camera and the field of view angle in the Y direction of the visible light, and the relationship between the focal length of the infrared camera and the field of view angle in the Y direction of the infrared light on the ground, record the relationship between the field of view angle and the focal length as a file, and store it on the camera control board;
[0085] The measurement and calibration method specifically includes: Using a folding visible-infrared dual-light co-visual field optical tube (with a crosshair inside the optical tube and a step length of 0.1 m) to gradually measure the field of view angle in the Y direction of the visible light camera and the field of view angle in the Y direction of the infrared camera;
[0086] Specifically, the optoelectronic turret is internally equipped with a measured azimuth and elevation encoder; Assume that the focal length of the visible light camera lens at this time is CCD_F, the elevation angle of the optoelectronic turret is locked at 0, align the upper edge of the visible light camera with the crosshair of the optical tube, and read the azimuth angle CCDup of the optoelectronic turret at this time; Rotate the optoelectronic turret so that the lower edge of the visible light camera aligns with the crosshair of the optical tube, and read the azimuth angle CCDdown of the optoelectronic turret at this time. Then, when the focal length of the visible light camera is CCD_F, the field of view angle in the Y direction of the visible light is CCD_Y = CCDup - CCDdown; Record the corresponding relationship between the focal length of the visible light lens and the field of view angle into the array c; Similarly, measure the field of view angle in the Y direction of the infrared camera, IR_Y = IRup - IRdown; Record the corresponding relationship between the focal length of the infrared lens and the field of view angle into the array d, and store it in the recorder of the airborne optoelectronic turret of the UAV.
[0087] S2. Coarse matching of field of view angles: The camera control board controls the movement of the visible light camera or the infrared camera through the PID algorithm to make the field of view angle of a single pixel of the visible light image equal to the field of view angle of a single pixel of the infrared image; The specific operation is as follows:
[0088] If the field of view angles of each pixel of the visible light image and the infrared image are to be the same, the derivation of the focal length relationship between the visible light lens and the infrared lens is as follows:
[0089] Visible light pixel size: ; Infrared pixel size: ; The focal length of the visible light lens is fvi ; The focal length of the infrared lens is fir ;
[0090] The angular resolution θ can be expressed by the following formula: θ = pixel size / focal length;
[0091] The angular resolution of the visible light image is θ vis = dvis / fvis ; The angular resolution of the infrared image is ;
[0092] To make θ vis equal to θ ir , the following equation can be established: ; ; That is, in the case of a visible light pixel size of 2.5 μm and an infrared pixel size of 15 μm, that is, when the visible light focal length is 1 / 6 of the infrared focal length, the field of view angle of a single pixel in the visible light image is equal to the field of view angle of a single pixel in the infrared image.
[0093] The visible light camera and the infrared camera adopt an external trigger method to ensure that the visible light and infrared are exposed at the same time.
[0094] Adjust the field of view angle of a single pixel in the visible light image to be equal to the field of view angle of a single pixel in the infrared image. The method of adjusting the visible light focal length to be 1 / 6 of the infrared focal length is as follows:
[0095] When the drone performs a mission, if the visible light lens is the main one and the infrared is fused, the following steps are executed: Taking the visible light lens as the main lens, when the ground station sends commands for a large field of view and a small field of view, only the visible light lens responds; after the ground station control ends, the visible light lens is in a stationary state. At this time, read out the focal length value g of the visible light CCD ; According to the focal length value g of the visible light CCD , the visible pixel size and the infrared pixel size, calculate the infrared focal length under the condition of the same field of view for a single pixel. The camera control board controls the repeated movement of the infrared camera through the PID algorithm until the visible light focal length is 1 / 6 of the infrared focal length;
[0096] When the drone performs a mission, if the infrared lens is the main one and the visible light lens is fused, the following steps are executed: Taking the infrared lens as the main lens, when the ground station sends commands for a large field of view and a small field of view, only the infrared lens responds; after the ground station control ends, the infrared lens is in a stationary state. At this time, read out the focal length value g of the infrared lens IR ; According to the focal length value g of the infrared lens IR , according to the visible pixel size and the infrared pixel size, calculate the visible light focal length under the condition of the same field of view for a single pixel. The camera control board controls the repeated movement of the visible light camera through the PID algorithm until the visible light focal length is 1 / 6 of the infrared focal length.
[0097] S3. Image descriptor extraction and registration: Extract heterologous image descriptors of visible light images and infrared images, match the descriptors, and store the projection relationship of each pixel coordinate of the visible light image and the infrared image as matrix A; call matrix A to complete the registration.
[0098] The heterologous image descriptors of visible light images and infrared images include phase consistency descriptors and gradient descriptors; the method of matching the descriptors is to combine the phase consistency descriptors and gradient descriptors, extract the general feature descriptors of visible light images and infrared images, and use the method of weighted correlation distance (WCD) for matching; the specific method is as follows:
[0099] S301. Calculate the phase consistency descriptor of the heterologous image to measure the phase of specific frequency components in the heterologous image; for each pixel, the phase can be calculated by the following formula:
[0100] ;
[0101] where is the pixel coordinate, w is different scales, is the phase at scale w, and M represents the number of sizes (usually taking the value of 10);
[0102] S302. Calculate the gradient descriptor of the heterologous image;
[0103] The gradient direction descriptor of the heterologous image is as follows:
[0104] ; ;
[0105] where and are the gradients in the x - direction and y - direction at scale m respectively, I is the brightness of the image; thus, the gradient descriptor of the heterologous image is expressed as:
[0106]
[0107] ;
[0108] In the formula, represents the direction of the gradient, that is, it represents the direction of the brightness change of the image at the point (x,y);
[0109] S303. Set the statistical scale as m, and the gradient histogram of 10×10 pixels around the pixel (x,y) is represented as follows:
[0110] ;
[0111] S304. Based on the image target size, take M / 2 scales upward and M / 2 scales downward, for a total of M scales. The m-th scale The phase consistency descriptor , the gradient descriptor (gradient magnitude) , and the gradient histogram constitute the descriptor vector at the m-th scale :
[0112] ;
[0113] Both the visible light image and the infrared image are operated according to the above method, and a rich multi-scale descriptor vector and are generated for each feature point in the visible light image and the infrared image obtained. This descriptor effectively describes the local structure information of the visible image and the infrared image and provides a solid foundation for feature matching.
[0114] S305. Use the method of weighted correlation distance (WCD) for matching; the expression is as follows:
[0115] ;
[0116] In the formula, WCD represents the weighted correlation distance, that is, the weighted sum of the differences between the descriptor vectors of the visible light image and the infrared image from scale 1 to scale M; m is the scale of the descriptor (the value range of m is from 1 to M); M represents the number of sizes (usually taking the value of 10); and respectively represent the multi-scale descriptor vectors at the m-th scale in the visible light image and the infrared image; is the weight of the m-th scale, expressed as: ; is the descriptor vector of the m-th scale and is the variance;
[0117] When the WCD value is greater than 0.5, it is considered that and are successfully matched; the projection relationship of each pixel coordinate in the visible light image and the infrared image is stored as matrix A.
[0118] Principle: After step S2, the visible and infrared fields of view are basically consistent in the Y direction in theory. However, due to the field of view angle calibration error in step S1 and the focus control error in step S2, there are still slight differences in the visible and infrared fields of view, so registration is required. The calculation of the common feature descriptors of visible light images and infrared images is a key step in the matching of visible light images and infrared images. It provides a unique vector representation for each feature point of the visible light image and the infrared image.
[0119] The weighted correlation distance (WCD) method not only considers the Euclidean distance between feature vectors, but also the correlation between feature dimensions. This method is particularly suitable for situations where there are intrinsic connections between feature dimensions.
[0120] Because visible infrared uses external triggering to ensure that visible infrared is exposed at the same time, after exposure, a new frame of visible light image and infrared image arrive at the onboard computer. Only by calling matrix A can the real-time registration of visible infrared be completed, so that the real-time performance of the registration is effectively guaranteed.
[0121] S4. Image fusion: According to the registered image, convert the visible light image format from RGB to YUV, and superimpose the corresponding brightness value of the infrared image on the Y channel, calculate the local contrast and information entropy of the visible light image and the infrared image, dynamically adjust the fusion ratio of the visible light image and the infrared image, and generate a registered fused image; specifically, it includes the following sub-steps:
[0122] S401. Calculate the local contrast of the visible light image and the infrared image; normalize the local contrast of the visible light image and the infrared image to obtain the weight of the local contrast of the visible light image and the infrared image; specifically include the following sub-steps:
[0123] S4011. Determine a local area centered on the feature, apply a mean filter to the local area to obtain the local average brightness, calculate the standard deviation of the deviation between the pixel values in the local area and the local average brightness, and take the maximum value of the standard deviation as the local contrast.
[0124] Local standard deviation It is a local area The measure of the internal pixel value fluctuation, the local standard deviation calculation formulas for visible light images and infrared images are as follows:
[0125] ;
[0126] ;
[0127] In the formula, is the visible image at coordinates ( x , y ), is the pixel value of the infrared image at the coordinate ( x , y ); is the average pixel value of the local region Ω in the visible light image or the infrared image f . |Ω f | represents the total number of pixels in the local region Ω in the visible light image or the infrared image f .
[0128] The local standard deviation is calculated by computing the sum of the squares of the deviations of the pixel values within the local region Ω f from its average value μ f , and then taking the square root. This value reflects the degree of dispersion of the pixel intensities within the feature region and is an indicator of the local texture complexity.
[0129] The local contrast of the visible light image or the infrared image is defined as the maximum value of its local standard deviation:
[0130] CCDLocalContrast = max( );
[0131] IRLocalContrast = max( );
[0132] S4012. Normalize the local contrasts of the visible light image and the infrared image such that the sum of CCDLocalContrast and IRLocalContrast is 1; obtain the weights of the local contrasts of the visible light image and the infrared image:
[0133] ;
[0134] wherein, represents the weight of the visible light image based on the local contrast, represents the weight of the infrared image based on the local contrast.
[0135] S402. Calculate the information entropy of the visible light image and the information entropy of the infrared image, normalize the information entropies of the visible light image and the infrared image, and obtain the weights of the information entropies of the visible light image and the infrared image; specifically including the following sub-steps:
[0136] S4021. Use a histogram to represent the number of occurrences of the gray value i in the visible light image or the infrared image; for the gray value of each pixel point in the gray image or , perform the following operations:
[0137] ;
[0138] Wherein, and represent the visible histogram and the infrared histogram; i represents the gray value; represents the indicator function; when ; equals i, is 1, otherwise 0; when equals i, is 1, otherwise 0; represents the gray value of the pixel at coordinates (x, y) in the visible image, represents the gray value of the pixel at coordinates (x, y) in the infrared image;
[0139] S4022. Perform normalization to convert the histogram into a probability distribution by dividing the frequency of each gray value by the total number of pixels in the image:
[0140] ;
[0141] ;
[0142] Wherein, and are the probability distributions of each gray value of the visible light image and the infrared image respectively, and are the width and height of the visible light image respectively, and are the width and height of the infrared image respectively.
[0143] S4023. Use the Shannon information entropy formula to multiply the probability of each gray value by its logarithm to the base 2 and accumulate over all 256 possible gray values to calculate the visible light image information entropy H ccd and the infrared image information entropy H ir;
[0144] ;
[0145] ;
[0146] The range of the information entropy is from 0 to the maximum value of 8 bits, and this range reflects the amount of information in the image from completely disordered (the gray value of each pixel is random and has equal probability) to completely ordered (all pixels have the same gray value); when all pixels have the same gray value (i.e., the image is completely uniform), H ccd = 0, H ir = 0; when the probability of each gray value appearing is equal, the information entropy H ccd = 8, H ir = 8, reaching the maximum value.
[0147] S4024. Normalize the information entropy of the visible light image and the infrared image so that H ccd and H the sum of ir is 1 to obtain the weights of the information entropy of the visible light image and the infrared image:
[0148] ;
[0149] In the formula, represents the weight of the visible light image based on the information entropy, represents the weight of the infrared image based on the information entropy.
[0150] S403. Convert the RGB format file of the visible light image to YUV, and make the Y-channel brightness value of the visible light image Y ccd be superimposed with the Y-channel brightness value of the infrared image Y ir, and perform pixel-by-pixel fusion to generate a registered fusion image Y fused , and the expression is as follows:
[0151] ;
[0152] In the formula, Y fused represents the registered fusion image, Y ccd represents the Y-channel brightness value of the visible light image, Y ir is the Y-channel brightness value of the infrared image; represents the weight of the visible light image based on the local contrast, represents the weight of the infrared image based on the local contrast; represents the weight of the visible light image based on the information entropy, represents the weight of the infrared image based on the information entropy.
[0153] S5. Input the registered fusion image into the relevant tracking algorithm for target tracking; the relevant tracking algorithm specifically includes the following steps:
[0154] S501. Target area initialization: In the fusion tracking image, determine the initial area of the target through the target detection algorithm, and use this initial area as the starting point of the tracking algorithm;
[0155] S502. Template matching and correlation calculation: Use the initial area of the target as the template, and search for the target by calculating the correlation between the template and each possible position in the image; the correlation can be calculated by the following formula:
[0156] ;
[0157] Among them, R (x , y ) is the correlation score at the position of the registered and fused image ( x , y ), T ( i , j ) is the pixel value of the fused template image, I ( x + i , y + j ) is the pixel value of the target fused image, μT and μI are the means of the template and target regions respectively;
[0158] S503. Peak Detection and Target Location: In the correlation map, the new position of the target is determined by finding the local maximum; if the peak is higher than a preset threshold, it is considered that the target is successfully tracked at this position, and the model of the target is updated according to the tracking result.
[0159] In summary, the present invention simplifies the ground calibration process. Only the field of view angles and imaging parameters of the visible light camera and the infrared are required to be calibrated. After the zoom or parameter change, the registration algorithm is run on the on-board computer, the heterologous image descriptors are extracted, the descriptors of the visible light and infrared images are matched, and the projection relationship of each pixel coordinate of the visible light and infrared is stored as matrix A. When the focal length or parameter configuration remains unchanged, matrix A only needs to be calculated once and stored in the on-board computer.
[0160] Since the visible light and infrared use the synchronous trigger mode to ensure that the visible light and infrared acquire data at the same moment. After the data is acquired, a new frame of visible light image and infrared image arrive at the on-board computer simultaneously. Only by calling matrix A can the real-time registration of the visible light and infrared be completed, effectively guaranteeing the real-time performance of the registration. After each frame of image is registered by matrix A, the visible light video format is converted from RGB to YUV, and the corresponding brightness value of the infrared is superimposed on the Y channel. The local contrast and information entropy of the visible light and infrared are calculated, and the fusion ratio of the visible light Y and infrared brightness values is dynamically adjusted according to the local contrast and information entropy to maximize the information amount of the fused image.
[0161] The method of the present invention is applicable to image registration under different resolutions and conditions where the visible light camera and the infrared have variable focal lengths or variable parameter configurations, especially applicable to fields such as the registration of variable focal length visible light cameras and infrared images of the integrated payload of an unmanned aerial vehicle's optoelectronic radar, etc. It has a wide range of application prospects and practical application values, and significantly reduces the workload of registration calibration.
[0162] It should be understood that the various forms of processes shown above can be used, with steps reordered, added or deleted. For example, the steps described in the disclosure of the present invention can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solution disclosed in the present invention can be achieved, and no limitation is imposed herein.
[0163] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A fusion tracking method of visible light and infrared images in dynamic flight of unmanned aerial vehicle, characterized in that: The specific steps include: S1. Ground calibration: measure and calibrate the relationship between the focal length of the visible light camera and the visible light Y-direction field angle and the relationship between the focal length of the infrared camera and the infrared Y-direction field angle on the ground, record the relationship between the field angle and the focal length as a file, and store it on the camera control board; the measurement and calibration method specifically includes: using a reentrant visible infrared dual-light common field light tube to gradually measure the visible light camera Y-direction field angle and the infrared camera Y-direction field angle; the reentrant visible infrared dual-light common field light tube contains a crosshair, and the step length is 0.1m; S2. Coarse matching of field of view angle: The camera control board controls the movement of the visible light camera or infrared camera through the PID algorithm, so that the field of view angle of a single pixel of the visible light image is equal to the field of view angle of a single pixel of the infrared image; S3. Image descriptor extraction and registration: extract heterogeneous image descriptors of visible light image and infrared image, match the descriptors, store the projection relationship of each pixel coordinate of visible light image and infrared image as matrix A; call matrix A to complete the registration; S4. Image fusion: According to the registered image, convert the visible light image format from RGB to YUV, and superimpose the corresponding brightness value of the infrared image on the Y channel, calculate the local contrast and information entropy of the visible light image and the infrared image, dynamically adjust the fusion ratio of the visible light image and the infrared image, and generate a registered fused image; specifically, it includes the following sub-steps: S401. Calculate the local contrast of the visible light image and the infrared image; normalize the local contrast of the visible light image and the infrared image to obtain the weight of the local contrast of the visible light image and the infrared image; S402. Calculate the information entropy of the visible light image and the information entropy of the infrared image, normalize the information entropy of the visible light image and the infrared image, and obtain the weights of the information entropy of the visible light image and the infrared image; S403. Convert the RGB format file of the visible light image to YUV, and make the Y channel brightness value of the visible light image Y CCD and infrared image Y channel brightness value Y IR superposition, pixel-by-pixel fusion, to generate a registered fused image Y fused , the expression is as follows: ; In the formula, Y fused represents the registered fusion image, Y CCD represents the Y channel brightness value of the visible light image. Y ir is the Y channel brightness value of the infrared image; represents the visible light image weight based on local contrast, represents the infrared image weight based on local contrast; represents the weight of the visible light image based on information entropy, represents the weight of infrared image based on information entropy; S5. Input the registered fused image into the relevant tracking algorithm to perform target tracking.
2. The method for fusion tracking of visible light and infrared images in dynamic flight of a UAV according to claim 1 is characterized in that: The method for adjusting the field of view angle of a single pixel of the visible light image to be equal to the field of view angle of a single pixel of the infrared image in step S2 is specifically as follows: When the UAV performs a mission, if the visible lens is used as the main lens and infrared is integrated, the following steps are performed: when the visible lens is used as the main lens and the ground station sends a large field of view or small field of view command, only the visible lens responds; after the ground station control ends, the visible lens is in a static state. At this time, the focal length value g of the visible light is read out. CCD ; According to the focal length of visible light g CCD , visible pixel size and infrared pixel size, calculate the infrared focal length of a single pixel under the same field of view, and the camera control board controls the repeated movement of the infrared camera through the PID algorithm until the visible light focal length is 1 / 6 of the infrared focal length; When the UAV is performing a mission, if the infrared lens is used as the main lens and the visible lens is integrated, the following steps are performed: when the infrared lens is used as the main lens and the ground station sends a large field of view or a small field of view command, only the infrared lens responds; after the ground station control ends, the infrared lens is in a static state. At this time, the focal length value g of the infrared lens is read out. IR ; According to the focal length value g of the infrared lens IR , according to the visible pixel size and the infrared pixel size, the visible light focal length of a single pixel under the same field of view is calculated. The camera control board controls the repeated movement of the visible light camera through the PID algorithm until the visible light focal length is 1 / 6 of the infrared focal length.
3. The method for fusion tracking of visible light and infrared images in dynamic flight of a UAV according to claim 1 is characterized in that: In step S3, the heterogeneous image descriptors of the visible light image and the infrared image include a phase consistency descriptor and a gradient descriptor; the weighted correlation distance method is used for matching; the specific method is as follows: S301. Calculate the phase consistency descriptor of the heterogeneous image, which is used to measure the phase of a specific frequency component in the heterogeneous image; for each pixel, the phase can be calculated by the following formula: ; in, are pixel coordinates, w are different scales, is the phase at scale w, M represents the number of sizes; S302. Calculate the gradient descriptor of the heterogeneous image; The gradient direction descriptor of the heterogeneous image is as follows: ; ; in, and They are the gradients in the x and y directions when the scale is m, I is the brightness of the image; thus, the gradient descriptor of the heterogeneous image is expressed as: ; In the formula, Indicates the direction of the gradient, that is, the direction of the brightness change of the image at the point (x, y); S303. Set the statistical scale to m, and take the (x, y) pixel as the center and the gradient histogram of the surrounding 10×10 pixels as follows: ; S304. Based on the image target size, take M / 2 scales upward and M / 2 scales downward, for a total of M scales; ; Phase consistency descriptor at the mth scale , gradient descriptor , gradient histogram It constitutes the descriptor vector at the mth scale : ; S305. Use the weighted correlation distance method to perform matching; the expression is as follows: ; Where WCD represents the weighted correlation distance, which is the weighted sum of the differences between the descriptor vectors of the visible light image and the infrared image from scale 1 to scale M; m is the scale of the descriptor; M represents the number of sizes; and Represent the multi-scale descriptor vector at the mth scale in the visible light image and infrared image respectively; is the weight of the mth scale, expressed as: ; is the descriptor vector of the mth scale and The variance of When the WCD value is greater than 0.5, it is considered and The match is successful; the projection relationship between the coordinates of each pixel in the visible light image and the infrared image is stored as matrix A.
4. The method for fusion tracking of visible light and infrared images in dynamic flight of a UAV according to claim 1 is characterized in that: The step S401 specifically includes the following sub-steps: S4011. Determine a local area centered on the feature, apply a mean filter to the local area to obtain the local average brightness, calculate the standard deviation of the deviation between the pixel value in the local area and the local average brightness, and take the maximum value of the standard deviation as the local contrast; The calculation formulas for the local standard deviation of visible light images and infrared images are as follows: ; ; In the formula, is the visible image at coordinates ( x , y ), is the infrared image at coordinates ( x , y ), is the local area Ω in the visible light image or infrared image f The average pixel value, |Ω f ∣ represents the local area Ω in the visible light image or infrared image f The total number of pixels; The local contrast expression of a visible light image or infrared image is as follows: CCDLocalContrast=max( ); IRLocalContrast=max( ); S4012. Normalize the local contrast of the visible light image and the infrared image so that the sum of CCDLocalContrast and IRLocalContrast is 1; obtain the weight of the local contrast of the visible light image and the infrared image: ; In the formula, represents the visible light image weight based on local contrast, Represents the infrared image weight based on local contrast.
5. The method for fusion tracking of visible light and infrared images in dynamic flight of a UAV according to claim 4, characterized in that: The step S402 specifically includes the following sub-steps: S4021. Use a histogram to represent the number of times the gray value i appears in the visible light image or infrared image; for each pixel in the gray image, the gray value or , do the following: ; In the formula, and Represents visible histogram and infrared histogram; i represents grayscale value; represents the indicator function; when ; when it is equal to i, is 1, otherwise it is 0; When it is equal to i, is 1, otherwise it is 0; Represents the grayscale value of the pixel with coordinates (x, y) in the visible image. Represents the grayscale value of the pixel with coordinates (x, y) in the infrared image; S4022. Perform normalization to convert the histogram into a probability distribution by dividing the frequency of each gray value by the total number of pixels in the image: ; ; In the formula, and are the probability distribution of each gray value of the visible light image and infrared image, and are the width and height of the visible light image, respectively. and are the width and height of the infrared image respectively; S4023. Using the Shannon information entropy formula, multiply the probability of each gray value by its logarithm with base 2, and accumulate all 256 possible gray values to calculate the information entropy of the visible light image. H CCD and infrared image information entropy H ir; the calculation formula is as follows: ; ; S4024. Normalize the information entropy of the visible light image and the infrared image so that H CCD and H ir is 1, and the weights of the information entropy of the visible light image and the infrared image are obtained: ; In the formula, represents the weight of the visible light image based on information entropy, Represents the infrared image weight based on information entropy.
6. The method for fusion tracking of visible light and infrared images in dynamic flight of a UAV according to claim 5, characterized in that: The value range of the information entropy is 0 to 8 bits; when all pixels have the same grayscale value, H CCD = 0, H ir = 0; when the probability of each gray value appearing is equal, the information entropy H CCD = 8, H ir =8.
7. The method for fusion tracking of visible light and infrared images in dynamic flight of a UAV according to claim 1, characterized in that: The correlation tracking algorithm in step S5 specifically includes the following steps: S501. Target region initialization: In the fused tracking image, the initial region of the target is determined by the target detection algorithm, and the initial region is used as the starting point of the tracking algorithm; S502. Template matching and correlation calculation: Use the initial area of the target as a template and search for the target by calculating the correlation between the template and each possible position in the image; the correlation can be calculated by the following formula: ; in, R ( x , y ) is at the position of the registered fusion image ( x , y ), T ( i , j ) is the pixel value of the fused template image, I ( x + i , y + j ) is the pixel value of the target fused image, μT and μI are the means of the template and target regions, respectively; S503. Peak detection and target positioning: In the correlation graph, the new position of the target is determined by finding the local maximum; if the peak value is higher than the preset threshold, the target is considered to be successfully tracked at this position, and the target model is updated according to the tracking result.
Citation Information
Patent Citations
Optimized registration method and system for different-source images
CN113793372A