Unmanned aerial vehicle infrared light and visible light fusion automatic calibration method

Through the combination of SIFT+KNN algorithm and spatial structure similarity algorithm, the problems of single information and sparse point clouds in the infrared visible light fusion calibration of drone are solved, achieving more efficient image registration and error reduction.

CN120278926APending Publication Date: 2025-07-08DEHONG POWER SUPPLY BUREAU OF YUNNAN POWER GRID CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311699159.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-12
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The traditional drone infrared visible light fusion calibration method has single information, lacks texture characteristics, and some targets are incomplete and the distance point clouds are sparse, making it difficult to achieve the completeness and density of point clouds.

Method used

The SIFT+KNN algorithm is used to obtain the two-dimensional key points of infrared light and visible light images, and the depth information of the two-dimensional key points is back-projected to the three-dimensional space. The three-dimensional space key points are obtained through the spatial structure similarity algorithm, and the three-dimensional space key points are connected in series to achieve automatic image calibration of infrared light and visible light.

Benefits of technology

It significantly improves the operating speed and registration effect of the automatic calibration of infrared visible light fusion of drones, reducing rotation and translation errors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120278926A_ABST
    Figure CN120278926A_ABST
Patent Text Reader

Abstract

The invention provides an automatic calibration method for infrared light and visible light fusion of an unmanned aerial vehicle. The automatic calibration method comprises the following steps: acquiring two-dimensional key points of infrared light and visible light images by using an SIFT + KNN algorithm; according to depth camera information of the two-dimensional key points, back-projecting to a three-dimensional space; obtaining three-dimensional space key points according to a spatial structure similarity algorithm; and the three-dimensional space key points are connected in series to realize automatic image calibration of infrared light and visible light The invention provides a point cloud registration fusion algorithm based on visual key point space structure similarity for solving the problems of low operation speed and poor registration effect when a traditional point cloud registration algorithm is used for processing a power grid configuration scene point cloud, and the method comprises the following steps: firstly, obtaining a two-dimensional key point pair of a visual image by using an SIFT + KNN algorithm; then estimating depth information of the two-dimensional key points and back-projecting the depth information to a three-dimensional space, and then obtaining several groups of three-dimensional key point pairs with the highest confidence degree through a space structure similarity algorithm to calculate rigid transformation parameters;
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of infrared and visible light fusion optics, and particularly relates to an automatic calibration method for infrared and visible light fusion of drones. Background Technique

[0002] In recent years, to meet the demand for electricity in national economic and social development, new requirements have been put forward for the stability of power grid equipment operation. However, the traditional calibration method has two main defects:

[0003] (1) Single information and lack of texture features. Ordinary optical cameras obtain visual images with dense color information, and it is impossible to fuse point cloud-visual information at the pixel level through calibration technology to break through the limitations of data;

[0004] (2) Some targets are incomplete and the far-view point cloud is sparse. It is difficult to fuse multiple frames of point clouds through point cloud registration technology by the traditional calibration method to achieve the integrity and densification of the point cloud.

[0005] To solve the above problems, an automatic calibration method for infrared and visible light fusion of drones is specifically proposed. Summary of the Invention

[0006] The embodiment of this application provides an automatic calibration method for infrared and visible light fusion of drones, which is a point cloud registration and fusion algorithm based on the spatial structure similarity of visual key points. First, use the SIFT+KNN algorithm to obtain the two-dimensional key point pairs of the visual image, then estimate the depth information of the two-dimensional key points and back-project them into the three-dimensional space, and then obtain several groups of three-dimensional key point pairs with the highest confidence through the spatial structure similarity algorithm to calculate the rigid transformation parameters. Experiments prove that the method in this paper has achieved a large improvement in both running speed and registration effect, and solves the problems of single information, lack of texture features, and incomplete and sparse far-view point clouds in the existing automatic calibration methods for infrared and visible light fusion of drones.

[0007] The embodiment of this application provides an automatic calibration method for infrared and visible light fusion of drones, including:

[0008] Use the SIFT and KNN algorithms to obtain the two-dimensional key points of the infrared and visible light images;

[0009] Back-project according to the depth camera information of the two-dimensional key points into the three-dimensional space;

[0010] Obtain the key points in the three-dimensional space according to the spatial structure similarity algorithm;

[0011] Connect the key points in the three-dimensional space in series to realize the automatic calibration of the infrared and visible light images.

[0012] In a feasible implementation manner, the method for obtaining two-dimensional key points of infrared light and visible light images using the SIFT and KNN algorithms includes:

[0013] Select the center point, and search for non-zero value points within the 8×8 neighborhood; calculate the two-dimensional Euclidean distance and direction from all neighborhoods to the center point, and save them.

[0014] According to the angle threshold:

[0015] [(-22.5° - 22.5°], (22.5° - 67.5°], (67.5° - 112.5°], (112.5° - 157.5°], (157.5° - 202.5°], (202.5° - 247.5°], (247.5° - 292.5°], (292.5° - 337.5°]], classify all neighborhood points according to these eight threshold intervals.

[0016] Record the point with the minimum Euclidean distance in each category as the only neighborhood point within the threshold, and finally obtain eight uniformly angular neighborhood points, and save their Euclidean distances and direction angles to the center point.

[0017] In a feasible implementation manner, the three-dimensional space coordinate system adopts Cartesian coordinates, expressed as (Xw, Yw, Zw), where the X-axis and Y-axis are parallel to the X-axis and Y-axis of the image coordinate axes, and the optical axis of the camera is the Z-axis;

[0018] For a two-dimensional space point, (x, y) is its coordinate point, and the conversion relationship between the camera coordinate system and the three-dimensional coordinate system is:

[0019]

[0020] where: f is the camera focal length, and cx and cy are the coordinates of the center of the imaging plane, belonging to the depth camera information parameters.

[0021] In a feasible implementation manner, the relationship between the projection coordinate system and the camera coordinate system is:

[0022]

[0023] where: R and T are rotation and translation parameters.

[0024] In a feasible implementation manner, the method for obtaining three-dimensional space key points according to the spatial structure similarity algorithm is:

[0025] First, use the uniform angle search algorithm to obtain the specific coordinates of eight neighborhood points; in the clockwise direction, compare the pixel value relationship between these eight neighborhood points and the center, record those larger than the center pixel value as 0, otherwise record as 1, and obtain an eight-bit binary number.

[0026] Shift this eight - bit binary number to the right until the least significant bit is not 0. Complement the shifted higher bits with 0, and convert the above - mentioned binary number into a decimal number as the eigenvalue of the central pixel point;

[0027] The right - shift of the binary number is to endow the feature with rotational invariance.

[0028] In a feasible implementation, the three - dimensional space point projection method is as follows:

[0029] First, use the uniform - angle search algorithm to obtain eight neighborhood points;

[0030] Reconstruct the two - dimensional pixel point p(ai, vi) into a three - dimensional space point P(xi, yi, zi);

[0031] Select three points that can form an approximate triangle in a clockwise direction to obtain a virtual plane, and calculate the unit normal vector;

[0032] Finally, obtain eight groups of unit normal vectors to get the three - dimensionally calibrated image;

[0033] The calculation method of the normal vector n is as follows:

[0034]

[0035] Where: PA, PB, and PC are the three points of the approximate triangle.

[0036] An infrared - light and visible - light fusion test was carried out. In order to achieve end - to - end training, the work of point - cloud projection, visual depth estimation, and homogeneous - feature calculation was added to the model's convolutional module in the form of non - weight layers and cuda operations were used, which greatly improved the calculation speed. During model training, the hardware was NVidia GTX1080ti, the BatchSize was set to 4, the weights of rotation and translation losses were set to 0.2 and 0.8, the maximum number of training Epochs was set to 60, the learning rate was initialized using an adaptive method, starting from 0.0015, and decreasing by 0.00015 every 10 training Epochs until 0.00005 and then no longer decreasing. During the model - training process, the rotation / translation error changes of the training set and the validation set are shown in the figure. It can be seen from the figure that when the training reaches 40 rounds, it has converged. The 45th - round model is selected as the optimal model. At this time, the loss of the training set is 0.051, and the loss of the test set is 0.059.

[0037] Set the initial errors: θ = 10, D = 2m, generate the training set and the validation set, and compare with the existing deep - learning - based point - cloud - image calibration methods: RegNet, CalibNet, and CMRNet. The corrected rotation error and translation error are shown in the above table.

[0038] The results show that the method of the present invention is superior to the past methods in both rotational and translational errors. Especially in the translational result, the error after calibration is significantly reduced compared with other methods.

[0039] An automatic calibration method for fusing infrared light and visible light of an unmanned aerial vehicle provided by an embodiment of the present application obtains two-dimensional key points of infrared light and visible light images by using the SIFT+KNN algorithm; back-projects according to the depth camera information of the two-dimensional key points into three-dimensional space; obtains three-dimensional space key points according to the spatial structure similarity algorithm; and cascades the three-dimensional space key points to realize automatic calibration of infrared light and visible light images. In view of the problems of slow running speed and poor registration effect of traditional point cloud registration algorithms when processing point clouds in power grid configuration scenarios, this paper proposes a point cloud registration and fusion algorithm based on the spatial structure similarity of visual key points. First, use the SIFT+KNN algorithm to obtain pairs of two-dimensional key points of visual images, then estimate the depth information of the two-dimensional key points and back-project them into three-dimensional space, and then obtain several pairs of three-dimensional key points with the highest confidence through the spatial structure similarity algorithm to calculate the rigid transformation parameters. Experiments prove that the method in this paper has achieved a large improvement in both running speed and registration effect. Description of the Drawings

[0040] Figure 1 is a flowchart of the automatic calibration method for fusing infrared light and visible light of an unmanned aerial vehicle provided by the present application;

[0041] Figure 2 is a schematic diagram for determining two-dimensional space points in an automatic calibration method for fusing infrared light and visible light of an unmanned aerial vehicle;

[0042] Figure 3 is a registration diagram of three-dimensional space key points in an automatic calibration method for fusing infrared light and visible light of an unmanned aerial vehicle;

[0043] Figure 4 is an angular error diagram in an automatic calibration method for fusing infrared light and visible light of an unmanned aerial vehicle;

[0044] Figure 5 is a translational error diagram in an automatic calibration method for fusing infrared light and visible light of an unmanned aerial vehicle. Detailed Embodiments

[0045] In order to enable those skilled in the art to better understand the technical solutions in the present application, the following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0046] Drones can capture infrared images and visible light images by using dual-light cameras or by carrying infrared cameras and visible light cameras.

[0047] Infrared cameras can capture the infrared radiation emitted by objects and convert it into images. This technology is commonly used for shooting at night or in low-light conditions because it can detect the heat emitted by objects, thus showing the outlines and shapes of the objects.

[0048] Visible light cameras, on the other hand, can capture the visible light reflected by objects and convert it into images. Such cameras are usually used for shooting during the day or in high-light conditions and can provide clearer and more vivid images.

[0049] In practical applications, drones can carry both infrared cameras and visible light cameras simultaneously in order to obtain more comprehensive information under different lighting conditions. For example, in search and rescue missions, drones can use infrared cameras to search for missing persons at night or in low-light conditions, and then use visible light cameras to obtain more detailed images during the day or in high-light conditions.

[0050] The infrared and visible light fusion of drones can fuse the data from the infrared light and visible light sensors on the drones and automatically calibrate this data to improve the quality and accuracy of the images.

[0051] This technology usually uses computer vision and machine learning algorithms to achieve. It can analyze and process infrared light and visible light images to determine the corresponding relationships between them and calibrate them. This can eliminate errors caused by factors such as sensor position, angle, and environmental factors, thereby improving the quality and accuracy of the images.

[0052] The automatic calibration technology for the infrared and visible light fusion of drones has extensive applications in fields such as military, security, fire protection, and agriculture. It can help users obtain more accurate information about the targets and improve the accuracy and efficiency of decision-making.

[0053] However, there are two main defects in traditional calibration methods:

[0054] (1) Single information and lack of texture features. Ordinary optical cameras obtain visual images with dense color information, and it is impossible to fuse point cloud-visual information at the pixel level through calibration technology to break through the limitations of the data;

[0055] (2) Some targets are incomplete and the far-view point clouds are sparse. It is very difficult for traditional calibration methods to fuse multiple frames of point clouds through point cloud registration technology to achieve the integrity and densification of the point clouds.

[0056] The automatic calibration method for the fusion of infrared light and visible light of the drone provided by this application is a point cloud registration and fusion algorithm based on the spatial structure similarity of visual key points. First, the SIFT+KNN algorithm is used to obtain the two-dimensional key point pairs of the visual image, then the depth information of the two-dimensional key points is estimated and back-projected into the three-dimensional space, and then several groups of three-dimensional key point pairs with the highest confidence are obtained through the spatial structure similarity algorithm to calculate the rigid transformation parameters.

[0057] The following will describe in detail the specific structure of the automatic calibration method for the fusion of infrared light and visible light of the drone provided by this application with reference to the accompanying drawings.

[0058] Example 1:

[0059] Refer to Figure 1 As shown, the embodiment of this application provides an automatic calibration method for the fusion of infrared light and visible light of a drone, including:

[0060] Use the SIFT+KNN algorithm to obtain the two-dimensional key points of the infrared light and visible light images;

[0061] Back-project according to the depth camera information of the two-dimensional key points into the three-dimensional space;

[0062] Obtain the three-dimensional space key points according to the spatial structure similarity algorithm;

[0063] Connect the three-dimensional space key points in series to realize the automatic calibration of the infrared light and visible light images.

[0064] This application obtains the two-dimensional key points of the infrared light and visible light images by using the SIFT+KNN algorithm; back-projects according to the depth camera information of the two-dimensional key points into the three-dimensional space; obtains the three-dimensional space key points according to the spatial structure similarity algorithm; connects the three-dimensional space key points in series to realize the automatic calibration of the infrared light and visible light images. Aiming at the problems of slow running speed and poor registration effect of traditional point cloud registration algorithms when processing point clouds in the power grid configuration scenario, this paper proposes a point cloud registration and fusion algorithm based on the spatial structure similarity of visual key points. First, the SIFT+KNN algorithm is used to obtain the two-dimensional key point pairs of the visual image, then the depth information of the two-dimensional key points is estimated and back-projected into the three-dimensional space, and then several groups of three-dimensional key point pairs with the highest confidence are obtained through the spatial structure similarity algorithm to calculate the rigid transformation parameters. Experiments prove that the method in this paper has achieved a large improvement in both running speed and registration effect.

[0065] In some embodiments, the algorithm steps of using the SIFT+KNN algorithm to obtain the two-dimensional key points of the infrared light and visible light images are:

[0066] Select the center point, and find the non-zero value points within the range of 8×8 in the neighborhood; calculate the two-dimensional Euclidean distance and direction from all neighborhoods to the center point, and save them;

[0067] As Figure 2 shown, according to the angular threshold:

[0068] [(-22.5° - 22.5°], (22.5° - 67.5°], (67.5° - 112.5°], (112.5° - 157.5°], (157.5° - 202.5°], (202.5° - 247.5°], (247.5° - 292.5°], (292.5° - 337.5°]], all neighborhood points are classified according to these eight threshold intervals;

[0069] The point with the minimum Euclidean distance in each category is denoted as the unique neighborhood point within the threshold. Finally, eight angularly uniform neighborhood points are obtained, and their Euclidean distances and direction angles to the center point are saved.

[0070] SIFT (Scale - Invariant Feature Transform) and KNN (K - Nearest Neighbors) algorithms are commonly used techniques in computer vision and image processing.

[0071] SIFT is an algorithm for image feature extraction. It can extract key points (also known as feature points) in an image and calculate their feature descriptors. SIFT feature descriptors have scale invariance and rotation invariance, so they can be used to identify objects of different sizes and orientations.

[0072] KNN is an algorithm for classification or regression. It is based on the proximity principle, that is, the similarity between samples is judged according to the distance between them. In the KNN algorithm, first, the distance between the test sample and each sample in the training set is calculated, then K nearest samples are selected as neighbors, and finally, the class or attribute of the test sample is predicted according to the classes or attributes of these neighbors.

[0073] SIFT and KNN algorithms are usually used in combination for tasks such as image recognition and classification. First, the SIFT algorithm is used to extract key points and feature descriptors in the image, and then the KNN algorithm is used to classify or regress these features. This method has been widely used in computer vision and image processing, such as object recognition, image matching, image classification, etc.

[0074] In some embodiments, for the three - dimensional space coordinate system, Cartesian coordinates are adopted and expressed as (Xw, Yw, Zw), where the X - axis and Y - axis are parallel to the X - axis and Y - axis of the image coordinate axes, and the optical axis of the camera is the Z - axis;

[0075] For two - dimensional space points, (x, y) is their coordinate point, and the conversion relationship between the camera coordinate system and the three - dimensional coordinate system is:

[0076]

[0077] Among them, f is the focal length of the camera, and cx and cy are the coordinates of the center of the imaging plane, which belong to the depth camera information parameters.

[0078] The Cartesian coordinate system is a coordinate system used to describe the position of points on a plane. It consists of two perpendicular coordinate axes, usually represented by the x-axis and the y-axis.

[0079] In the Cartesian coordinate system, each point can be represented by an ordered pair (x, y), where x represents the coordinate of the point on the x-axis and y represents the coordinate of the point on the y-axis. The values of these two coordinates can be any real numbers, so the Cartesian coordinate system can represent any point on the plane.

[0080] The Cartesian coordinate system is commonly used in fields such as mathematics, physics, and engineering to describe geometric shapes, motion trajectories, vectors, etc. on a plane. It is also widely used in fields such as computer graphics and computer vision to represent the pixel positions in a two-dimensional image, the positions of graphic objects, etc.

[0081] In some embodiments, the relationship between the projection coordinate system and the camera coordinate system is:

[0082]

[0083] Among them, R and T are rotation and translation parameters.

[0084] In some embodiments, the method for obtaining three-dimensional space key points according to the spatial structure similarity algorithm is;

[0085] First, use the uniform angle search algorithm to obtain the specific coordinates of eight neighborhood points; in a clockwise direction, compare the pixel value relationship between these eight neighborhood points and the center, and those larger than the center pixel value are recorded as 0, otherwise recorded as 1, to obtain an eight-bit binary number;

[0086] Shift this eight-bit binary number to the right until the lowest bit is not 0, and the shifted high bits are filled with 0, and convert the above binary number to a decimal number as the eigenvalue of the center pixel point;

[0087] The right shift of the binary number is to endow the feature with rotational invariance.

[0088] In some embodiments, as Figure 3 The three-dimensional space point projection method shown is:

[0089] First, use the uniform angle search algorithm to obtain eight neighborhood points;

[0090] Reconstruct the two-dimensional pixel point p(ai, vi) into a three-dimensional space point P(xi, yi, zi);

[0091] Select three points that can form an approximate triangle in the clockwise direction to obtain a virtual plane, and calculate the unit normal vector;

[0092] Finally, obtain eight groups of unit normal vectors to obtain the three-dimensionally calibrated image;

[0093] The calculation method of the normal vector n is:

[0094]

[0095] where PA, PB, and PC are the three points of the approximate triangle.

[0096] Example 2:

[0097] As Figure 4 shown, an infrared light and visible light fusion test was carried out. In order to achieve end-to-end training, the point cloud projection, visual depth estimation, and homogeneous feature calculation work were all added to the convolutional module of the model in the form of non-weight layers, and cuda operations were used, which greatly improved the calculation speed. During model training, the hardware was NVidia GTX1080ti, the BatchSize was set to 4, the weights of the rotation and translation losses were set to 0.2 and 0.8, the maximum number of training Epochs was set to 60, and the learning rate was initialized using an adaptive method, starting from 0.0015. Every 10 Epochs of training, it was reduced by 0.00015 until it reached 0.00005 and then stopped decreasing. During the model training process, the rotation / translation error changes of the training set and the validation set are shown in the figure. It can be seen from the figure that when the training reached 40 rounds, it had converged. The 45th round model was selected as the optimal model. At this time, the loss of the training set was 0.051, and the loss of the test set was 0.059. Shown Figure 4 is the angular loss curve, shown Figure 5 is the translation loss curve.

[0098] At the same time, the present invention carried out a comparison of the automatic calibration accuracy with other algorithms:

[0099]

[0100] Set the initial error: θ = 10, D = 2m, generate the training set and the validation set, and compare with the existing deep learning-based point cloud-image calibration methods: RegNet, CalibNet, and CMRNet. The corrected rotation error and translation error are shown in the above table.

[0101] The results show that the method of the present invention is superior to the past methods in both rotation and translation errors. Especially in the translation result, the error after calibration is greatly reduced compared with other methods.

[0102] It is easy to understand that those skilled in the art can combine, split, reorganize, etc. the embodiments provided in this application to obtain other embodiments on the basis of several embodiments provided in this application, and these embodiments do not exceed the protection scope of this application.

[0103] The above specific implementation manners further elaborate in detail the purpose, technical solutions and beneficial effects of the embodiments of this application. It should be understood that the above are only the specific implementation manners of the embodiments of this application, and are not used to limit the protection scope of the embodiments of this application. Any modifications, equivalent replacements, improvements, etc. made on the basis of the technical solutions of the embodiments of this application shall be included within the protection scope of the embodiments of this application.

Claims

1. An automatic calibration method for the fusion of infrared light and visible light of an unmanned aerial vehicle, characterized in that: including; Using the SIFT and KNN algorithms to obtain two-dimensional key points of infrared and visible light images; Back-projecting the depth camera information of the two-dimensional key points into three-dimensional space; Obtaining three-dimensional space key points according to the spatial structure similarity algorithm; Connecting the three-dimensional space key points in series to realize automatic calibration of infrared and visible light images.

2. The method for automatically calibrating the fusion of infrared light and visible light of an unmanned aerial vehicle according to claim 1, wherein: The method of using the SIFT and KNN algorithms to obtain two-dimensional key points of infrared and visible light images includes: Selecting a center point and finding non-zero value points within a neighborhood of 8×8; calculating the two-dimensional Euclidean distance and direction from all neighborhoods to the center point and saving them; According to the angle threshold: [(-22.5° - 22.5°], (22.5° - 67.5°], (67.5° - 112.5°], (112.5° - 157.5°], (157.5° - 202.5°], (202.5° - 247.5°], (247.5° - 292.5°], (292.5° - 337.5°]], classifying all neighborhood points according to these eight threshold intervals; Denoting the point with the minimum Euclidean distance in each class as the only neighborhood point within the threshold, and finally obtaining eight uniformly angled neighborhood points, saving the Euclidean distance domain direction angle from the uniformly angled neighborhood points to the center point.

3. The method for automatically calibrating the fusion of infrared light and visible light of an unmanned aerial vehicle according to claim 1, wherein: The three-dimensional space coordinate system adopts Cartesian coordinates and is expressed as (Xw, Yw, Zw), where the X-axis and Y-axis are parallel to the X-axis and Y-axis of the image coordinate axes, and the optical axis of the camera is the Z-axis; For a two-dimensional space point, (x, y) is its coordinate point, and the conversion relationship between the camera coordinate system and the three-dimensional coordinate system is: where: f is the camera focal length, and cx and cy are the coordinates of the center of the imaging plane, which belong to the depth camera information parameters.

4. The method for automatically calibrating the fusion of infrared light and visible light of an unmanned aerial vehicle according to claim 3, wherein: The relationship between the projection coordinate system and the camera coordinate system is: where: R and T are rotation and translation parameters.

5. The method for automatically calibrating the fusion of infrared light and visible light of an unmanned aerial vehicle according to claim 3, wherein: The method of obtaining three-dimensional space key points according to the spatial structure similarity algorithm includes: Using the uniform angle search algorithm to obtain the specific coordinates of eight neighborhood points; comparing the pixel value relationship between the eight neighborhood points and the center in a clockwise direction, denoting those larger than the center pixel value as 0, otherwise as 1, to obtain an eight-bit binary number; Shifting the eight-bit binary number to the right until the lowest bit is not 0, filling the shifted high bits with 0, and converting the above binary number into a decimal number as the characteristic value of the center pixel point; The right shift of the binary number is used to endow the feature with rotational invariance.

6. The method for automatically calibrating the fusion of infrared light and visible light of an unmanned aerial vehicle according to claim 3, wherein: The three-dimensional space point projection method includes: Using the uniform angle search algorithm to obtain eight neighborhood points; Reconstruct the two-dimensional pixel point p(ai, vi) into a three-dimensional space point P(xi, yi, zi); Select three points that can form an approximate triangle in the clockwise direction to obtain a virtual plane, and calculate the unit normal vector; Obtain eight groups of unit normal vectors to obtain a three-dimensionally calibrated image; The calculation method of the normal vector n is as follows: Where: PA, PB, and PC are the three points of the approximate triangle.