A method and system for spatial registration of visible light images and infrared images

By calculating the rigid transformation matrix through time alignment and inverse perspective projection transformation, the complexity of spatial registration between visible light and infrared images is solved, achieving efficient and low-cost image registration.

CN119784803BActive Publication Date: 2026-05-01BEIJING SINOITS TECH
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING SINOITS TECH
Filing Date
2024-12-16
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing technologies for spatial registration of visible light and infrared images suffer from problems such as complex feature extraction and matching, and the need for a large amount of training data, resulting in a complex and costly registration process.

Method used

By aligning visible light and infrared cameras in time, and utilizing target recognition and inverse perspective projection transformation, a rigid transformation matrix is ​​calculated to achieve spatial registration of images, simplifying the conversion process and eliminating the need for a large number of training images.

Benefits of technology

It achieves precise registration between visible light camera and infrared camera images, reducing costs and improving the practicality and accuracy of registration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119784803B_ABST
    Figure CN119784803B_ABST
Patent Text Reader

Abstract

The application discloses a kind of spatial registration method and system of visible light image and infrared image, it is related to spatial registration technical field, the position coordinate in the position coordinate in the coordinate system of second bird's-eye view obtained by visible light camera is registered with the position coordinate in the position coordinate in the coordinate system of second bird's-eye view obtained by infrared camera, more accurate conversion relationship can be obtained, and the acquisition process of conversion relationship is simpler than prior art, does not need a large number of training images, low in cost, strong practicality.
Need to check novelty before this filing date? Find Prior Art

Description

A method and system for spatial registration of visible light and infrared images Technical Field

[0001] This invention relates to the field of spatial registration technology, and in particular to a method and system for spatial registration of visible light images and infrared images. Background Technology

[0002] Dual-spectrum integrated cameras include a visible light camera and an infrared camera (specifically, a far-infrared camera), providing day and night complementarity and also offering target temperature information. They are crucial components of vehicle-to-everything (V2X) and digital twin systems. Therefore, it is necessary to spatially register the images captured by the visible light camera and infrared camera of the dual-spectrum integrated camera. Currently, the following existing technologies exist:

[0003] 1) The invention patent with publication number "CN103337077A" and subject title "A method for registration of visible light and infrared images based on multi-scale segmentation and SIFT" discloses that SIFT feature matching has achieved great success in the field of visible light, but there are complex parameter and engineering adjustment problems between it and feature extraction and matching of far-infrared images.

[0004] 2) The invention patent with publication number "CN113628261A" and subject title "A Method for Registering Infrared and Visible Light Images in a Power Inspection Scenarios" discloses the following: A SuperPoint feature extraction network is used to detect feature points in two edge images and calculate descriptors; based on the feature points of the two edge images, a SuperGlue feature matching network is used to match the feature points, selecting correct feature point matching pairs and discarding unmatched feature points; affine transformation model parameters are calculated based on the matched feature point pairs, and bilinear interpolation is used to perform spatial coordinate transformation on the images to be registered, thus achieving image registration. This invention achieves accurate registration of infrared and visible light images of power equipment, obtaining temperature information of the power equipment in the background of the visible light image, but requires a large number of images for training.

[0005] 3) The invention patent with publication number “CN109448035A” and subject name “Infrared image and visible light image registration method based on deep learning” discloses that the affine or homography matrix of the two is directly given by the deep learning model. This requires a large number of samples to train the model, but in actual post-installation applications, these data are often not guaranteed. Summary of the Invention

[0006] The technical problem to be solved by the present invention is to address the shortcomings of the prior art, specifically by providing a spatial registration method and system for visible light images and infrared images, as detailed below:

[0007] 1) In a first aspect, the present invention provides a spatial registration method for visible light images and infrared images, the specific technical solution of which is as follows:

[0008] After time alignment of the visible light camera and the infrared camera, the first visible light image and the first infrared image are acquired at the same time using the visible light camera and the infrared camera, respectively.

[0009] Target recognition is performed on the first visible light image to obtain the first position coordinates of each first preset target in the coordinate system used by the visible light camera; target recognition is performed on the first infrared image to obtain the second position coordinates of each first preset target in the coordinate system used by the infrared camera.

[0010] Based on the first transformation matrix, the first visible light image is subjected to inverse perspective projection transformation to obtain the first bird's-eye view. The first transformation matrix is ​​then used to transform each first position coordinate to obtain the third position coordinate of each first position coordinate in the coordinate system used by the first bird's-eye view.

[0011] Based on the second transformation matrix, the first infrared image is subjected to inverse perspective projection transformation to obtain the second bird's-eye view. The second transformation matrix is ​​then used to transform each second position coordinate to obtain the fourth position coordinate of each second position coordinate in the coordinate system used in the second bird's-eye view.

[0012] Calculate the rigid transformation matrix between the coordinate system used in the first bird's-eye view and the coordinate system used in the second bird's-eye view;

[0013] Using a rigid transformation matrix, the fifth position coordinates of each third position coordinate in the coordinate system used in the second bird's-eye view are obtained;

[0014] Calculate the error distance between all fourth position coordinates and all fifth position coordinates; determine whether the error distance is less than a preset error distance threshold, and obtain the judgment result;

[0015] When the judgment result is yes, the conversion relationship between the images captured by the visible light camera and the infrared camera at the same time is determined based on all the fourth position coordinates and all the fifth position coordinates.

[0016] Using transformation relationships, spatial registration is performed on images captured by a visible light camera and an infrared camera at the same time.

[0017] The beneficial effects of the spatial registration method for visible light images and infrared images provided by this invention are as follows:

[0018] The position coordinates of the first preset target obtained by the visible light camera are coordinates in the coordinate system used by the visible light camera, and the first position coordinates of the first preset target obtained by the infrared camera are coordinates in the coordinate system used by the infrared camera. In practical applications, it is necessary to convert the second position coordinates of the target obtained by the visible light camera and the infrared camera. Therefore, the accuracy of the conversion relationship between the images captured by the visible light camera and the infrared camera at the same time determines whether the conversion between the first position coordinates and the second position coordinates is accurate. This invention registers the position coordinates of the position coordinates obtained by the visible light camera in the coordinate system used by the second bird's-eye view with the position coordinates of the position coordinates obtained by the infrared camera in the coordinate system used by the second bird's-eye view. This can obtain a more accurate conversion relationship, and the process of obtaining the conversion relationship is simpler than that of the prior art. It does not require a large number of training images, has low cost, and is highly practical.

[0019] Based on the above scheme, the spatial registration method for visible light images and infrared images of the present invention can be further improved as follows.

[0020] Furthermore, the process of obtaining the first transformation matrix includes:

[0021] Target tracking is performed using a time-aligned visible light camera to obtain the trajectory of each second preset target and perform straight line fitting to fit the first straight line corresponding to each second preset target. The first transformation matrix is ​​determined based on the two outermost first straight lines.

[0022] Furthermore, the process of obtaining the second transformation matrix includes:

[0023] Target tracking is performed using a time-aligned infrared camera to obtain the trajectory of each third preset target and perform straight line fitting to fit the second straight line corresponding to each third preset target. The second transformation matrix is ​​determined based on the two outermost second straight lines.

[0024] Furthermore, the rigid transformation matrix between the coordinate system used in the first bird's-eye view and the coordinate system used in the second bird's-eye view is calculated, including:

[0025] The rigid transformation matrix between the coordinate system used in the first bird's-eye view and the coordinate system used in the second bird's-eye view is calculated using either the ICP algorithm or the CPD algorithm.

[0026] 2) In a second aspect, the present invention also provides a spatial registration system for visible light images and infrared images, the specific technical solution of which is as follows:

[0027] It includes an image acquisition module, a target recognition module, a bird's-eye view acquisition module, a position coordinate transformation module, a rigid transformation matrix calculation module, an error distance calculation and judgment module, a transformation relationship determination module, and a spatial registration module;

[0028] The image acquisition module is used to: time-align the visible light camera and the infrared camera, and then use the visible light camera and the infrared camera to acquire the first visible light image and the first infrared image at the same time.

[0029] The target recognition module is used to: perform target recognition on the first visible light image to obtain the first position coordinates of each first preset target in the coordinate system used by the visible light camera; and perform target recognition on the first infrared image to obtain the second position coordinates of each first preset target in the coordinate system used by the infrared camera.

[0030] The bird's-eye view acquisition module is used to: perform inverse perspective projection transformation on the first visible light image based on the first transformation matrix to obtain the first bird's-eye view;

[0031] The position coordinate transformation module is used to: transform each first position coordinate using the first transformation matrix to obtain the third position coordinate of each first position coordinate in the coordinate system used in the first bird's-eye view;

[0032] The bird's-eye view acquisition module is also used to: perform inverse perspective projection transformation on the first infrared image based on the second transformation matrix to obtain a second bird's-eye view;

[0033] The position coordinate transformation module is also used to: transform each second position coordinate using the second transformation matrix to obtain the fourth position coordinate of each second position coordinate in the coordinate system used in the second bird's-eye view;

[0034] The rigid transformation matrix calculation module is used to calculate the rigid transformation matrix between the coordinate system used in the first bird's-eye view and the coordinate system used in the second bird's-eye view.

[0035] The position coordinate transformation module is also used to: obtain the fifth position coordinate of each third position coordinate in the coordinate system used in the second bird's-eye view using a rigid transformation matrix;

[0036] The error distance calculation and judgment module is used to: calculate the error distance between all fourth position coordinates and all fifth position coordinates; determine whether the error distance is less than a preset error distance threshold, and obtain the judgment result;

[0037] The conversion relationship determination module is used to: when the judgment result is yes, determine the conversion relationship between the images captured by the visible light camera and the infrared camera at the same time based on all the fourth position coordinates and all the fifth position coordinates;

[0038] The spatial registration module is used to spatially register images captured by a visible light camera and an infrared camera at the same time using a transformation relationship.

[0039] Based on the above scheme, the spatial registration system for visible light images and infrared images of the present invention can be further improved as follows.

[0040] Furthermore, it also includes a first transformation matrix acquisition module, which is used for:

[0041] Target tracking is performed using a time-aligned visible light camera to obtain the trajectory of each second preset target and perform straight line fitting to fit the first straight line corresponding to each second preset target. The first transformation matrix is ​​determined based on the two outermost first straight lines.

[0042] Furthermore, it also includes a second transformation matrix acquisition module, which is used for:

[0043] Target tracking is performed using a time-aligned infrared camera to obtain the trajectory of each third preset target and perform straight line fitting to fit the second straight line corresponding to each third preset target. The second transformation matrix is ​​determined based on the two outermost second straight lines.

[0044] Furthermore, the rigid transformation matrix calculation module is specifically used to: calculate the rigid transformation matrix between the coordinate system used in the first bird's-eye view and the coordinate system used in the second bird's-eye view using the ICP algorithm or the CPD algorithm.

[0045] 3) In a third aspect, the present invention also provides an electronic device, the electronic device including a processor coupled to a memory, the memory storing at least one computer program, the at least one computer program being loaded and executed by the processor, so as to enable the electronic device to implement any of the above-mentioned spatial registration methods for visible light images and infrared images.

[0046] 4) In a fourth aspect, the present invention also provides a computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, implements any of the above-described spatial registration methods for visible light images and infrared images.

[0047] It should be noted that the beneficial effects of the technical solutions of the second to fourth aspects of the present invention and their corresponding possible implementations can be found in the above description of the technical effects of the first aspect and its corresponding possible implementations, and will not be repeated here. Attached Figure Description

[0048] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments of the present invention will be briefly introduced below:

[0049] Figure 1 is a schematic flowchart of a spatial registration method for visible light images and infrared images according to an embodiment of the present invention.

[0050] Figure 2 is a geometric diagram used to determine the first transformation matrix;

[0051] Figure 3 is a schematic diagram of a rectangle preset in the coordinate system used in the first bird's-eye view;

[0052] Figure 4 shows the geometric diagram used to determine the second transformation matrix.

[0053] Figure 5 is a schematic diagram of a rectangle preset in the coordinate system used in the second bird's-eye view;

[0054] Figure 6 is a schematic diagram of the structure of an electronic device according to an embodiment of the present invention. Detailed Implementation

[0055] The principles and features of the present invention are described below. The examples given are only for explaining the present invention and are not intended to limit the scope of the present invention.

[0056] The technical solution of the present invention and how the technical solution of the present invention solves the above-mentioned technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of the present invention will now be described with reference to the accompanying drawings.

[0057] As shown in Figure 1, a spatial registration method for visible light images and infrared images according to an embodiment of the present invention includes the following steps:

[0058] S1. After time-aligning the visible light camera and the infrared camera, the first visible light image and the first infrared image are acquired at the same time using the visible light camera and the infrared camera, respectively.

[0059] In this process, a visible light camera and an infrared camera are used to capture images of the same preset area, resulting in a first visible light image and a first infrared image acquired at the same time. The preset area can be a designated road area or a preset area can be specified according to the actual situation.

[0060] S2. Perform target recognition on the first visible light image to obtain the first position coordinates of each first preset target in the coordinate system used by the visible light camera; perform target recognition on the first infrared image to obtain the second position coordinates of each first preset target in the coordinate system used by the infrared camera, specifically:

[0061] Target recognition can be performed on the first visible light image using algorithms such as YOLO or ViD, detecting N1 targets. Target recognition can then be performed on the first infrared image using the same algorithms, detecting N2 targets. Multiple identical targets can be selected from the N1 and N2 targets as the first preset targets, denoted as n. When N1 = N2, all targets can be used as the first preset targets, in which case N1 = N2 = n. It should be noted that other single-stage detection algorithms can also be used to perform target recognition on the first visible light image.

[0062] Since the detection boxes of the first preset target are output by algorithms such as YOLO or ViD, the coordinates of the center point of the bottom edge of the detection box of each first preset target in the coordinate system used by the visible light camera are used as the corresponding first position coordinates. Furthermore, since the detection boxes of the target are output by algorithms such as YOLO or ViD, the coordinates of the center point of the bottom edge of the detection box of each target in the coordinate system used by the infrared camera are used as the corresponding second position coordinates.

[0063] The coordinate system used by the visible light camera is a rectangular coordinate system established with a designated position on the visible light image captured by the camera as the origin. The designated position can be the center point of the visible light image or any position on the bottom edge of the visible light image. The straight line containing the bottom edge of the visible light image can be the x-axis of this rectangular coordinate system. It should be noted that the origin of the coordinate system used by the visible light camera can be set according to the actual situation, and the x-axis and y-axis of this rectangular coordinate system can also be set according to the actual situation.

[0064] The coordinate system used by the infrared camera is a rectangular coordinate system established with a designated position on the infrared image captured by the infrared camera as the origin. The designated position can be the center point of the infrared image or any position on the bottom edge of the infrared image. The straight line containing the bottom edge of the infrared image can be the x-axis of the rectangular coordinate system. It should be noted that the origin of the coordinate system used by the infrared camera can be set according to the actual situation, and the x-axis and y-axis of the rectangular coordinate system can also be set according to the actual situation.

[0065] Among them, multiple identical targets can be selected from N1 targets and N2 targets as the first preset targets, and the number of the first preset targets is denoted as n. When N1=N2, all targets can be used as the first preset targets, and at this time N1=N2=n.

[0066] S3. Based on the first transformation matrix, perform inverse perspective projection transformation on the first visible light image to obtain the first bird's-eye view, and use the first transformation matrix to transform each first position coordinate to obtain the third position coordinate of each first position coordinate in the coordinate system used by the first bird's-eye view.

[0067] The process of obtaining the first transformation matrix includes:

[0068] Target tracking is performed using a time-aligned visible light camera to obtain the trajectory of each second preset target and perform straight line fitting to fit the first straight line corresponding to each second preset target. The first transformation matrix is ​​determined based on the two outermost first straight lines.

[0069] In this process, target recognition can be performed on the second visible light image captured by the visible light camera using algorithms such as YOLO or ViD. If N3 targets are detected, a portion of the N3 targets can be selected as the second preset targets, or all N3 targets can be used as the second preset targets. Since the YOLO or ViD algorithm will output the detection box of the target, the coordinates of the center point of the bottom edge of the detection box of each target are used as the corresponding sixth position coordinates. The sixth position coordinates are the coordinates in the coordinate system used by the visible light camera. It should be noted that other single-stage detection algorithms can also be used to perform target recognition on the first visible light image.

[0070] The system uses a visible light camera to track each second preset target, obtaining the trajectory of each second preset target. The trajectory of each second preset target is then fitted with a straight line to obtain multiple first straight lines. The intersection of all the first straight lines is then determined to obtain the first hidden point. The two outermost first straight lines are selected and denoted as L1 and L2 respectively. When the second preset target is a vehicle, the first hidden point is the hidden point of the road. The specific meaning of the first hidden point will also be different when the target is different.

[0071] As shown in Figure 2, calculate the position coordinates of the first intersection point A1 and the second intersection point B1. The first intersection point A1 is the intersection of the first straight line L1 with the x-axis in the coordinate system where the visible light camera is located. The second intersection point B1 is the intersection of the second straight line L2 with the x-axis in the coordinate system where the visible light camera is located. The position coordinates of the first intersection point A1 and the second intersection point B1 are both position coordinates in the coordinate system where the visible light camera is located.

[0072] Calculate the position coordinates of the first hidden point VP1 in the coordinate system of the visible light camera;

[0073] Let the midpoint of the line segment between the first hidden point VP1 and the first intersection point A1 be denoted as the first midpoint A2, and the midpoint of the line segment between the first hidden point VP1 and the second intersection point B1 be denoted as the second midpoint B2. Then obtain the position coordinates of the first midpoint A2 and the second midpoint B2. At this time, the first intersection point A1, the second intersection point B1, the first midpoint A2 and the second midpoint B2 form a trapezoid.

[0074] As shown in Figure 3, in the coordinate system used in the first bird's-eye view, a rectangle is pre-defined. The four corner points of the rectangle are denoted as: first corner point a1, second corner point b1, third corner point a2, and fourth corner point b2. The trapezoid formed by the first intersection point A1, the second intersection point B1, the first midpoint A2, and the second midpoint B2 is mapped to the rectangle formed by the same four corner points. Specifically, the first intersection point A1 corresponds to the first corner point a1, the second intersection point B1 corresponds to the second corner point b1, the first midpoint A2 corresponds to the third corner point a2, and the second midpoint B2 corresponds to the fourth corner point b2. Based on these four pairs of point correspondences, the transformation matrix between the visible light coordinate system and the coordinate system used in the first bird's-eye view, i.e., the first transformation matrix, is obtained, denoted as H. vis_bev,1 .

[0075] The coordinate system used in the first bird's-eye view is a rectangular coordinate system with the origin at a specified position on the first bird's-eye view. The specified position can be the center point of the first bird's-eye view or any position on the bottom edge of the first bird's-eye view. The straight line containing the bottom edge of the first bird's-eye view can be the x-axis of the rectangular coordinate system. It should be noted that the origin of the first bird's-eye view can be set according to the actual situation, and the x-axis and y-axis of the rectangular coordinate system can also be set according to the actual situation.

[0076] Specifically, each first position coordinate is transformed using a first transformation matrix to obtain the third position coordinates of each first position coordinate in the coordinate system used in the first bird's-eye view, including:

[0077] Using the first transformation formula, the third position coordinates of each first position coordinate in the coordinate system used in the first bird's-eye view are obtained. The first transformation formula is: ,in, express: The first in The first position coordinates, The first in The first position coordinate is the first The first position coordinates of the first preset target Represents the set of all first position coordinates. express: The third position coordinate in the coordinate system used in the first bird's-eye view. , Let the set of coordinates of the third position be a positive integer. .

[0078] S4. Based on the second transformation matrix, perform inverse perspective projection transformation on the first infrared image to obtain the second bird's-eye view, and use the second transformation matrix to transform each second position coordinate to obtain the fourth position coordinate of each second position coordinate in the coordinate system used by the second bird's-eye view.

[0079] The process of obtaining the second transformation matrix includes:

[0080] Target tracking is performed using a time-aligned infrared camera to obtain the trajectory of each third preset target and perform straight line fitting to fit the second straight line corresponding to each third preset target. The second transformation matrix is ​​determined based on the two outermost second straight lines.

[0081] In this process, target recognition can be performed on the second infrared image captured by the infrared camera using algorithms such as YOLO or ViD. If N4 targets are detected, a portion of the N4 targets can be selected as the third preset targets, or all N4 targets can be selected as the third preset targets. Since the YOLO or ViD algorithm will output the target detection box, the coordinates of the center point of the bottom edge of the detection box of each target are used as the corresponding sixth position coordinates. The sixth position coordinates are the coordinates in the coordinate system used by the infrared camera. It should be noted that other single-stage detection algorithms can also be used to perform target recognition on the first visible light image.

[0082] Infrared cameras are used to track each third preset target to obtain the trajectory of each third preset target. Straight line fitting is performed on the trajectory of each third preset target to obtain multiple second straight lines. The intersection of all the second straight lines is obtained, and the second disappearance point is calculated. The two outermost second straight lines are selected and denoted as L3 and L4 respectively. When the second preset target is a vehicle, the second disappearance point is the disappearance point of the road. The specific meaning of the disappearance point will also be different when the target is different.

[0083] As shown in Figure 4, calculate the position coordinates of the third intersection point A3 and the fourth intersection point B3. The third intersection point A3 is the intersection of the third line L3 with the x-axis in the coordinate system where the infrared camera is located, and the fourth intersection point B3 is the intersection of the fourth line L4 with the x-axis in the coordinate system where the infrared camera is located. The position coordinates of the third intersection point A3 and the fourth intersection point B3 are both position coordinates in the coordinate system where the infrared camera is located.

[0084] Calculate the position coordinates of the second blanking point VP2 in the coordinate system of the infrared camera;

[0085] Let the midpoint of the line segment between the second hidden point VP2 and the third intersection point A3 be denoted as the third midpoint A4, and the midpoint of the line segment between the second hidden point VP2 and the fourth intersection point B3 be denoted as the fourth midpoint B4. Then obtain the position coordinates of the third midpoint A4 and the fourth midpoint B4. At this time, the third intersection point A3, the fourth intersection point B3, the third midpoint A4, and the fourth midpoint B4 form a trapezoid.

[0086] As shown in Figure 5, in the coordinate system used in the second bird's-eye view, a rectangle is pre-defined. The four corner points of the rectangle are denoted as: fifth corner point a3, sixth corner point b3, seventh corner point a4, and eighth corner point b4. The trapezoid formed by the third intersection point A3, the fourth intersection point B3, the third midpoint A4, and the fourth midpoint B4 is mapped to the rectangle formed by the fifth corner point a3, the sixth corner point b3, the seventh corner point a4, and the eighth corner point b4. Specifically, the third intersection point A3 corresponds to the fifth corner point a3, the third midpoint A4 corresponds to the seventh corner point a4, the fourth intersection point B4 corresponds to the sixth corner point b3, and the fourth midpoint B4 corresponds to the eighth corner point b4. Based on the correspondence of these four pairs of points, the transformation matrix between the coordinate system used for infrared and the coordinate system used in the second bird's-eye view, i.e., the second transformation matrix, can be obtained, denoted as H. vis_bev,2 .

[0087] The coordinate system used in the second bird's-eye view is a rectangular coordinate system established with the specified position on the second bird's-eye view as the origin. The specified position can be the center point of the second bird's-eye view or any position on the bottom edge of the second bird's-eye view. The straight line containing the bottom edge of the second bird's-eye view can be the x-axis of this rectangular coordinate system. It should be noted that the origin of the second bird's-eye view can be set according to the actual situation, and the x-axis and y-axis of this rectangular coordinate system can also be set according to the actual situation.

[0088] Specifically, the second transformation matrix is ​​used to transform each second position coordinate to obtain the fourth position coordinates of each second position coordinate in the coordinate system used in the second bird's-eye view, including:

[0089] Using the second transformation formula, we obtain the fourth position coordinates of each second position coordinate in the coordinate system used in the second bird's-eye view. The first transformation formula is: ,in, express: The first in The second position coordinates, The first in The second position coordinate is: the first The second position coordinates of the first preset target Represents: the set of all second position coordinates. express: The fourth position coordinate in the coordinate system used in the second bird's-eye view , Let the set of the fourth position coordinates be positive integers. .

[0090] S5. Calculate the rigid transformation matrix between the coordinate system used in the first bird's-eye view and the coordinate system used in the second bird's-eye view;

[0091] Specifically, the ICP (Iterative Closest Point) algorithm or the CPD (Coherent Point Drift) algorithm can be used to calculate the rigid transformation matrix between the coordinate system used in the first bird's-eye view and the coordinate system used in the second bird's-eye view.

[0092] S6. Using a rigid transformation matrix, obtain the fifth position coordinates of each third position coordinate in the coordinate system used in the second bird's-eye view. This is specifically achieved through the third transformation formula, which is: ,in, Represents the rigid transformation matrix. express: The fifth position coordinates in the coordinate system used in the second bird's-eye view, the set of all fifth position coordinates is denoted as... .

[0093] S7. Calculate the error distance between all fourth position coordinates and all fifth position coordinates; determine whether the error distance is less than the preset error distance threshold, and obtain the judgment result;

[0094] The error distance can be Euclidean distance, or error distance (Euclidean distance). It is calculated using the following formula:

[0095]

[0096] The preset error distance threshold can be set according to the actual situation.

[0097] S8. When the judgment result is yes, determine the conversion relationship between the images captured by the visible light camera and the infrared camera at the same time based on all the fourth position coordinates and all the fifth position coordinates.

[0098] S9. Using the transformation relationship, spatial registration is performed on the images captured by the visible light camera and the infrared camera at the same time.

[0099] Here, the set of all fourth position coordinates is denoted as Let the set of all fourth position coordinates be denoted as ,pass and The one-to-one correspondence between the positional coordinates can determine and The one-to-one correspondence between the positional coordinates of the two coordinates allows for the calculation of the homography transformation matrix used to characterize the transformation relationship. :

[0100]

[0101] in, , , , , , , , and The homography transformation matrix The elements in.

[0102] At this point, using time-aligned visible light and infrared cameras for target tracking, the position coordinates of the same target detected at any given time in the corresponding coordinate system satisfy:

[0103]

[0104] The position coordinates in the coordinate system used by the visible light camera are: The position coordinates in the coordinate system used by the infrared camera are: ,and and This refers to the position coordinates of the same target detected by a visible light camera and an infrared camera at any given time after time alignment. The coefficients represent homogeneous coordinates.

[0105] Through homography transformation matrix It can accurately convert the position coordinates in the coordinate system used by the visible light camera to the coordinate system used by the infrared camera, and it can also accurately convert the position coordinates in the coordinate system used by the infrared camera to the coordinate system used by the visible light camera.

[0106] Optionally, the above technical solution also includes: when the judgment result is negative, manual registration can be performed. and The one-to-one correspondence between the positional coordinates of the two can be used to calculate the homography transformation matrix used to characterize the transformation relationship. Alternatively, the homography transformation matrix used to characterize the transformation relationship can be obtained through other calculation methods.

[0107] In the above embodiments, although the steps are numbered S1, S2, etc., they are only specific embodiments given by the present invention. Those skilled in the art can adjust the execution order of S1, S2, etc. according to the actual situation, which is also within the protection scope of the present invention. It can be understood that in some embodiments, some or all of the above embodiments may be included.

[0108] A spatial registration system for visible light and infrared images according to an embodiment of the present invention includes an image acquisition module, a target recognition module, a bird's-eye view acquisition module, a position coordinate transformation module, a rigid transformation matrix calculation module, an error distance calculation and judgment module, a transformation relationship determination module, and a spatial registration module.

[0109] The image acquisition module is used to: time-align the visible light camera and the infrared camera, and then use the visible light camera and the infrared camera to acquire the first visible light image and the first infrared image at the same time.

[0110] The target recognition module is used to: perform target recognition on the first visible light image to obtain the first position coordinates of each first preset target in the coordinate system used by the visible light camera; and perform target recognition on the first infrared image to obtain the second position coordinates of each first preset target in the coordinate system used by the infrared camera.

[0111] The bird's-eye view acquisition module is used to: perform inverse perspective projection transformation on the first visible light image based on the first transformation matrix to obtain the first bird's-eye view;

[0112] The position coordinate transformation module is used to: transform each first position coordinate using the first transformation matrix to obtain the third position coordinate of each first position coordinate in the coordinate system used in the first bird's-eye view;

[0113] The bird's-eye view acquisition module is also used to: perform inverse perspective projection transformation on the first infrared image based on the second transformation matrix to obtain a second bird's-eye view;

[0114] The position coordinate transformation module is also used to: transform each second position coordinate using the second transformation matrix to obtain the fourth position coordinate of each second position coordinate in the coordinate system used in the second bird's-eye view;

[0115] The rigid transformation matrix calculation module is used to calculate the rigid transformation matrix between the coordinate system used in the first bird's-eye view and the coordinate system used in the second bird's-eye view.

[0116] The position coordinate transformation module is also used to: obtain the fifth position coordinate of each third position coordinate in the coordinate system used in the second bird's-eye view using a rigid transformation matrix;

[0117] The error distance calculation and judgment module is used to: calculate the error distance between all fourth position coordinates and all fifth position coordinates; determine whether the error distance is less than a preset error distance threshold, and obtain the judgment result;

[0118] The conversion relationship determination module is used to: when the judgment result is yes, determine the conversion relationship between the images captured by the visible light camera and the infrared camera at the same time based on all the fourth position coordinates and all the fifth position coordinates;

[0119] The spatial registration module is used to spatially register images captured by a visible light camera and an infrared camera at the same time using a transformation relationship.

[0120] Optionally, the above technical solution further includes a first transformation matrix acquisition module, which is used for:

[0121] Target tracking is performed using a time-aligned visible light camera to obtain the trajectory of each second preset target and perform straight line fitting to fit the first straight line corresponding to each second preset target. The first transformation matrix is ​​determined based on the two outermost first straight lines.

[0122] Optionally, the above technical solution further includes a second transformation matrix acquisition module, which is used for:

[0123] Target tracking is performed using a time-aligned infrared camera to obtain the trajectory of each third preset target and perform straight line fitting to fit the second straight line corresponding to each third preset target. The second transformation matrix is ​​determined based on the two outermost second straight lines.

[0124] Optionally, in the above technical solution, the rigid transformation matrix calculation module is specifically used to: calculate the rigid transformation matrix between the coordinate system used in the first bird's-eye view and the coordinate system used in the second bird's-eye view using the ICP algorithm or the CPD algorithm.

[0125] It should be noted that the beneficial effects of the spatial registration system for visible light and infrared images provided in the above embodiments are the same as those of the spatial registration method for visible light and infrared images described above, and will not be repeated here. Furthermore, the system provided in the above embodiments is only illustrated by the division of the above functional modules. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the system can be divided into different functional modules according to the actual situation to complete all or part of the functions described above. In addition, the system and method embodiments provided in the above embodiments belong to the same concept, and their specific implementation process is detailed in the method embodiments, and will not be repeated here.

[0126] The spatial registration system for visible light and infrared images of the present invention can be a computer program (including program code) running on a computer device. For example, the spatial registration system for visible light and infrared images of the present invention is an application software that can be used to execute the corresponding steps in the spatial registration method for visible light and infrared images of the present invention.

[0127] In some embodiments, the spatial registration system for visible light and infrared images of the present invention can be implemented using a combination of hardware and software. As an example, the spatial registration system for visible light and infrared images of the present invention can be a processor in the form of a hardware decoding processor, which is programmed to execute the spatial registration method for visible light and infrared images of the present invention. For example, the processor in the form of a hardware decoding processor can be one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), or other electronic components.

[0128] The modules described in the embodiments of this invention can be implemented in software or hardware. The names of the modules are not, in some cases, limiting the scope of the module itself.

[0129] An electronic device according to an embodiment of the present invention includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements any of the above-mentioned spatial registration methods for visible light images and infrared images. That is, an electronic device according to an embodiment of the present invention may include, but is not limited to: a processor and a memory; the memory is used to store the computer program; the processor is used to execute a spatial registration method for visible light images and infrared images shown in any embodiment of the present invention by calling the computer program.

[0130] In one optional embodiment, an electronic device is provided, as shown in FIG6. The electronic device 4000 shown in FIG6 includes a processor 4001 and a memory 4003. The processor 4001 and the memory 4003 are connected, for example, via a bus 4002. Optionally, the electronic device 4000 may further include a transceiver 4004, which can be used for data interaction between the electronic device and other electronic devices, such as sending and / or receiving data. It should be noted that in practical applications, the transceiver 4004 is not limited to one type, and the structure of the electronic device 4000 does not constitute a limitation on the embodiments of the present invention.

[0131] Processor 4001 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this invention. Processor 4001 may also be a combination that implements computational functions, such as including one or more microprocessor combinations, a combination of a DSP and a microprocessor, etc.

[0132] Bus 4002 may include a pathway for transmitting information between the aforementioned components. Bus 4002 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. Bus 4002 can be divided into address bus, data bus, control bus, etc. For ease of illustration, Figure 6 uses only one thick line to represent bus 4002, but this does not indicate that there is only one bus or one type of bus.

[0133] The memory 4003 may be ROM (Read Only Memory) or other types of static storage devices capable of storing static information and instructions, RAM (Random Access Memory) or other types of dynamic storage devices capable of storing information and instructions, or EEPROM (Electrically Erasable Programmable Read Only Memory), CD-ROM (Compact Disc Read Only Memory) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but not limited thereto.

[0134] The memory 4003 stores application code (computer program) for executing the present invention, and its execution is controlled by the processor 4001. The processor 4001 executes the application code stored in the memory 4003 to implement the content shown in the foregoing method embodiments.

[0135] Among them, electronic devices can also be terminal devices. Terminal devices can be any terminal device that can install applications, including at least one of smartphones, tablets, laptops, desktop computers, smart speakers, smartwatches, smart TVs, and smart in-vehicle devices.

[0136] It should be noted that the electronic device shown in Figure 6 is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of the present invention.

[0137] An embodiment of the present invention provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements any of the above-described spatial registration methods for visible light images and infrared images.

[0138] Alternatively, the computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CD-ROM), magnetic tape, a floppy disk, and an optical data storage device, etc.

[0139] In an exemplary embodiment, a computer program product or computer program is also provided, which includes computer instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the electronic device to perform any of the above-described spatial registration methods for visible light and infrared images.

[0140] Computer program code for performing the operations of this invention can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as C or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0141] It should be understood that the flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of methods and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0142] The computer-readable storage medium provided in this invention can be, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this invention, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0143] The aforementioned computer-readable storage medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to perform the method shown in the above embodiments.

[0144] The above description is merely a preferred embodiment of the present invention and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of disclosure in this invention is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-disclosed concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features with similar functions disclosed in this invention.

[0145] It should be noted that the terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and represent a limitation on a specific order or sequence. Where appropriate, the order of use for similar objects can be interchanged so that the embodiments of this application described herein can be implemented in an order other than that shown or described.

[0146] Those skilled in the art will recognize that this invention can be implemented as a system, method, or computer program product. Therefore, this invention can be specifically implemented in the following forms: it can be entirely hardware, entirely software (including firmware, resident software, microcode, etc.), or a combination of hardware and software, generally referred to herein as a "circuit," "module," or "system." Furthermore, in some embodiments, this invention can also be implemented as a computer program product contained in one or more computer-readable media, which includes computer-readable program code.

[0147] Although embodiments of the present invention have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention.

Claims

1. A method for spatial registration of visible light images and infrared images, characterized in that, include: After time alignment of the visible light camera and the infrared camera, the first visible light image and the first infrared image are acquired at the same time using the visible light camera and the infrared camera, respectively. Target recognition is performed on the first visible light image to obtain the first position coordinates of each first preset target in the coordinate system used by the visible light camera; Target recognition is performed on the first infrared image to obtain the second position coordinates of each first preset target in the coordinate system used by the infrared camera; Based on the first transformation matrix, the first visible light image is subjected to inverse perspective projection transformation to obtain a first bird's-eye view. The first transformation matrix is ​​then used to transform each first position coordinate to obtain the third position coordinate of each first position coordinate in the coordinate system used in the first bird's-eye view. Based on the second transformation matrix, the first infrared image is subjected to inverse perspective projection transformation to obtain a second bird's-eye view. The second transformation matrix is ​​then used to transform each second position coordinate to obtain the fourth position coordinate of each second position coordinate in the coordinate system used by the second bird's-eye view. The rigid transformation matrix between the coordinate system used by the first bird's-eye view and the coordinate system used by the second bird's-eye view is calculated. Using the rigid transformation matrix, the fifth position coordinates of each third position coordinate in the coordinate system used in the second bird's-eye view are obtained; the error distance between all fourth position coordinates and all fifth position coordinates is calculated; it is determined whether the error distance is less than a preset error distance threshold, and a judgment result is obtained; when the judgment result is yes, the transformation relationship between the images captured by the visible light camera and the infrared camera at the same time is determined based on all fourth position coordinates and all fifth position coordinates; using the transformation relationship, spatial registration is performed on the images captured by the visible light camera and the infrared camera at the same time.

2. The spatial registration method for visible light images and infrared images according to claim 1, characterized in that, The process of obtaining the first transformation matrix includes: using a time-aligned visible light camera to track targets, obtaining the trajectory of each second preset target and performing straight line fitting, fitting the first straight line corresponding to each second preset target, and determining the first transformation matrix based on the two outermost first straight lines.

3. The spatial registration method for visible light images and infrared images according to claim 1, characterized in that, The process of obtaining the second transformation matrix includes: using a time-aligned infrared camera to track targets, obtaining the trajectory of each third preset target and performing straight line fitting, fitting the second straight line corresponding to each third preset target, and determining the second transformation matrix based on the two outermost second straight lines.

4. A spatial registration method for visible light images and infrared images according to any one of claims 1 to 3, characterized in that, Calculating the rigid transformation matrix between the coordinate system used in the first bird's-eye view and the coordinate system used in the second bird's-eye view includes: using the ICP algorithm or the CPD algorithm to calculate the rigid transformation matrix between the coordinate system used in the first bird's-eye view and the coordinate system used in the second bird's-eye view.

5. A spatial registration system for visible light images and infrared images, characterized in that, The system includes an image acquisition module, a target recognition module, a bird's-eye view acquisition module, a position coordinate transformation module, a rigid transformation matrix calculation module, an error distance calculation and judgment module, a transformation relationship determination module, and a spatial registration module. The image acquisition module is used to: after time alignment of a visible light camera and an infrared camera, simultaneously acquire a first visible light image and a first infrared image using the visible light camera and the infrared camera, respectively. The target recognition module is used to: perform target recognition on the first visible light image to obtain the first position coordinates of each first preset target in the coordinate system used by the visible light camera; and perform target recognition on the first infrared image to obtain the second position coordinates of each first preset target in the coordinate system used by the infrared camera. The bird's-eye view acquisition module is used to: perform inverse perspective projection transformation on the first visible light image based on a first transformation matrix to obtain a first bird's-eye view. The position coordinate transformation module is used to: transform each first position coordinate using the first transformation matrix to obtain a third position coordinate for each first position coordinate in the coordinate system used by the first bird's-eye view; the bird's-eye view acquisition module is further used to: perform inverse perspective projection transformation on the first infrared image based on the second transformation matrix to obtain a second bird's-eye view; the position coordinate transformation module is further used to: transform each second position coordinate using the second transformation matrix to obtain a fourth position coordinate for each second position coordinate in the coordinate system used by the second bird's-eye view; the rigid transformation matrix calculation module is used to: calculate the rigid transformation matrix between the coordinate system used by the first bird's-eye view and the coordinate system used by the second bird's-eye view; the position coordinate transformation... The transformation module is further configured to: use the rigid transformation matrix to obtain the fifth position coordinates of each third position coordinate in the coordinate system used in the second bird's-eye view; the error distance calculation and judgment module is configured to: calculate the error distance between all fourth position coordinates and all fifth position coordinates; determine whether the error distance is less than a preset error distance threshold, and obtain a judgment result; the transformation relationship determination module is configured to: when the judgment result is yes, determine the transformation relationship between the images captured by the visible light camera and the infrared camera at the same time based on all fourth position coordinates and all fifth position coordinates; the spatial registration module is configured to: use the transformation relationship to perform spatial registration on the images captured by the visible light camera and the infrared camera at the same time.

6. A spatial registration system for visible light and infrared images according to claim 5, characterized in that, It also includes a first transformation matrix acquisition module, which is used to: use a time-aligned visible light camera to perform target tracking, obtain the trajectory of each second preset target and perform straight line fitting, fit the first straight line corresponding to each second preset target, and determine the first transformation matrix based on the two outermost first straight lines.

7. A spatial registration system for visible light and infrared images according to claim 5, characterized in that, It also includes a second transformation matrix acquisition module, which is used to: use a time-aligned infrared camera to track targets, obtain the trajectory of each third preset target and perform straight line fitting, fit the second straight line corresponding to each third preset target, and determine the second transformation matrix based on the two outermost second straight lines.

8. A spatial registration system for visible light images and infrared images according to any one of claims 5 to 7, characterized in that, The rigid transformation matrix calculation module is specifically used to: calculate the rigid transformation matrix between the coordinate system used in the first bird's-eye view and the coordinate system used in the second bird's-eye view using the ICP algorithm or the CPD algorithm.

9. An electronic device, characterized in that, The device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the spatial registration method for visible light images and infrared images as described in any one of claims 1 to 4.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the spatial registration method for visible light images and infrared images as described in any one of claims 1 to 4.

Citation Information

Patent Citations

  • Registration method for visible light and infrared images based on multi-scale segmentation and SIFT (Scale Invariant Feature Transform)

    CN103337077A

  • An infrared image and visible light image registration method based on deep learning

    CN109448035A

  • Infrared and visible light image registration method in electric power inspection scene

    CN113628261A

  • Infrared image and visible light image registration method and device and readable storage medium

    CN111667520A

  • Infrared and visible light image adaptive fusion alignment method and system

    CN114255197A