Infrared and visible light image registration method fusing depth estimation
By integrating depth estimation technology, depth information is added to infrared and visible images, and a hierarchical screening strategy is used to filter feature points, solving the error problem caused by feature instability in infrared and visible images registration, achieving more accurate and reliable image registration.
Patent Information
- Application Number
- CN202411962820.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-05-02
- Estimated Expiration
- 2044-12-30
AI Technical Summary
The significant modal difference between infrared and visible light images leads to challenges in feature extraction, matching and geometric correction, especially in dynamic scenes or weak texture backgrounds, which may lead to registration errors due to feature instability.
By fusion depth estimation, depth information is added to the image to be registered, and a hierarchical filtering strategy is used to filter feature points to reduce mismatch and achieve image registration.
The amount of prior information for matching feature point selection tasks is effectively increased, making feature point matching more accurate, and the registration results are more reliable, and matching accuracy and robustness are improved.
Smart Images

Figure CN119919463A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to an infrared and visible light image registration method integrating depth estimation, and belongs to the technical field of image registration. Background Art
[0002] With the rapid development of artificial intelligence and computer vision technology, heterogeneous image registration has been widely used in the fields of multimodal fusion, target detection, autonomous driving, and night vision enhancement. The registration technology of infrared images and visible light images is an important research direction. Infrared images are mainly based on the thermal radiation characteristics of the target, with strong anti-interference and all-weather adaptability, while visible light images provide rich texture and detail information. The effective combination of the two can achieve complementary advantages and provide important support for multimodal analysis in complex scenes.
[0003] However, due to the significant modal differences between infrared and visible light images, such as inconsistencies in spectral response, texture features, and resolution, traditional registration algorithms face many challenges in feature extraction, matching, and geometric correction. In dynamic scenes or weakly textured backgrounds, existing methods may cause registration errors due to feature instability. In addition, the high computational cost of complex algorithms limits their application in scenes with high real-time requirements. Summary of the invention
[0004] The technical problem to be solved by the present invention is to provide an infrared and visible light image registration method that integrates depth estimation, adds depth information to the image to be registered by fusing the depth map, and screens the feature points matched by the deep learning network through a hierarchical screening strategy, effectively reducing mismatches and achieving the final image registration.
[0005] The present invention adopts the following technical solutions to solve the above technical problems:
[0006] A method for registering infrared and visible light images by integrating depth estimation comprises the following steps:
[0007] Step 1, preprocessing the infrared image and visible light image to be registered, and obtaining depth maps of the preprocessed infrared and visible light images;
[0008] Step 2: Fusing the infrared image to be registered with its corresponding depth map to obtain an infrared image containing depth information; fusing the visible light image to be registered with its corresponding depth map to obtain a visible light image containing depth information;
[0009] Step 3, segmenting the infrared image and the visible light image containing the depth information according to a preset number of rows and columns, so that the infrared image blocks obtained by segmentation have the same size as the visible light image blocks, and treating the infrared image blocks and the visible light image blocks at the same position in the respective images as a pair of image pairs;
[0010] Step 4: extract and match feature points for each pair of images to obtain an initial set of matching point pairs;
[0011] Step 5, set the D0 value, perform a first screening on the initial set of matching point pairs according to D0, delete the matching point pairs that fail the screening, and obtain the matching point pair set after the first screening;
[0012] Step 6, setting the D1 value, performing a second screening on the set of matching point pairs after the first screening according to D1, deleting the matching point pairs that fail the screening, and obtaining the set of matching point pairs after the second screening;
[0013] Step 7, for the image pairs that do not obtain matching point pairs after feature point extraction and matching in step 4, and the image pairs that do not have matching point pairs after the second screening in step 6, matching point compensation is performed through a compensation mechanism so that each pair of images has at least one pair of matching point pairs, thereby obtaining a final set of matching point pairs;
[0014] Step 8: According to the final set of matching point pairs, the infrared image to be registered is subjected to thin-sample strip transformation to obtain the registered infrared image.
[0015] Compared with the prior art, the present invention adopts the above technical solution and has the following technical effects:
[0016] 1. The present invention adds depth information to the image to be registered by fusing the original image with the depth map, which can effectively increase the amount of prior information for the matching feature point selection task, making the feature point matching more accurate and the registration result more reliable.
[0017] 2. The present invention can effectively check the matching errors that occur in the feature point extraction and matching stages through a multi-level screening mechanism, making the registration result more stable and reliable and improving the matching accuracy.
[0018] 3. The present invention performs matching feature point compensation on the area where the matching points are missing according to the successfully matched feature points, which can effectively remedy the situation where the multispectral image features do not correspond, so that the registration algorithm has higher robustness.
[0019] 4. The method of the present invention has a wide range of practical value as a prerequisite for any image processing task that requires registration of multispectral images, such as target detection and target tracking. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 It is a flow chart of a method for registering infrared and visible light images by integrating depth estimation according to the present invention;
[0021] Figure 2 It is a schematic diagram of the deep feature extraction and fusion process;
[0022] Figure 3 It is a flowchart of twice screening matching points. DETAILED DESCRIPTION
[0023] The embodiments of the present invention are described in detail below, and examples of the embodiments are shown in the accompanying drawings. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and cannot be interpreted as limiting the present invention.
[0024] The present invention provides an infrared and visible light image registration method integrating depth estimation, which can effectively realize the infrared and visible light image registration task by integrating the depth image information into the original image and performing hierarchical screening on the matching points. Figure 1 As shown, the following steps are included:
[0025] Step 1: Perform monocular depth estimation on the infrared and visible light images respectively to obtain respective depth maps. In the present invention, the monocular depth estimation method uses Depth-Anything-V2, and uses its official pre-training model to obtain the relative depth image required by the present invention.
[0026] It should be noted that the depth estimation method is not limited to Depth-Anything-V2, and any monocular depth estimation method can be used as the depth estimation method adapted by the present invention.
[0027] Step 2: Fuse the visible light image and infrared image with the depth map generated by them, such as Figure 2 As shown, infrared and visible light images containing depth information are generated.
[0028] It should be noted that the present invention realizes image fusion by first extracting contours from the depth map and then weighted fusion of the original image, depth map and contour map. This is a fusion method with lower complexity. In fact, this module can also be replaced by various image fusion methods such as wavelet transform image fusion, multi-scale pyramid image fusion and even a special image fusion method based on deep learning network.
[0029] Step 3: Segment the visible light and infrared images to generate a series of image pairs of suitable sizes for feature point extraction and matching.
[0030] Specifically, the specified cols and rows values are set to determine the number of image pairs to be segmented, where cols represents the number of columns into which the original image is segmented, and rows represents the number of rows into which the original image is segmented. cols*rows represents the total number of image pairs to be segmented.
[0031] Step 4: Extract and match feature points for each pair of images to obtain an initial set of matching points.
[0032] It should be noted that the present invention adopts the LoFTR deep learning network to collect feature matching points, denoises each pair of images, and then converts the array in numpy format into a tensor tensor, adopts the LoFTR feature matching deep learning network to extract the initial feature point set, and converts the result back to numpy format.
[0033] In fact, classic feature point extraction and matching methods such as SIFT feature point extraction and matching, Harris corner detection, etc. can replace this module, but the cost is that there will be differences in accuracy.
[0034] Step 5: Perform the first matching point screening: perform D0 screening on the matching points in each pair of images, update the matching point set, and delete the matching point pairs that do not pass the screening.
[0035] Specifically, first set the value of D0, the calculation formula is:
[0036] D0=w / (8*cols)
[0037] Among them, w represents the width of the infrared and visible light images, and cols represents the number of columns into which the original image is divided. Then calculate the length of the matching line of the matching point pair in each pair of images, select the first three matching point pairs with the shortest matching line length, and for these three matching point pairs, calculate the distance between any two matching points in each image. If there are matching points with a distance less than D0, delete the matching point pair with a longer matching line and update the matching point set.
[0038] Step 6: Perform the second matching point screening: Perform D1 screening on the matching points that have passed the first matching point screening, update the matching point set, and delete the matching point pairs that have not passed the screening.
[0039] Specifically, the positions of the matching points that have passed the first matching point screening are mapped to the original visible light and infrared images to generate a new matching point set, and the value of D1 is set. The calculation formula is:
[0040] D1=h / (4*rows)
[0041] Among them, h represents the height of the infrared and visible light images, and rows represents the number of rows into which the original image is divided.
[0042] Calculate the absolute value of the slope of the matching line. The calculation formula is:
[0043] grad=|(y vi -y ir ) / (x vi -x ir)|
[0044] Among them, y vi Represents the ordinate of the matching point in the visible light image, y ir Represents the ordinate of the matching point in the infrared image, x vi Represents the horizontal coordinate of the matching point in the visible light image, x ir Represents the horizontal coordinate of the matching point in the infrared image. If the absolute value of the slope of the matching point is greater than the preset threshold thres, the matching point (x i ,y i ) and the shortest distance L between other matching points i , the shortest distance calculation formula is:
[0045]
[0046] If the distance is within the specified distance range D1, the matching point pair is deleted.
[0047] The two matching point screening processes involved in step 5 and step 6 are as follows: Figure 3 As shown, specifically, the first matching point screening in step 5 focuses on the selection of local feature points, while the second matching point screening in step 6 focuses on the global matching point screening.
[0048] Step 7, compensation of matching points: retrieve the segmented image pairs without matching points after screening to obtain their areas in the original image, add new matching points to the corresponding areas through the compensation mechanism, and update the matching point set. The specific method is as follows: for visible light images, find the original image area corresponding to the image pairs without matching points after the second matching point screening, find the matching point (x1, y1) closest to its center point (a1, b1), and the calculation formula is:
[0049]
[0050] Among them, x i ,y i Respectively represent the horizontal and vertical coordinates of the matching points that remain after two screenings. For infrared images, the matching point of the center point (a1, b1) on the infrared image is (a2, b2), and a2 = a1 + x2 - x1, b2 = b1 + y2 - y1; finally, (a1, b1) and (a2, b2) are added to the matching point set as the compensated matching point pair.
[0051] It should be noted that step 7 sets an anchor point for the area without matching points to avoid abnormal distortion of the infrared image when using a thin sample strip in the subsequent step 8, so as to make the registration error smaller.
[0052] Step 8, perform registration transformation: According to the final set of matching points, perform thin-sample strip transformation on the infrared image to obtain the registered infrared image. Regarding the thin-sample strip transformation, specifically, with the help of the correspondence between the matching points, a smooth function is established to minimize the deformation energy, thereby achieving accurate mapping of the source image to the target image. The thin-sample strip transformation has both global smoothness and local flexibility, can effectively handle nonlinear distortion, and is particularly suitable for solving complex deformation problems caused by modal differences between infrared images and visible light images.
[0053] Based on the same inventive concept, an embodiment of the present application provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the aforementioned infrared and visible light image registration method for fusion depth estimation are implemented.
[0054] Based on the same inventive concept, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the steps of the aforementioned infrared and visible light image registration method for fusion depth estimation are implemented.
[0055] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0056] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0057] These computer program instructions may also be stored in a computer readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture including an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0058] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process in the computer or other programmable device. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0059] The above embodiments are only for illustrating the technical idea of the present invention, and cannot be used to limit the protection scope of the present invention. Any changes made on the basis of the technical solution in accordance with the technical idea proposed by the present invention shall fall within the protection scope of the present invention.
Claims
1. A method for infrared and visible light image registration integrating depth estimation, characterized in that: The steps include: Step 1, preprocessing the infrared image and visible light image to be registered, and obtaining depth maps of the preprocessed infrared and visible light images; Step 2: Fusing the infrared image to be registered with its corresponding depth map to obtain an infrared image containing depth information; fusing the visible light image to be registered with its corresponding depth map to obtain a visible light image containing depth information; Step 3, segmenting the infrared image and the visible light image containing the depth information according to a preset number of rows and columns, so that the infrared image blocks obtained by segmentation have the same size as the visible light image blocks, and treating the infrared image blocks and the visible light image blocks at the same position in the respective images as a pair of image pairs; Step 4: extract and match feature points for each pair of images to obtain an initial set of matching point pairs; Step 5, set the D0 value, perform a first screening on the initial set of matching point pairs according to D0, delete the matching point pairs that fail the screening, and obtain the matching point pair set after the first screening; Step 6, setting the D1 value, performing a second screening on the set of matching point pairs after the first screening according to D1, deleting the matching point pairs that fail the screening, and obtaining the set of matching point pairs after the second screening; Step 7, for the image pairs that do not obtain matching point pairs after feature point extraction and matching in step 4, and the image pairs that do not have matching point pairs after the second screening in step 6, matching point compensation is performed through a compensation mechanism so that each pair of images has at least one pair of matching point pairs, thereby obtaining a final set of matching point pairs; Step 8: According to the final set of matching point pairs, the infrared image to be registered is subjected to thin-sample strip transformation to obtain the registered infrared image.
2. The infrared and visible light image registration method integrating depth estimation according to claim 1, characterized in that: The specific process of step 1 is as follows: The infrared image and visible light image to be registered are preprocessed so that the resolution of the preprocessed infrared and visible light images are the same. The Depth-Anything-V2 model is used to perform monocular depth estimation on the preprocessed infrared and visible light images respectively to obtain the depth maps corresponding to the preprocessed infrared and visible light images.
3. The infrared and visible light image registration method integrating depth estimation according to claim 1, characterized in that: The specific process of step 2 is as follows: When fusing the infrared image to be registered with its corresponding depth map, firstly extract the contour of the depth map to obtain a contour map; then fuse the infrared image to be registered, the depth map and the contour map in a weighted fusion manner to obtain an infrared image containing depth information; When fusing the visible light image to be registered with its corresponding depth map, the depth map is firstly contour extracted to obtain a contour map; then the visible light image to be registered, the depth map and the contour map are fused in a weighted fusion manner to obtain a visible light image containing depth information.
4. The infrared and visible light image registration method integrating depth estimation according to claim 1, characterized in that: In step 4, the LoFTR deep learning network is used to extract and match feature points for each pair of images.
5. The infrared and visible light image registration method integrating depth estimation according to claim 1, characterized in that: The specific process of step 5 is as follows: Set the D0 value, the calculation formula is: D0=w / (8*cols) Wherein, w represents the width of the infrared or visible light image to be registered, and cols represents the number of columns into which the infrared image or visible light image containing depth information is divided; Calculate the length of the matching line corresponding to each pair of matching point pairs in the initial set of matching point pairs, sort all the matching line lengths in ascending order, and select the three matching point pairs corresponding to the first three matching line lengths; For the three matching points on the infrared image among the three pairs of matching points, calculate the distance between any two matching points. If there is a distance less than D0, delete the matching point with a longer matching line between the two matching points corresponding to the distance less than D0, and delete the matching point corresponding to the deleted matching point on the visible light image; if it does not exist, do not delete it; If the remaining matching point pairs are greater than or equal to two pairs, then for several matching points on the visible light image, calculate the distance between any two matching points. If there is a distance less than D0, delete the matching point with a longer matching line between the two matching points corresponding to the distance less than D0, and delete the matching point corresponding to the deleted matching point on the infrared image; if it does not exist, do not delete it; and obtain the set of matching point pairs after the first screening.
6. The infrared and visible light image registration method integrating depth estimation according to claim 1, characterized in that: The specific process of step 6 is as follows: Mapping each matching point pair in the matching point pair set after the first screening to the infrared image and the visible light image to be registered to generate a new matching point pair set; Set the D1 value, the calculation formula is: D1=h / (4*rows) Wherein, h represents the height of the infrared or visible light image to be registered, and rows represents the number of rows into which the infrared image or visible light image containing depth information is divided; Calculate the absolute value of the slope of the matching line corresponding to each pair of matching points in the new set of matching point pairs. The calculation formula is: grad=|(and vi -and ir ) / (x vi -x ir )| Among them, y vi Represents the ordinate of the matching point in the visible light image, y ir Represents the ordinate of the matching point in the infrared image, x vi Represents the horizontal coordinate of the matching point in the visible light image, x ir Represents the horizontal coordinate of the matching point in the infrared image; Determine whether there is a matching line whose absolute value of slope is greater than the preset threshold thres. If so, calculate the matching point pair (x i ,y i ) and other matching points (x j ,y j )The shortest distance L between i , the shortest distance calculation formula is: If the shortest distance is less than D1, the matching point pair (x i ,y i ) is deleted, otherwise it is not deleted, and the set of matching point pairs after the second screening is obtained.
7. The infrared and visible light image registration method integrating depth estimation according to claim 1, characterized in that: The specific process of step 7 is as follows: For the image pairs that do not obtain matching point pairs after feature point extraction and matching in step 4, and the image pairs that do not have matching point pairs after the second screening in step 6; For the visible light image block A, the center point (a1, b1) of A is used as the matching point on A; the matching point a (x1, y1) closest to the center point (a1, b1) of the visible light image block A is found on the visible light image to be registered, and the matching point b (x2, y2) corresponding to the matching point a (x1, y1) is found on the infrared image to be registered; For the infrared image block B, the matching point of the center point (a1, b1) on the infrared image block B is (a2, b2), and a2 = a1 + x2 - x1, b2 = b1 + y2 - y1; (a1, b1) and (a2, b2) are taken as the matching point pairs of the visible light image block A and the infrared image block B, and added to the matching point pair set after the second screening to obtain the final matching point pair set.
8. A computer device comprising a memory, a processor, and a computer program stored in the memory and capable of running on the processor, characterized in that: When the processor executes the computer program, the steps of the infrared and visible light image registration method for integrating depth estimation as described in any one of claims 1 to 7 are implemented.
9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the infrared and visible light image registration method for integrating depth estimation as described in any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Registration method of infrared image and visible light image and related device
CN112102380A
Infrared and visible light image registration method based on time-space domain depth information completion
CN114926515A