Infrared image and LiDAR point cloud registration method based on hypergraph convolutional network

The hypergraph structure and learning deformation field are constructed through hypergraph convolutional networks, which solves the problem of insufficient alignment accuracy of existing methods in dynamic scenarios, and realizes high-precision alignment of infrared images with LiDAR point clouds, enhancing the robustness of the model and the ability to adapt to complex dynamic environments.

CN120259379APending Publication Date: 2025-07-04KUNMING UNIV OF SCI & TECH
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510198001.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-21
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The existing cross-modal data registration methods have limited effects in dynamic scenarios, and it is difficult to adapt to the differences in LiDAR and infrared imaging due to acquisition frequency, dynamic changes of objects and non-rigid deformation, resulting in insufficient alignment accuracy.

Method used

The hypergraph convolutional network is used to build a hypergraph structure, extract cross-modal consistency features, and learn deformation fields through convolutional gated loop network to achieve high-precision alignment of infrared images and LiDAR point clouds. The hypergraph structure and hypergraph convolutional network are used to enhance the correlation of cross-modal features, estimate non-rigid deformation, and compensate for the differences in multimodal data under dynamic perturbation.

Benefits of technology

The correlation and alignment accuracy of cross-modal data features are improved, the model's adaptability to dynamic scenes is enhanced, and high-precision and dense alignment is achieved in complex dynamic environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259379A_ABST
    Figure CN120259379A_ABST
Patent Text Reader

Abstract

The invention relates to an infrared image and LiDAR point cloud registration method based on a hypergraph convolutional network, and belongs to the technical field of point cloud and image registration. The method comprises the following steps: firstly, establishing a hypergraph structure on infrared images and LiDAR point cloud data, extracting cross-modal consistency features by using an improved hypergraph convolutional network, and estimating rigid coarse transformation between cameras; superpixels and superpoints are mapped to original pixels and points, point clouds are projected after being transformed, a deformation field is estimated through correlation features and a convolution gating network so as to perform non-rigid deformation on a projection image, and a correction image aligned with an infrared image space is generated; and finally, indexing and establishing a pixel-point bidirectional dense mapping relation. According to the method, the inconsistency of infrared radiation features and point cloud geometric features in an embedding space is overcome, the mismatching phenomenon of an existing model in a complex scene is relieved, the limitation that a static registration model based on rigid scene prior is difficult to adapt to dynamic disturbance is broken through, and accurate alignment of 2D-3D heterogeneous data in the complex dynamic scene is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for registering infrared images and LiDAR point clouds based on a hypergraph convolutional network, belonging to the technical field of point cloud and image registration. Background Art

[0002] In the technical field of point cloud and image registration, fusing different sensor data is the key to solving the problems of incomplete perception or false detection caused by factors such as viewing angle, illumination, weather, and vegetation occlusion in a single sensor. As an important link before data fusion, the main challenge of cross-modal data registration lies in how to effectively achieve accurate alignment of features between different modalities, laying a foundation for subsequent information fusion and analysis. However, the existing cross-modal alignment methods have limited effects in dynamic scenes, mainly limited by the differences in acquisition frequency, object dynamic changes, and non-rigid deformations between lidar (LiDAR) and infrared imaging. These differences make it difficult for existing methods to adapt to the perception challenges in complex dynamic environments, and there is an urgent need for targeted improvement.

[0003] The aligned data can be widely applied to tasks such as object detection, semantic segmentation, and object tracking. By combining the advantages of infrared images and LiDAR point clouds, under the premise of dense alignment, the robustness of object detection and tracking can be significantly enhanced, so as to better cope with the perception challenges in complex dynamic environments. Summary of the Invention

[0004] The technical problem to be solved by the present invention is to propose a method for registering infrared images and LiDAR point clouds based on a hypergraph convolutional network, aiming to achieve high-precision and robust alignment of infrared images and LiDAR point clouds in complex dynamic environments through hypergraph structure construction, cross-modal feature extraction, and deformation field learning, overcome the problem of insufficient alignment accuracy caused by data heterogeneity, environmental changes, and noise interference in traditional methods in dynamic scenes, and provide an effective and accurate solution for the application of multi-source sensor data registration in complex environments.

[0005] The technical solution adopted by the present invention is: a method for registering infrared images and LiDAR point clouds based on a hypergraph convolutional network, which constructs a data set by acquiring heterogeneous data, constructs a hypergraph structure and extracts cross-modal consistent features by using a feature extractor, estimates the rough camera transformation, extracts correlation features after restoring superpixels and superpoints, inputs a convolutional gated recurrent network to learn the deformation field, and obtains a dense alignment result of point-to-pixel after deformation through index matching, so as to achieve precise alignment of infrared images and LiDAR point clouds in complex dynamic scenes.

[0006] The specific steps are as follows:

[0007] Step1: Construct a dual-channel data loader to process infrared images and LiDAR point cloud data in complex dynamic scenarios respectively. Generate superpixels and superpoints through spatial clustering, and establish superedges based on spatial proximity constraints and feature similarity metrics to generate a hypergraph structure;

[0008] Step2: Input the hypergraph structure into the spatial domain hypergraph convolutional network to extract cross-modal consistent features, and learn vertex features through the superedge propagation mechanism; then use convolution to establish the probability correspondence between superpixels and superpoints, remove the mismatched point pairs below the threshold, and obtain the feature map with feature consistency in the overlapping area;

[0009] Step3: Calculate the feature similarity matrix of superpoints and superpixels in the overlapping area according to the consistent feature map, screen the matching points and use the algorithm to solve the optimal rigid transformation parameters to obtain the rotation matrix and translation vector of the rigid transformation between cross-modal data;

[0010] Step4: Map the superpixels and superpoints to the original points and pixels, project the points onto the two-dimensional plane through the rigid transformation matrix, and interpolate according to the resolution of the infrared image to obtain the projected image; extract the similarity features of the infrared image and the point cloud projected image through the dual-branch encoder, splice and input them into the convolutional gated recurrent unit to iteratively update the deformation field;

[0011] Step5: Use the updated deformation field to transform the projected point cloud image, compare the transformed result with the infrared image pixel by pixel to obtain the dense alignment result of pixel pairs, thereby constructing an alignment network for infrared images and LiDAR point clouds in complex dynamic scenarios;

[0012] Step6: Iteratively train the alignment network in combination with the loss function designed for the task to obtain the final alignment model of infrared images and LiDAR point clouds in dynamic scenarios.

[0013] The specific content of Step1 is as follows:

[0014] Step1.1: Construct superpixels by dividing the image with the original resolution into image patches of size 1 / a, and construct superpoints using a regular voxel grid, such that the density of superpoints is 1 / a of the original point cloud density;

[0015] For image nodes, each superpixel is regarded as a node, denoted as I i , and each superpoint in the point cloud is regarded as a node, denoted as P z , connect the superpixel nodes in the image using the nearest neighbor principle, and the similarity metric D pixel (I i , I j ) is as follows:

[0016]

[0017] Among them, σ1 is an adjustable parameter of the similarity metric, and D pixel (I i , I j ) is controlled between [0, 1]. I i and I j refer to different superpixel nodes, and d pixel (I i , I j ) is the coordinate distance metric between two superpixels. exp is the exponential function with base e;

[0018] For each point P i in the point cloud:

[0019]

[0020] Among them, D point (P i , P j ) is the similarity metric of the spatial distance, σ2 is an adjustable parameter of the similarity metric, and P i and P j are different superpoints, and d point (P i , P j ) is the coordinate distance metric between two superpoints;

[0021] Step1.2: The point cloud is input into the shared multi-layer perceptron MLP network, and the superpoint features are extracted through the fully connected layer. The image is input into the convolutional network CNN with a convolutional kernel size of 1*1 to extract the features of the superpixel nodes in the image. For each pair of image superpixel nodes I i or point cloud superpoint nodes P i in the spatial proximity constraint domain, the distance in the feature space is used as a metric to calculate the feature similarity:

[0022]

[0023] Among them, S point (P i , P j ) represents the feature similarity of the superpoint, S pixel (I i , I j ) represents the feature similarity of the superpixel, σ3 and σ4 are adjustable parameters of the similarity metric, represents the feature of the i-th superpoint, represents the feature of the j-th superpoint, represents the feature of the i-th superpixel, Denote the j-th superpixel feature, where i≠j; select and retain the k points with the highest feature similarity, and establish the feature similarity hyperedge ∈ between the superpixel and the superpoint ik and ∈ pk :

[0024]

[0025] where, W feature (∈ ik ) is the weight of the superpixel feature similarity hyperedge, |ε k | represents the number of nodes in the hyperedge, C(|∈ k |, 2) represents the combination number of selecting 2 elements from |∈ k | elements, Σ is the summation symbol, W feature (∈ pk ) is the weight of the superpoint feature similarity hyperedge, and the constructed hypergraph structure is G(I i , ∈ ik , W feature (∈ ik )) and G(P i , ∈ pk , W feature (∈ pk ))), where G represents the graph structure.

[0026] Specifically, Step 2 is as follows:

[0027] Step 2.1: Calculate the incidence matrix, node degree matrix, and hyperedge degree matrix. Adopt the hypergraph convolutional network in the spatial domain. In the first stage, the hyperedge features are aggregated through the vertex features they connect. In the second stage, the vertex features are updated by aggregating the features of their affiliated hyperedges;

[0028] Step 2.2: Repeat and expand the global image features according to the number of superpoints to make their dimensions consistent with the superpoint features. Concatenate the expanded global features and the features of each superpoint element-wise to form a fused feature representation. At the same time, repeat and expand the global features of the point cloud according to the number of superpixels and concatenate them with the features of each superpixel. Then apply a neural network to the concatenated features to output the probability features of each superpixel or each superpoint in the overlapping area, and screen the superpoint and superpixel features in the overlapping area that are higher than the threshold.

[0029] Specifically, Step 4 is as follows:

[0030] Step 4.1: For each superpixel, extract all the pixel sets it contains to obtain the infrared image; for each superpoint, extract all the point sets it contains; apply the rough transformation matrix to each point to obtain the transformed point cloud point set, and calculate the projection coordinates of each point on the image through the camera internal parameter focal length and principal point to obtain the projection point set;

[0031] Step 4.2: Create a blank image with the same size as the infrared image. Apply nearest neighbor interpolation to each projection point to draw points on the blank image. Input the blank image and the infrared image into the dual-branch encoder. Use a residual network with shared weights to extract multi-scale features respectively. Calculate the cross-modal hybrid correlation through dot product, construct a 4D pyramid. Given an initial displacement field estimate, input the correlation features indexed from the pyramid and the extracted context features, and use convolutional GRU units and residual connections to update iteratively.

[0032] The specific content of Step 6 is as follows:

[0033] The loss function includes an overlapping region detection loss, a gradient loss, and a deformation loss;

[0034] The overlapping region detection loss L ovlp is:

[0035]

[0036] where N in represents the number of superpoint and superpixel pairs sampled in the overlapping region, and N out represents the number of superpoint and superpixel pairs sampled in the non-overlapping region, represents the probability that the j-th superpixel is located in the overlapping region, represents the probability that the j-th superpoint is located in the overlapping region, represents the probability that the i-th superpixel is located in the non-overlapping region, represents the probability that the i-th superpoint is located in the non-overlapping region;

[0037] The gradient loss L grad is:

[0038]

[0039] where is the gradient of the projected image after deformation;

[0040] The deformation loss L def is:

[0041] L def = ||d(I proj ) - d inv (I proj )|| 2

[0042] where d(I proj ) represents the displacement field from I proj to the deformed image, and d inv (I proj ) represents the inverse transformation of the deformed image to Iproj The calculated displacement field;

[0043] Integrate the above losses to obtain the final loss L total :

[0044] L total = λ1L ovlp + λ2L grad + λ3L def

[0045] where λ1, λ2, and λ3 are the weight coefficients of each loss.

[0046] Existing static registration models based on rigid scene priors are difficult to adapt to the non-rigid deformation field changes caused by dynamic disturbances (such as moving objects and vegetation deformation), and the topological inconsistency of infrared radiation features and point cloud geometric features in the embedding space affects the robustness of heterogeneous data alignment in complex scenes. Through the present invention, the correlation of cross-modal data features can be improved, and at the same time, the application of 2D-3D data alignment methods in dynamic scenes can be expanded, enhancing the robustness while improving the accuracy of dense alignment.

[0047] The beneficial effects of the present invention are as follows: The present invention proposes an infrared image and LiDAR point cloud registration method based on a hypergraph convolutional network. By applying the hypergraph structure and hypergraph convolutional network, the correlation of cross-modal matching features is enhanced. The application of correlation features and gated convolutional recurrent networks enables the model to accurately estimate the non-rigid deformation between the projected image of the point cloud and the infrared image, compensates for the multi-modal data differences under dynamic disturbances, and improves the accuracy and robustness of matching. The above design can not only improve the accuracy of heterogeneous data alignment, but also enhance the adaptability of the model to dynamic scenes. The present invention provides an efficient, accurate and robust solution for the cross-modal data alignment task in complex dynamic scenes. Description of the Drawings

[0048] Figure 1 is the overall step flow chart of the present invention;

[0049] Figure 2 is the hypergraph construction and hypergraph convolution flow chart of the present invention. Detailed Embodiments

[0050] In order to enable those skilled in the art to better understand the solution of the present invention, the present invention will be further described below in conjunction with the drawings and specific embodiments, but the protection scope of the present invention is not limited to the described scope.

[0051] Example 1: As Figure 1 shown, an infrared image and LiDAR point cloud registration method based on a hypergraph convolutional network, the specific steps are as follows:

[0052] Step1: Construct a dual-channel data loader to process infrared images and LiDAR point cloud data in complex dynamic scenarios respectively. Generate superpixels and superpoints through spatial clustering, and establish superedges based on spatial proximity constraints and feature similarity metrics to generate a hypergraph structure;

[0053] Step1.1: Input the infrared image and LiDAR point cloud data of the mountain jungle scene captured by the drone. Construct superpixels by dividing the image with the original resolution into image patches of 1 / 8 size, and construct superpoints using a regular voxel grid, so that the density of the superpoints is 1 / 8 of the original point cloud density. Define the voxel size as:

[0054]

[0055] where, l voxel is the voxel side length, and D pc is the density of the original point cloud, that is, the number of points per unit volume;

[0056] For image nodes, each superpixel is regarded as a node, denoted as I i , and each superpoint in the point cloud is regarded as a node, denoted as P z . Adopt the k-hop strategy to connect the superpixel nodes in the image. The similarity metric D pixel (I i ,I j ) is:

[0057]

[0058] where, σ1 is an adjustable parameter of the similarity metric, and D pixel (I i ,I j ) is controlled between [0,1]. I i and I j refer to different superpixel nodes, and d pixel (I i ,I j ) is the coordinate distance metric of two superpixels, and exp is the exponential function with e as the base;

[0059] For each point P i in the point cloud:

[0060]

[0061] where, D point (P i ,P j ) is the similarity metric of spatial distance, σ2 is an adjustable parameter of the similarity metric, P i and P j are different superpoints, and dpoint (P i , P j ) is the coordinate distance metric between two super points;

[0062] Step1.2: The point cloud is input into the shared multi-layer perceptron MLP network, and the super point features are extracted through the fully connected layer. The image is input into the convolutional network CNN with a convolutional kernel size of 1*1 to extract the features of the superpixel nodes in the image. For each pair of image superpixel nodes I i or point cloud super point nodes P i , the distance in the feature space is used as a metric to calculate the feature similarity:

[0063]

[0064]

[0065] where S point (P i , P j ) represents the feature similarity of the super point, S pixel (I i , I j ) represents the feature similarity of the superpixel, σ3 and σ4 are adjustable parameters of the similarity metric, represents the feature of the i-th super point, represents the feature of the j-th super point, represents the feature of the i-th superpixel, represents the feature of the j-th superpixel, i≠j; Keep the k points with the highest feature similarity, and establish the feature similarity hyperedges ∈ ik and ∈ pk :

[0066]

[0067]

[0068] where W feature (∈ ik ) is the weight of the superpixel feature similarity hyperedge, |ε k | represents the number of nodes in the hyperedge, D(|∈ k |, 2) represents the combination number of selecting 2 elements from |∈ k | elements, Σ is the summation symbol, W feature (∈ pk ) is the weight of the super point feature similarity hyperedge, and the constructed hypergraph structure is G(I i , ∈ ik , W feature (∈ ik )) and G(P i, ∈ pk , W feature ( ∈ pk ), where G represents the graph structure.

[0069] Step2: Input the hypergraph structure into the spatial domain hypergraph convolutional network to extract cross-modal consistent features, and learn vertex features through the hyperedge propagation mechanism; then use convolution to establish the probability correspondence between superpixels and superpoints, remove the mismatched point pairs below the threshold, and obtain the feature map with feature consistency in the overlapping region;

[0070] Step2.1: Calculate the incidence matrix H ∈ R |v|×|∈| :

[0071]

[0072] where v represents a graph node, ∈ represents a hyperedge, and v ∈ ∈ means that node v is connected by hyperedge ∈;

[0073] Calculate the node degree matrix D v and the hyperedge degree matrix D ∈ :

[0074]

[0075] where diag{·} represents the operation of converting a vector into a diagonal matrix, and H(i, j) is the incidence matrix;

[0076] Adopt the HGNN + hypergraph convolutional network in the spatial domain. In the first stage, the hyperedge features are aggregated through the vertex features they connect: Y (l) = H T X (l) , where X (l) is the node feature of the l-th layer, H T is the transpose of the incidence matrix H, representing the transfer from nodes to hyperedges, and Y (l) is the feature of the hyperedge; in the second stage, the vertex features are updated by aggregating the features of the hyperedges they belong to: where X (l+1) is the node feature of the (l + 1)-th layer, and are the inverse matrices of the node degree matrix and the hyperedge degree matrix respectively, used for normalization;

[0077] Step2.2: Repeat and expand the global feature G 2D of the image according to the number of superpoints to make its dimension consistent with the superpoint features. The expanded global feature is concatenated element-wise with the feature f 3D of each superpoint to form a fused feature representation. At the same time, the global feature G 3D of the point cloud is repeated and expanded according to the number of superpixels and is combined with each superpixel feature f2D Stitch them together, then apply CNN or MLP to the stitched features, and output the probability features of each superpixel or each superpoint in the overlapping region. Filter the superpoint and superpixel features in the overlapping region that are higher than the threshold. The threshold of the present invention is set to 0.9.

[0078] Step3: Calculate the superpoint and superpixel feature similarity matrix according to the consistency feature map in the overlapping region, filter the matching points and use the algorithm to solve the optimal rigid transformation parameters to obtain the rotation matrix and translation vector of the rigid transformation between cross-modal data;

[0079] Calculate for each superpixel node I i The superpoint P with the largest feature similarity z to establish a 2D-3D rough matching relationship:

[0080] M = {(P x ∈P, I y ∈I) | y = argmin||F P (P x ) - F I (I y )||}

[0081] where M is the pair of superpoints and superpixels of the rough matching, P is the set of superpoints in the overlapping region, P x is a certain superpoint, I is the set of superpixels in the overlapping region, I y is a certain superpixel, F P is the corresponding superpoint feature, F I is the corresponding superpixel feature, argmin is the parameter value that makes the objective function obtain the minimum value, ||*|| is the distance metric of the L2 norm. According to the rough matching relationship, input the pose solution algorithm to obtain the rough transformation matrix T = R|t of the infrared image superpixel and the point cloud superpoint, where T is the rough transformation matrix, R is an n×n rotation matrix, t is an n×1 translation vector, and | represents horizontal stitching.

[0082] Step4: Map the superpixels and superpoints to the original points and pixels, project the points onto the two-dimensional plane through the rigid transformation matrix, and interpolate according to the resolution of the infrared image to obtain the projected image; extract the similarity features of the infrared image and the point cloud projected image through the double-branch encoder, stitch and input them into the convolutional gated recurrent unit to iteratively update the deformation field;

[0083] Step4.1: For each superpixel, extract all the pixel sets i = {i x} it contains to obtain the infrared image I inf ; for each superpoint, extract all the point sets p = {p y} it contains; apply the rough transformation matrix to each point: p' y = R·py +t to obtain the transformed point cloud point set p' = {p' y}, through the camera internal parameter focal lengths (f x , f y ) and the principal point (c x , c y ), for each point p' y = (x, y, z), calculate its projection coordinates on the image:

[0084]

[0085] to obtain the projection point set U = {u y , v y}, where u y is the horizontal coordinate of the projection point, and v y is the vertical coordinate of the projection point;

[0086] Step4.2: Create a blank image I proj with the same size as the infrared image, draw points on I proj for each projection point using nearest neighbor interpolation, input I proj and I inf to the dual-branch encoder, use the residual network with shared weights to extract multi-scale features respectively, the output feature map has the original resolution of 1 / 4, 1 / 8, 1 / 16, 1 / 32, calculate the cross-modal hybrid correlation through dot product, construct a 4D pyramid, given an initial displacement field estimate, input the correlation features indexed from the pyramid and extract context features, and use the convolutional GRU cell and residual connection method to iteratively update:

[0087]

[0088] F t+1 = F t + ΔF

[0089] where, h t is the hidden state of the current iteration, h t+1 is the updated hidden state, F t is the current displacement field estimate, F t+1 is the updated displacement field estimate, is the displacement field gradient, ΔF is the displacement increment, C is the correlation feature, Context is the context feature, ConvGRU is the convolutional gated recurrent unit, [] is the concatenation operation, and the estimated displacement field is the deformation field of the dynamic scene projection points to the pixels.

[0090] Step 5: Use the updated deformation field to transform the projected point cloud image. Compare the result of the transformation pixel by pixel with the infrared image to obtain the dense alignment result of pixel-to-point pairs, thereby constructing an alignment network for infrared images and LiDAR point clouds in complex dynamic scenes;

[0091] Apply the deformation field to I proj image to obtain the deformed image, which corresponds pixel by pixel to the infrared image I inf Index the projected points corresponding to each pixel of I proj That is, the matching result of pixel by pixel and point by point is obtained, realizing the dense alignment of infrared image pixels to LiDAR point cloud points in a dynamic scene.

[0092] Step 6: Iteratively train the alignment network by combining the loss function designed for the task to obtain the final alignment model of infrared images and LiDAR point clouds in dynamic scenes.

[0093] The loss function includes overlapping region detection loss, gradient loss, and deformation loss;

[0094] The overlapping region detection loss L ovlp is:

[0095]

[0096] where N in represents the number of superpoint and superpixel pairs sampled in the overlapping region, and N out represents the number of superpoint and superpixel pairs sampled in the non-overlapping region, represents the probability that the j-th superpixel is located in the overlapping region, represents the probability that the j-th superpoint is located in the overlapping region, represents the probability that the i-th superpixel is located in the non-overlapping region, represents the probability that the i-th superpoint is located in the non-overlapping region;

[0097] The gradient loss L grad is:

[0098]

[0099] where is the gradient of the projected image after deformation;

[0100] The deformation loss L def is:

[0101] L def = ||d(I proj ) - d inv (I proj )|| 2

[0102] Among them, d(I proj ) represents the displacement field from I proj to the deformed image, and d inv (I proj ) represents the displacement field calculated by inversely transforming the deformed image to I proj ;

[0103] Integrate the above losses to obtain the final loss L total :

[0104] L total = λ1L ovlp + λ2L grad + λ3L def

[0105] Among them, λ1, λ2, and λ3 are the weight coefficients of each loss.

[0106] The specific embodiments of the present invention have been described in detail above in conjunction with the accompanying drawings. However, the present invention is not limited to the above embodiments, and various changes can be made without departing from the spirit of the present invention within the scope of knowledge possessed by those of ordinary skill in the art.

Claims

1. An infrared image and LiDAR point cloud registration method based on a hypergraph convolutional network, characterized in that Specifically, it includes the following steps: Step1: Construct a dual-channel data loader to process infrared images and LiDAR point cloud data in complex dynamic scenarios respectively. Generate superpixels and superpoints through spatial clustering, and establish superedges based on spatial proximity constraints and feature similarity metrics to generate a hypergraph structure; Step2: Input the hypergraph structure into the spatial-domain hypergraph convolutional network to extract cross-modal consistent features, and learn vertex features through the superedge propagation mechanism; then use convolution to establish the probability correspondence between superpixels and superpoints, remove the mismatched point pairs below the threshold, and obtain a feature map with feature consistency in the overlapping area; Step3: Calculate the feature similarity matrix of superpoints and superpixels in the overlapping area according to the consistent feature map, screen the matching points and use the algorithm to solve the optimal rigid transformation parameters to obtain the rotation matrix and translation vector of the rigid transformation between cross-modal data; Step4: Map the superpixels and superpoints to the original points and pixels, project the points onto the two-dimensional plane through the rigid transformation matrix, and interpolate according to the resolution of the infrared image to obtain the projected image; extract the similarity features of the infrared image and the point cloud projected image through the dual-branch encoder, splice and input them into the convolutional gated recurrent unit to iteratively update the deformation field; Step5: Use the updated deformation field to transform the projected point cloud image, compare the transformed result with the infrared image pixel by pixel to obtain the dense alignment result of pixel-to-point pairs, thereby constructing an alignment network for infrared images and LiDAR point clouds in complex dynamic scenarios; Step6: Iteratively train the alignment network in combination with the loss function designed for the task to obtain the final alignment model of infrared images and LiDAR point clouds in dynamic scenarios.

2. The infrared image and LiDAR point cloud registration method based on a hypergraph convolutional network according to claim 1, characterized in that, Specifically, Step1 is as follows: Step1.1: Construct superpixels by dividing the image with the original resolution into image patches of size 1 / a, and construct superpoints using a regular voxel grid, so that the density of superpoints is 1 / a of the original point cloud density; For the image nodes, each superpixel is regarded as a node, denoted as I i , and each superpoint in the point cloud is regarded as a node, denoted as P z . The nearest neighbor principle is used to connect the superpixel nodes in the image, and the similarity metric D pixel (I i , I j ) is as follows: Among them, σ1 is an adjustable parameter of the similarity measure, and D pixel (I i ,I j ) is controlled to be between [0, 1]. I i and I j refer to different superpixel nodes, and d pixel (I i ,I j ) is the coordinate distance measure between two superpixels. exp is the exponential function with base e; For each point P in the point cloud i : Among them, D point (P i , P j ) is the similarity measure of the spatial distance, σ2 is the adjustable parameter of the similarity measure, P i and P j are different superpoints, d point (P i , P j ) is the coordinate distance measure of the two superpoints; Step1.2: The point cloud is input into a shared multi-layer perceptron (MLP) network, and hyperpoint features are extracted through a fully connected layer. The image is input into a convolutional neural network (CNN) with a convolutional kernel size of 1*1 to extract the features of superpixel nodes in the image. For each pair of image superpixel nodes I i or point cloud hyperpoint nodes I i in the spatial proximity constraint domain, the feature similarity is calculated using the distance in the feature space: Among them, S point (P i , P j ) represents the feature similarity of superpoints, and S pixel (I i , I j ) represents the feature similarity of superpixels. σ3 and σ4 are adjustable parameters for similarity measurement. represents the feature of the i-th superpoint. represents the feature of the j-th superpoint. represents the feature of the i-th superpixel. represents the feature of the j-th superpixel, where i ≠ j; keep the k points with the highest feature similarity, and establish the feature similarity hyperedges ∈ ik and ∈ pk : Among them, W feature (∈ ik ) is the hyperedge weight of the superpixel feature similarity. |ε k | represents the number of nodes in the hyperedge. C(|∈ k |, 2) represents the combination number of selecting 2 elements from |∈ k | elements. Σ is the summation symbol. W feature (∈ pk ) is the hyperedge weight of the superpoint feature similarity. The constructed hypergraph structures are G(I i , ∈ ik , W feature (∈ ik )) and G(P i , ∈ pk , W feature (∈ pk )) where G represents the graph structure.

3. A method for registering infrared images and LiDAR point clouds based on a hypergraph convolutional network according to claim 1, characterized in that, Specifically, Step2 is as follows: Step2.1: Calculate the incidence matrix, node degree matrix, and superedge degree matrix, and use the spatial-domain hypergraph convolutional network. In the first stage, the superedge features are aggregated through the vertex features connected to them, and in the second stage, the vertex features are updated by aggregating the features of the superedges to which they belong; Step2.2: Repeat and expand the global image features according to the number of superpoints to make their dimensions consistent with the superpoint features. The expanded global features are concatenated element-wise with the features of each superpoint to form a fused feature representation. At the same time, the global features of the point cloud are repeated and expanded according to the number of superpixels and concatenated with the features of each superpixel, and then the concatenated features are applied to a neural network to output the probability features of each superpixel or each superpoint in the overlapping area, and screen the superpoint and superpixel features in the overlapping area above the threshold.

4. A method for registering infrared images and LiDAR point clouds based on a hypergraph convolutional network according to claim 1, characterized in that, Specifically, Step4 is as follows: Step4.1: For each superpixel, extract the set of all pixels it contains to obtain an infrared image; for each superpoint, extract the set of all points it contains; apply the coarse transformation matrix to each point to obtain the set of transformed point cloud points, and calculate the projected coordinates of each point on the image through the focal length and principal point of the camera's internal parameters to obtain the set of projected points; Step4.2: Create a blank image with the same size as the infrared image, draw points on the blank image for each projected point using nearest neighbor interpolation, input the blank image and the infrared image into a two-branch encoder, use a residual network with shared weights to extract multi-scale features respectively, calculate the cross-modal mixing correlation through dot product, construct a 4D pyramid, given an initial displacement field estimate, input the correlation features indexed from the pyramid and the extracted context features, and iteratively update using convolutional GRU units and residual connection methods.

5. A method for registering infrared images and LiDAR point clouds based on a hypergraph convolutional network according to claim 1, characterized in that, The specific content of Step6 is as follows: The loss function includes an overlapping region detection loss, a gradient loss, and a deformation loss; Overlapping region detection loss L ovlp is as follows: Among them, N in represents the number of sampled superpoints and superpixels pairs in the overlapping region, N out represents the number of sampled superpoints and superpixels pairs in the non - overlapping region, represents the probability that the j - th superpixel is located in the overlapping region, represents the probability that the j - th superpoint is located in the overlapping region, represents the probability that the i - th superpixel is located in the non - overlapping region, represents the probability that the i - th superpoint is located in the non - overlapping region; Gradient loss L grad is as follows: Among them, is the gradient of the projected image after deformation; Deformation loss L def is as follows: L def = ||d(I proj ) - d inv (I proj )|| 2 Among them, d(I proj ) represents the displacement field from I proj to the deformed image, and d inv (I proj ) represents the displacement field calculated by inversely transforming the deformed image to I proj and the displacement field obtained by calculation; Integrate the above losses to obtain the final loss L total : L total = λ1L ovlp + λ2L grad + λ3L def Among them, λ1, λ2, and λ3 are the weight coefficients of each loss.

Citation Information

Cited By

  • Non-rigid cross-modal space correspondence method based on diffusion probability flow guidance

    CN121861127A

  • Intelligent access control system

    CN122244983A