A motion consistency based multi-modality imaging data spatio-temporal registration method

By constructing cross-modal invariant feature maps and applying the motion consistency principle for temporal compensation, combined with affine models and dense deformation correction, the registration deviation problem of multimodal imaging data in UAV dynamic flight is solved, achieving high-precision spatiotemporal consistency registration, which is suitable for multimodal imaging data processing on UAV platforms.

CN122134773APending Publication Date: 2026-06-02XIAN UNVERSITY OF ARTS & SCI

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
XIAN UNVERSITY OF ARTS & SCI
Filing Date
2026-05-07
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Under the dynamic flight conditions of UAVs, the registration deviation caused by time asynchrony and parallax effect of multimodal imaging data makes it difficult for existing technologies to achieve high-precision spatiotemporal consistency registration.

Method used

By constructing cross-modal invariant feature maps, using the motion consistency principle for temporal compensation, and combining an affine model for global spatial coarse registration, and further using region of interest masks and dense deformation correction of infrared images, precise alignment of multimodal images is achieved.

Benefits of technology

It significantly improves the registration accuracy and robustness of multimodal imaging data under the high dynamic conditions of UAVs, providing highly consistent input data for subsequent image fusion and target detection, and is suitable for engineering applications on UAV platforms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122134773A_ABST
    Figure CN122134773A_ABST
Patent Text Reader

Abstract

This invention relates to the field of multimodal imaging technology, and more particularly to a spatiotemporal registration method for multimodal imaging data based on motion consistency. This method effectively solves the registration deviation problem caused by time asynchrony and parallax effects between visible light and infrared images during dynamic flight of a UAV. Under the high-dynamic flight conditions of a UAV, stable alignment of multimodal images is achieved through temporal compensation driven by motion consistency constraints, cross-modal geometric coarse registration, and dense deformation correction weighted by the region of interest. The method outputs overlay maps, difference maps, deformation fields, and geometric alignment indices for verification and engineering deployment. This invention does not require complex coaxial hardware design or strict time synchronization conditions, and can significantly improve the registration accuracy and robustness of multimodal imaging data in moving scenarios, providing highly consistent input data for subsequent image fusion, target detection, and recognition, and has good engineering applicability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of multimodal imaging technology, and in particular to a spatiotemporal registration method for multimodal imaging data based on motion consistency. Background Technology

[0002] Multimodal imaging, encompassing visible light imaging and infrared thermal imaging, has significant applications in fields such as UAV reconnaissance, target detection, and situational awareness. However, due to sensor characteristics and the influence of UAV motion, the registration of visible light imaging with infrared images faces the following challenges:

[0003] 1. Timing discrepancies caused by time asynchrony: Visible light cameras and infrared cameras often have different refresh rates and exposure times. During the dynamic flight of a drone, the two sensors cannot guarantee that they will acquire images at the same moment simultaneously, resulting in a discrepancy in acquisition time. If images from different times are directly aligned, the positions of moving targets in the scene will be misaligned. Especially when the drone is moving at high speed or the observed target is moving rapidly, asynchronous sampling will exacerbate the registration error.

[0004] 2. Spatial misalignment caused by imaging parallax: Visible light cameras and infrared cameras typically have different focal lengths and fields of view, and their installation positions differ in relative pose. When the optical axes of the two cameras do not coincide, depth parallax effects occur for objects at different depths of field: the positional deviation of near-distance targets in the two images is larger, while the deviation of far-distance targets is smaller. Traditional pixel-level registration methods that assume a one-to-one correspondence between objects in the two images struggle to simultaneously align targets at different distances. They often only work well at long distances, where parallax is negligible, and registration fails at close distances.

[0005] In existing technologies, to alleviate the problem of asynchronous sampling times among multiple sensors, methods such as interpolation, extrapolation, or global clock synchronization are commonly used to map multi-source data to a unified time axis. For example, least-squares curve fitting is used to align the time reference of high-frequency sampling sensors to that of low-frequency sampling sensors. However, these methods are essentially based on the assumption of time-series interpolation and do not incorporate platform motion into the imaging modeling process. Under the conditions of highly dynamic platforms such as UAVs, it is difficult to effectively compensate for image-level temporal misalignment caused by rapid pose changes of the platform, which can easily lead to spatial position shifts of moving targets in cross-modal images.

[0006] To address the parallax problem in imaging, existing technologies typically rely on precise camera calibration parameters to calculate the spatial mapping relationship between images. One approach involves introducing optical structures such as beam splitters to allow visible light and infrared imaging to share the same optical axis, thus eliminating parallax at the hardware level. However, this type of coaxial optical imaging scheme depends on a highly integrated, customized optical system, which has high requirements for size, weight, and assembly accuracy, making it difficult to meet the engineering constraints of UAV platforms in terms of payload size, weight, and modular deployment.

[0007] At the pure algorithm level, assuming the scene is a single plane or a distant target, homography transformation is used to achieve geometric registration between thermal infrared and visible light images. However, its spatial mapping model cannot characterize the parallax changes between targets at different depths, and local registration failures are prone to occur in close-range observation or complex 3D scenes. Furthermore, the concept of stereo vision is introduced, and scene depth is estimated through binocular stereo matching to assist in multispectral image registration, which alleviates the depth parallax problem to some extent. However, such methods are usually based on the premise of synchronous image acquisition or negligible platform motion, making them difficult to directly apply to asynchronous multimodal imaging scenarios under dynamic flight conditions of UAVs.

[0008] Furthermore, existing multi-source spatiotemporal registration technologies primarily align time and space at the level of sensor state data or sparse features, mainly for positioning and environmental perception applications on low-dynamic platforms such as vehicles. Under the conditions of high-speed movement of UAVs, the imaging timing misalignment introduced by differences in refresh rate, exposure mechanism, and data path between different imaging sensors, coupled with the depth parallax caused by inconsistent focal length and field of view, significantly amplifies image-level registration errors. This makes it difficult for traditional methods based on global geometric transformation or feature similarity matching to obtain stable and reliable alignment results.

[0009] In summary, current technologies lack a method for achieving high-precision spatiotemporal registration at the image level by uniformly modeling the effects of temporal differences and depth parallax on multimodal imaging data under motion platform conditions. Therefore, in UAV multi-sensor fusion applications, it is necessary to propose a motion-consistent spatiotemporal registration method that fully utilizes UAV motion information to perform temporal and spatial collaborative correction of images from different modalities, thereby improving the accuracy and robustness of multimodal data fusion. Summary of the Invention

[0010] This invention provides a spatiotemporal registration method and system for multimodal imaging data based on motion consistency, which can effectively solve the registration deviation problem caused by time asynchrony and parallax effect between visible light images and infrared images during the dynamic flight of UAVs. Under the high dynamic flight conditions of UAVs, stable alignment of multimodal images is achieved through temporal compensation driven by motion consistency constraints, cross-modal geometric coarse registration, and ROI-weighted dense deformation correction. The system outputs overlay maps, difference maps, deformation fields, and geometric alignment indices for verification and engineering deployment.

[0011] This invention provides a spatiotemporal registration method for multimodal imaging data based on motion consistency, applicable to a motion platform equipped with a multimodal imaging sensor, comprising the following steps:

[0012] S1. Obtain visible light image sequences and infrared image sequences of the same scene, and construct an initial set of image pairs;

[0013] S2. Extract features from visible light and infrared images respectively, and construct cross-modal invariant feature maps to reduce the impact of photometric differences;

[0014] S3. Based on the principle of motion consistency, the temporal translation deviation between visible light images and infrared images is estimated using cross-modal invariant feature maps, and temporal compensation is performed on the visible light images.

[0015] S4. In the feature domain, perform global spatial coarse registration based on an affine model on the temporally compensated visible light image and infrared image to correct scale, rotation and shearing differences.

[0016] S5. Extract thermally significant regions from infrared images and construct a region of interest mask and corresponding spatial weight map;

[0017] S6. Using the spatial weight map, perform dense deformation fine registration with region of interest weighting on the coarsely registered image in the feature domain, and solve the dense displacement field to compensate for local parallax and non-rigid deformation.

[0018] S7. Perform geometric correction on the visible light image based on the dense displacement field, and output the registration result that is spatiotemporally aligned with the infrared image.

[0019] According to the motion-consistency-based spatiotemporal registration method for multimodal imaging data provided by the present invention, the formula for constructing the cross-modal invariant feature map in step S2 is as follows:

[0020] ,

[0021] Where I represents the image. This represents the magnitude of the Sobel gradient. Let α be the amplitude of the discrete Laplace operator response, and α be the weighting coefficient.

[0022] According to the spatiotemporal registration method for multimodal imaging data based on motion consistency provided by the present invention, the temporal compensation method in S3 is as follows: the translation vector from the visible light feature map to the infrared feature map is calculated in the feature domain by using the phase correlation method, the peak-to-sidelobe ratio of the correlation peak is calculated, if the peak-to-sidelobe ratio is lower than a preset threshold, the matching is determined to be unreliable, the temporal compensation is abandoned, and a translation transformation matrix is ​​constructed using the reliable translation vector to perform geometric remapping on the visible light image to achieve temporal deviation compensation.

[0023] According to the spatiotemporal registration method for multimodal imaging data based on motion consistency provided by the present invention, the method for constructing the region of interest mask in S5 is as follows: threshold segmentation is performed on the infrared image to obtain the initial thermally significant region; morphological closing operation, hole filling and area filtering are performed on the initial region to obtain a stable connected region as the region of interest mask; if the area of ​​the region of interest mask is too small or does not exist, the entire image region is used as the region of interest mask to ensure the robustness of the algorithm.

[0024] According to the spatiotemporal registration method for multimodal imaging data based on motion consistency provided by the present invention, the calculation formula for the dense deformation fine registration weighted by the region of interest in S6 is as follows:

[0025] ;

[0026] Where D(x) is the dense displacement field to be determined, w (x) F is a spatial weight map constructed based on a region of interest mask. vis and F ir λ represents the feature maps of the visible light and infrared images, respectively, λ is the regularization coefficient, and Ω is the image domain.

[0027] According to the spatiotemporal registration method for multimodal imaging data based on motion consistency provided by the present invention, the dense deformation fine registration is implemented using the multi-resolution Demons algorithm, and the number of iterations of the resolution level is gradually reduced during the iteration process, and the displacement field is smoothed and attenuated at the mask boundary of the region of interest.

[0028] According to the motion-consistency-based spatiotemporal registration method for multimodal imaging data provided by the present invention, S7, while outputting the registration result spatiotemporally aligned with the infrared image, also includes generating a visualization output, wherein the visualization output includes at least one of the following:

[0029] An overlay of an infrared image and a registered visible light image;

[0030] Absolute difference plots before and after registration;

[0031] The displacement field amplitude map obtained in the dense deformation fine registration step.

[0032] The present invention also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the above-described method for spatiotemporal registration of multimodal imaging data based on motion consistency.

[0033] The beneficial effects of this invention are:

[0034] This invention provides a spatiotemporal registration method for multimodal imaging data based on motion consistency. Addressing the spatiotemporal registration error caused by different acquisition times, imaging parameter variations, and parallax effects in dynamic platform conditions, this method constructs a collaborative registration framework that integrates temporal compensation and spatial correction. The method constructs cross-modal invariant features to estimate and compensate for temporal deviations in different modal images based on the principle of motion consistency. Building upon this, an affine model is used to perform global spatial coarse registration to correct for scale, rotation, and shearing differences. Furthermore, by combining thermally salient regions in infrared images to construct region-of-interest weights, a weighted dense deformation fine registration is performed on the coarse registration results to compensate for local parallax and non-rigid deformation, achieving accurate alignment of multimodal images under a unified temporal reference and spatial coordinate system.

[0035] This invention does not require complex coaxial hardware design or strict time synchronization conditions. It can significantly improve the registration accuracy and robustness of multimodal imaging data in motion scenarios, providing highly consistent input data for subsequent image fusion, target detection and recognition, and has good engineering applicability.

[0036] Additional aspects and advantages of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description

[0037] To more clearly illustrate the technical solutions in this invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0038] Figure 1 This is a flowchart of a spatiotemporal registration method for motion-consistent multimodal imaging data;

[0039] Figure 2 It is a comparison of raw visible light and infrared images;

[0040] Figure 3 It is an overlay of an uncorrected infrared image and a visible light image;

[0041] Figure 4It is an image of infrared and visible light superimposed after spatiotemporal registration;

[0042] Figure 5 This is a diagram showing the absolute difference between the infrared image and the visible light image before spatiotemporal registration;

[0043] Figure 6 This is a diagram showing the absolute difference between the infrared image and the visible light image after spatiotemporal registration.

[0044] Figure 7 It is a dense deformation field amplitude diagram;

[0045] Figure 8 It is a display image of a spatiotemporally registered visible light image and an infrared image fused together. Detailed Implementation

[0046] Example 1

[0047] This embodiment provides a spatiotemporal registration method for multimodal imaging data based on motion consistency, and the specific steps are as follows:

[0048] S1: Data Acquisition and Input Definition:

[0049] Acquire a sequence of visible light images of the same scene With infrared image sequence For each image pair, the visible light image is converted to grayscale, and the infrared and visible light images are scaled to the same size. Each frame contains an image matrix, which supports image formats such as PNG / JPG / BMP / TIFF, a timestamp for sequence pairing, sensor basic information (resolution, frame rate, and lens parameters), and constructs an initial set of paired image pairs.

[0050] ;

[0051] Where I refers to the two-dimensional pixel matrix corresponding to a single frame image, vis means visible light, ir means infrared light, t is time, and P refers to the cross-modal paired image set, that is, the set of all time-synchronized visible light and infrared image pairs.

[0052] S2: Cross-modal invariant feature construction, gradient-Laplacian fusion:

[0053] To reduce the impact of the difference in luminance between visible light and infrared images, a feature map with relatively invariant modes is constructed:

[0054]

[0055] in This refers to the Sobel gradient (a method for edge detection in digital image processing), which corresponds to the edge features of the image. The discrete Laplacian operator response amplitude corresponds to the two-dimensional texture features of the image to supplement weak edge scenes, α∈(0,1), where α is the weighting coefficient, yielding infrared and cross-modal infrared invariant features F. ir =G(I ir ) and visible light invariant characteristic F vis =G(I vis ).

[0056] S3: Timing compensation based on motion consistency:

[0057] Estimating the visible-to-infrared shift in the characteristic domain using phase correlation :

[0058]

[0059] Where F and F -1 Representing Fourier transform and inverse Fourier transform, To represent complex conjugate, the peak position of r is taken to obtain a translation estimate and periodic reversal processing is performed. To suppress mismatches, the peak-to-sidelobe ratio (PSLR) is introduced:

[0060]

[0061] Where r peak The maximum value of the correlation peak at the finest scale, µ r and σ r Here, represents the mean and standard deviation of the entire correlation plane at this scale, respectively, and ε is a numerical stability term used to avoid zero denominators and enhance numerical stability; it is typically taken as 10. -6 ~10 -8 A smaller constant;

[0062] When PSLR < θ (θ is the PSLR discrimination threshold, usually set empirically or tuned on a validation set to achieve a balance between matching accuracy and robustness), set... Without performing time-series compensation, ensuring that time-series compensation does not introduce additional errors, the final time-series compensation transformation (homogeneous coordinate T) is constructed. time ):

[0063]

[0064] Then, applying this to visible light and performing translation compensation on the camera image yields:

[0065]

[0066] This represents an image geometric remapping operator used to spatially rearrange the original image based on predicted motion transformation relationships, thereby achieving data consistency across time and sensors. ir,1 Indicates infrared image Through timing transformation After geometric remapping, an infrared image aligned in a unified spatiotemporal coordinate system is obtained.

[0067] S4: Coarse spatial registration based on affine model:

[0068] Since visible light and infrared are separate lenses, focal length, field of view, installation angle and slight baseline deviations will cause global scale, rotation and shear differences. If they directly enter dense deformation, the displacement field is prone to bear too much global error and cause instability.

[0069] Therefore, this step first performs coarse affine registration in the feature domain to "linearly absorb" the main global differences, and then performs affine registration on the visible light image in the feature domain, with the affine model T... aff :

[0070]

[0071] With F ir For a fixed graph, with G(I) vis,1 Given a moving graph, find the T that optimizes the similarity metric over the feature domain. aff :

[0072]

[0073] D uses a similarity metric from a single-modal placement optimizer. If the affine solution fails, it degenerates into a translation model to ensure robustness.

[0074] Get T aff Perform geometric transformations on visible light images:

[0075]

[0076] First, align the global differences and compress the remaining errors into local residuals (parallax / depth difference) to construct suitable initial values ​​for the next step of dense deformation.

[0077] S5: Construction and Weighting of Regions of Interest (ROIs) for Parallax and Moving Targets

[0078] In drone scenarios, backgrounds such as grass and sky exhibit significant differences in visible light / infrared response and possess rich textures, easily inducing mismatches. Simultaneously, tasks often focus on thermal targets, such as the drone body, ground heat sources, and their surroundings. Therefore, this step uses infrared image I... ir,0 Constructing a thermally significant ROI mask Ω roi The algorithm uses quantile thresholds to form candidate hot regions and obtains stable connected regions through closing operations, hole filling, and area filtering. When the ROI is too small or unstable, it degenerates into the whole graph to ensure the robustness of the algorithm.

[0079] Get Ω roi Construct a weighted graph:

[0080]

[0081] Where w roi The target region weight is used to enhance target region alignment. bg The weights for the background region are used to suppress background interference, w roi >w bg This indicates that the alignment of structures within the ROI is prioritized during optimization. This weight map will serve as the weight input for the cost function in the dense deformation stage, so that the dense displacement field is mainly used to correct the local parallax of the target region, rather than being driven by background noise across the entire map.

[0082] S6: ROI-weighted dense deformation Demons fine registration

[0083] Even after temporal compensation and affine coarse registration, local parallax still exists due to variations in focal length / field of view and scene depth, manifesting as position-dependent non-rigid misalignment. A dense displacement field D(x)=[u(x),v(x)] is introduced. T The weighted energy is solved over the feature domain as follows:

[0084]

[0085] Where λ is the regularization coefficient, the first term indicates that structural consistency is prioritized within the ROI; the second term is the smoothing regularization term, which suppresses displacement field oscillations, folding and noise amplification, making the estimation results continuous and interpretable.

[0086] The engineering implementation uses multi-resolution Demons, with the number of iterations decreasing layer by layer, and attenuation is introduced at the ROI boundary to avoid discontinuous displacement at the mask boundary.

[0087] After obtaining D(x), dense remapping is performed on the coarsely registered visible light to obtain the final aligned output. :

[0088]

[0089] S7: Output Image Generation

[0090] For each image pair, the output includes an overlay map, a difference map, and a deformation field amplitude map. The overlay map maps infrared and visible light to different color channels respectively.

[0091]

[0092] If geometric alignment improves, the red-green separation between the target outline and the structure edge should decrease, and the overlapping area should show a stronger yellow tint. Overlay images can verify the improvement in geometric consistency.

[0093] Difference plots are used to observe whether the residual distribution converges in the target region:

[0094]

[0095] The deformation field amplitude map represents the intensity of the local geometric correction required to achieve alignment.

[0096] .

[0097] The process includes generating a visualization output along with the registration result spatiotemporally aligned with the infrared image. The visualization output includes at least one of the following:

[0098] An overlay of an infrared image and a registered visible light image;

[0099] Absolute difference plots before and after registration;

[0100] The displacement field amplitude map obtained in the dense deformation fine registration step.

[0101] The flowchart of the spatiotemporal registration method for multimodal imaging data based on motion consistency provided in this embodiment is as follows: Figure 1 As shown.

[0102] Example 2

[0103] This embodiment uses a set of visible light images and infrared thermal imaging images collected in a UAV ground test scenario as an example to illustrate the implementation process and effect of the spatiotemporal registration method for multimodal imaging data proposed in Embodiment 1.

[0104] In this embodiment, the visible light imaging device and the infrared imaging device are installed on the same UAV platform or adjacent observation positions to perform synchronous or quasi-synchronous imaging of the same ground scene. The visible light image is a three-channel color image, which is used for registration calculation after grayscale processing, and its original resolution is approximately M×N pixels; the infrared image is a single-channel grayscale thermal image, with a different resolution than the visible light image, resulting in differences in field of view, imaging scale, and imaging distortion. To facilitate subsequent processing, the visible light image is first resampled to the same size as the infrared image, and the infrared image is used as the reference coordinate system.

[0105] Figure 2The visible light (b) and infrared (a) images input in this embodiment are shown. The characteristic targets in the infrared image include a human target in the center of the image and a drone target in front of it. Both exhibit obvious high-brightness thermal features in the infrared image. The background area mainly consists of the ground, buildings, and distant trees, with a relatively smooth overall texture but low-frequency brightness variations. The visible light image acquired at the corresponding time, after grayscale processing, clearly shows the human outline, drone structure, and ground texture. However, due to the difference in imaging modality, the target brightness distribution is significantly inconsistent with that of the infrared image.

[0106] Due to time delays and platform movement during the acquisition process of visible light and infrared imaging equipment, obvious geometric misalignment can be observed when two images are directly superimposed. Figure 3 The superposition effect of uncorrected infrared and visible light images is presented. The positions of human and drone targets cannot be accurately superimposed in the two modes, and the background structure also shows overall offset and local stretching.

[0107] To address the aforementioned issues, this embodiment first performs motion consistency analysis on the two images based on cross-modal invariance features. By estimating the overall translational relationship between the two images in the feature domain, temporal compensation is performed on the visible light image to eliminate the main translational errors caused by inconsistent acquisition times. After completing temporal compensation, affine coarse registration is performed on the visible light and infrared images in the feature domain to correct rotation, scale, and shearing deviations caused by differences in viewing angle and overall attitude changes.

[0108] Even after completing the global geometric correction described above, local misalignment and non-rigid distortion can still be observed in the target area and at the edges of the image. Therefore, this embodiment further extracts thermally significant regions from the infrared image as regions of interest (ROIs). These ROIs primarily cover the human and UAV targets and their surrounding areas. Based on this, spatial weighted constraints are constructed to give the target region a higher matching priority in the subsequent registration process.

[0109] Subsequently, dense deformation correction is performed in the weighted feature domain to compensate for non-rigid deformation in the visible light image, correcting local misalignments caused by depth parallax, lens distortion, and non-planar structures in the scene. The dense displacement field output by this step reflects the amount of geometric correction required for each pixel position of the visible light image relative to the infrared reference image.

[0110] Figure 4 The image shows the superimposed result of the infrared and visible light images after all spatiotemporal registration steps have been completed, and... Figure 3 In comparison, it can be clearly observed that the human body outline and drone structure are more accurately aligned in both modes, the target boundary misalignment is significantly reduced, and the overall structural consistency of the background area is also significantly improved.

[0111] To further illustrate the registration effect more intuitively, this embodiment also provides difference images before and after registration.

[0112] Figure 5 The absolute difference results between the infrared image and the visible light image before registration are shown. There are large areas of high difference values ​​in the human body and drone regions, indicating that the geometric correspondence between the two images is poor in these regions.

[0113] Figure 6 The absolute difference results between the registered infrared image and the visible light image are shown. It can be observed that the difference intensity between the target area and the background area is significantly reduced, indicating that the two images are more consistent in spatial position after registration.

[0114] also, Figure 7 The diagram shows the dense deformation field amplitude obtained in this embodiment, which reflects the distribution of displacement magnitude of each pixel position during the registration process of the visible light image. It can be seen from the diagram that larger displacements are mainly concentrated in the human body, the drone, and their surrounding areas, while the displacement changes in the background area are relatively gradual. This indicates that the dense deformation correction is mainly used to compensate for target-related local parallax, while maintaining the continuity and stability of the overall geometric structure.

[0115] After registration is completed, this embodiment also fuses the registered visible light image with the infrared image for display, resulting in... Figure 8 The fused image shown is an example of this process. This fusion result preserves the structural details of the visible light image while accurately superimposing infrared thermal information onto the corresponding target location, thus providing more consistent and reliable multimodal input data for subsequent target detection, recognition, and situational awareness.

[0116] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention in any way. Any simple modifications, alterations, and equivalent changes made to the above embodiments based on the inventive essence shall still fall within the protection scope of the present invention.

Claims

1. A spatiotemporal registration method for multimodal imaging data based on motion consistency, applied to a motion platform equipped with a multimodal imaging sensor, characterized in that, Includes the following steps: S1. Obtain visible light image sequences and infrared image sequences of the same scene, and construct an initial set of image pairs; S2. Extract features from visible light and infrared images respectively, and construct cross-modal invariant feature maps to reduce the impact of photometric differences; S3. Based on the principle of motion consistency, the temporal translation deviation between visible light images and infrared images is estimated using cross-modal invariant feature maps, and temporal compensation is performed on the visible light images. S4. In the feature domain, perform global spatial coarse registration based on an affine model on the temporally compensated visible light image and infrared image to correct scale, rotation and shearing differences. S5. Extract thermally significant regions from infrared images and construct a region of interest mask and corresponding spatial weight map; S6. Using the spatial weight map, perform dense deformation fine registration with region of interest weighting on the coarsely registered image in the feature domain, and solve the dense displacement field to compensate for local parallax and non-rigid deformation. S7. Perform geometric correction on the visible light image based on the dense displacement field, and output the registration result that is spatiotemporally aligned with the infrared image.

2. The spatiotemporal registration method for multimodal imaging data based on motion consistency according to claim 1, characterized in that, The formula for constructing the cross-modal invariant feature map described in S2 is: , Where I represents the image. This represents the magnitude of the Sobel gradient. Let α be the amplitude of the discrete Laplace operator response, and α be the weighting coefficient.

3. The spatiotemporal registration method for multimodal imaging data based on motion consistency according to claim 1, characterized in that, The timing compensation method described in S3 is as follows: the translation vector from the visible light feature map to the infrared feature map is calculated in the feature domain using the phase correlation method, the peak-to-sidelobe ratio of the correlation peak is calculated, and if the peak-to-sidelobe ratio is lower than a preset threshold, the matching is determined to be unreliable, and the timing compensation is abandoned. A translation transformation matrix is ​​constructed using the reliable translation vector, and the visible light image is geometrically remapped to achieve timing deviation compensation.

4. The spatiotemporal registration method for multimodal imaging data based on motion consistency according to claim 1, characterized in that, The method for constructing the region of interest mask described in S5 is as follows: threshold segmentation is performed on the infrared image to obtain the initial thermally significant region, and morphological closing operation, hole filling and area filtering are performed on the initial region to obtain a stable connected region as the region of interest mask. If the area of ​​the region of interest mask is too small or does not exist, the entire image region is used as the region of interest mask to ensure the robustness of the algorithm.

5. The spatiotemporal registration method for multimodal imaging data based on motion consistency according to claim 1, characterized in that, The calculation formula for the dense deformation fine registration weighted by the region of interest described in S6 is as follows: ; Where D(x) is the dense displacement field to be determined, w (x) F is a spatial weight map constructed based on a region of interest mask. vis and F ir λ represents the feature maps of the visible light and infrared images, respectively, λ is the regularization coefficient, and Ω is the image domain.

6. The spatiotemporal registration method for multimodal imaging data based on motion consistency according to claim 5, characterized in that, The dense deformation fine registration is implemented using the multi-resolution Demons algorithm, and the number of iterations of the resolution level is gradually reduced during the iteration process. Furthermore, the displacement field is smoothly attenuated at the mask boundary of the region of interest.

7. The spatiotemporal registration method for multimodal imaging data based on motion consistency according to claim 1, characterized in that, S7, while outputting the registration result spatiotemporally aligned with the infrared image, also includes generating a visualization output, which includes at least one of the following: An overlay of an infrared image and a registered visible light image; Absolute difference plots before and after registration; The displacement field amplitude map obtained in the dense deformation fine registration step.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the method as described in any one of claims 1 to 7.