An unsupervised infrared-visible light optical flow estimation method based on cross-modal alignment

CN122597745APending Publication Date: 2026-08-18NORTHWESTERN POLYTECHNICAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610643538.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-11
Publication Date
2026-08-18

AI Technical Summary

Technical Problem

现有无监督光流估计方法均以单一可见光模态为输入,依托光度一致性和平滑先验完成模型训练,然而该类方法在逆光、夜间、雾天等复杂城市场景中存在显著技术缺陷:可见光图像易受光照变化、低对比度、纹理退化影响,导致光度一致性假设失效,光流估计出现大量错误匹配、光流漂移问题,端点误差显著升高;同时模型在常规白天数据集训练后,在夜间或恶劣天气场景中的泛化能力急剧下降,无法满足自动驾驶的实际感知需求

Benefits of technology

一、光流估计鲁棒性大幅提升,有效弥补了可见光模态在低光照、恶劣天气下的感知缺陷,充分发挥红外模态的全天候光照鲁棒性与恶劣天气穿透性,在 KAIST 夜间场景中平均角误差(AAE)较单红外模型降低 83%,较单可见光模型降低 55%,且在遮挡区域能保持高 IoU 与 SSIM,彻底解决了单模态方法在复杂场景下的错误匹配、光流漂移问题;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122597745A_ABST
    Figure CN122597745A_ABST
Patent Text Reader

Abstract

The application discloses an unsupervised infrared-visible light optical flow estimation method based on cross-modal alignment. Firstly, a cross-dataset verification strategy is relied on to adapt to KITTI labeled data and KAIST unlabeled data mixed scenes, and reliable pseudo-supervision signals are generated. Secondly, a modal complementary alignment module based on cross-modal motion alignment is designed. Multi-scale deep features of infrared and visible light modalities are synchronously acquired through a shared pyramid extractor, and then independent motion information propagation modules are used to generate specific optical flow and motion features of each modality, and output aligned motion features, optical flow and joint occlusion map. Finally, a multi-dimensional alignment loss is introduced, and the optical flow estimation accuracy is greatly optimized through supervision of three dimensions of global adversarial distribution, local statistical consistency and cross-modal structure reconstruction. In combination with the optical flow estimation network, a fusion enhanced optical flow estimation system is formed, and finally high-precision optical flow field output is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of computer vision and multimodal perception technology, specifically relating to an unsupervised infrared-visible optical flow estimation method based on cross-modal alignment. Background Technology

[0002] Optical flow estimation is a core technology for autonomous driving perception, providing support for scene understanding, obstacle avoidance, and trajectory prediction by extracting pixel-level motion cues. Existing unsupervised optical flow estimation methods all use a single visible light modality as input and rely on photometric consistency and smoothness priors to complete model training. However, these methods have significant technical shortcomings in complex urban scenes such as backlighting, nighttime, and foggy weather: visible light images are easily affected by changes in illumination, low contrast, and texture degradation, causing the photometric consistency assumption to fail, resulting in a large number of mismatches and optical flow drift problems in optical flow estimation, and a significant increase in endpoint error; at the same time, after training on a regular daytime dataset, the model's generalization ability drops sharply in nighttime or inclement weather scenarios, failing to meet the actual perception needs of autonomous driving.

[0003] To compensate for the limitations of visible light modes, infrared images, with their all-weather robustness and penetration in adverse weather conditions, have become an ideal complementary mode to visible light. Infrared-visible light fusion technology has been applied in tasks such as image classification and object detection. However, existing fusion strategies mostly remain at the image level or shallow feature stitching stage, without exploring the motion consistency alignment mechanism required for optical flow tasks, and thus failing to effectively tap the complementary value of cross-modal motion features. In addition, existing technologies lack ground truth labeled datasets for infrared-visible dual-modal optical flow. Related cross-modal optical flow research relies on complex ground truth label selection or downstream task verification, which is cumbersome and lacks robustness.

[0004] In summary, existing unsupervised optical flow estimation methods suffer from three major problems: poor robustness in single-modal scenarios, lack of motion alignment mechanism in cross-modal fusion, and difficulty in model validation and training in dual-modal scenarios without ground truth. These problems result in low optical flow estimation accuracy and weak generalization ability in complex urban scenarios, making it difficult to implement in practical engineering applications. Summary of the Invention

[0005] To overcome the shortcomings of existing technologies, this invention provides an unsupervised infrared-visible optical flow estimation method based on cross-modal alignment. First, it utilizes a cross-dataset verification strategy adapted to mixed scenarios of KITTI labeled data and KAIST unlabeled data to generate reliable pseudo-supervisory signals. Second, it designs a modal complementary alignment module based on cross-modal motion alignment. This module first acquires multi-scale deep features of infrared and visible light modes simultaneously through a shared pyramid extractor, and then generates modality-specific optical flow and motion features through an independent motion information propagation module, outputting aligned motion features, optical flow, and a joint occlusion map. Finally, it introduces a multi-dimensional alignment loss, significantly optimizing optical flow estimation accuracy through three dimensions: supervised global adversarial distribution, local statistical consistency, and cross-modal structural reconstruction. This loss is combined with an optical flow estimation network to form a fusion-enhanced optical flow estimation system, ultimately achieving high-precision optical flow field output.

[0006] The technical solution adopted by this invention to solve its technical problem is as follows: Step 1: Cross-dataset validation; It provides reliable pseudo-supervisory signals for unsupervised model training, adapts to mixed training scenarios of KITTI labeled single-modal dataset and KAIST unlabeled infrared-visible dual-modal dataset, and forms a closed-loop verification system through KITTI ground truth quantization screening and KAIST unlabeled semi-quantitative consistency verification. Step 2: Modal complementary alignment based on cross-modal motion alignment; Taking four consecutive frames of infrared-visible light images as input, the system goes through three sub-steps: shared pyramid feature extraction, motion information propagation, and cross-modal motion alignment. The output is aligned motion features, optical flow field, and joint occlusion mask, providing robust cross-modal feature input for subsequent optical flow estimation. Step 3: Fuse the enhanced optical flow estimation network; The optical flow estimation network employs a multi-scale architecture, recursively predicting and correcting the initial optical flow; the network input includes aligned features. , Alignment optical flow and by joint masking Weighted contextual information; at each scale The network calculates the enhanced feature correlation through cost volume and outputs the optical flow increment. ; Step 4: Composite loss function; The composite loss function consists of multidimensional alignment loss and unsupervised optical flow loss. The multidimensional alignment loss is based on three dimensions: global adversarial distribution constraint, local statistical consistency constraint, and cross-modal structural reconstruction constraint. The unsupervised optical flow loss relies on the photometric consistency assumption to ensure the smoothness of the optical flow field and the consistency between frames. The composite loss function realizes end-to-end unsupervised training, and the total loss is as shown in formula (12).

[0007] in For traditional unsupervised optical flow loss, unflowLoss For multidimensional alignment loss, To align the loss weights.

[0008] Preferably, step 1 specifically comprises: Step 1-1: KITTI truth quantification screening; By utilizing the high-precision optical flow ground truth of the KITTI dataset, occluded regions are filtered out through forward and backward consistency checks, and noise predictions are filtered out by combining photometric consistency constraints. Finally, reliable pseudo-labels are selected for model training. Step 1-1-1: For the t-th frame image pair in the KITTI dataset With the Frame Image Pair Calculate the forward optical flow With backward optical flow The occlusion mask is obtained by performing a forward and backward consistency check using formula (1). Filter out occluded pixels with inconsistent optical flow:

[0009] in , For the occlusion threshold, Indicates the first The coordinates of a pixel on a frame image; Step 1-1-2: In the effective pixel mask Calculate and predict optical flow With truth value The endpoint error EPE, as shown in formula (2), is used as a quantitative evaluation standard for optical flow accuracy:

[0010] in Total number of effective pixels, It is the Euclidean norm; Step 1-1-3: Introduce photometric consistency constraints to filter noise prediction; Photometric consistency is based on the assumption of constant brightness and is used to calculate the reconstructed image. With the original image L1 loss between:

[0011] in , This indicates a warping operation based on optical flow. Steps 1-2: KAIST semi-quantitative consistency verification without truth values; Based on the assumption of local smoothness of optical flow, i.e., the direction and amplitude of adjacent pixels in the real optical flow field are highly continuous, three complementary evaluation indicators are designed: mean angular error (AAE), intersection-over-union ratio (IoU) of moving regions, and structural similarity index (SSIM). This forms a semi-quantitative verification system without ground truth, enabling effective performance evaluation of the model in complex bimodal scenarios. The calculation methods for each indicator are as follows: Step 1-2-1: IoU: Perform 5×5 Gaussian smoothing (σ=1) on the optical flow amplitude map to suppress isolated noise; generate a motion mask using Otsu adaptive threshold binarization. Reference motion mask The IoU value is generated from the 3×3 neighborhood mean of the optical flow amplitude map and is finally calculated using formula (4). The higher the value, the more accurate the localization of the motion region.

[0012] Step 1-2-2: AAE: Calculate the predicted unit direction vector within the motion region. relative to the reference unit direction vector The included angle error is shown in formula (5):

[0013] in The total number of pixels in the moving region, ε=1e 8. To prevent the value from becoming an unstable minimum, the included angle is limited to the range of [0,π]. Steps 1-2-3: SSIM; The optical flow amplitude map and the reference amplitude map are normalized to [0, 255], and the mean, variance and covariance of the local region are calculated to quantify the similarity of the optical flow field structure. The higher the value, the better the structural integrity of the optical flow field. Preferably, step 2 specifically comprises: Step 2-1: Shared pyramid feature extraction; A pyramid feature extractor with shared weights is used to extract features from the input image, as shown in formula (6):

[0014] in , Infrared and visible light modes are respectively located in the pyramid. Feature map of the layer For the shared weight parameters of the extractor, , For the first The pyramid consists of 5 layers, which can extract infrared and visible light input images to achieve multi-scale extraction from shallow texture features to deep semantic features. Step 2-2: To address the differences in modal characteristics between infrared and visible light, the multi-scale features extracted from the shared pyramid are propagated independently to achieve decoupling of modality-specific dynamic features and enhance the continuity of features in the time dimension, laying the foundation for cross-modal alignment. Step 2-2-1: Analyze the features of adjacent frames for each modality. and Perform feature correlation calculations to capture inter-frame motion cues and generate initial motion features. ; Step 2-2-2: Introduce feature dimensionality reduction operation The high-dimensional initial motion features are compressed to a fixed number of channels to generate mode-specific motion descriptors. , To ensure dimensional uniformity of infrared and visible light motion descriptors; Steps 2-3: Cross-modal motion alignment; Through four steps—cross-modal coupled feature extraction, adaptive channel attention weight allocation, learnable residual alignment, optical flow residual correction, and joint occlusion estimation—precise alignment of bimodal motion features and unified propagation of global motion information are achieved, while correcting motion deviations of single modality in challenging scenarios. Step 2-3-1: Cross-modal coupling feature extraction; Initial motion characteristics of infrared and visible light modes and Cross-modal coupling features are extracted through cascaded convolutional layers. As shown in formula (7):

[0015] Step 2-3-2: Adaptive channel attention weight allocation; Based on cross-modal coupling features Adaptive channel attention weights are generated through 1×1 convolution and the Sigmoid activation function to dynamically adjust the contribution of infrared and visible light modes in motion information propagation, thereby achieving adaptive fusion of modal information in different scenarios, as shown in formula (8):

[0016] Step 2-3-3: Learnable residual alignment; Introduce a learnable residual scaling factor with an initial value of 0.1. A residual alignment mechanism is constructed to achieve robust alignment of dual-modal motion features, while simultaneously promoting the unified propagation of global motion information. The aligned motion features and As in formula (9):

[0017] in This represents element-wise multiplication; Steps 2-3-4: Optical flow residual correction and joint occlusion estimation; Based on motion feature alignment, the initial optical flow of each mode is... and Perform residual correction and generate a joint occlusion mask. This guides subsequent optical flow estimation to focus on a reliable, unobstructed region, resulting in aligned optical flow. The calculation is as shown in formula (10):

[0018] By utilizing the global statistical properties of the aligned dual-modal optical flow, a joint occlusion mask is generated using formula (11). : (11).

[0019] Preferably, step 4 specifically comprises: Step 4-1: Multidimensional Alignment Loss ; Step 4-1-1: Global Adversarial Constraints ; Using a discriminator Perform domain alignment in the feature space, utilizing unaligned features. , The discriminator is trained to identify the modality source; subsequently, the features are forcibly aligned. , Mapping to the intermediate domain, the adversarial loss is as shown in formula (13):

[0020] Step 4-1-2: Local Statistical Matching Constraints ; Calculate the local mean of the alignment feature within a k×k sliding window. With variance For example, in formula (14):

[0021] By minimizing the second-order statistical difference between the two modes using formula (15), the consistency constraint of local motion characteristics is achieved:

[0022] Step 4-1-3: Exchange and reconfigure constraints ; Through a cross-decoding mechanism, using an infrared decoder Recovering infrared images from visible light alignment features using a visible light decoder The visible light image is recovered from the infrared alignment features. The L1 loss between the reconstructed image and the original image is calculated using formula (16), thus achieving the constraint for cross-modal structure reconstruction.

[0023] Step 4-2: Traditional Unsupervised Optical Flow Loss ; Based on the photometric consistency assumption, this method combines pixel-level L1 loss, structural similarity loss, and ternary loss, and introduces a second-order edge-aware smoothing loss. It also utilizes a joint occlusion mask. Invalid constraints in the shielded area, as shown in formula (17):

[0024] in For pixel-level L1 luminance loss, The loss is ternary, and α1, α2, α3, and α4 are the loss weights. The second gradient of the final optical flow field. It is a second-order edge-aware smoothing loss.

[0025] Preferably, the occlusion threshold Set to 1 pixel.

[0026] Preferably, the alignment loss weight The value is 0.5.

[0027] Preferably, k×k is 7×7.

[0028] An electronic device includes: a processor and a memory; the memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory to enable the electronic device to perform the above-described unsupervised infrared-visible optical flow estimation method.

[0029] A computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements the above-described unsupervised infrared-visible optical flow estimation method.

[0030] A chip includes a processor for calling and running a computer program from a memory, causing a device on which the chip is mounted to perform the above-described unsupervised infrared-visible optical flow estimation method.

[0031] A computer program product includes a computer storage medium storing a computer program, the computer program including instructions executable by at least one processor, which, when executed by the at least one processor, implement the above-described unsupervised infrared-visible optical flow estimation method.

[0032] The beneficial effects of this invention are as follows: I. The robustness of optical flow estimation is greatly improved, effectively compensating for the perception defects of the visible light mode in low light and severe weather, and giving full play to the all-weather robustness and penetration in severe weather of the infrared mode. In the KAIST night scene, the mean angular error (AAE) is reduced by 83% compared with the single infrared model and by 55% compared with the single visible light model. Moreover, it can maintain high IoU and SSIM in the occluded area, and completely solve the problems of mismatch and optical flow drift of the single mode method in complex scenes. Second, the model has excellent generalization ability. Relying on the design idea of ​​cross-dataset verification strategy and modality-specific dynamic decoupling, the model learns modality-independent general motion features. It shows a significantly leading cross-domain generalization ability in zero-shot evaluation of the M3FD dataset. It can achieve high-precision optical flow estimation without domain fine-tuning and can adapt to complex motion perception scenarios with different urban backgrounds and sensor parameters. Third, the computational efficiency takes into account engineering requirements. While achieving a significant performance improvement, it only increases the number of parameters and computational overhead. The model FLOPs are consistent with the single-modal baseline, and the single-frame inference speed reaches 58.03ms, which fully meets the real-time requirements of scenarios such as autonomous driving. Fourth, the training cost is significantly reduced. Through cross-dataset pseudo-supervision and unsupervised constraints of multidimensional alignment loss, there is no need to collect and label infrared-visible dual-modal optical flow ground truth data, which greatly reduces the labeling cost and engineering implementation difficulty of model training. V. Precise and efficient motion feature alignment: Based on the modal complementary alignment module for cross-modal motion alignment, it achieves all-round alignment of dual modes from the feature layer to the motion field, effectively decoupling the specific dynamic features of different modes, promoting the unified propagation of global motion information, avoiding motion deviation of single modes in backlight, low texture and other scenarios, and generating sharp, smooth and continuous optical flow field boundaries without artifacts and breaks. Attached Figure Description

[0033] Figure 1 This is a schematic diagram of the overall architecture of CrossFlow; Figure 2 This is a schematic diagram of a cross-modal motion alignment module; Figure 3 A schematic diagram of the three-branch constraint for alignment loss; Figure 4 This is a qualitative comparison diagram of optical flow estimation results on the KAIST dataset under harsh conditions, according to an embodiment of the present invention. Figure 5 This is a schematic diagram comparing the quantitative performance of models A, B, and C in various challenging scenarios according to embodiments of the present invention. Figure 6 This is a zero-sample qualitative result on the M3FD dataset of this invention – examples of foggy and rainy days. Detailed Implementation

[0034] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0035] To address the core technical challenges of existing unsupervised optical flow estimation methods in complex urban scenarios such as backlighting, nighttime, and foggy conditions—including insufficient robustness of single visible light modes, lack of motion consistency alignment mechanisms in infrared-visible cross-modal fusion, and high difficulty in model training and validation in ground truth dual-modal scenarios—this invention aims to provide an unsupervised infrared-visible optical flow estimation method based on cross-modal alignment. Through multi-component collaborative design, it achieves deep alignment and complementarity of motion features between infrared and visible light modes. Without requiring additional dual-modal optical flow ground truth annotation, it significantly improves the accuracy, robustness, and generalization ability of optical flow estimation in complex urban scenarios, meeting the high reliability and real-time requirements of optical flow perception in engineering scenarios such as autonomous driving and robot navigation.

[0036] To achieve the aforementioned objectives, this invention constructs an end-to-end CrossFlow unsupervised optical flow estimation framework. Using four consecutive frames of infrared-visible dual-modal images as core input, it integrates three core components to form a complete technical solution. Unsupervised model training is achieved through precise alignment of cross-modal motion features and multi-dimensional loss constraints. The core design logic is as follows: First, relying on a cross-dataset verification strategy adapted to mixed scenarios of KITTI labeled data and KAIST unlabeled data, reliable pseudo-supervisory signals are generated. This strategy, through a closed-loop design of ground truth quantization screening and semi-quantitative consistency verification, completely solves the technical bottleneck of model training and verification in ground truth-less dual-modal scenarios, providing a stable and reliable supervisory basis for unsupervised model learning. Second, a modal complementary alignment module based on cross-modal motion alignment is designed. This module, through the design idea of ​​decoupling modal-specific dynamic features and promoting the unified propagation of global motion information, firstly acquires multi-scale deep features of infrared and visible modes synchronously through a shared pyramid extractor, and then generates modal-specific optical flow and motion features through an independent motion information propagation module. A cross-modal motion alignment mechanism is used to achieve accurate alignment of bimodal motion features, outputting aligned motion features, optical flow, and joint occlusion map. This fully leverages the modal complementarity value of infrared and visible light, correcting motion deviations in challenging scenarios. Finally, a multi-dimensional alignment loss is introduced. This loss serves as an important supplement to traditional unsupervised loss. By supervising global adversarial distribution, local statistical consistency, and cross-modal structural reconstruction, it comprehensively enhances cross-modal feature consistency from global to local and from feature distribution to structural representation, significantly optimizing optical flow estimation accuracy. Combined with the optical flow estimation network, it forms a fusion-enhanced optical flow estimation system, ultimately achieving high-precision optical flow field output.

[0037] This invention, through the aforementioned technical solution, overcomes the limitations of traditional single-modal unsupervised optical flow estimation, achieving a comprehensive improvement in optical flow estimation performance in complex urban scenarios. It possesses significant technical advantages and engineering application value: First, the robustness of optical flow estimation is greatly improved, effectively compensating for the perception deficiencies of the visible light mode under low light and inclement weather conditions, while fully leveraging the all-weather robustness and penetration capabilities of the infrared mode in adverse weather conditions. In KAIST nighttime scenarios, the mean angular error (AAE) is reduced by 83% compared to the single infrared model and by 55% compared to the single visible light model. Furthermore, it maintains high IoU and SSIM in occluded areas, completely solving the problems of mismatch and optical flow drift in complex scenarios caused by single-modal methods. Second, the model exhibits excellent generalization ability. Relying on a cross-dataset validation strategy and a design approach of modality-specific dynamic decoupling, the model learns modality-independent general motion features, achieving significant performance improvements in M3FD. The model demonstrates significantly superior cross-domain generalization capabilities in zero-shot evaluations of the dataset, achieving high-precision optical flow estimation without domain fine-tuning, and adapting to complex motion perception scenarios with varying urban backgrounds and sensor parameters. Thirdly, its computational efficiency balances engineering requirements, achieving substantial performance improvements with minimal increase in parameters and computational overhead. The model's FLOPs remain consistent with the single-modal baseline, and its single-frame inference speed reaches 58.03ms, fully meeting the real-time requirements of scenarios such as autonomous driving. Fourthly, training costs are significantly reduced through unsupervised constraints using cross-dataset pseudo-supervision and multi-dimensional alignment loss, eliminating the need for infrared data acquisition and labeling. The ground truth data of visible light dual-modal optical flow significantly reduces the annotation cost and engineering implementation difficulty of model training; fifth, the motion feature alignment is accurate and efficient. Based on the modal complementary alignment module of cross-modal motion alignment, it realizes the all-round alignment of dual modes from the feature layer to the motion field, effectively decouples the specific dynamic features of different modes, promotes the unified propagation of global motion information, avoids the motion deviation of single mode in backlight, low texture and other scenes, and the generated optical flow field boundary is sharp, smooth and continuous, without artifacts and breakage.

[0038] Overall, this invention successfully addresses the robustness, generalization, and training / validation challenges of existing unsupervised optical flow estimation methods in complex urban scenarios. Through a cross-dataset validation strategy, a modal complementary alignment module based on cross-modal motion alignment, and a multi-dimensional alignment loss, it achieves end-to-end unsupervised training for infrared-visible cross-modal optical flow estimation, providing an efficient and reliable technical solution for multimodal motion perception in fields such as autonomous driving and robot navigation. Furthermore, this method can be fused with data from other sensors such as LiDAR and millimeter-wave radar to further enhance overall perception robustness in complex environments, demonstrating broad engineering application prospects and expansion value.

[0039] This invention presents an unsupervised infrared-visible optical flow estimation method based on cross-modal alignment. The core of this method is the construction of an end-to-end CrossFlow framework. This framework takes four consecutive frames of infrared (IR) and visible (VIS) image sequences as input. It generates reliable pseudo-supervision through a cross-dataset verification strategy adapted to mixed KITTI-annotated and KAIST-unannotated scenarios. It achieves modal-specific dynamic decoupling and unified propagation of global motion information through a modal complementary alignment module based on cross-modal motion alignment. A composite loss constraint is formed by combining multi-dimensional alignment loss and traditional unsupervised optical flow loss. A fused and enhanced optical flow estimation network outputs a high-precision optical flow field, completing end-to-end unsupervised training. The overall technical solution revolves around three core components and includes six core implementation steps. These steps work together to achieve accurate estimation of cross-modal optical flow in complex scenarios. Specific technical details are as follows: I. Cross-dataset validation strategy; This strategy provides reliable pseudo-supervision signals for unsupervised model training and is specifically adapted to mixed training scenarios involving the KITTI labeled unimodal dataset and the KAIST unlabeled infrared-visible bimodal dataset. A closed-loop validation system is formed through KITTI ground truth quantization screening and KAIST unlabeled semi-quantitative consistency verification, ensuring both high-quality pseudo-labels and effective model performance evaluation in the unlabeled bimodal scenario. The specific implementation steps and mathematical expressions are as follows: 1. KITTI truth quantification screening; Using high-precision ground truth optical flow values ​​from the KITTI dataset, occluded regions are filtered out through forward and backward consistency checks, and noise predictions are filtered out by combining photometric consistency constraints. Finally, reliable pseudo-labels are selected for model training. The specific process is as follows: 1.1 For the KITTI dataset, the first... Frame Image Pair With the Frame Image Pair Calculate the forward optical flow With backward optical flow The occlusion mask is obtained by performing a forward and backward consistency check using formula (1). Filter out occluded pixels with inconsistent optical flow:

[0040] in , The occlusion threshold is typically set to 1 pixel.

[0041] 1.2 In the effective pixel mask Calculate and predict optical flow With truth value The endpoint error EPE, as shown in formula (2), is used as a quantitative evaluation standard for optical flow accuracy:

[0042] in Total number of effective pixels, It is the Euclidean norm.

[0043] 1.3 Introducing photometric consistency constraints to filter noise prediction. Photometric consistency is based on the assumption of constant brightness and is used to calculate the reconstructed image. With the original image L1 loss between:

[0044] in , This represents a warping operation based on optical flow. This invention retains the requirement to simultaneously satisfy... The predictions with low luminance error serve as reliable pseudo-labels, providing high-quality supervision for subsequent cross-modal model training.

[0045] 2. KAIST lacks true value semi-quantitative consistency verification; To address the issue of lacking a true value for the infrared-visible dual-modal optical flow in KAIST, this paper proposes three complementary evaluation metrics based on the assumption of local smoothness in optical flow—that is, the direction and amplitude of adjacent pixels in the true optical flow field have high continuity. These metrics include Average Absolute Error (AAE), Intersection over Union (IoU), and Structural Similarity (SSIM). This forms a semi-quantitative verification system without a true value, enabling effective performance evaluation of the model in complex dual-modal scenarios. The calculation methods for each metric are as follows: 2.1 IoU: Measures the consistency of the model's spatial localization of moving targets. The optical flow amplitude map is smoothed with 5×5 Gaussian (σ=1) to suppress isolated noise, and a motion mask is generated by Otsu adaptive threshold binarization. Reference motion mask The IoU value is generated from the 3×3 neighborhood mean of the optical flow amplitude map and is finally calculated using formula (4). The higher the value, the more accurate the localization of the motion region.

[0046] 2.2 AAE: Focuses on the continuity of optical flow direction, specifically addressing the problem of inaccurate optical flow direction estimation in low-light and low-texture scenes. It calculates and predicts the unit direction vector within the moving region. relative to the reference unit direction vector The included angle error is shown in formula (5):

[0047] in The total number of pixels in the moving region, ε=1e 8. To prevent the value from becoming an unstable minimum, the included angle is limited to the range of [0,π]. 2.3 SSIM: To further measure the structural and textural consistency of the optical flow field, the optical flow amplitude map and the reference amplitude map are normalized to [0, 255]. The mean, variance and covariance of the local region are calculated to quantify the structural similarity of the optical flow field. The higher the value, the better the structural integrity of the optical flow field.

[0048] The three metrics work together: IoU provides an objective basis for spatial localization of the motion region, AAE highlights the continuity of optical flow direction, and SSIM supplements the consistency of structure and texture, enabling a comprehensive and reliable evaluation of model performance in dual-modal scenarios without ground truth.

[0049] II. Modal complementary alignment module based on cross-modal motion alignment; This module is the core technology supporting the accurate cross-modal optical flow estimation of this invention. Its core design goal is to decouple the modal-specific dynamic features of infrared and visible light, promote the unified propagation of global motion information, and achieve omnidirectional alignment of dual-modal motion features from the feature layer to the motion field. The module takes four consecutive frames of infrared-visible light images as input, and through three sub-stages—shared pyramid feature extraction, motion information propagation, and cross-modal motion alignment—outputs aligned motion features, optical flow field, and joint occlusion mask, providing robust cross-modal feature input for subsequent optical flow estimation. The specific implementation of each sub-stage is as follows: 1. Shared pyramid feature extraction; To simultaneously capture multi-scale deep features of infrared and visible light while maintaining computational efficiency, and to ensure the dimensionality consistency of the dual-modal features, a pyramid feature extractor with shared weights is used to extract features from the input image, as shown in formula (6):

[0050] in , Infrared and visible light modes are respectively located in the pyramid. Feature map of the layer For the shared weight parameters of the extractor, , For the first The input images are infrared and visible light, and the pyramid has 5 layers to achieve multi-scale extraction from shallow texture features to deep semantic features.

[0051] 2. Addressing the differences in modal characteristics between infrared and visible light, independent motion information propagation is performed on the multi-scale features extracted from the shared pyramid. This decouples modality-specific dynamic features and enhances the continuity of features over time, laying the foundation for cross-modal alignment. 2.1 Features of adjacent frames for each modality and Perform feature correlation calculations to capture inter-frame motion cues and generate initial motion features. ; 2.2 Introducing Feature Dimensionality Reduction Operation The high-dimensional initial motion features are compressed to a fixed number of channels to generate mode-specific motion descriptors. , This ensures the dimensional uniformity of infrared and visible light motion descriptors, providing a consistent feature basis for subsequent cross-modal motion alignment.

[0052] 3. Cross-modal motion alignment; This step is the core of the modal complementarity alignment module. Through four steps—cross-modal coupled feature extraction, adaptive channel attention weight allocation, learnable residual alignment, optical flow residual correction, and joint occlusion estimation—it achieves accurate alignment of bimodal motion features and unified propagation of global motion information. Simultaneously, it corrects motion deviations in challenging scenarios for single-modal systems. The specific implementation of each step is as follows: 3.1 Cross-modal coupling feature extraction: Extracting the initial motion features of infrared and visible light modes. and Cross-modal coupling features are extracted through cascaded convolutional layers. As shown in formula (7):

[0053] 3.2 Adaptive Channel Attention Weight Allocation: Based on Cross-Modal Coupling Features Adaptive channel attention weights are generated through 1×1 convolution and the Sigmoid activation function to dynamically adjust the contribution of infrared and visible light modes in motion information propagation, thereby achieving adaptive fusion of modal information in different scenarios, as shown in formula (8):

[0054] 3.3 Learnable Residual Alignment: To prevent optical flow oscillations caused by the instability of the alignment module in the early stages of training, a learnable residual scaling factor with an initial value of 0.1 is introduced. A residual alignment mechanism is constructed to achieve robust alignment of dual-modal motion features, while simultaneously promoting the unified propagation of global motion information. The aligned motion features and As in formula (9):

[0055] in This indicates element-wise multiplication.

[0056] 3.4 Optical Flow Residual Correction and Joint Occlusion Estimation: Based on motion feature alignment, the initial optical flow of each mode is adjusted. and Perform residual correction and generate a joint occlusion mask. This guides subsequent optical flow estimation to focus on a reliable, unobstructed region, resulting in aligned optical flow. The calculation is as shown in formula (10):

[0057] By utilizing the global statistical properties of the aligned dual-modal optical flow, a joint occlusion mask is generated using formula (11). This shields unreliable regions with excessively large differences in cross-modal optical flow, providing effective constraints for subsequent optical flow estimation.

[0058] III. Fusion-enhanced optical flow estimation network; The optical flow estimation network employs a multi-scale architecture, recursively predicting and correcting the initial optical flow. The network input contains aligned features. , Alignment optical flow and by joint masking Weighted contextual information. At each scale The network calculates the enhanced feature relevance using the cost volume and outputs the optical flow increment. Because the input features have undergone motion compensation and alignment by the Cross-Modal Motion Alignment module, the network can more effectively focus on the actual geometric displacement, thus maintaining the continuity of the motion trajectory on roads under strong light or at night.

[0059] IV. Design of Composite Loss Function; The composite loss function of this invention consists of multidimensional alignment loss and traditional unsupervised optical flow loss. Multidimensional alignment loss is the core innovation, serving as an important supplement to traditional unsupervised optical flow loss. It comprehensively strengthens cross-modal feature consistency through three dimensions: global adversarial distribution constraints, local statistical consistency constraints, and cross-modal structure reconstruction constraints, significantly optimizing optical flow estimation accuracy. Traditional unsupervised optical flow loss relies on the photometric consistency assumption to ensure the smoothness and inter-frame consistency of the optical flow field. The composite loss function enables end-to-end unsupervised training, with the total loss as shown in formula (12):

[0060] in For traditional unsupervised optical flow loss, unflowLoss For multidimensional alignment loss, To align the loss weights, experiments have shown that the optimal value is 0.5, achieving the best balance between modal complementarity and optical flow estimation accuracy.

[0061] 1. Multidimensional alignment loss ; 1.1 Global Adversarial Constraints From the perspective of global feature distribution, the forced alignment of cross-modal motion features is mapped to a mode-independent intermediate domain, eliminating the modal distribution differences between infrared and visible light, and achieving a unified expression of global motion features. A discriminator is then used. Perform domain alignment in the feature space. This utilizes unaligned features. , The discriminator is trained to identify the modality source; subsequently, the features are forcibly aligned. , Mapping to the intermediate domain, the adversarial loss is as shown in formula (13):

[0062] 1.2 Local Statistical Matching Constraints From the perspective of local feature statistics, to ensure the synchronicity of aligned cross-modal motion features in spatial details, the local mean of the aligned features is calculated within a k×k (7×7 in experiments) sliding window. With variance For example, in formula (14):

[0063] By minimizing the second-order statistical difference between the two modes using formula (15), the consistency constraint of local motion characteristics is achieved:

[0064] 1.3 Exchange and Reconfiguration Constraints From the perspective of feature structure representation, the completeness of aligned cross-modal motion features is verified, ensuring that the aligned features are both modally independent and retain the core motion semantics of the original mode. This is achieved through a cross-decoding mechanism using an infrared decoder. Recovering infrared images from visible light alignment features using a visible light decoder The visible light image is recovered from the infrared alignment features. The L1 loss between the reconstructed image and the original image is calculated using formula (16), thus achieving the constraint for cross-modal structure reconstruction.

[0065] 2. Traditional unsupervised optical flow loss ; Following the classic unsupervised loss design in the field of optical flow estimation, and relying on the photometric consistency assumption, this paper combines pixel-level L1 loss, structural similarity loss, and ternary loss, and introduces a second-order edge-aware smoothing loss. Simultaneously, it utilizes a joint occlusion mask. Invalid constraints in the shielded area, as shown in formula (17):

[0066] in For pixel-level L1 luminance loss, The loss is ternary, and α1, α2, α3, and α4 are loss weights, which are classic empirical values ​​in the field of optical flow. The second gradient of the final optical flow field. The second-order edge-aware smoothing loss ensures both the smoothness of the optical flow field and the sharpness of the moving target's boundary, thus avoiding blurring of the optical flow field.

[0067] The above six technical aspects revolve around three core components: cross-dataset validation strategy provides reliable pseudo-supervision for the model; modal complementarity alignment module realizes accurate alignment of cross-modal motion features and modal-specific dynamic decoupling; multidimensional alignment loss and traditional unsupervised optical flow loss form a composite loss that provides comprehensive constraints for model training; and finally, a fusion-enhanced optical flow estimation network outputs a high-precision and highly robust cross-modal optical flow field.

[0068] Example: This invention presents an unsupervised infrared-visible optical flow estimation method based on cross-modal alignment. It constructs an end-to-end CrossFlow framework, integrating three core components: a cross-dataset validation strategy, a modal complementarity alignment module based on cross-modal motion alignment, and a multidimensional alignment loss. This achieves high-precision infrared-visible dual-modal optical flow estimation in complex urban scenarios. The overall model framework is as follows: Figure 1As shown in the figure. This method, based on the traditional unsupervised optical flow framework, innovatively introduces a cross-modal motion alignment mechanism and a multi-dimensional alignment loss. It uses the KITTI labeled dataset for ground truth quantization screening and the KAIST unlabeled bimodal dataset for semi-quantitative consistency verification, generating reliable pseudo-supervisory signals. Simultaneously, it achieves deep fusion of bimodal features through shared pyramid feature extraction, modality-specific dynamic decoupling, and cross-modal motion alignment. Combined with a three-branch alignment loss of global adversarial, local statistical, and cross-modal reconstruction, it strengthens feature consistency. This results in an 83% reduction in mean angular error (AAE) compared to a single infrared model and a 55% reduction compared to a single visible light model in KAIST nighttime scenes. It also demonstrates excellent cross-domain generalization ability in zero-shot tests on the M3FD dataset. This method can be widely applied to vision tasks such as autonomous driving and robot navigation, and can also be fused with sensors such as LiDAR and millimeter-wave radar to further improve perception robustness in complex environments. The specific implementation process of this invention is as follows.

[0069] This embodiment provides an unsupervised infrared-visible optical flow estimation method based on the CrossFlow framework and combining the KITTI and KAIST datasets. A multi-dimensional alignment loss module is introduced after the modal complementarity alignment module to further enhance the model's ability to express the consistency of cross-modal motion features. The specific implementation steps are described below with reference to the accompanying drawings: Step 1: Analyze the training parameters; The Python standard library argparse is used to parse command-line arguments and obtain hyperparameter configuration information, providing a complete configuration entry point for subsequent training. Step 2: Load and preprocess the KAIST and KITTI datasets; KITTI 2015 is a labeled unimodal dataset used for ground truth quantization and model accuracy evaluation. KAIST is an unlabeled infrared-visible dual-modal dataset. Multispectral pedestrian detection video sequences from KAIST were selected for unsupervised model training and robustness verification in complex scenes. Four consecutive frames were extracted as a training sample, each containing an infrared (IR) image and a visible light (VIS) image. The loaded datasets underwent uniform preprocessing: infrared and visible light images were scaled to a target size of 384×128; infrared images were normalized to the [0,1] interval; and visible light images were normalized according to ImageNet mean-variance. The image arrays were converted to tensor format. Simultaneously, the KITTI dataset was filtered using the ground truth quantization method of the technical solution, retaining endpoint errors (EPE) < 5px and light loss. For prediction results with a value <0.1, reliable pseudo-labels are generated for model training.

[0070] Step 3: Create a data loader and online enhancements; Use `torch.utils.data.DataLoader` to create a training data loader with `BatchSize=4`. Real-time augmentation strategies include: Geometric consistency transformations: RandomVerticalFlip (vflip=True), RandomHorizontalFlip (hflip=True), and RandomSwap (swap=True). These transformations are applied synchronously to image pairs and instance masks to ensure geometric consistency. Appearance enhancement transformations: Color Jitter (brightness=0.5, contrast=0.0, saturation=0.0, hue=0.0, applied only to image pairs) and Random Gaussian Blur (p=0.5, radius range 0~3, applied only to image pairs). Input transformation: the array is converted into a tensor and uniformly scaled to the target size.

[0071] Step 4: Create the CrossFlow model framework; The model modules are constructed sequentially according to the technical solution of this invention and end-to-end splicing is completed. Irrelevant constraints of traditional single-modal optical flow are disabled, a custom CrossModalLoss module is registered, and preparation is made to access the multidimensional alignment loss based on cross-modal alignment, as shown in Figure 1 (network structure) and Figure 2 (details of the cross-modal motion alignment module). The core modules of the model include: a shared pyramid feature extractor, a motion information propagation module, a modal complementary alignment module based on cross-modal motion alignment, a fusion-enhanced optical flow estimation network, and a multidimensional alignment loss module.

[0072] Step 5: Configure the optimizer and learning rate scheduler; Using the Adam optimizer, with an initial learning rate of 2×10⁻⁶. -4 The ExponentialLR scheduler is used, decaying every 1000 steps. Gradient clipping (clip=1.0) is enabled to prevent explosion. All trainable parameters participate in optimization, with no frozen layers.

[0073] Step 6: Train the model; A complete training loop is executed for 100,000 iterations: At each iteration step, a batch of training samples is retrieved from the DataLoader, including four consecutive frames of infrared images I0^ir~I3^ir and four consecutive frames of visible light images. The model also includes reliable pseudo-labels for the KITTI dataset used for evaluation and selection. During the model's forward propagation, the following steps are performed sequentially: shared pyramid multi-scale feature extraction, motion information propagation and modality-specific dynamic decoupling, cross-modal motion alignment and joint occlusion mask generation, and fusion-enhanced optical flow estimation. The outputs are multi-scale optical flow prediction results and the final global optical flow field. Then the total loss is calculated, which consists of the traditional unsupervised optical flow loss. With multidimensional alignment loss Weighted composition, i.e. .

[0074] Specifically, in the process of calculating the multidimensional alignment loss, the aligned motion features output by the modal complementarity alignment module are first trained with global adversarial distribution constraints. A discriminator identifies the modal origin of the unaligned features, forcing the aligned features to map to a modality-independent intermediate domain, as shown in Figure 3 (three-branch constraint diagram of alignment loss). Then, the local mean and variance within a 7×7 sliding window are calculated to minimize the second-order statistical difference between the infrared and visible light dual modes, achieving local statistical consistency constraints. Next, a cross-decoding mechanism is used, where the infrared decoder recovers the infrared image from the visible light alignment features, and the visible light decoder recovers the visible light image from the infrared alignment features. The L1 loss between the reconstructed image and the original image is calculated, completing the cross-modal structure reconstruction constraint. The weighted sum of the three-branch losses yields the multidimensional alignment loss, which, together with the traditional unsupervised optical flow loss combined with joint occlusion masks, constitutes the total loss.

[0075] After calculating the total loss, the Adam optimizer is used for backpropagation and to update the model parameters. The TensorBoard training log is recorded every 1000 steps, including indicators such as the loss value of each branch, endpoint error EPE, and mean angular error AAE. Every 10000 steps, the model checkpoint is saved, and the model weights with the best performance on the validation set are retained. The entire training loop is shown in Figure 1, which is the overall framework diagram of this invention, including data loading, model forward propagation, loss calculation, and optimization update.

[0076] Step 7: Verification and Testing; During training, model performance was evaluated on the validation set every 10,000 steps. Table 1 shows the evaluation results of the CrossFlow framework trained on the KAIST dataset and on the KITTI 2015 visible light single-modality dataset. The performance of the single infrared model A, the single visible light baseline model C, and the dual-modality fusion model B in this paper were compared. The core metrics include EPE, F1-all, NOC, and OCC. The smaller the value, the better the performance. The single visible light baseline C performs best in the texture-rich KITTI daytime scene, while the single infrared model A performs the worst due to the lack of texture details. Our proposed dual-modal fusion model B achieves a comprehensive improvement over A, with relative reductions in EPE, F1-all, NOC, and OCC of 12.93%, 10.53%, 15.15%, and 12.25%, respectively. The largest improvement is observed in the non-occluded region, effectively validating that the modal complementarity fusion and cross-modal motion alignment mechanism can fully combine the strong robustness of the infrared modality with the texture detail advantage of the visible light modality, thereby improving the matching accuracy of motion features and the reliability of optical flow estimation. Model B's performance is still inferior to C, which stems from the data distribution difference resulting from training the model in the KAIST dual-modal nighttime scene but testing it in the KITTI single-modal daytime scene. This aligns with our design goal of focusing on robustness in complex nighttime scenes, rather than pursuing the best ranking on the KITTI daytime benchmark.

[0077] Table 1. Quantitative Comparison of Optical Flow Estimation Performance on the KITTI Dataset

[0078] According to the cross-dataset validation strategy mentioned in the technical solution of this invention, namely KAIST truth-free semi-quantitative consistency validation, the mean angular error (AAE), the region of motion (IoU), and the structural similarity (SSIM) of the above-mentioned models A, B, and C are calculated on the test set partitioned by KAIST, respectively. Figure 4 And provide a qualitative demonstration of some examples.

[0079] Table 2. Semi-quantitative comparison of optical flow estimation in complex scenes using KAIST

[0080] The semi-quantitative analysis in Table 2 fully confirms that the CrossFlow framework exhibits significant robustness advantages in extremely challenging scenarios such as low illumination, severe weather, and extremely dark nights. Compared with single infrared model A and single visible light model C, the proposed fusion model B achieves the best performance across all metrics. Particularly in backlit, cloudy, and nighttime scenes, B reduces AAE by 70.78%, 79.46%, and 78.39% compared to A, reaching extremely low levels of 0.225°, 0.281°, and 0.222°, respectively, while maintaining a leading IoU and SSIM. In contrast, model C, in extremely dark environments, not only shows a significant increase in AAE, but also a significant decrease in SSIM and IoU, reaching a performance bottleneck. This directly highlights the core advantage of CrossFlow in overcoming texture loss and illumination limitations through cross-modal complementary alignment.

[0081] The trend comparison in Figure 5 further reveals that Model B can maintain optimal and stable performance under various operating conditions, effectively avoiding the drastic performance fluctuations common in single-mode models. Combined with... Figure 5 The qualitative visualization results clearly show that the optical flow field generated by CrossFlow has sharp, smooth, and continuous boundaries with almost no artifacts or breaks, which significantly improves the model's perception accuracy and generalization ability in real driving environments.

[0082] Step 8: Ablation experiment of fusion strategy; To verify the necessity of the proposed modal complementarity fusion strategy, Table 3 shows multiple ablation control experiments. While maintaining the network structure and training strategy, only the modal fusion type was changed. The results show that the proposed infrared + visible light modal complementarity fusion scheme (B) achieves optimal performance on both the KITTI and KAIST datasets: KITTI-EPE as low as 17.785 px and KAIST-AAE as low as 0.225°. In contrast, schemes replacing it with infrared + random noise (B1), infrared + doubled infrared features (B2), or infrared + textureless visible light low-pass (B3) all exhibited significant performance degradation. The AAE of B1 and B3 increased to 0.890° and 1.212°, respectively, demonstrating that only the complementarity of real visible light texture and infrared modality can effectively improve the robustness and accuracy of optical flow estimation. This result fully illustrates that cross-modal complementarity fusion is the core mechanism for the CrossFlow framework to maintain its advantage in complex scenarios, rather than simply feature overlay or data augmentation.

[0083] Table 3 Comparison of ablation experimental performance of fusion strategies

[0084] Step 9, Alignment Loss Step-by-step verification experiment; The ablation experiment results in Table 4 show that the KAIST-AAE index continuously decreases after progressively integrating global adversarial, local statistical, and structural reconstruction losses, verifying the complementary effect of each alignment branch. This progressive improvement confirms that applying multidimensional feature consistency constraints can effectively bridge the gap between modal domains, thereby significantly improving the motion estimation accuracy in complex target scenarios.

[0085] Table 4. Comparison of Alignment Loss Ablation

[0086] Step 10 Hyperparameter analysis experiments were conducted to verify the cross-modal alignment loss weights. Impact on model performance; Table 5 compares different The KITTI-EPE index under various values. The results show that when... The model achieves optimal performance with a weighting of 0.5, resulting in an EPE as low as 17.78 px. If the weights are too small, the cross-modal constraints are insufficient, and the EPE rises to 19.15 px. If the weights are too large, the feature distribution will be over-constrained, leading to a decrease in optical flow estimation accuracy, and the EPE will rise to 22.02 px. This result indicates that a moderate cross-modal alignment constraint can effectively balance modal complementarity and feature representation ability.

[0087] Table 5 Comparison of Hyperparameter Analysis

[0088] Step 11: Robustness and Generalization Analysis Experiments. To further verify the reliability of the CrossFlow framework under extreme conditions and its cross-scene transferability, this invention conducted robustness tests in extremely complex scenarios and zero-shot validation across datasets. In the KAIST project comparison shown in Table 6, Model B (Ours) exhibits an overwhelming robustness advantage: under severe weather conditions, its AEE is only 1.327, achieving an accuracy improvement of nearly an order of magnitude compared to the 11.645 of the single visible light model C; in low-light nighttime scenarios, the AEE of Model B drops to 0.443, and the IoU in the occluded area is as high as 99.28%. This fully demonstrates that through the cross-modal motion alignment mechanism, the model can effectively utilize the penetration and stability of the infrared mode, compensating for the deficiency of photometric consistency failure of visible light in extreme environments.

[0089] Table 6 Comparison of KAIST Robustness in Extremely Complex Scenarios

[0090] Regarding generalization validation, as shown in Table 7, this invention directly applies model B, trained only on the KAIST dataset, to the M3FD dataset for zero-shot testing. Figure 6 The qualitative demonstration intuitively reflects the model's zero-shot performance on the M3FD dataset. Experimental results show that, without any domain fine-tuning, model B achieves an AEE of 0.206, significantly outperforming single-modal models A (0.838) and C (0.809), while maintaining the best performance in both IoU and SSIM metrics. This outstanding cross-dataset performance demonstrates that the motion features extracted by CrossFlow possess good modality invariance and domain generalization, not only adapting to differences in sensor parameters but also effectively addressing the complex motion perception challenges in different urban contexts.

[0091] Table 7. Zero-shot cross-dataset generalization evaluation on the M3FD multimodal dataset.

Claims

1. An unsupervised infrared-visible optical flow estimation method based on cross-modal alignment, characterized in that, Includes the following steps: Step 1: Cross-dataset validation; It provides reliable pseudo-supervisory signals for unsupervised model training, adapts to mixed training scenarios of KITTI labeled single-modal dataset and KAIST unlabeled infrared-visible dual-modal dataset, and forms a closed-loop verification system through KITTI ground truth quantization screening and KAIST unlabeled semi-quantitative consistency verification. Step 2: Modal complementary alignment based on cross-modal motion alignment; Taking four consecutive frames of infrared-visible light images as input, the system goes through three sub-steps: shared pyramid feature extraction, motion information propagation, and cross-modal motion alignment. The output is aligned motion features, optical flow field, and joint occlusion mask, providing robust cross-modal feature input for subsequent optical flow estimation. Step 3: Fuse the enhanced optical flow estimation network; The optical flow estimation network adopts a multi-scale architecture and corrects the initial optical flow through recursive prediction; The network input contains aligned features , Alignment optical flow and by joint masking Weighted contextual information; at each scale The network calculates the enhanced feature correlation through cost volume and outputs the optical flow increment. ; Step 4: Composite loss function; The composite loss function consists of multidimensional alignment loss and unsupervised optical flow loss. The multidimensional alignment loss is based on three dimensions: global adversarial distribution constraint, local statistical consistency constraint, and cross-modal structural reconstruction constraint. The unsupervised optical flow loss relies on the photometric consistency assumption to ensure the smoothness of the optical flow field and the consistency between frames. The composite loss function realizes end-to-end unsupervised training, and the total loss is as shown in formula (12). in For traditional unsupervised optical flow loss, unflowLoss For multidimensional alignment loss, To align the loss weights.

2. The unsupervised infrared-visible optical flow estimation method based on cross-modal alignment according to claim 1, characterized in that, Step 1 specifically involves: Step 1-1: KITTI truth quantification screening; By utilizing the high-precision optical flow ground truth of the KITTI dataset, occluded regions are filtered out through forward and backward consistency checks, and noise predictions are filtered out by combining photometric consistency constraints. Finally, reliable pseudo-labels are selected for model training. Step 1-1-1: For the t-th frame image pair in the KITTI dataset With the Frame Image Pair Calculate the forward optical flow With backward optical flow The occlusion mask is obtained by performing a forward and backward consistency check using formula (1). Filter out occluded pixels with inconsistent optical flow: in , For the occlusion threshold, Indicates the first The coordinates of a pixel on a frame image; Step 1-1-2: In the effective pixel mask Calculate and predict optical flow With truth value The endpoint error EPE, as shown in formula (2), is used as a quantitative evaluation standard for optical flow accuracy: in Total number of effective pixels, It is the Euclidean norm; Step 1-1-3: Introduce photometric consistency constraints to filter noise prediction; Photometric consistency is based on the assumption of constant brightness and is used to calculate the reconstructed image. With the original image L1 loss between: in , This indicates a warping operation based on optical flow. Steps 1-2: KAIST semi-quantitative consistency verification without truth values; Based on the assumption of local smoothness of optical flow, i.e., the direction and amplitude of adjacent pixels in the real optical flow field are highly continuous, three complementary evaluation indicators are designed: mean angular error (AAE), intersection-over-union ratio (IoU) of moving regions, and structural similarity index (SSIM). This forms a semi-quantitative verification system without ground truth, enabling effective performance evaluation of the model in complex bimodal scenarios. The calculation methods for each indicator are as follows: Step 1-2-1: IoU: Perform 5×5 Gaussian smoothing (σ=1) on the optical flow amplitude map to suppress isolated noise; generate a motion mask using Otsu adaptive threshold binarization. Reference motion mask The IoU value is generated from the 3×3 neighborhood mean of the optical flow amplitude map and is finally calculated using formula (4). The higher the value, the more accurate the localization of the motion region. Step 1-2-2: AAE: Calculate the predicted unit direction vector within the motion region. relative to the reference unit direction vector The included angle error is shown in formula (5): in The total number of pixels in the moving region, ε=1e 8. To prevent the value from becoming an unstable minimum, the included angle is limited to the range of [0,π]. Steps 1-2-3: SSIM; The optical flow amplitude map and the reference amplitude map are normalized to [0, 255]. The mean, variance and covariance of the local region are calculated to quantify the structural similarity of the optical flow field. The higher the value, the better the structural integrity of the optical flow field.

3. The unsupervised infrared-visible optical flow estimation method based on cross-modal alignment according to claim 2, characterized in that, Step 2 specifically involves: Step 2-1: Shared pyramid feature extraction; A pyramid feature extractor with shared weights is used to extract features from the input image, as shown in formula (6): in , Infrared and visible light modes are respectively located in the pyramid. Feature map of the layer For the shared weight parameters of the extractor, , For the first The pyramid consists of 5 layers, which can extract infrared and visible light input images to achieve multi-scale extraction from shallow texture features to deep semantic features. Step 2-2: To address the differences in modal characteristics between infrared and visible light, the multi-scale features extracted from the shared pyramid are propagated independently to achieve decoupling of modality-specific dynamic features and enhance the continuity of features in the time dimension, laying the foundation for cross-modal alignment. Step 2-2-1: Analyze the features of adjacent frames for each modality. and Perform feature correlation calculations to capture inter-frame motion cues and generate initial motion features. ; Step 2-2-2: Introduce feature dimensionality reduction operation The high-dimensional initial motion features are compressed to a fixed number of channels to generate mode-specific motion descriptors. , To ensure dimensional uniformity of infrared and visible light motion descriptors; Steps 2-3: Cross-modal motion alignment; Through four steps—cross-modal coupled feature extraction, adaptive channel attention weight allocation, learnable residual alignment, optical flow residual correction, and joint occlusion estimation—precise alignment of bimodal motion features and unified propagation of global motion information are achieved, while correcting motion deviations of single modality in challenging scenarios. Step 2-3-1: Cross-modal coupling feature extraction; Initial motion characteristics of infrared and visible light modes and Cross-modal coupling features are extracted through cascaded convolutional layers. As shown in formula (7): Step 2-3-2: Adaptive channel attention weight allocation; Based on cross-modal coupling features Adaptive channel attention weights are generated through 1×1 convolution and the Sigmoid activation function to dynamically adjust the contribution of infrared and visible light modes in motion information propagation, thereby achieving adaptive fusion of modal information in different scenarios, as shown in formula (8): Step 2-3-3: Learnable residual alignment; Introduce a learnable residual scaling factor with an initial value of 0.

1. A residual alignment mechanism is constructed to achieve robust alignment of dual-modal motion features, while simultaneously promoting the unified propagation of global motion information. The aligned motion features and As in formula (9): in This represents element-wise multiplication; Steps 2-3-4: Optical flow residual correction and joint occlusion estimation; Based on motion feature alignment, the initial optical flow of each mode is... and Perform residual correction and generate a joint occlusion mask. This guides subsequent optical flow estimation to focus on a reliable, unobstructed region, resulting in aligned optical flow. The calculation is as shown in formula (10): By utilizing the global statistical properties of the aligned dual-modal optical flow, a joint occlusion mask is generated using formula (11). : (11)。 4. The unsupervised infrared-visible optical flow estimation method based on cross-modal alignment according to claim 3, characterized in that, Step 4 specifically involves: Step 4-1: Multidimensional Alignment Loss ; Step 4-1-1: Global Adversarial Constraints ; Using a discriminator Perform domain alignment in the feature space, utilizing unaligned features. , Train the discriminator to identify the modality source; Subsequently, the features after forced alignment , Mapping to the intermediate domain, the adversarial loss is as shown in formula (13): Step 4-1-2: Local Statistical Matching Constraints ; Calculate the local mean of the alignment feature within a k×k sliding window. With variance For example, in formula (14): By minimizing the second-order statistical difference between the two modes using formula (15), the consistency constraint of local motion characteristics is achieved: Step 4-1-3: Exchange and reconfigure constraints ; Through a cross-decoding mechanism, using an infrared decoder Recovering infrared images from visible light alignment features using a visible light decoder The visible light image is recovered from the infrared alignment features. The L1 loss between the reconstructed image and the original image is calculated using formula (16), thus achieving the constraint for cross-modal structure reconstruction. Step 4-2: Traditional Unsupervised Optical Flow Loss ; Based on the photometric consistency assumption, this method combines pixel-level L1 loss, structural similarity loss, and ternary loss, and introduces a second-order edge-aware smoothing loss. It also utilizes a joint occlusion mask. Invalid constraints in the shielded area, as shown in formula (17): in For pixel-level L1 luminance loss, The loss is ternary, and α1, α2, α3, and α4 are the loss weights. The second gradient of the final optical flow field. It is a second-order edge-aware smoothing loss.

5. The unsupervised infrared-visible optical flow estimation method based on cross-modal alignment according to claim 4, characterized in that, The occlusion threshold Set to 1 pixel.

6. The unsupervised infrared-visible optical flow estimation method based on cross-modal alignment according to claim 5, characterized in that, The alignment loss weight The value is 0.

5.

7. The unsupervised infrared-visible optical flow estimation method based on cross-modal alignment according to claim 6, characterized in that, The k×k value is 7×7.

8. An electronic device, characterized in that, include: Processor and memory; The memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory to cause the electronic device to perform the method as described in any one of claims 1 to 7.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1 to 7.

10. A chip, characterized in that, include: A processor for retrieving and running a computer program from memory, causing a device on which the chip is mounted to perform the method as described in any one of claims 1 to 7.