Multimodal image fusion inspection method and system for complex environments of stations

CN122574576APending Publication Date: 2026-08-14SICHUAN XINCHUAN FENGYUE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-27
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

[0003]当前市面主流场站巡检技术大多依赖可见光或热成像单一模态感知,少数简易多模态方案仅做基础数据叠加,普遍缺失三维点云几何信息支撑,存在信息维度单一、环境适应性差的核心短板,难以适配场站光照多变、设备密集、工况复杂的场景

Benefits of technology

[0051] By collaboratively acquiring multi-source data from multiple sensors and performing refined preprocessing, combined with hardware synchronization and 3D spatial constraints, multimodal data spatiotemporal alignment and pixel-level precise registration are achieved, effectively eliminating multi-source data misalignment errors. Relying on parallel feature networks and cross-modal attention mechanisms, image texture semantics, thermal radiation features, and 3D geometric structure features are adaptively fused, resulting in richer dimensions of reconstructed fused image information and significantly improving target identification and anomaly detection capabilities. Simultaneously, this method considers both static equipment faults and instrument over-limit detection, as well as dynamic target situational analysis. Through spatiotemporal topology modeling, it achieves accurate identification of abnormal behavior. Coupled with tiered alarms and structured report outputs with multimodal evidence retention, it realizes comprehensive, high-precision, and traceable intelligent inspection, significantly improving the coverage, accuracy, and timeliness of risk warnings for site inspections.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122574576A_ABST
    Figure CN122574576A_ABST
Patent Text Reader

Abstract

This invention discloses a multimodal image fusion inspection method and system for complex environments at airfields, relating to the field of image processing. The method includes: collaboratively scheduling optical, thermal, and depth sensors to simultaneously acquire RGB images, thermal images, and point cloud data, and performing preprocessing and spatiotemporal alignment on each to construct pixel-level registered multimodal data pairs; fusing texture, thermal radiation, and 3D geometric features through parallel networks and cross-modal attention mechanisms to generate enhanced fused images; based on this, multi-target detection and recognition of equipment, instruments, personnel, and vehicles are performed, analyzing equipment temperature anomalies and instrument reading overruns, and tracking dynamic target trajectories; finally, a station-level spatiotemporal topology map is constructed to analyze behavioral relationships and anomaly patterns, enabling tiered alarms and generating structured early warning reports. The advantages of this invention are: fusing visible light, thermal imaging, and 3D point cloud multimodal data, combined with precise spatiotemporal registration and cross-modal intelligent fusion technology, adapting to the complex and ever-changing environments of airfields.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing, and in particular to a multimodal image fusion inspection method and system for complex environments at stations. Background Technology

[0002] The complex environment of industrial sites, with its vast space, numerous equipment, and frequent interference from strong light, weak light, smoke, and obstructions, makes it difficult for single visible light images to capture critical information in all weather conditions and from all angles. This can easily lead to missed detections of equipment thermal failures, minor defects, or hidden hazards. Therefore, given the severe challenges faced by traditional manual or single-vision inspection methods in large industrial sites such as power, petrochemical, and energy sectors, integrating the advantages of multimodal sensors has become an inevitable direction for technological breakthroughs.

[0003] Most current mainstream site inspection technologies rely on single-modal sensing using visible light or thermal imaging. A few simplified multimodal solutions only overlay basic data, generally lacking 3D point cloud geometric information support. This results in core shortcomings such as limited information dimensions and poor environmental adaptability, making them unsuitable for scenarios with variable lighting, dense equipment, and complex operating conditions. Existing methods lack a systematic multi-source data preprocessing mechanism, failing to effectively eliminate lighting interference, sensor color difference, and temperature measurement errors. Furthermore, they do not rely on hardware synchronization and 3D spatial constraints to achieve pixel-level accurate registration, leading to misalignment and mismatch issues in multimodal data. Simultaneously, they often involve shallow feature stitching and fusion, lacking cross-modal attention complementarity mechanisms, failing to adaptively mine the advantageous features of each modality, resulting in poor fusion effects and significant loss of detail. In addition, traditional inspection functions are relatively limited, only capable of detecting abnormal temperatures in basic equipment, unable to simultaneously complete the judgment of exceeding instrument reading limits, dynamic target tracking, and site spatiotemporal situation analysis, making it difficult to identify various abnormal behaviors. Furthermore, its anomaly detection relies heavily on fixed thresholds, lacks historical baseline comparison and a hierarchical alarm system, has a high rate of false positives and false negatives, and cannot generate structured reports with multimodal evidence. The accuracy, robustness and intelligent traceability of its inspections are all significantly deficient. Summary of the Invention

[0004] To improve existing methods and systems, a multimodal image fusion inspection method and system for complex environments at stations is provided. This method integrates visible light, thermal imaging, and 3D point cloud multimodal data, and combines precise spatiotemporal registration and cross-modal intelligent fusion technology to adapt to the complex and ever-changing environment of stations.

[0005] To achieve the above objectives, the technical solution adopted by the present invention is as follows:

[0006] A multimodal image fusion inspection method for complex environments at airfields includes:

[0007] The system coordinates visible light, thermal imaging, and depth sensors to simultaneously acquire RGB images, thermal radiation images, and point cloud data of the scene's 3D structure. It performs illumination enhancement and color correction on RGB images, temperature calibration on thermal radiation images, and filtering and normalization on point cloud data.

[0008] Based on hardware synchronization and sensor calibration parameters, RGB images, thermal radiation images and scene 3D structure point cloud data are aligned in time and space. By combining feature matching with the spatial constraints of 3D point clouds, pixel-level registration of visible light and thermal imaging images is performed to construct spatially corresponding multimodal data pairs.

[0009] The texture semantic features and thermal radiation pattern features of the image are extracted by parallel network branches respectively. The cross-modal attention mechanism guides the complementary selection and fusion of features. The geometric structure features extracted from the 3D point cloud are further fused to reconstruct and generate an enhanced fused image containing texture, temperature information and depth perception cues.

[0010] Multi-target detection and recognition of equipment, instruments, personnel and vehicles are performed on fused images. For equipment targets, the temperature distribution of the thermal imaging area is analyzed and compared with historical baselines to identify overheating anomalies. For instrument targets, their readings are read and it is determined whether they exceed the limits. For dynamic targets, their appearance and motion features are extracted.

[0011] The system continuously tracks dynamic targets across inspection cycles, generates continuous trajectories, constructs a site-level spatiotemporal topology map, analyzes the spatial proximity and behavioral interaction relationships between targets, identifies abnormal behavior patterns, and issues graded alarms based on the severity and development trend of anomalies, generating structured early warning reports containing multimodal evidence.

[0012] Preferably, the steps of enhancing and correcting the illumination of the RGB image, calibrating the temperature of the thermal radiation image, and filtering and normalizing the point cloud data specifically include:

[0013] Adaptive illumination enhancement and color consistency correction are performed on RGB images to eliminate the effects of drastic changes in ambient light and color differences between different sensors; non-uniformity correction and temperature calibration are performed on thermal radiation images to map grayscale values ​​to absolute temperature or relative temperature difference; noise reduction, outlier filtering, and coordinate system normalization are performed on point cloud data.

[0014] Adaptive illumination enhancement and color consistency correction employ an image enhancement algorithm based on Retinex theory. This algorithm separates the illumination and reflection components of the image, adaptively adjusts the illumination component to balance the overall brightness, and enhances the color fidelity of the reflection component. For multiple cameras, a method based on color chart calibration or feature point color transfer is used to unify the color responses of different cameras to a reference color space.

[0015] Preferably, the multimodal data pairs corresponding to the constructed space specifically include:

[0016] Based on hardware synchronization signals and high-frequency timestamps, time alignment is performed on RGB images, thermal radiation images and point cloud data acquired at the same time.

[0017] For fixed-point sensors, modal data are transformed to a unified world coordinate system through a pre-calibrated extrinsic parameter matrix.

[0018] For mobile sensors, the precise pose of each sensor is calculated and recorded in real time at each acquisition moment, and the point cloud data is spatially projected and aligned with the RGB image and thermal radiation image.

[0019] For the RGB-thermal image pairs after projection alignment, ORB feature points and descriptors are extracted from the RGB image and the thermal radiation image, respectively. Hamming distance is used for preliminary matching. Using the depth information provided by the point cloud data, the matched point pairs are back-projected to the three-dimensional space. Mismatched point pairs with a distance exceeding a preset threshold in the three-dimensional space are removed. Then, the remaining correct matched point pairs are used to estimate the affine transformation matrix through the RANSAC algorithm to generate multimodal image pairs with corresponding relationships.

[0020] Preferably, the reconstruction to generate an enhanced fused image containing texture, temperature information, and depth-aware cues specifically includes:

[0021] The registered multimodal image is input into a feature extraction network, which includes parallel RGB feature branches and thermal radiation feature branches.

[0022] In the RGB feature branch, deep semantic features of color, texture, and shape are extracted using a convolutional neural network. In the thermal radiation feature branch, deep semantic features of temperature distribution, thermal gradient, and abnormal thermal zone patterns are extracted using a convolutional neural network.

[0023] A cross-modal attention fusion module is introduced, which uses the feature map of one modality as the query vector and the feature map of another modality as the key and value vector. By calculating attention weights, the network is guided to focus on complementary regions in the other modality that are highly discriminative for the current task, and adaptive selection and weighted fusion are performed at the feature level.

[0024] The fused deep semantic features are concatenated with the 3D geometric features extracted from point cloud data after voxelization. The concatenated multimodal joint features are then input into the decoder network, and upsampling and reconstruction are performed to generate a fused image containing RGB texture details, temperature anomaly information, and depth perception cues.

[0025] Preferably, the workflow of the cross-modal attention fusion module is as follows:

[0026] RGB feature maps are obtained from the convolutional layers of the RGB feature branch, and thermal radiation feature maps are obtained from the corresponding layers of the thermal radiation feature branch.

[0027] The RGB feature map is reshaped into a two-dimensional feature matrix, and a query vector is generated through a linear projection layer. The thermal radiation feature map is reshaped into a two-dimensional feature matrix, and a key vector and a value vector are generated through two independent linear projection layers.

[0028] Calculate the dot product of the query vector and the key vector to measure the correlation between each position in the RGB feature space and all positions in the thermal radiation feature space. After scaling the dot product result, apply the Softmax function for normalization to generate an attention weight map and obtain the thermal radiation feature region that is most discriminative for the current RGB feature position.

[0029] The value vector is weighted and summed using an attention weight map to generate a context feature vector that incorporates significant thermal radiation information. This vector is then added to the original RGB feature map using residuals to output an enhanced fused feature map.

[0030] Preferably, the multi-target detection and recognition specifically includes:

[0031] The generated fused image is input into the target detection model based on the key point detection architecture. The model is pre-trained and fine-tuned for typical targets at the site, and the target's location bounding box, category label, and confidence score are output simultaneously.

[0032] For detected equipment targets, their thermal radiation channel data is extracted, the temperature distribution histogram is analyzed, and it is compared with the historical temperature baseline model of normal operation. If there are consecutive frames where the local temperature exceeds the preset dynamic threshold, it is judged as an overheating anomaly.

[0033] For detected instrument-type targets, the dial image is corrected using perspective transformation in the RGB channel, the reading is read using a digital character recognition algorithm, and compared with the preset safety threshold range to determine whether it exceeds the limit.

[0034] For detected dynamic targets such as people and vehicles, appearance and motion features are extracted based on the temporal 3D trajectory provided by point cloud data.

[0035] Preferably, the identification of abnormal behavior patterns specifically includes:

[0036] During a continuous inspection cycle, a multi-target tracking algorithm is used for all detected dynamic targets. The algorithm predicts the position of the target in the next frame of the fused image based on Kalman filtering.

[0037] By using the appearance fusion features and motion features of the target to associate data, a unique and continuous trajectory ID is generated for each target;

[0038] Construct a site-level spatiotemporal topology graph, where nodes are target instances and edges represent spatial proximity, temporal co-occurrence, or behavioral interaction relationships between targets;

[0039] Based on the target trajectory and topological relationship, abnormal behavior patterns are identified, including: personnel entering unauthorized areas, vehicles abnormally staying, multiple targets gathering near the source of danger, and target movement trajectories deviating significantly from the preset inspection route or workflow.

[0040] Preferably, generating a structured early warning report containing multimodal evidence specifically includes:

[0041] The identified equipment overheating, instrument over-limit, and abnormal behavior events are classified into multiple levels of alarms, warnings, and severe alarms based on their severity, scope of impact, and development trend.

[0042] Based on the alarm level and type, a structured early warning report is generated. The report includes at least the anomaly type, location of occurrence, target image, thermal image screenshot, associated multimodal fusion image, confidence level, timestamp, and historical records of similar events.

[0043] Furthermore, a multimodal image fusion inspection system for complex environments at stations is proposed, including:

[0044] Data acquisition and preprocessing module: synchronously schedules multiple types of sensors to acquire images and point cloud data, and completes RGB image enhancement and correction, thermal imaging temperature calibration, and point cloud filtering and normalization;

[0045] Spatiotemporal registration and alignment module: Relying on hardware synchronization signals and calibration parameters, it realizes spatiotemporal alignment of multimodal data, and completes pixel-level registration by combining feature matching and three-dimensional spatial constraints to build accurate corresponding data pairs;

[0046] Feature fusion and reconstruction module: Extracts image texture and temperature features through a parallel network, fuses features with the help of a cross-modal attention mechanism, and reconstructs a high-quality enhanced fused image by combining three-dimensional geometric features;

[0047] Multi-target detection and recognition module: Detects multiple types of targets in fused images, and performs feature extraction and analysis for abnormal equipment temperature, excessive instrument readings, and dynamic targets respectively;

[0048] Spatiotemporal situation analysis module: continuously tracks the dynamic target generation trajectory, constructs a spatiotemporal topology map of the site, analyzes the target interaction relationship, and identifies various abnormal behavior patterns of the site;

[0049] Tiered Alarm and Report Generation Module: Based on abnormal situations, it issues tiered early warnings, integrates multimodal evidence, and automatically generates structured inspection and early warning reports.

[0050] Compared with the prior art, the advantages of the present invention are:

[0051] By collaboratively acquiring multi-source data from multiple sensors and performing refined preprocessing, combined with hardware synchronization and 3D spatial constraints, multimodal data spatiotemporal alignment and pixel-level precise registration are achieved, effectively eliminating multi-source data misalignment errors. Relying on parallel feature networks and cross-modal attention mechanisms, image texture semantics, thermal radiation features, and 3D geometric structure features are adaptively fused, resulting in richer dimensions of reconstructed fused image information and significantly improving target identification and anomaly detection capabilities. Simultaneously, this method considers both static equipment faults and instrument over-limit detection, as well as dynamic target situational analysis. Through spatiotemporal topology modeling, it achieves accurate identification of abnormal behavior. Coupled with tiered alarms and structured report outputs with multimodal evidence retention, it realizes comprehensive, high-precision, and traceable intelligent inspection, significantly improving the coverage, accuracy, and timeliness of risk warnings for site inspections. Attached Figure Description

[0052] Figure 1 This is a schematic diagram of the multimodal image fusion inspection method for complex environments at stations proposed in this invention;

[0053] Figure 2 This is a schematic diagram of the reconstructed, enhanced, and fused image proposed in this invention;

[0054] Figure 3 This is a schematic diagram of the multi-target detection and recognition method proposed in this invention. Detailed Implementation

[0055] The following description is intended to disclose the invention and enable those skilled in the art to implement it. The preferred embodiments described below are merely examples, and other obvious variations will occur to those skilled in the art.

[0056] A multimodal image fusion inspection system for complex environments at airfields includes:

[0057] Data acquisition and preprocessing module: synchronously schedules multiple types of sensors to acquire images and point cloud data, and completes RGB image enhancement and correction, thermal imaging temperature calibration, and point cloud filtering and normalization;

[0058] Spatiotemporal registration and alignment module: Relying on hardware synchronization signals and calibration parameters, it realizes spatiotemporal alignment of multimodal data, and completes pixel-level registration by combining feature matching and three-dimensional spatial constraints to build accurate corresponding data pairs;

[0059] Feature fusion and reconstruction module: Extracts image texture and temperature features through a parallel network, fuses features with the help of a cross-modal attention mechanism, and reconstructs a high-quality enhanced fused image by combining three-dimensional geometric features;

[0060] Multi-target detection and recognition module: Detects multiple types of targets in fused images, and performs feature extraction and analysis for abnormal equipment temperature, excessive instrument readings, and dynamic targets respectively;

[0061] Spatiotemporal situation analysis module: continuously tracks the dynamic target generation trajectory, constructs a spatiotemporal topology map of the site, analyzes the target interaction relationship, and identifies various abnormal behavior patterns of the site;

[0062] Tiered Alarm and Report Generation Module: Based on abnormal situations, it issues tiered early warnings, integrates multimodal evidence, and automatically generates structured inspection and early warning reports.

[0063] See Figure 1 As shown, a multimodal image fusion inspection method for complex environments at airfields includes:

[0064] The system coordinates visible light, thermal imaging, and depth sensors to simultaneously acquire RGB images, thermal radiation images, and point cloud data of the scene's 3D structure. It performs illumination enhancement and color correction on RGB images, temperature calibration on thermal radiation images, and filtering and normalization on point cloud data.

[0065] Based on hardware synchronization and sensor calibration parameters, RGB images, thermal radiation images and scene 3D structure point cloud data are aligned in time and space. By combining feature matching with the spatial constraints of 3D point clouds, pixel-level registration of visible light and thermal imaging images is performed to construct spatially corresponding multimodal data pairs.

[0066] The texture semantic features and thermal radiation pattern features of the image are extracted by parallel network branches respectively. The cross-modal attention mechanism guides the complementary selection and fusion of features. The geometric structure features extracted from the 3D point cloud are further fused to reconstruct and generate an enhanced fused image containing texture, temperature information and depth perception cues.

[0067] Multi-target detection and recognition of equipment, instruments, personnel and vehicles are performed on fused images. For equipment targets, the temperature distribution of the thermal imaging area is analyzed and compared with historical baselines to identify overheating anomalies. For instrument targets, their readings are read and it is determined whether they exceed the limits. For dynamic targets, their appearance and motion features are extracted.

[0068] The system continuously tracks dynamic targets across inspection cycles, generates continuous trajectories, constructs a site-level spatiotemporal topology map, analyzes the spatial proximity and behavioral interaction relationships between targets, identifies abnormal behavior patterns, and issues graded alarms based on the severity and development trend of anomalies, generating structured early warning reports containing multimodal evidence.

[0069] Specifically, the ambient lighting conditions at the site are complex and variable, with phenomena such as direct sunlight, equipment shadows, and uneven nighttime lighting, leading to problems such as localized overexposure, loss of detail, and poor overall contrast in visible light images. Therefore, adaptive lighting enhancement and color consistency correction are performed on RGB images to eliminate the effects of drastic changes in ambient lighting and color differences between different sensors. This adaptive lighting enhancement and color consistency correction employs an image enhancement algorithm based on Retinex theory, separating the illumination and reflection components of the image. The illumination component is adaptively adjusted to balance overall brightness, and the reflection component undergoes color fidelity enhancement. For multiple cameras, a method based on color chart calibration or feature point color transfer is used to unify the color responses of different cameras to a reference color space. Specifically, this includes:

[0070] Adaptive illumination enhancement is performed using a multi-scale Retinex-based algorithm to enhance the observed images. Decomposed into incident light components and reflection component The model is represented as:

[0071]

[0072] Where (x, y) are pixel coordinates. The goal of enhancement is to estimate and eliminate the uneven illumination component L, thereby recovering the color-constant and detail-clear reflection component R. A logarithmic transformation converts the multiplicative relationship into an additive one:

[0073]

[0074] In practice, light component The value can be approximated by convolving the original image with a Gaussian wrapping function G(x,y). To preserve both detail and color fidelity, this invention employs a multi-scale Gaussian kernel for processing. The single-scale Retinex output is:

[0075]

[0076] in, The standard deviation is The k-th Gaussian kernel, This represents the convolution operation. The final multi-scale reflection components. It is a weighted sum of multiple single-scale results:

[0077]

[0078] Using three scales and assigning equal weights, for Exponential calculations and adaptive contrast stretching are performed to obtain an enhanced image with balanced illumination.

[0079] A color correction method based on a standard color chart is adopted. A standard 24-color chart is placed in the public field of view, and images from a reference camera and a camera to be calibrated are acquired. The color value of each color block in the standard sRGB color space is extracted, and the pixel value of each frame of the image acquired by the camera to be calibrated is corrected in real time.

[0080] Non-uniformity correction and temperature calibration of thermal radiation images, mapping grayscale values ​​to absolute temperature or relative temperature difference, specifically includes:

[0081] A two-point temperature calibration method was used for correction and physical calibration, employing two blackbody radiation sources with known temperatures, designated as the low-temperature reference source and the high-temperature reference source. Thermal images of the two blackbody sources were acquired, obtaining the raw digital output value corresponding to each pixel. Using these two points, a linear grayscale-temperature response model was established for each pixel. After calibration, the absolute temperature value corresponding to the raw output value of any pixel in any thermal image acquired during the inspection process was calculated.

[0082] The specific process for denoising, outlier removal filtering, and coordinate system normalization of point cloud data is as follows:

[0083] Raw point clouds acquired by depth sensors typically contain a large number of noisy points and isolated outliers. A two-stage filtering strategy is employed. First, statistical outlier filtering is performed. For each query point in the point cloud, its distance to all nearest neighbors is calculated, and the average distance and standard deviation are determined; a global distance threshold multiplier is set. If a query point meets the following conditions, it is identified as an outlier and removed:

[0084]

[0085] in, This is the global distance threshold multiplier. and These are all points The global mean and standard deviation.

[0086] Secondly, radius filtering is performed to remove residual sparse noise. For each point in the statistically filtered point cloud, the number of neighboring points within a spherical neighborhood centered at that point is counted. If this number is less than a preset threshold, the point is considered an isolated point and is removed. For fixed-position sensors, transformation is performed using pre-calibrated extrinsic parameters. For mobile platforms, the pose calculated from the point cloud acquired at each moment is converted to the global coordinate system in real time. Finally, all point cloud data are aligned and fused under a unified spatiotemporal reference.

[0087] Based on hardware synchronization signals and high-frequency timestamps, RGB images, thermal radiation images, and point cloud data acquired at the same time are time-aligned. For fixed-point sensors, the modal data are transformed to a unified world coordinate system using a pre-calibrated extrinsic parameter matrix. For mobile sensors, the precise pose of each sensor is calculated and recorded in real time at each acquisition moment, and the point cloud data is spatially projected and aligned with the RGB and thermal radiation images. The specific process is as follows:

[0088] By employing a hardware synchronization triggering or timestamp alignment scheme, synchronous trigger pulses are sent to all sensors to ensure that data acquisition occurs at the same physical moment, fundamentally eliminating motion blur or misalignment caused by scene changes due to asynchronous sampling periods. Spatially, fixed-point sensors obtain the extrinsic parameters of each sensor relative to the world coordinate system, i.e., the rotation matrix, through pre-calibrated joint calibration. Translation vector For a collected 3D point cloud point Its world coordinates Calculated using the following formula:

[0089]

[0090] The mobile sensors are mounted on a mobile platform, with each sensor's position relative to the platform body fixed. The platform calculates its own pose in the world coordinate system in real time using a tightly coupled visual-inertial odometry system. Converting any sensor data to world coordinates requires two steps: first, conversion to the platform coordinate system, and then conversion to the world coordinate system.

[0091] After initial alignment, subpixel-level offsets may still exist between the RGB image and the thermal image due to lens distortion, resolution differences, and slight calibration errors. A feature matching method incorporating three-dimensional spatial constraints is employed to achieve accurate registration.

[0092] For the projected and aligned RGB-thermal image pairs, ORB feature points and descriptors are extracted from the RGB image and the thermal radiation image, respectively. Preliminary matching is performed using Hamming distance. Using the depth information provided by the point cloud data, the matched point pairs are back-projected into 3D space. Mismatched point pairs with distances exceeding a preset threshold in 3D space are removed. Then, the remaining correctly matched point pairs are used to estimate the affine transformation matrix using the RANSAC algorithm, generating corresponding multimodal image pairs. The specific steps are as follows:

[0093] On the initially projected and aligned RGB and thermal radiation images, ORB feature points and their binary descriptors are extracted, respectively. Let the feature point set extracted from the RGB image be... Its descriptor is Extracted from thermal images The descriptor is Hamming distance is used for descriptor matching, and two binary descriptors are calculated. and Differences between them:

[0094]

[0095] in, This represents the XOR operation, where n is the descriptor length. For Each feature point in, Find the nearest and second nearest corresponding points in the Hamming distance. If the ratio of the nearest distance to the second nearest distance is less than a preset threshold, accept the nearest neighbor as a preliminary matching point pair to obtain the initial matching set.

[0096] The initial matching set contains a large number of mismatches. For each matching pair, using camera intrinsics and pose, they are back-projected into 3D space to obtain their corresponding 3D position estimates. The Euclidean distance between the two 3D points is calculated. If the distance is greater than a preset geometric consistency threshold, the match is considered a mismatch and is discarded. All matching pairs that meet a preset geometric consistency threshold are retained to form a refined matching set.

[0097] Using the refined matching set, a two-dimensional affine transformation matrix from the thermal image to the RGB image is applied to correct for remaining subtle deformations. This transformation is represented in homogeneous coordinates as follows:

[0098]

[0099] in,[ , ] is a point in the thermal image, [ , [] represents the point in the RGB image coordinate system after the transformation. Matrix Includes translation, rotation, scaling, and shearing parameters.

[0100] To robustly estimate the matrix To avoid the influence of residual false matches, the RANSAC algorithm is used to randomly extract a minimum sample set from the refined matching set. This sample set is then used to compute a candidate match. .use Test all matching point pairs and calculate the projection error. If the projection error of a pair of matching points is less than a threshold, mark it as an interior point. Repeat the above process multiple times, and finally select the pair with the most interior points. As the optimal transformation matrix, the optimal transformation matrix is ​​refined using the least squares method with all interior points to obtain the final high-precision affine transformation matrix. The thermal radiation image is then resampled and precisely twisted into the coordinate system of the RGB image to generate pixel-aligned RGB-thermal image pairs.

[0101] See Figure 2 As shown, the reconstruction and generation of the enhanced fused image specifically includes:

[0102] The registered multimodal image is input into a feature extraction network, which includes parallel RGB feature branches and thermal radiation feature branches.

[0103] In the RGB feature branch, deep semantic features of color, texture, and shape are extracted using a convolutional neural network. In the thermal radiation feature branch, deep semantic features of temperature distribution, thermal gradient, and abnormal thermal zone patterns are extracted using a convolutional neural network.

[0104] A cross-modal attention fusion module is introduced, which uses the feature map of one modality as the query vector and the feature map of another modality as the key and value vector. By calculating attention weights, the network is guided to focus on complementary regions in the other modality that are highly discriminative for the current task, and adaptive selection and weighted fusion are performed at the feature level.

[0105] The fused deep semantic features are concatenated with the 3D geometric features extracted from point cloud data after voxelization. The concatenated multimodal joint features are then input into the decoder network, and upsampling and reconstruction are performed to generate a fused image containing RGB texture details, temperature anomaly information, and depth perception cues.

[0106] Specifically, the registered multimodal image is input into the feature extraction network, which is a dual-branch encoder network, to process the registered RGB image and the thermal radiation image respectively.

[0107] Both branches employ convolutional neural networks with residual connections as their backbone, but do not share weights after the initial layer to learn modality-specific feature representations. The RGB branch aims to extract rich appearance-semantic features, including color, texture, and shape, with its feature maps encoding hierarchical information from low-level edges and textures to high-level semantics. The thermal radiation branch focuses on extracting features related to thermophysical states, including temperature distribution, thermal gradients, and anomalous thermal patterns, with its feature maps encoding temperature distribution patterns, thermal gradients, and potential anomalous high or low temperature regions.

[0108] A cross-modal attention fusion module is introduced, using the feature map of one modality as the query vector and the feature map of another modality as the key-value vector. By calculating attention weights, the network is guided to focus on complementary regions in the other modality that are highly discriminative for the current task, performing adaptive selection and weighted fusion at the feature level. The specific steps are as follows:

[0109] This allows a feature from one modality to actively "query" complementary information from another modality. Taking RGB features as the query and hot features as the key and value, the calculation process is as follows:

[0110] For the input feature map and Perform 1×1 convolutions on each matrix, project them onto the same feature space, and reshape them into two-dimensional matrices to generate the query matrix Q, the key matrix K, and the value matrix V:

[0111]

[0112] in, , , This represents a learnable linear projection operation.

[0113] Calculate the attention weight matrix A, where each element of the matrix... This indicates the degree of correlation between RGB feature position i and thermal radiation feature position j:

[0114]

[0115] in, It is the dimension of the key vector. This is a scaling factor used to prevent the gradient of the Softmax function from vanishing due to an excessively large dot product result.

[0116] The value matrix V is weighted and summed using attention weights to generate a context feature matrix C that incorporates significant thermal radiation information: C = AV. Finally, the obtained context feature C is reshaped back into its spatial format and compared with the original RGB features. Perform residual connections and output the enhanced fused features.

[0117] The fused deep semantic features are concatenated with the 3D geometric features extracted from point cloud data after voxelization. The concatenated multimodal joint features are then input into the decoder network, and upsampling and reconstruction are performed to generate a fused image containing RGB texture details, temperature anomaly information, and depth perception cues. Specifically, this includes:

[0118] The point cloud aligned to the world coordinate system is processed to extract geometric features. First, voxelization is performed, dividing the 3D space into a uniform mesh. For each non-empty voxel, a set of points is contained. The principal orientation of the point cloud within the voxel is calculated to obtain the surface normal vector. By constructing the covariance matrix and performing eigenvalue decomposition, the eigenvector corresponding to the smallest eigenvalue of the covariance matrix is ​​approximately the surface normal vector. Simultaneously, the curvature of the voxel is calculated as a measure of geometric complexity, which can be approximated as the ratio of the smallest eigenvalue to the sum of the three eigenvalues. The local height difference is defined as the difference between the maximum and minimum coordinates of points within the voxel. Each voxel is ultimately represented as a geometric eigenvector.

[0119] A 3D voxel mesh is projected onto a 2D image plane to generate a 3D geometric feature map. Finally, the visual fusion features and the geometric feature map are concatenated along the channel dimension to obtain a unified feature representation of joint appearance, thermal state, and 3D geometry.

[0120] The joint features are input into a decoder network consisting of multiple upsampling modules. Each upsampling module typically contains a transposed convolution or pixel shuffling operation, followed by a regular convolutional layer and an activation function to progressively restore spatial resolution. The final layer of the decoder uses a 1×1 convolution to map the number of channels to 3 and employs a sigmoid activation function to output a fused image of the same size as the input RGB image.

[0121] See Figure 3 As shown, multi-target detection and recognition specifically includes:

[0122] The generated fused image is input into the target detection model based on the key point detection architecture. The model is pre-trained and fine-tuned for typical targets at the site, and the target's location bounding box, category label, and confidence score are output simultaneously.

[0123] For detected equipment targets, their thermal radiation channel data is extracted, the temperature distribution histogram is analyzed, and it is compared with the historical temperature baseline model of normal operation. If there are consecutive frames where the local temperature exceeds the preset dynamic threshold, it is judged as an overheating anomaly.

[0124] For detected instrument-type targets, the dial image is corrected using perspective transformation in the RGB channel, the reading is read using a digital character recognition algorithm, and compared with the preset safety threshold range to determine whether it exceeds the limit.

[0125] For detected dynamic targets such as people and vehicles, appearance and motion features are extracted based on the temporal 3D trajectory provided by point cloud data.

[0126] Specifically, the generated fused image is input into a target detection model based on a keypoint detection architecture. Pre-training and fine-tuning are performed on typical targets at the site, and the target's location bounding box, category label, and confidence score are output simultaneously. The steps are as follows:

[0127] The generated fused image is input into a pre-trained target detection network fine-tuned for the field scene. A single-stage detector based on anchor boxes is preferred. The network outputs a set of detection results, each represented as a tuple containing the location bounding box, class label, and confidence parameter. Only detection results with a confidence level higher than a preset threshold are accepted for subsequent analysis. For each target detected as a "device," a corresponding temperature matrix region is extracted from the original thermal radiation image based on its bounding box. The temperature distribution within this region is analyzed to diagnose overheating anomalies.

[0128] For detected equipment targets, their thermal radiation channel data is extracted, the temperature distribution histogram is analyzed, and compared with the historical temperature baseline model of normal operation. If the local temperature of consecutive frames exceeds the preset dynamic threshold, it is judged as an overheating anomaly. For detected instrument targets, the dial image is corrected using perspective transformation in the RGB channel, the reading is read using a digit character recognition algorithm, and compared with the preset safety threshold range to determine whether it exceeds the limit. For detected dynamic targets such as personnel and vehicles, appearance and motion features are extracted based on the temporal 3D trajectory provided by point cloud data. Specifically:

[0129] The temperature values ​​of all valid pixels within the region are statistically analyzed to construct a temperature histogram. Key statistics are calculated, including average temperature, maximum temperature, and temperature standard deviation. This temperature distribution is compared with a historical temperature baseline model for normal device operation. The baseline model can be a dynamic threshold model obtained by statistically analyzing historical temperature data of the device under similar environmental conditions. One criterion for determining overheating anomalies is to check whether there are consecutive frames where the local maximum temperature consistently exceeds the historical baseline average temperature. The slope is calculated by linearly fitting the average temperature sequence of the most recent frames. If the slope is greater than a positive threshold, it is determined to be an abnormal temperature rise.

[0130] For "dashboard" type targets, the dial area image is extracted from the original RGB image using its bounding box. Perspective transformation based on feature points is used for correction. A homography matrix is ​​calculated by detecting the four vertices of the dial area. Images from a tilted perspective Transform into a front view image :

[0131]

[0132] For pointer-type instruments, The pointer line segment is detected using Hough transform or a deep learning model, and the angle of the pointer relative to the zero mark is calculated. If the instrument range is [ , ], the corresponding angle range is [ , Then the reading V is:

[0133]

[0134] For digital instruments, an optical character recognition engine is used to directly identify the displayed digit V. Finally, the reading V is compared with the instrument's preset safety threshold range to determine if the instrument reading exceeds the limit.

[0135] For dynamic targets such as people and vehicles, at each time step, the centroid position of the target in the 3D world coordinate system is calculated based on the position of the target detection box in the image, combined with camera parameters and point cloud data. By tracking the target's ID in consecutive frames, its motion trajectory sequence can be obtained.

[0136] Within a continuous inspection cycle, a multi-target tracking algorithm is employed for all detected dynamic targets. This algorithm predicts the target's position in the next frame of the fused image based on Kalman filtering. Data association is performed using the target's appearance fusion features and motion features to generate a unique and continuous trajectory ID for each target. A site-level spatiotemporal topology map is constructed, where nodes represent target instances, and edges represent spatial proximity, temporal co-occurrence, or behavioral interaction relationships between targets. Based on the target trajectory and topological relationships, abnormal behavior patterns are identified, including: personnel entering unauthorized areas, vehicles abnormally lingering, multiple targets clustering near a hazard source, and target movement trajectories significantly deviating from the preset inspection route or workflow. The specific steps are as follows:

[0137] In continuous video frames or inspection cycles, for each detected dynamic target, a tracking framework based on Kalman filtering motion prediction and multi-feature fusion is adopted to maintain a state vector for each active target k at time t. ,in( ) represents the center coordinates of the bounding box in the image plane. Its speed, The bounding box has a height and width. Kalman filtering consists of two steps: based on a uniform motion model, it predicts the target's state in the next frame t+1. and state covariance matrix :

[0138]

[0139] Where F is the state transition matrix and Q is the process noise covariance matrix.

[0140] When new detection data associated with the target is obtained in frame t+1 Then, update the state estimate. Covariance .

[0141] To match the detection set of a new frame with the existing trajectory set, a cost matrix is ​​constructed. This cost matrix integrates motion consistency cost and appearance similarity cost. The motion cost uses the Mahalanobis distance between the predicted state and the detection state, and the appearance cost uses the cosine distance to measure the similarity between the deep learning feature vectors extracted from the target fused image patch and the target. The Hungarian algorithm is used to solve for the optimal allocation of this cost matrix, thereby assigning the detection to the existing trajectory, initializing it as a new trajectory, or determining it as the end of the trajectory.

[0142] To perform higher-level situational analysis, all targets and their trajectories are modeled as a dynamic spatiotemporal topology graph. .

[0143] Each node represents a target instance, with attributes including target type, current position, current speed, and historical trajectory fragments. Edges represent a relationship between targets and are weighted to indicate the strength of that relationship. Three main types of edges are constructed: Spatial proximity edges: An edge is created if the Euclidean distance between two targets in 3D space is less than a threshold. The weight is inversely proportional to the distance. Temporal co-occurrence edges: An edge is created if two targets frequently appear in the same local area in the recent past, even if they are not currently adjacent. The weight is based on the co-occurrence frequency. Behavioral interaction edges: For example, person A pointing at device B for an extended period, or vehicle C following vehicle D. These edges can be created by analyzing the relative speed direction and line-of-sight direction of the trajectory.

[0144] The specific process for generating structured early warning reports containing multimodal evidence is as follows: Based on the constructed trajectory and topology map, intrusion into unauthorized areas, abnormal loitering, multi-target aggregation in dangerous areas, and trajectory deviation from preset paths are defined and identified. The identified events are then classified according to a predefined rule base. Alert reports are automatically generated, including: anomaly type, world coordinates, timestamp, confidence level, associated original / thermal / fused image snapshots, target trajectory fragments, topological relationship screenshots, and links to historical similar event records, forming a structured log for subsequent tracing and decision support.

[0145] It should be noted that the order of the above embodiments of the present invention is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. Furthermore, the above description focuses on specific embodiments of this specification. Additionally, the processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired results. In some embodiments, multitasking and parallel processing are possible or may be advantageous.

[0146] The various embodiments in this specification are described in a progressive manner. The same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on describing the differences from other embodiments.

[0147] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A multimodal image fusion inspection method for complex environments at stations, characterized in that, include: The system coordinates visible light, thermal imaging, and depth sensors to simultaneously acquire RGB images, thermal radiation images, and point cloud data of the scene's 3D structure. It performs illumination enhancement and color correction on RGB images, temperature calibration on thermal radiation images, and filtering and normalization on point cloud data. Based on hardware synchronization and sensor calibration parameters, RGB images, thermal radiation images and scene 3D structure point cloud data are aligned in time and space. By combining feature matching with the spatial constraints of 3D point clouds, pixel-level registration of visible light and thermal imaging images is performed to construct spatially corresponding multimodal data pairs. The texture semantic features and thermal radiation pattern features of the image are extracted by parallel network branches respectively. The cross-modal attention mechanism guides the complementary selection and fusion of features. The geometric structure features extracted from the 3D point cloud are further fused to reconstruct and generate an enhanced fused image containing texture, temperature information and depth perception cues. Multi-target detection and recognition of equipment, instruments, personnel and vehicles are performed on fused images. For equipment targets, the temperature distribution of the thermal imaging area is analyzed and compared with historical baselines to identify overheating anomalies. For instrument targets, their readings are read and it is determined whether they exceed the limits. For dynamic targets, their appearance and motion features are extracted. The system continuously tracks dynamic targets across inspection cycles, generates continuous trajectories, constructs a site-level spatiotemporal topology map, analyzes the spatial proximity and behavioral interaction relationships between targets, identifies abnormal behavior patterns, and issues graded alarms based on the severity and development trend of anomalies, generating structured early warning reports containing multimodal evidence.

2. The multimodal image fusion inspection method for complex environments of stations according to claim 1, characterized in that, The specific steps of enhancing and correcting the illumination of RGB images, calibrating the temperature of thermal radiation images, and filtering and normalizing point cloud data include: Adaptive illumination enhancement and color consistency correction are performed on RGB images to eliminate the effects of drastic changes in ambient light and color differences between different sensors; non-uniformity correction and temperature calibration are performed on thermal radiation images to map grayscale values ​​to absolute temperature or relative temperature difference; noise reduction, outlier filtering, and coordinate system normalization are performed on point cloud data. Adaptive illumination enhancement and color consistency correction employ an image enhancement algorithm based on Retinex theory. This algorithm separates the illumination and reflection components of the image, adaptively adjusts the illumination component to balance the overall brightness, and enhances the color fidelity of the reflection component. For multiple cameras, a method based on color chart calibration or feature point color transfer is used to unify the color responses of different cameras to a reference color space.

3. The multimodal image fusion inspection method for complex environments of stations according to claim 1, characterized in that, The multimodal data pairs corresponding to the constructed space specifically include: Based on hardware synchronization signals and high-frequency timestamps, time alignment is performed on RGB images, thermal radiation images and point cloud data acquired at the same time. For fixed-point sensors, modal data are transformed to a unified world coordinate system through a pre-calibrated extrinsic parameter matrix. For mobile sensors, the precise pose of each sensor is calculated and recorded in real time at each acquisition moment, and the point cloud data is spatially projected and aligned with the RGB image and thermal radiation image. For the RGB-thermal image pairs after projection alignment, ORB feature points and descriptors are extracted from the RGB image and the thermal radiation image, respectively. Hamming distance is used for preliminary matching. Using the depth information provided by the point cloud data, the matched point pairs are back-projected to the three-dimensional space. Mismatched point pairs with a distance exceeding a preset threshold in the three-dimensional space are removed. Then, the remaining correct matched point pairs are used to estimate the affine transformation matrix through the RANSAC algorithm to generate multimodal image pairs with corresponding relationships.

4. The multimodal image fusion inspection method for complex environments of stations according to claim 1, characterized in that, The reconstruction generates an enhanced fused image containing texture, temperature information, and depth-aware cues, specifically including: The registered multimodal image is input into a feature extraction network, which includes parallel RGB feature branches and thermal radiation feature branches. In the RGB feature branch, deep semantic features of color, texture, and shape are extracted using a convolutional neural network. In the thermal radiation feature branch, deep semantic features of temperature distribution, thermal gradient, and abnormal thermal zone patterns are extracted using a convolutional neural network. A cross-modal attention fusion module is introduced, which uses the feature map of one modality as the query vector and the feature map of another modality as the key and value vector. By calculating attention weights, the network is guided to focus on complementary regions in the other modality that are highly discriminative for the current task, and adaptive selection and weighted fusion are performed at the feature level. The fused deep semantic features are concatenated with the 3D geometric features extracted from point cloud data after voxelization. The concatenated multimodal joint features are then input into the decoder network, and upsampling and reconstruction are performed to generate a fused image containing RGB texture details, temperature anomaly information, and depth perception cues.

5. The multimodal image fusion inspection method for complex environments of stations according to claim 4, characterized in that, The workflow of the cross-modal attention fusion module is as follows: RGB feature maps are obtained from the convolutional layers of the RGB feature branch, and thermal radiation feature maps are obtained from the corresponding layers of the thermal radiation feature branch. The RGB feature map is reshaped into a two-dimensional feature matrix, and a query vector is generated through a linear projection layer. The thermal radiation feature map is reshaped into a two-dimensional feature matrix, and a key vector and a value vector are generated through two independent linear projection layers. Calculate the dot product of the query vector and the key vector to measure the correlation between each position in the RGB feature space and all positions in the thermal radiation feature space. After scaling the dot product result, apply the Softmax function for normalization to generate an attention weight map and obtain the thermal radiation feature region that is most discriminative for the current RGB feature position. The value vector is weighted and summed using an attention weight map to generate a context feature vector that incorporates significant thermal radiation information. This vector is then added to the original RGB feature map using residuals to output an enhanced fused feature map.

6. The multimodal image fusion inspection method for complex environments of stations according to claim 1, characterized in that, The multi-target detection and recognition specifically includes: The generated fused image is input into the target detection model based on the key point detection architecture. The model is pre-trained and fine-tuned for typical targets at the site, and the target's location bounding box, category label, and confidence score are output simultaneously. For detected equipment targets, their thermal radiation channel data is extracted, the temperature distribution histogram is analyzed, and it is compared with the historical temperature baseline model of normal operation. If there are consecutive frames where the local temperature exceeds the preset dynamic threshold, it is judged as an overheating anomaly. For detected instrument-type targets, the dial image is corrected using perspective transformation in the RGB channel, the reading is read using a digital character recognition algorithm, and compared with the preset safety threshold range to determine whether it exceeds the limit. For detected dynamic targets such as people and vehicles, appearance and motion features are extracted based on the temporal 3D trajectory provided by point cloud data.

7. The multimodal image fusion inspection method for complex environments of stations according to claim 1, characterized in that, The identification of abnormal behavior patterns specifically includes: During a continuous inspection cycle, a multi-target tracking algorithm is used for all detected dynamic targets. The algorithm predicts the position of the target in the next frame of the fused image based on Kalman filtering. By using the appearance fusion features and motion features of the target to associate data, a unique and continuous trajectory ID is generated for each target; Construct a site-level spatiotemporal topology graph, where nodes are target instances and edges represent spatial proximity, temporal co-occurrence, or behavioral interaction relationships between targets; Based on the target trajectory and topological relationship, abnormal behavior patterns are identified, including: personnel entering unauthorized areas, vehicles abnormally staying, multiple targets gathering near the source of danger, and target movement trajectories deviating significantly from the preset inspection route or workflow.

8. The multimodal image fusion inspection method for complex environments of stations according to claim 1, characterized in that, The generation of a structured early warning report containing multimodal evidence specifically includes: The identified equipment overheating, instrument over-limit, and abnormal behavior events are classified into multiple levels of alarms, warnings, and severe alarms based on their severity, scope of impact, and development trend. Based on the alarm level and type, a structured early warning report is generated. The report includes at least the anomaly type, location of occurrence, target image, thermal image screenshot, associated multimodal fusion image, confidence level, timestamp, and historical records of similar events.

9. A multimodal image fusion inspection system for complex environments of stations, used to implement the multimodal image fusion inspection method for complex environments of stations as described in any one of claims 1-8, characterized in that, include: Data acquisition and preprocessing module: synchronously schedules multiple types of sensors to acquire images and point cloud data, and completes RGB image enhancement and correction, thermal imaging temperature calibration, and point cloud filtering and normalization; Spatiotemporal registration and alignment module: Relying on hardware synchronization signals and calibration parameters, it realizes spatiotemporal alignment of multimodal data, and completes pixel-level registration by combining feature matching and three-dimensional spatial constraints to build accurate corresponding data pairs; Feature fusion and reconstruction module: Extracts image texture and temperature features through a parallel network, fuses features with the help of a cross-modal attention mechanism, and reconstructs a high-quality enhanced fused image by combining three-dimensional geometric features; Multi-target detection and recognition module: Detects multiple types of targets in fused images, and performs feature extraction and analysis for abnormal equipment temperature, excessive instrument readings, and dynamic targets respectively; Spatiotemporal situation analysis module: continuously tracks the dynamic target generation trajectory, constructs a spatiotemporal topology map of the site, analyzes the target interaction relationship, and identifies various abnormal behavior patterns of the site; Tiered Alarm and Report Generation Module: Based on abnormal situations, it issues tiered early warnings, integrates multimodal evidence, and automatically generates structured inspection and early warning reports.