A furnace online temperature measurement target positioning method based on double light fusion
By employing deep learning-based image correction, intelligent fusion, and target tracking prediction technologies, the challenges of image distortion and tracking caused by thermal disturbances and rotation in rotary kilns have been solved. This has enabled high-precision and robust positioning and prediction of thermal anomalies on the surface of rotary kilns, thereby enhancing the intelligence of the monitoring system and production safety.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANGHAI JINYI INSPECTION TECH
- Filing Date
- 2026-04-20
- Publication Date
- 2026-07-10
AI Technical Summary
Existing dual-light fusion systems face challenges in industrial settings such as rotary kilns, including geometric distortion caused by heat wave disturbances and dynamic tracking difficulties caused by furnace rotation. This results in insufficient accuracy and reliability of image registration and thermal anomaly detection, making it impossible to achieve efficient and stable monitoring.
Image correction is performed using a deep learning-based thermal disturbance perception and hierarchical registration network (H2R-Net), intelligent fusion is performed by combining an environment-adaptive cross-modal fusion and segmentation network (EACF-Net), and target tracking and prediction are performed using a kinematically-aware spatiotemporal graph Kalman network (K-STGKN), thereby achieving accurate localization and prediction of thermal anomalies on the surface of a rotary kiln.
It enables high-precision and robust location and prediction of thermal anomalies in complex industrial environments, improves the level of intelligent monitoring, reduces manual intervention, avoids unplanned downtime, and enhances production safety and equipment lifespan.
Smart Images

Figure CN122368003A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of industrial automation monitoring technology, specifically to a method for online temperature measurement target positioning in furnaces and kilns based on dual-light fusion. Background Technology
[0002] Industrial furnaces, especially rotary kilns, are core production equipment in basic industries such as steel, cement, and chemicals. They operate under harsh conditions of high temperature, heavy load, and continuous operation for extended periods. The integrity and stability of the furnace structure are crucial to production safety, product quality, and equipment lifespan. One of the main failure modes of furnaces is the erosion, spalling, or damage to their internal refractory materials. This internal damage directly manifests as areas of abnormally high temperature on the outer surface of the furnace shell, i.e., thermal anomaly points or hot spots. Failure to detect and address these thermal anomalies in a timely manner can lead to burn-through and deformation of the furnace shell steel plates, or even major production accidents such as kiln shutdown for maintenance, resulting in significant economic losses.
[0003] To mitigate the aforementioned risks, the industry commonly employs temperature monitoring to assess the health status of furnaces and kilns. Traditional monitoring methods primarily rely on manual, periodic inspections using handheld infrared thermometers or thermal imagers. However, this approach has several inherent drawbacks: First, it is inefficient and lacks real-time accuracy; the long inspection cycles prevent the capture of instantaneous temperature changes, resulting in significant data delays. Second, the test results are highly susceptible to human error, easily leading to missed or false detections and untimely warnings. Finally, the harsh working environment surrounding the furnaces and kilns, with its high-temperature radiation and dust hazard, not only threatens the health of inspection personnel but also contributes to recruitment difficulties and continuously rising labor costs.
[0004] To overcome the limitations of manual monitoring, early automated systems primarily employed single infrared thermal imaging technology, using fixed thermal imagers to continuously scan the temperature field of the furnace surface 24 / 7. While these systems achieved continuous monitoring, their limitations were significant: lacking the rich texture and contextual information provided by visible light images, they struggled to effectively distinguish genuine thermal anomalies caused by internal refractory material damage from false hotspots caused by external factors such as stains, ash accumulation, and temporary deposits on the furnace shell surface. Operators had to expend considerable effort on manual identification, resulting in limited levels of automation and intelligence.
[0005] To address these issues, dual-light fusion monitoring systems integrating visible light cameras and infrared thermal imagers have become the mainstream development direction. However, existing dual-light fusion systems still face a series of severe and unresolved technical challenges when applied to special industrial scenarios such as rotary kilns: 1. Geometric Distortion Problem Caused by Thermal Wave Disturbance: The enormous heat radiated from the furnace shell surface heats the surrounding air, creating intense heat waves and turbulence. This non-uniform air medium acts like a dynamically changing lens, irregularly refracting light along its propagation path, thus introducing severe, nonlinear geometric distortions into visible light and thermal imaging video streams. This distortion causes object contours in the image to become distorted, jittery, and blurred. Traditional image registration algorithms based on rigid transformations (such as translation and rotation) or affine transformations assume only simple global geometric relationships between images, making them completely incapable of handling such complex, local, time-varying nonlinear distortions. Incorrect registration leads to severe spatial misalignment between visible light texture and thermal imaging temperature information, rendering subsequent fusion analysis meaningless; this is a core failure point of existing technologies.
[0006] 2. Dynamic Tracking Challenges Arising from Rotating Furnaces: The characteristic of rotary kilns is that the furnace body itself is constantly rotating. Therefore, the trajectory of a thermal anomaly point on the furnace shell surface within the fixed field of view of the monitoring equipment is not a simple two-dimensional linear motion, but rather a projection of a helical motion on the surface of a three-dimensional rotating cylinder onto a two-dimensional plane, exhibiting complex nonlinear periodic motion. When the thermal anomaly point rotates to the back of the furnace, it will experience prolonged occlusion. Although existing dual-light systems mention the function of "dynamic tracking of thermal image targets," they typically employ two-dimensional target tracking algorithms designed for conventional scenarios (such as tracking methods based on correlation filtering or deep learning detectors). These algorithms struggle to establish accurate motion models for rotating targets, and are highly susceptible to tracking loss or ID-Switch errors when the target undergoes deformation, rapid movement, or prolonged occlusion.
[0007] The fundamental flaw in existing technologies lies in their failure to address a coupled geometric-dynamic problem. Thermal distortion corrupts the spatial information required for accurate tracking, while the complex dynamics introduced by furnace rotation make learning and compensating for thermal distortion models from time series data exceptionally difficult. Consequently, any subsequent fusion and tracking algorithms are built upon contaminated and unstable data, compromising both reliability and accuracy. Summary of the Invention
[0008] In order to overcome the shortcomings of the existing technology, the purpose of this invention is to provide a method for online temperature measurement target localization of furnaces and kilns based on dual-light fusion. This method solves the problems of accurate registration, robust fusion and stable tracking and prediction of thermal anomaly targets in complex industrial environments such as strong heat wave disturbances and furnace rotation. This enables high-precision, robust and predictive localization and analysis of thermal anomalies on the surface of rotating industrial furnaces and kilns.
[0009] To achieve the above objectives, the technical solution adopted by the present invention is as follows: A method for online temperature measurement target localization in furnaces and kilns based on dual-light fusion includes the following steps; a. Data collection steps: The visible light image sequence and thermal imaging image sequence of the rotary furnace surface are acquired simultaneously through the dual-light data acquisition module; b. Distortion correction and non-rigid registration steps: The visible light image sequence and thermal imaging image sequence are processed using a pre-trained heat wave perturbation sensing and hierarchical registration network H2R-Net to generate a dense non-rigid deformation field. The non-rigid deformation field is then applied to obtain a pair of geometrically corrected and spatially precisely aligned corrected images. c. Fusion and segmentation steps: The corrected image pairs are fused using the pre-trained Environment Adaptive Cross-Modal Fusion and Segmentation Network EACF-Net. This mechanism can sense and adapt to the interference of the on-site environment (such as dust, smoke, and light) on different modal signals in real time, dynamically adjust the fusion strategy, and generate a binary segmentation mask that identifies the thermal anomaly area on the surface of the furnace. d. Tracking and prediction steps: Using a kinematically aware spatiotemporal graph Kalman network (K-STGKN), the state tracking and future state prediction of thermal anomaly targets identified in the segmentation mask are performed.
[0010] In step a, the dual-light data acquisition module is fixedly installed at an appropriate position directly in front of or diagonally above the rotary kiln to ensure complete coverage of the kiln surface area, and the hardware synchronous triggering mechanism ensures that the images of the two modes are strictly aligned in time; during acquisition, the visible light camera and the thermal imager acquire data synchronously at the same frame rate, and the timing consistency is ensured by a unified clock signal. The visible light image sequence contains nonlinear geometric distortion caused by heat wave disturbance; that is, the air density change above the high-temperature area on the surface of the furnace causes uneven refractive index of light, which causes the visible light image to produce non-rigid deformation similar to water ripples, such as distortion and shaking. Thermal imaging images are less affected by this due to their longer wavelength.
[0011] In step b, The H2R-Net is a deep neural network with a twin encoder and a hierarchical decoder structure. The twin encoder independently extracts features from visible light images and thermal imaging images to establish a dual-modal feature representation, providing a foundation for subsequent cross-modal registration. The hierarchical decoder optimizes the non-rigid deformation field at each level through a dual attention correction module (DACB). The dual attention correction module includes a self-attention mechanism for capturing spatial distortion features within a single modality and a cross-attention mechanism for establishing pixel correspondences across modalities. During processing, the visible light image and the thermal imaging image are first input into the twin encoder for feature extraction. Then, a dense non-rigid deformation field is gradually generated through a hierarchical decoder. Finally, the deformation field is applied to the visible light image for geometric transformation, resulting in a pair of corrected images that are geometrically corrected and spatially precisely aligned. The H2R-Net is trained in a self-supervised manner, and its loss function includes: a structural similarity loss to penalize the structural information differences between the corrected image pairs, a temporal consistency loss to ensure smooth transitions of the deformation field between consecutive frames, and a regularization loss to constrain the spatial smoothness of the deformation field itself.
[0012] In step c, the EACF-Net includes a two-stream feature encoder for extracting multi-scale features, an environment encoder for encoding the original visible light image into a conditional vector representing the current environmental state, and a fusion decoder. The dual-stream feature encoder extracts multi-scale features from the visible light image and the thermal imaging image respectively. The environment encoder extracts the global environment condition vector from the original visible light image. The fusion decoder integrates the multi-scale features output by the dual-stream feature encoder in a hierarchical manner and performs cross-modal fusion and upsampling under the guidance of the condition vector output by the environment encoder, finally generating a binary segmentation mask.
[0013] The fusion decoder performs cross-modal feature fusion at each level through the Environment Adaptive Fusion Module (EAFM), which uses the conditional vector to dynamically modulate the fusion weights of visible light and thermal imaging features.
[0014] Specifically, EAFM employs a cross-attention mechanism, using thermal imaging features as the query and visible light features as the key and value. This leverages the semantic localization information from thermal imaging to guide the network in extracting precise boundary details from the visible light features. Furthermore, it adaptively calibrates the fusion weights through conditional vectors to achieve feature complementarity between the two modalities. The fused features are then upsampled and skip connections are used to gradually restore spatial resolution. Finally, the segmentation head outputs pixel-level classification results, generating a binary segmentation mask for the thermal anomaly region.
[0015] The environment-adaptive cross-modal fusion and segmentation network is pre-trained in an end-to-end manner. It uses a dataset containing visible light-thermal imaging paired data and corresponding thermal anomaly label masks, with binary cross-entropy loss as the optimization objective. The parameters of the dual-stream encoder, environment encoder and fusion decoder are updated synchronously through backpropagation until the model converges.
[0016] In step d, the K-STGKN models each thermal anomaly target as a node in a dynamic graph network. The thermal anomaly target is extracted by connected component analysis through the binary segmentation mask obtained in step c. Each target corresponds to a node, and the node attributes include the target's position, appearance, and motion characteristics. Then, the state of each node is tracked, and its state (position, temperature, area, etc.) at the next time t+1 is predicted.
[0017] The tracking and prediction process for each node includes: i. Kalman prediction based on kinematics: Using a Kalman filter with a nonlinear kinematic model based on the known physical parameters of the furnace (including radius and angular velocity) as the state transition function, the node states are predicted a priori to estimate the position changes caused by the rotation of the furnace body. ii. Spatiotemporal graph network correction: The spatiotemporal graph neural network (STGNN) is used to learn the evolution of the non-kinematic states of the thermal anomaly target itself, such as size, shape, and temperature, and outputs the residual correction vector for the prior predicted state; iii. State fusion update: The prior predicted state is added to the residual correction vector to obtain the final posterior predicted state, and then the data is correlated and updated with the detection result at the next time step.
[0018] A furnace online temperature measurement target positioning system based on dual-light fusion includes a dual-light data acquisition module and a central data processing unit; A dual-light data acquisition module is used to simultaneously acquire visible light image sequences and thermal imaging image sequences of the surface of a rotary kiln; a central data processing unit is connected to the dual-light data acquisition module and is configured to execute the method described above.
[0019] The central data processing unit includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the method described above.
[0020] The beneficial effects of this invention are: 1. By actively learning and correcting thermal distortion through H2R-Net, this invention fundamentally eliminates the most significant source of error in existing systems, enabling sub-pixel-level precise registration between visible light and thermal imaging, thereby ensuring the accuracy of thermal anomaly localization.
[0021] 2. EACF-Net's environment adaptive fusion mechanism enables the system to consistently select and rely on more reliable modal information for decision-making in harsh industrial environments such as dust, smoke, and changes in lighting, greatly improving the stability and reliability of thermal anomaly detection.
[0022] 3. K-STGKN not only stably tracks existing thermal anomalies, but more importantly, it can accurately predict key state parameters such as location, temperature, and area of thermal anomalies over a future period based on physical models and learned evolutionary patterns. This predictive capability enables preventative maintenance, helping companies intervene before failures occur and avoid unplanned downtime. Simultaneously, its prediction-based tracking mechanism effectively addresses prolonged obstruction of thermal anomalies caused by furnace rotation, maintaining the continuity of the target's identity.
[0023] 4. The end-to-end deep learning framework proposed in this invention automatically completes distortion correction and accurate registration of dual-light images through a heat wave disturbance perception and hierarchical registration network (H2R-Net), eliminating geometric distortion caused by heat wave disturbance. Then, it uses an environment-adaptive cross-modal fusion and segmentation network (EACF-Net) to intelligently fuse and segment the registered dual-light images, generating accurate masks for thermal anomaly regions in real time. Finally, it uses a kinematically aware spatiotemporal graph Kalman network (K-STGKN) to continuously track and predict the status of thermal anomaly targets, automatically generating early warning level assessments and predictive analysis reports based on their motion trajectories and temperature evolution trends. This achieves full-process automation from raw video data input to final output of operable early warning information and predictive analysis reports, freeing operators from tedious data identification and status judgment, and significantly improving the intelligence level of furnace monitoring. Attached Figure Description
[0024] Figure 1 This is a schematic diagram of the system hardware architecture according to an embodiment of the present invention.
[0025] Figure 2 This is the overall flowchart of the three-stage algorithm proposed in this invention.
[0026] Figure 3 This is a detailed network structure diagram of the heat wave disturbance sensing and hierarchical registration network (H2R-Net) in this invention.
[0027] Figure 4 This is a detailed network structure diagram of the Environment Adaptive Cross-Modal Fusion and Segmentation Network (EACF-Net) in this invention.
[0028] Figure 5 This is a detailed structure and data flow diagram of the kinematic perception spatiotemporal graph Kalman network (K-STGKN) in this invention. Detailed Implementation
[0029] The present invention will now be described in further detail with reference to the accompanying drawings.
[0030] 1. System Overall Architecture and Process Reference Figure 1 and Figure 2This invention proposes an online temperature measurement target positioning system and method for furnaces and kilns based on dual-light fusion; System hardware architecture ( Figure 1 The system mainly includes a dual-light data acquisition module 101 and a central data processing unit 102. The dual-light data acquisition module 101 integrates a visible light (RGB) camera and a long-wave infrared (LWIR) thermal imager, both of which have undergone rigorous internal and external parameter calibration to ensure that the initial field of view is basically aligned.
[0031] This module can be installed in a fixed location or deployed in a portable manner to adapt to different monitoring needs. The central data processing unit 102 can be an edge computing device or a cloud server connected via a wireless communication network such as 5G, and is responsible for receiving the visible light image sequence synchronously transmitted by the dual-light data acquisition module 101. and thermal imaging image sequences And execute the three-stage algorithm proposed in this invention.
[0032] Overall algorithm flow ( Figure 2 )as follows: (1) Input: At each time step t, the dual-light data acquisition module acquires a pair of synchronized but geometrically distorted raw images. and .
[0033] (2) Stage 1 (H2R-Net): will and The input is fed into the heat wave disturbance sensing and hierarchical registration network 201. This network outputs a dense non-rigid deformation field. By applying this deformation field to the input image, a pair of geometrically corrected and spatially precisely aligned images is obtained. .
[0034] (3) Stage 2 (EACF-Net): Align the images and the original visible light image (For environmental perception) The input is fed into an environment-adaptive cross-modal fusion and segmentation network 202. This network outputs a binary segmentation mask. Pixels with a value of 1 represent detected thermal anomaly areas.
[0035] (4) Phase 3 (K-STGKN): For the segmentation mask Post-processing (such as connected component analysis) is performed to instantiate thermal anomaly targets. These targets are then fed into a kinematically aware spatiotemporal graph Kalman network 203. This network is responsible for state tracking, identity maintenance, and predicting the state (position, size, temperature, etc.) of each thermal anomaly target at the next time step t+1.
[0036] (5) Output: The system finally outputs the real-time status, historical trajectory and future status prediction of each thermal anomaly target, and can generate alarm information according to preset thresholds.
[0037] 2. Phase One: Heat-Haze-Aware Hierarchical Registration Network (H2R-Net); This network is a generative adversarial network whose core tasks are to simultaneously achieve two objectives: 1. Blindly correct nonlinear geometric distortions caused by heat wave turbulence; 2. Perform non-rigid registration between visible light and thermal imaging images at high fidelity. It outputs a pair of geometrically corrected and pixel-level precisely aligned visible light and thermal imaging images. The core objective of this phase is to generate a pair of geometrically distortion-free and spatially pixel-aligned visible light and thermal imaging images from a dual-light video stream containing severe heat wave distortion. .
[0038] The innovation of H2R-Net lies in its unification of two originally independent and challenging tasks—distortion correction and non-rigid registration—within a single deep learning framework for end-to-end joint optimization. The network learns to target a single, dense deformation field. The deformation field contains two types of transformation information: First, it is a transformation that reverses the nonlinear spatial distortion caused by the turbulence of heat waves; Secondly, it involves a non-rigid registration transformation that maps the visible light image coordinate system to the thermal imaging image coordinate system. This integrated modeling approach allows the network to utilize the structural consistency between the two modes as a supervision signal to simultaneously solve two coupled problems.
[0039] Network structure: Refer to Figure 3 The structural design of H2R-Net was inspired by the Hierarchical Vision Transformer (H-ViT) in the field of deformable image registration, and it was deeply modified to address the specific problem of blind geometric distortion correction.
[0040] (1) Siamese Encoder 301: The network adopts a weight-sharing Siamese encoder structure to process the input visible light image in parallel. and thermal imaging images Encoders typically consist of a series of convolutional layers used to extract multi-scale feature pyramids from the raw image. These feature maps capture different levels of information, from low-level texture to high-level semantics.
[0041] (2) Hierarchical Deformation Decoder 302: The decoder adopts a top-down hierarchical structure, refining the deformation field level by level. The estimation method starts with the coarsest feature map scale, generates a low-resolution initial deformation field, then upsamples layer by layer, and refines it at each scale using feature maps from the corresponding level of the encoder, ultimately generating a fine deformation field with the same resolution as the original image.
[0042] (3) Dual-Attention Correction Block (DACB) 303: This is the core component of each layer of the H2R-Net decoder. Unlike the standard H-ViT, which mainly focuses on registration, the DAB is designed to decouple and handle common-mode distortion (heat waves) and modal differences. Within a DAB: Self-Attention (401): Self-attention computation is performed on feature patches from a single modality (such as visible light). This allows the network to capture the spatial dependencies of features within that modality, effectively modeling the non-uniform spatial distribution of heat wave distortion.
[0043] Cross-Attention (402): This method uses features from one modality as the query and features from another modality as the key and value for cross-attention computation. This aims to establish the correlation between corresponding regions of the two modalities, which is crucial for achieving accurate registration.
[0044] Fusion and Update: The outputs of self-attention and cross-attention are fused to update the deformation field at the current scale. This dual attention mechanism works in concert, enabling the network to perceive and compensate for internal geometric distortions while searching for cross-modal correspondences.
[0045] Training Method: The training process of H2R-Net is fully self-supervised, requiring no ground truth labeled deformation fields or distortion-free image pairs. This is crucial for industrial applications, as obtaining such labeled data is virtually impossible. Its loss function... It consists of the following three weighted components: in, This represents the total loss during model training. Structural similarity loss is used to penalize differences in structural information between corrected image pairs. This is a time consistency loss used to ensure a smooth transition of deformation fields between consecutive frames. To smooth out the regularization loss, it is used to constrain the spatial smoothness of the deformation field itself. , , These are the weighting coefficients for the corresponding loss functions, used to balance the contribution of each loss to the total loss.
[0046] (1) Structural similarity loss This is the core supervisory signal driving the network's learning. Its basic assumption is that although visible light and thermal imaging differ significantly in radiation intensity, the physical object they observe (the furnace shell) shares a common underlying structure, such as the gaps in the refractory bricks and the outlines of the reinforcing ribs. These structures should exhibit high similarity in the gradient or edge maps of both modes. Therefore, the loss function is calculated using the deformation field. Distorted visible light image and thermal imaging images Structural similarity is used (e.g., by using Normalized Cross-Correlation (NCC) or Gradient Difference (SSIM)). By minimizing this structural difference, the network is forced to learn a deformation field that can eliminate common-mode distortion and align the underlying structure. .
[0047] (2) Time consistency loss This loss term assumes that the pattern of thermal turbulence evolves smoothly over time. Therefore, it penalizes the deformation field between consecutive frames. and Dramatic, discontinuous changes occur. This helps to generate correction results that are more stable over time and more in line with physical laws.
[0048] (3) Smoothing regularization loss The loss term affects the deformation field. The spatial gradient is penalized to encourage smoothness within local regions. This effectively prevents unrealistic artifacts such as tearing and wrinkling from appearing in the corrected image.
[0049] Through this self-supervised learning paradigm, H2R-Net learns an implicit "standard coordinate space." It doesn't simply distort one image into the shape of another; rather, during optimization, to maximize the structural similarity between the two, it is forced to find a transformation that simultaneously "undistorts" and aligns the two distorted images. The final product of this process is the corrected image pair. It can be regarded as the projection of the original scene onto a stable, distortion-free, and normalized coordinate system, laying a solid data foundation for subsequent accurate analysis.
[0050] 3. Phase Two: Environment-Adaptive Cross-Modal Fusion & Segmentation Network (EACF-Net); This network receives aligned image pairs output by H2R-Net and intelligently combines information from the two modalities through an innovative dynamic fusion mechanism. This mechanism can sense and adapt in real time to the interference level of the field environment (such as dust, smoke, and water vapor) on different modal signals, dynamically adjust the fusion strategy, and finally generate a pixel-level segmentation mask that accurately identifies all thermal anomaly regions.
[0051] Receive aligned image pairs from H2R-Net output And generate an accurate binary segmentation mask. It accurately identifies pixels in all thermal anomaly areas.
[0052] The core innovation of EACF-Net lies in the dynamic adaptability of its fusion mechanism. Unlike previous methods that employed fixed and unchanging fusion strategies, EACF-Net can assess the quality of the current monitoring environment in real time and dynamically adjust the weights and interaction methods of visible light and thermal imaging modal information during the fusion process. For example, when smoke and dust severely degrade the quality of visible light images, the network automatically increases its reliance on thermal imaging features; while in well-lit and detailed conditions, it utilizes visible light features more extensively to refine the boundaries of thermal anomalies.
[0053] Network structure: Refer to Figure 4 The overall architecture of EACF-Net is a variant of U-Net with a dual encoder and a single decoder, and its design incorporates the latest advances in the fields of RGB-T salient object detection and multimodal semantic segmentation.
[0054] (1) Dual-Stream Encoder 501: The network contains two parallel encoder branches, which are used to extract aligned visible light images. and thermal imaging images Multiscale features.
[0055] (2) Environment Encoder 502: This is a lightweight convolutional neural network branch whose input is raw, uncorrected visible light images. It compresses the scene information of the entire image into a low-dimensional condition vector (ct) through several layers of convolution and global average pooling operations. This vector is a compact representation of the scene's environmental state, and during training, the network spontaneously learns to map different visual conditions (such as smoke concentration and light intensity) to different regions of this vector space. This design draws inspiration from the idea of conditional fusion.
[0056] (3) FusionDecoder 503: The decoder is responsible for fusing and upsampling the multi-scale features extracted by the encoder, gradually restoring the spatial resolution and generating the final segmentation mask.
[0057] (4) Environment-Adaptive Fusion Module (EAFM) 504: EAFM is a key component of each layer of the decoder, responsible for performing dynamic cross-modal feature fusion. Within an EAFM: Cross-Modal Attention (601): This module employs a cross-attention mechanism to facilitate effective information exchange between modalities. Specifically, thermal imaging features are used as the query, and visible light features as the key and value. The physical meaning of this setup is that it utilizes the strong prior information provided by thermal imaging features regarding "where the hotspots are" to "query" detailed information about the "hotspot boundaries and textures" at the corresponding locations in the visible light feature map. This allows the network to effectively combine the semantic localization capabilities of thermal imaging with the high-frequency detail depiction capabilities of visible light.
[0058] Conditional Modulation (602): The output features of cross-modal attention undergo a modulation process controlled by a conditional vector `ct` before being passed to the next layer. For example, `ct` is mapped to a set of channel-level scaling factors and bias terms (similar to a FiLM layer) through a small fully connected network and then applied to the feature map. This process allows environmental information to directly influence the feature fusion result: when `ct` represents a harsh environment (such as visible light blur), the modulation process will correspondingly suppress feature contributions from the visible light branch, and vice versa.
[0059] Feature aggregation: Modulated cross-modal features are concatenated or added with upsampled features from the previous layer of the decoder, and then the features are finally fused through a convolutional layer and passed to the next level.
[0060] From an information processing perspective, the fusion process of EACF-Net can be understood as a kind of "guided denoising." The thermal imaging modality provides a "noisy signal" with high signal intensity, although the boundaries are blurred, for the location of "salient targets" (i.e., hot spots). The visible light modality provides high-resolution structural and boundary information for the scene, but it is itself "noise" for judging temperature. The cross-attention mechanism in EACF-Net uses the strong signal from thermal imaging to guide the network to accurately extract effective details related to the hot spot boundary from the high-resolution information stream of visible light, thereby "denoising" the "blurred spots" in thermal imaging and obtaining a segmentation result with clear contours and accurate positioning. The conditional vector ct acts as a "signal-to-noise ratio regulator" in this guided process. When the signal-to-noise ratio of the visible light information stream itself is too low (such as when it is obscured by smoke), it reduces the intensity of the guidance to avoid introducing more noise.
[0061] 4. Phase Three: Kinematics-Informed Spatio-Temporal Graph Kalman Network (K-STGKN); This network is a predictive tracking network that abstracts segmented thermal anomaly targets as nodes in a dynamic graph network. This network innovatively integrates a physics-based Kalman filter with a data-driven spatio-temporal graph neural network. The Kalman filter utilizes known kinematic parameters of the furnace rotation as a powerful prior, while the spatio-temporal graph neural network is responsible for learning the complex evolutionary patterns of the thermal anomaly itself (such as growth, decay, and splitting) and their potential interactions.
[0062] Robust, time-series tracking of thermal anomaly targets segmented by EACF-Net is performed, and their future state evolution is predicted.
[0063] The core idea of K-STGKN is to construct a hybrid model that deeply integrates physical information and data-driven approaches to achieve accurate tracking of targets on the surface of a rotating furnace. It cleverly combines the advantages of the classical Kalman filter in handling motions with explicit physical models (i.e., furnace rotation) with the powerful ability of the Spatiotemporal Graph Neural Network (STGNN) in learning complex, nonlinear dynamic relationships (i.e., the self-evolution of thermal anomalies). This design avoids the large amount of training data required by simply using a data-driven model and the potential for predictions that violate physical laws, while also overcoming the limitation that a purely physical model cannot describe the intrinsic changes of the target.
[0064] Network structure and process: Refer to Figure 5The K-STGKN processing flow consists of three steps: graph construction, state prediction, and state update and association, which are performed iteratively. This design combines the application of STGNN in trajectory prediction with the emerging trend of integrating learnable Kalman filters into deep networks.
[0065] (1) Dynamic graph construction (GraphConstruction) 701: In each frame t, the system first performs segmentation on the mask. Connectivity analysis is performed to identify each independent connected component (i.e., a thermal anomaly) as a target. The system constructs a dynamic graph of all tracked targets in the current frame. Each node ∈ This represents a thermal anomaly target i. Node eigenvectors The code encodes the complete state of the target at the current moment, including: geometric state (centroid coordinates (x, y), area A, major and minor axes, orientation angles, and other shape descriptors) and thermodynamic state (maximum temperature). average temperature (Temperature gradient, etc.). (Graph edges) ∈ It can be defined based on the spatial proximity or other physical relationships between nodes.
[0066] (2) Hybrid State Prediction: This is the core of K-STGKN. For each node i in the graph, the prediction of its state from time t to time t+1 is divided into two steps: Step A: Kinematic Kalman Prediction 702: First, a Kalman filter based on a physical model makes an initial prediction of the state of each node. Its state transition equation is as follows: ; in It is the state vector of node i at time t. It is its prior predicted state at time t+1. The key lies in the state transition function. It is not a simple linear model (such as uniform velocity or uniform acceleration), but a model based on the known physical parameters of the furnace. (e.g., furnace body radius R, rotational angular velocity) The nonlinear kinematic function is precisely derived. This function accurately describes the two-dimensional projected trajectory of a point attached to the surface of a rotating cylinder from a fixed camera viewpoint. This step provides an extremely powerful and reliable physical prior for prediction, accurately estimating the positional changes caused by the furnace rotation. wt represents process noise.
[0067] Step B: Spatiotemporal Graph Network Correction (STGNN) 703: Kalman prediction can only describe rigid body motion and cannot capture the evolution of thermal anomalies themselves, such as area growth, shape changes, or temperature increases. These dynamic changes are considered "residuals" of the pure kinematic model. The task of STGNN is to learn this complex spatiotemporal residual. STGNN takes the entire graph Gt as input and aggregates neighborhood information and captures temporal dependencies through multi-layer graph convolution and temporal attention mechanisms (borrowing from advanced models such as STAR). It learns the interactions between thermal anomalies (e.g., the expansion of a hotspot may indicate the appearance of a new hotspot nearby) and the evolutionary trend of each thermal anomaly itself. The output of STGNN is the residual correction vector for each node. .
[0068] Final prediction: Combining the results of the two steps yields the final posterior predicted state. .
[0069] (3) Data Association and State Update 704: When the data of the next frame t+1 arrives and is processed by EACF-Net to obtain the new target detection result. Then, the system uses either the Hungarian algorithm or a greedy algorithm to predict the state. With new detection targets Matching is performed between targets, and the cost function for matching can be distance in the state vector space. For successfully matched targets, their state is updated with new observations (the update step of Kalman filtering); unmatched predicted trajectories are marked as "occluded" and prediction continues for several frames; newly detected unmatched trajectories are initialized as new tracking trajectories. This tracking paradigm based on "joint reasoning of the past and future" allows the tracker to smoothly handle long-term occlusion and maintain the stability of the target ID.
[0070] In K-STGKN, the Kalman filter and STGNN are symbiotic. On one hand, the Kalman filter greatly simplifies the learning task of STGNN by accurately stripping away the main, deterministic rotational motions. STGNN no longer needs to learn complex dynamics from mixed motions, but only needs to focus on learning the purer "residual" dynamics about the evolution of the thermal anomaly itself. This makes model training more efficient and requires less data. On the other hand, the residuals learned by STGNN... This can be viewed as a dynamic, data-driven model of the noise in the Kalman filter process. It enables the entire system to adapt to complex dynamic changes that traditional Kalman filter assumptions cannot cover. This closed-loop mechanism of physics guiding data and data refining physics is key to the high-precision predictive tracking achieved in this invention, embodying a hybrid modeling approach similar to models like GKNet.
[0071] To verify the effectiveness of this symbiotic mechanism, we conducted simulation experiments: in a dynamic rotational scene superimposed with complex thermal anomalies, the prediction accuracy (RMSE) of K-STGKN significantly outperformed that of the purely physical Kalman filter and the purely data-driven STGNN model. This result intuitively demonstrates that the Kalman filter, by stripping away the rotational motion of the main body, provides the STGNN with cleaner learning samples, enabling it to focus more on capturing residual anomalies; conversely, the dynamic residuals learned by the STGNN compensate for the shortcomings of the Kalman filter's fixed noise assumption. The closed loop formed by these two factors jointly contributes to the model's superior predictive performance.
Claims
1. A method for online temperature measurement target localization in furnaces and kilns based on dual-light fusion, characterized in that, Includes the following steps; a. Simultaneously acquire visible light image sequences and thermal imaging image sequences of the rotary furnace surface through a dual-light data acquisition module; b. Using a pre-trained heat wave disturbance sensing and hierarchical registration network H2R-Net, the visible light image sequence and the thermal imaging image sequence are processed to generate a dense non-rigid deformation field, and the non-rigid deformation field is applied to obtain a pair of corrected images that are geometrically corrected and spatially precisely aligned. c. Using the pre-trained Environment Adaptive Cross-Modal Fusion and Segmentation Network EACF-Net, the corrected image pairs are fused to sense and adapt to the interference level of the field environment on different modal signals in real time, dynamically adjust the fusion strategy, and generate a binary segmentation mask that identifies the thermal anomaly area on the surface of the furnace. d. Using a kinematically aware spatiotemporal graph Kalman network (K-STGKN), the thermal anomaly targets identified in the segmentation mask are used for state tracking and future state prediction.
2. The method for online temperature measurement target positioning of furnaces and kilns based on dual-light fusion according to claim 1, characterized in that, In step a, the dual-light data acquisition module is fixedly installed at an appropriate position directly in front of or diagonally above the rotary kiln to ensure complete coverage of the kiln surface area, and the hardware synchronous triggering mechanism ensures that the images of the two modes are strictly aligned in time; during acquisition, the visible light camera and the thermal imager acquire data synchronously at the same frame rate, and the timing consistency is ensured by a unified clock signal. The visible light image sequence contains nonlinear geometric distortions caused by heat wave disturbances.
3. The method for online temperature measurement target positioning of furnaces and kilns based on dual-light fusion according to claim 2, characterized in that, In step b, the H2R-Net is a deep neural network with a twin encoder and a hierarchical decoder structure. The twin encoder independently extracts features from visible light images and thermal imaging images to establish a dual-modal feature representation, providing a basis for subsequent cross-modal registration. The hierarchical decoder optimizes the non-rigid deformation field at each level through a dual attention correction module (DACB). The dual attention correction module includes a self-attention mechanism for capturing spatial distortion features within a single modality and a cross-attention mechanism for establishing pixel correspondences across modalities. During processing, the visible light image and the thermal imaging image are first input into the twin encoder for feature extraction. Then, a dense non-rigid deformation field is gradually generated through a hierarchical decoder. Finally, the deformation field is applied to the visible light image for geometric transformation to obtain a pair of corrected images that are geometrically corrected and spatially aligned. The H2R-Net is trained using a self-supervised approach, and its loss function includes: structural similarity loss to penalize differences in structural information between corrected image pairs, temporal consistency loss to ensure smooth transition of deformation fields between consecutive frames, and regularization loss to constrain the spatial smoothness of the deformation field itself.
4. The method for online temperature measurement target positioning of furnaces and kilns based on dual-light fusion according to claim 3, characterized in that, In step c, the EACF-Net includes a two-stream feature encoder for extracting multi-scale features, an environment encoder for encoding the original visible light image into a conditional vector representing the current environmental state, and a fusion decoder. The dual-stream feature encoder extracts multi-scale features from the visible light image and the thermal imaging image respectively. The environment encoder extracts the global environment condition vector from the original visible light image. The fusion decoder integrates the multi-scale features output by the dual-stream feature encoder in a hierarchical manner and performs cross-modal fusion and upsampling under the guidance of the condition vector output by the environment encoder, finally generating a binary segmentation mask.
5. The method for online temperature measurement target positioning of furnaces and kilns based on dual-light fusion according to claim 4, characterized in that, The fusion decoder performs cross-modal feature fusion at each level through the Environment Adaptive Fusion Module (EAFM), which uses the conditional vector to dynamically modulate the fusion weights of visible light and thermal imaging features. Specifically, EAFM employs a cross-attention mechanism, using thermal imaging features as queries and visible light features as keys and values. It leverages the semantic localization information of thermal imaging to guide the network in extracting precise boundary details from visible light features. Furthermore, it adaptively calibrates the fusion weights through conditional vectors to achieve feature complementarity between the two modalities. The fused features are then upsampled and skip connections are used to gradually restore spatial resolution. Finally, the segmentation head outputs pixel-level classification results to generate a binary segmentation mask for the thermal anomaly region.
6. The method for online temperature measurement target positioning of furnaces and kilns based on dual-light fusion according to claim 5, characterized in that, The environment-adaptive cross-modal fusion and segmentation network is pre-trained in an end-to-end manner. It uses a dataset containing visible light-thermal imaging paired data and corresponding thermal anomaly label masks, with binary cross-entropy loss as the optimization objective. The parameters of the dual-stream encoder, environment encoder and fusion decoder are updated synchronously through backpropagation until the model converges.
7. The method for online temperature measurement target positioning of furnaces and kilns based on dual-light fusion according to claim 6, characterized in that, In step d, the K-STGKN models each thermal anomaly target as a node in a dynamic graph network. The thermal anomaly target is extracted by connected component analysis through the binary segmentation mask obtained in step c. Each target corresponds to a node, and the node attributes include the target's position, appearance, and motion characteristics. Then, the state of each node is tracked, and its state at the next time t+1 is predicted.
8. The method for online temperature measurement target positioning of furnaces and kilns based on dual-light fusion according to claim 1, characterized in that, The tracking and prediction process for each node includes: i. Kalman prediction based on kinematics: Using a Kalman filter with a nonlinear kinematic model based on the known physical parameters of the furnace as the state transition function, the node states are predicted a priori to estimate the position changes caused by the rotation of the furnace body. ii. Spatiotemporal graph network correction: The spatiotemporal graph neural network is used to learn the evolution law of the non-kinematic state of the size, shape and temperature of the thermal anomaly target itself, and outputs the residual correction vector of the prior predicted state; iii. State fusion update: The prior predicted state is added to the residual correction vector to obtain the final posterior predicted state, and then the data is correlated and updated with the detection result at the next time step.
9. A furnace online temperature measurement target positioning system based on dual-light fusion, characterized in that, Includes a dual-light data acquisition module and a central data processing unit; A dual-light data acquisition module is used to simultaneously acquire visible light image sequences and thermal imaging image sequences of the surface of a rotary kiln; a central data processing unit is connected to the dual-light data acquisition module, and the central data processing unit is configured to perform the method as described in any one of claims 1-8.
10. A furnace online temperature measurement target positioning system based on dual-light fusion according to claim 9, characterized in that, The central data processing unit includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the method described in any one of claims 1-8.