Dynamic temperature field reconstruction method based on multi-modal embedding and modal routing
By employing multimodal embedding and modal routing methods, the problems of rapid temperature evolution and insufficient texture in dynamic thermal field reconstruction are solved, achieving stable dynamic temperature field reconstruction and new perspective rendering, thus improving reconstruction quality and stability.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HANGZHOU DIANZI UNIV
- Filing Date
- 2026-06-23
- Publication Date
- 2026-07-24
AI Technical Summary
Existing technologies struggle to stably characterize the rapid temperature evolution over time in dynamic thermal field reconstruction. Furthermore, weak thermal imaging textures and insufficient salient features lead to unstable optimization, a lack of geometric consistency and thermal field realism under multimodal constraints, and a lack of a unified evaluation basis.
By employing a multimodal embedding and modal routing approach, the time-varying characteristics of scene structure, visible light appearance, and temperature field are jointly represented through a shared geometric representation framework. Camera parameters and poses are estimated using visible light image sequences, and a shared geometric multimodal Gaussian representation is constructed. The reconstruction process is optimized through an adaptive modal routing mechanism and a progressive dynamic deformation model, thereby achieving four-dimensional reconstruction of the dynamic temperature field.
It improves the stability of dynamic thermal field reconstruction and the rendering quality from a new perspective, and can accurately represent dynamic temperature distribution and scene structure while maintaining geometric consistency, thus enhancing the ability to represent dynamic thermal processes.
Smart Images

Figure CN122454068A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of computer vision and image processing technology, and relates to infrared thermal imaging and dynamic scene reconstruction and new perspective rendering, specifically to a dynamic temperature field reconstruction method based on multimodal embedding and modal routing. Background Technology
[0002] In the fields of infrared thermal imaging and 3D reconstruction, the spatial representation of temperature fields is an important and challenging technical direction. Infrared thermal imaging obtains thermal images by detecting infrared radiation information from the surface of an object, which are used to characterize temperature-related distribution features. Due to its irreplaceable role under low illumination, weak texture, or partial occlusion conditions, it is widely used in scenarios such as military monitoring, industrial inspection, building assessment, search and rescue, and medical auxiliary diagnosis. However, traditional thermal imaging results are usually presented in the form of two-dimensional heat maps with a fixed viewpoint, which makes it difficult to establish an accurate correspondence with the three-dimensional geometry of the scene, limiting the ability to perform refined spatial analysis of heat source locations, spatial scale of anomalous temperature regions, and the relationship between temperature distribution and structure.
[0003] To enhance the interpretability and spatial analysis capabilities of thermal information, existing technologies are gradually incorporating multi-view observation and 3D reconstruction concepts. These technologies attempt to recover scene geometry from multi-view thermal imaging images or multimodal images combining visible light and thermal imaging, and to express thermal distribution in 3D space. This supports more flexible viewing angles and more comprehensive spatial measurement analysis. The rapid development of novel perspective synthesis and 3D reconstruction methods, such as Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS), provides effective means for high-fidelity new perspective rendering and explicit 3D representation, and drives the evolution of thermal field reconstruction from traditional geometry pipelines towards differentiable rendering and learning-based modeling.
[0004] Against this backdrop, existing work has incorporated thermal modes into NeRF or 3DGS frameworks to achieve 3D thermal field reconstruction and thermal imaging rendering in static scenes. It's important to note that temperature distribution exhibits significant time dependence. Unlike visible light appearance, which is typically relatively stable under steady lighting conditions, thermal fields are constantly evolving due to the continuous influence of heat conduction, convection, radiation, and changes in external heat sources. Furthermore, in practical applications such as open-flame heating and high-temperature welding, temperature changes are often frequent and significant, making the static thermal field assumption difficult to uphold. However, most existing methods rely on static temperature distribution as a premise, and static 3D thermal field reconstruction alone cannot adequately characterize transient processes such as heat propagation, hotspot generation and decay, nor can it meet the application requirements for tracking and analyzing dynamic thermal phenomena. This makes extending thermal field modeling from three dimensions to four-dimensional spatiotemporal reconstruction that includes a time dimension a necessary development direction.
[0005] However, dynamic thermal field reconstruction faces several technical challenges. First, thermal imaging images often lack texture and significant features, making them prone to optimization instability. Second, in dynamic scenes, the visible light mode is better suited to providing stable geometric structures and motion cues, while the thermal mode focuses more on characterizing temperature distribution and its evolution over time. Therefore, although the visible light and thermal modes observe the same physical scene, their imaging mechanisms and temporal variation patterns are not consistent. Maintaining cross-modal geometric consistency and temporal continuity while simultaneously characterizing the dynamic changes of different modes is a key issue in multimodal dynamic thermal field reconstruction. Furthermore, existing publicly available data resources are relatively limited, especially the lack of benchmark data that can cover high-frequency temperature changes and provide simultaneous multi-view visible light and thermal imaging observations for four-dimensional evaluation. This, to some extent, restricts the research and validation of dynamic thermal field reconstruction methods.
[0006] In summary, while existing technologies have made some progress in static thermal field reconstruction, they still face challenges in reconstructing dynamic scenes with rapid and significant temperature changes. These challenges include difficulty in stably representing temperature evolution over time, difficulty in balancing geometric consistency and thermal field realism under multimodal constraints, and a lack of a unified evaluation basis. Therefore, how to achieve spatiotemporal temperature field reconstruction of dynamic scenes under multimodal observation conditions, and improve the rendering quality and optimization stability from new perspectives, remains a pressing technical problem to be solved in this field. Summary of the Invention
[0007] To address the shortcomings of existing technologies, this invention proposes a dynamic temperature field reconstruction method based on multimodal embedding and modal routing. This method uses simultaneously acquired visible light and thermal imaging images as inputs, jointly characterizing the scene structure, visible light appearance, and time-varying characteristics of the temperature field within a unified representation framework of shared geometry. Furthermore, it improves reconstruction convergence stability and new perspective rendering quality through multimodal collaborative constraints and spatiotemporal consistency optimization. This addresses the problems of existing technologies, such as reliance on static temperature distribution assumptions in thermal field reconstruction, difficulty in depicting rapid temperature evolution over time, and instability in dynamic optimization due to weak thermal imaging textures and insignificant features.
[0008] The dynamic temperature field reconstruction method based on multimodal embedding and modal routing specifically includes the following steps:
[0009] Step 1: Obtain visible light image sequences and thermal imaging image sequences of the same dynamic scene at consecutive moments, and complete time synchronization and basic preprocessing.
[0010] Step 2: Estimate camera parameters and camera pose using the visible light image sequence, and generate a sparse point cloud for initializing a 3D Gaussian ensemble based on the estimated camera pose. Use the estimated camera parameters, camera pose, and sparse point cloud as a geometric prior, and use this geometric prior as a shared reference for joint modeling of the visible light mode and the thermal imaging mode.
[0011] Step 3: Construct a multimodal Gaussian representation with shared geometry, setting spatial geometric parameters shared between the two modes and unique unimodal properties for each Gaussian.
[0012] The spatial geometric parameters include position, rotation, and scale. Visible light modal attributes include visible light appearance features, visible light opacity, and visible light modal embedding. Thermal imaging modal attributes include thermal modal appearance features, thermal modal opacity, and thermal modal embedding.
[0013] Simultaneously, a structural role routing variable is introduced for each Gaussian, including shared structure, visible light mode-specific structure, and thermal imaging mode-specific structure. The shared structure represents a Gaussian structure that participates in both visible light and thermal imaging mode modeling; the visible light mode-specific structure represents a Gaussian structure that participates only in visible light mode modeling; and the thermal imaging mode-specific structure represents a Gaussian structure that participates only in thermal imaging mode modeling.
[0014] Step 4: Construct a dynamic deformation model including an embedding layer, a visible light branch, and a thermal imaging branch to predict the residual update of Gaussian parameters relative to the gauge space under different time states.
[0015] The embedding layer is used to generate temporal embeddings and modal embeddings. The visible light branch, conditioned on temporal embeddings and visible light modal embeddings, uses a visible light branch deformation function to predict dynamic residual updates of geometric parameters, visible light appearance features, and visible light opacity. Subsequently, based on the update results of the visible light branch, the thermal imaging branch, conditioned on temporal embeddings and thermal modal embeddings, uses a thermal imaging branch deformation function to further predict dynamic residual updates of geometric parameters, thermal modal appearance features, and thermal modal opacity, thereby enhancing the dynamic modeling capability related to temperature distribution and mitigating cross-modal conflicts.
[0016] Step 5: Introduce an adaptive modal routing mechanism to generate a modal participation mask based on the Gaussian structure role routing variables, thereby enabling selective participation occlusion contribution control of Gaussian under different modalities. During training, a differentiable discrete selection strategy is adopted, and early over-specialization is suppressed through prior constraints to improve convergence stability while maintaining geometric consistency.
[0017] Step 6: Optimize the above parameters by using multimodal reconstruction loss, routing prior loss and spatiotemporal regularization term; output new perspective rendering results of visible light and thermal imaging at any time step to realize four-dimensional reconstruction and visualization of dynamic temperature field.
[0018] The present invention has the following beneficial effects:
[0019] 1. To address the problem that existing temperature field reconstructions are mostly based on static temperature distribution assumptions and are difficult to characterize the rapid evolution of temperature over time, this method proposes a four-dimensional temperature field reconstruction framework for dynamic scenes. Under a unified representation, it simultaneously models the time-varying processes of visible light appearance and temperature distribution, and supports new perspective rendering and visualization under different time states, thereby improving the ability to characterize dynamic thermal processes.
[0020] 2. To address the issues of weak texture and insufficient salient features in thermal imaging, which can lead to unstable optimization and insufficient geometric constraints when relying solely on thermal modes for modeling in dynamic scenes, this method proposes a multimodal modeling strategy that combines visible light and thermal imaging. First, visible light observations are used to complete pose estimation and geometric initialization, establishing a stable shared geometry. Then, corresponding Gaussian geometric parameters are shared across different modes, and visible light attributes, thermal mode attributes, and modal embeddings are set separately to achieve a unified representation of "geometric sharing and attribute separation." This design effectively enhances cross-modal geometric consistency, achieves multimodal information complementarity, and improves the convergence stability and rendering consistency of dynamic reconstruction.
[0021] 3. To address the conflict in coupled modeling caused by different dynamic change patterns in multimodal modes, this method designs a progressive, modal-separated dynamic deformation strategy: first, the visible light branch predicts the time-varying residual updates of shared geometric and visible light properties to stabilize geometric motion; then, the thermal branch further predicts the residual updates of geometric and thermal modal properties based on the update results of the previous stage, thereby enhancing the dynamic representation capability related to temperature distribution, realizing progressive decoupling modeling under shared geometric constraints, reducing cross-modal interference, and improving the dynamic thermal field expression capability.
[0022] 4. To address the problem that the same representation in multimodal joint optimization cannot simultaneously take into account both shared and modality-specific regions, and is prone to conflict and instability, this method proposes an adaptive modal routing mechanism. By learning shared / modality-specific structural roles and generating modality participation masks, Gaussian selective participation contribution control is achieved under different modalities. At the same time, combined with prior constraints on the routing distribution, early over-specialization is suppressed and stable learning of structural roles is promoted. Thus, while maintaining geometric consistency, cross-modal competition is alleviated, training is promoted to converge stably, and reconstruction quality is improved. Attached Figure Description
[0023] Figure 1 This is a flowchart of a dynamic temperature field reconstruction method based on multimodal embedding and modal routing.
[0024] Figure 2 Comparison of reconstruction results using the same scene, same perspective, but different methods.
[0025] Figure 3 Qualitative comparison of thermal imaging modal reconstruction results for different scenarios and methods.
[0026] Figure 4 A qualitative comparison of RGB modal and thermal imaging modal reconstruction results for the same scene using different methods. Detailed Implementation
[0027] The present invention will be further explained below with reference to the accompanying drawings;
[0028] The overall process of the dynamic temperature field reconstruction method based on multimodal embedding and modal routing is as follows: Figure 1 As shown, the specific steps include:
[0029] Step 1: Obtain the multimodal dynamic sequence and complete synchronization and preprocessing.
[0030] The system acquires a sequence of visible light images and thermal imaging images of the same dynamic scene at consecutive time points, ensuring synchronization between the two modalities as much as possible during acquisition. Then, the sequences of the two modalities are time-aligned. If there is a start-end frame offset, a one-to-one temporal index is established by aligning the start and end frames, unifying the frame rate, or extracting frames. Next, the images of the two modalities are cropped, downsampled, and matched to resolution, ensuring that the visible light images and thermal imaging images maintain spatial consistency, thereby guaranteeing that subsequent rendering and loss calculations are performed in a unified pixel coordinate system.
[0031] After preprocessing, time-aligned visible light image sequences and thermal imaging image sequences are obtained, along with the correspondence between the two sequences and a set of time indices.
[0032] Step 2: Initialize and obtain shared geometric priors.
[0033] Camera parameters and pose are estimated using visible light image sequences, and a sparse point cloud is generated based on the estimated camera pose to initialize a 3D Gaussian ensemble. The estimated camera parameters, camera pose, and sparse point cloud are used as geometric priors, which are then used as a shared reference for joint modeling of the visible light mode and the thermal imaging mode. This addresses the problem that thermal imaging images have weak texture and insufficient salient features, making direct pose estimation based on thermal imaging images prone to instability.
[0034] Step 3: Construct a multimodal Gaussian representation of the shared geometry and initialize the parameters.
[0035] The 3D Gaussian set is initialized based on the geometric priors obtained in step two, and differentiable rendering is achieved using Gaussian splashing. The camera parameters and camera pose in the geometric priors are used to determine the projection relationship of the 3D Gaussian set onto image planes from different viewpoints, and the sparse point cloud is used to initialize the position of the 3D Gaussian set. Each Gaussian set is configured with shared spatial geometric parameters between the two modalities, as well as unique unimodal attributes.
[0036] The spatial geometric parameters include position. Rotation and scale These are used to characterize a unified spatial structure within the same dynamic scene. Visible light modal properties include visible light appearance features. Visible light opacity and visible light mode embedding Thermal imaging modal attributes are used to characterize the appearance, visibility, and modal state of visible light modes. These attributes include thermal modal appearance features. Thermal mode opacity and thermal mode embedding It is used to characterize the temperature-dependent response, visibility, and modal state under thermal imaging modes.
[0037] For the A three-dimensional Gaussian, which is uniformly represented as:
[0038]
[0039] in, Representing visible light modes and thermal imaging modes, , and They represent the first A Gaussian in mode The appearance features, opacity, and modal embedding are included.
[0040] During rendering, a 3D Gaussian is projected onto the image plane of the target viewpoint to form a 2D Gaussian, and the transparency is accumulated and blended according to the depth order to obtain the rendering result of the target pixel. :
[0041]
[0042] in, Indicates the target pixel Contributing Gaussian sets Indicates the first The rendering properties of a Gaussian in the corresponding modality. Indicates the first Gaussian contribution to opacity.
[0043] At the same time, a structural role routing variable is introduced for each Gaussian. This is used to generate participation masks for different modalities during subsequent training, thereby forming a unified multimodal Gaussian representation with "geometric sharing, attribute separation, and controllable participation." The structural role routing variable... Including shared structures Visible light mode-specific structure and thermal imaging modal-specific structures Among them, the shared structure represents a Gaussian structure that participates in both visible light mode modeling and thermal imaging mode modeling; the visible light mode-specific structure represents a Gaussian structure that participates only in visible light mode modeling; and the thermal imaging mode-specific structure represents a Gaussian structure that participates only in thermal imaging mode modeling.
[0044] Step 4: Introduce time embedding and establish a dynamic deformation model.
[0045] An embedding-based dynamic deformation model is established, parameterized using a residual approach, predicting the increment of the relative normalized Gaussian parameters, and then applying it additively to the normalized parameters to characterize the continuous evolution of the dynamic scene over time, thereby improving training stability and temporal continuity. The dynamic deformation model includes an embedding layer, a visible light branch, and a thermal imaging branch.
[0046] The embedding layer provides each timestamp Introducing time embedding It is used to encode the global time state; modal embedding is introduced for each Gaussian in each modality. Used to encode the Gaussian mode The current state.
[0047] To avoid coupling conflicts caused by differences in the dynamic change patterns of the two modes, a two-stage progressive, mode-separated update strategy is adopted. The first stage involves the visible light branch. In input embedding conditions Residual updates for predicting shared geometric and visible light properties The first stage captures key geometrical motions and visible light appearance changes. The second stage, based on the geometry determined in the first stage, utilizes the thermal imaging branch... Input conditions Next, the residual update of the predicted geometric properties and thermal imaging modal properties will be performed. This enhances the dynamic modeling capabilities related to temperature distribution and mitigates mutual interference in multimodal joint optimization.
[0048] By employing the two-stage strategy described above, progressive decoupled modeling can be achieved under shared geometric constraints, balancing geometric consistency with the ability to dynamically represent thermal fields.
[0049] Step 5: Introduce an adaptive modal routing mechanism and generate a modal participation mask.
[0050] To achieve modal selectivity and suppress cross-modal conflicts under the premise of shared geometry, an adaptive modal routing mechanism is set up. For each Gaussian, a structural role routing variable is learned. ,
[0051] Since the structural role is a discrete variable, Gumbel-Softmax is used during training to perform a continuous approximation on the discrete choices, resulting in the... The soft-assignment probability of a Gaussian for the k-th structural role :
[0052]
[0053] Where k represents the structural role index, Indicates Gumbel noise, This represents the temperature parameter. The forward rendering stage uses hard allocation to determine the structural roles. :
[0054]
[0055] In the backpropagation phase, a Straight-Through estimator is used to propagate gradients via soft assignment, achieving compatibility between discrete structure selection and end-to-end training. Participation masks for visible light and thermal imaging modes are generated based on the structure's role. , :
[0056]
[0057]
[0058] By combining the participation mask with the opacity of the corresponding modality, the effective rendering contribution under different modalities can be obtained:
[0059]
[0060] in, Indicates the first The opacity of a Gaussian at time t and mode m This indicates the corresponding effective opacity.
[0061] The adaptive modal routing mechanism only adjusts the participation and occlusion contributions of Gaussians in different modalities, without modifying shared geometric parameters; geometric evolution and dynamic deformation are still defined in the shared canonical space, thus maintaining spatial consistency and temporal continuity. A phased, progressive strategy is adopted during training: in the initial stage, Gaussians are preferentially treated as shared structures to stabilize geometric and dynamic deformation learning; in the intermediate stage, routing variables are gradually activated, allowing Gaussians to adaptively differentiate into shared and modality-specific structures; in subsequent stages, the structure role assignment results are fixed to improve training convergence and evaluation stability.
[0062] Step 6: Construct a joint loss function and optimize its parameters to output the reconstruction results from a new perspective at any time step.
[0063] After completing shared geometry modeling, dynamic deformation modeling, and modal routing modeling, a joint loss function is constructed. The system performs unified optimization on shared geometric parameters, modal-specific properties, routing variables, embedding vectors, and two-stage dynamic deformation network parameters, and outputs new perspective reconstruction results for visible light and thermal modes at any time step.
[0064]
[0065] in, Indicates the multimodal reconstruction loss. Indicates the prior loss of the route. This represents the modal embedding local smoothing regularization term. This indicates a time-embedded smoothing regularization term. , and These represent the weight coefficients of the corresponding loss terms. During the training phase, gradient descent is used to jointly optimize these parameters.
[0066] The multimodal reconstruction loss This is used to constrain the consistency between the rendered results and actual observations. For visible light modes, pixel-level... Loss can be applied, and structural similarity loss can be optionally added to better constrain high-frequency structures and sensing quality; for thermal imaging modalities, pixel-level loss is used. The loss is adapted to its relatively smooth imaging characteristics:
[0067]
[0068] in, This represents the pixel-level reconstruction loss of the visible light mode. This represents the structural similarity loss of visible light modes. This represents the pixel-level reconstruction loss of the thermal imaging modality. and This represents the weighting coefficients used to balance the magnitudes of losses between different modes, through... and The settings can adjust the relative contributions of the two modal reconstruction terms to the total loss, preventing one modality from dominating the joint optimization and thus improving the stability of multimodal joint training.
[0069] In one embodiment, a multimodal reconstruction regularization term is introduced, which balances the training contributions of the two modes in joint optimization by setting weight coefficients for the visible light mode reconstruction loss and the thermal imaging mode reconstruction loss.
[0070]
[0071]
[0072] The routing prior loss This is used to constrain the distribution of soft routes in the optimization objective, making the route learning process more stable.
[0073]
[0074] in, This represents the average soft assignment probability of the k-th structural role across all Gaussians. Indicates the prior distribution of the preset route. This represents the Gaussian quantity. By using the route prior loss, the route distribution can be made more inclined to share structural roles in the early stage of training, thereby avoiding premature structural specialization; as the training process progresses, the route prior distribution is gradually adjusted to promote the stable differentiation between shared structures and modality-specific structures.
[0075] The modal embedding local smoothing regularization term and the temporal embedding smoothing regularization term are used to enhance the stability of dynamic modeling:
[0076]
[0077]
[0078] in, This represents the Gaussian set participating in the local smoothing constraint. Indicates the first The nearest neighbor set of Gaussians. and They represent the first The and the first Modal embedding of a Gaussian under mode m Indicates time Temporal embedding. ||||2 represents the L2 norm.
[0079] After training, inputting arbitrary timestamps and arbitrary new viewpoint camera parameters, the visible light branch first generates residual updates of shared geometric and visible light attributes; then, the thermal mode branch further generates residual updates of geometric and thermal mode attributes based on the results of the previous stage; then, according to the mode participation mask obtained in step five, the participation rendering contribution of each Gaussian in different modes is controlled; finally, the visible light new viewpoint result and the thermal mode new viewpoint result at that time step are rendered and output respectively through Gaussian splashing, thereby realizing the four-dimensional reconstruction and visualization of the dynamic temperature field.
[0080] To demonstrate the improved reconstruction performance of the proposed dynamic temperature field reconstruction method based on multimodal embedding and modal routing in dynamic scenes, quantitative comparative experiments were conducted. D-3DGS (Dynamic 3D Gaussian Splash), 4DGaussians (Four-Dimensional Gaussian Splash), E-D3DGS (Deformable 3D Gaussian Splash Based on Gaussian Embedding), ThermalGaussian (T-GS, Thermal Imaging 3D Gaussian Splash), and MMone (Multimodal Representation in a Single Scene) were selected as comparison methods. Reconstruction results in visible light and thermal imaging modes were evaluated on the DynamicRGBT-Scenes dataset. Peak Signal-to-Noise Ratio (PSNR), Structural Similarity (SSIM), and Perceptual Difference (LPIPS) were used to assess the results. The results are shown in Tables 1 and 2.
[0081] Table 1
[0082]
[0083] Table 2
[0084]
[0085] Wherein, "this method*" indicates that in the joint loss function The method of introducing a multimodal reconstruction regularization term is proposed.
[0086] As shown in Tables 1 and 2, our proposed method achieves superior reconstruction results in both the visible light and thermal imaging modes, with overall performance improved compared to existing comparative methods. This indicates that our method can more effectively characterize appearance and temperature changes in dynamic scenes while maintaining shared geometric consistency. Furthermore, the addition of a multimodal reconstruction regularization term further improves the reconstruction performance in the thermal imaging mode, demonstrating that this regularization term helps balance the training contributions of the visible light and thermal imaging modes in joint optimization and enhances the quality of dynamic temperature field reconstruction.
[0087] To further illustrate the reconstruction quality of the novel perspective method of this invention from a visual perspective, reconstruction results from different methods under the same scene, time step, and observation perspective in the test set were qualitatively compared. The results are as follows: Figure 2 As shown. By Figure 2 It can be seen that this method can preserve scene structure and local details well in the visible light mode, and can recover the thermal response area more accurately in the thermal imaging mode. Furthermore, the thermal imaging reconstruction results of different methods in multiple test scenarios are compared, and the results are as follows... Figure 3 As shown. By Figure 3It can be seen that this method can more stably recover the heat source region, temperature boundary, and local high-temperature details, indicating that it has a better characterization ability for dynamic temperature distribution. To illustrate the dynamic consistency in the time dimension, the reconstruction results of the visible light mode and thermal imaging mode at multiple time steps in the same dynamic scene are compared. The results are as follows: Figure 4 As shown. By Figure 4 It can be seen that the proposed method can maintain the consistency of visible light structure and the continuity of thermal imaging temperature distribution at multiple time steps, indicating that the proposed method has good temporal stability.
[0088] To illustrate the impact of shared Gaussian representation, multimodal embedding, adaptive modal routing mechanism, and multimodal reconstruction regularization on reconstruction performance, ablation experiments were conducted, with a stepwise comparison using the single-modal version of E-D3DGS as a baseline. The results are shown in Table 3.
[0089] Table 3
[0090]
[0091] The results show that introducing shared Gaussian representation improves the reconstruction performance of the visible light mode, while the performance of the thermal mode fluctuates slightly, indicating that while shared geometric representation enhances cross-modal structural consistency, it may still introduce some modal coupling effects. Further introduction of multimodal embedding and modal routing improves the reconstruction performance of both modes, demonstrating that collaborative modeling of shared geometry and modal-specific properties can effectively mitigate cross-modal interference and enhance the representation of dynamic temperature fields. Adding a multimodal reconstruction regularization term further improves the reconstruction performance of the thermal imaging mode, indicating that this regularization term helps balance the training contributions of the visible light mode and the thermal imaging mode in joint optimization and improves the quality of dynamic temperature field reconstruction.
[0092] The above description, in conjunction with specific and preferred embodiments, provides a further detailed explanation of the present invention. It should not be construed that the specific implementation of the present invention is limited to these descriptions. For those skilled in the art, various substitutions or modifications can be made to these described embodiments without departing from the concept of the present invention, and all such substitutions or modifications should be considered within the scope of protection of the present invention.
[0093] The parts of this invention not described in detail are well-known to those skilled in the art.
Claims
1. A dynamic temperature field reconstruction method based on multimodal embedding and modal routing, which utilizes visible light image sequences and thermal imaging image sequences of the same dynamic scene at consecutive time steps to output new perspective rendering results of visible light and thermal imaging at arbitrary time steps, and reconstructs the dynamic temperature field, characterized by: A multimodal Gaussian representation with shared geometry is constructed, and each Gaussian is assigned spatial geometric parameters shared between two modes, as well as mode-specific properties; at the same time, a structural role routing variable is introduced for each Gaussian to distinguish the modeling process in which the Gaussian structure participates. A dynamic deformation model including an embedding layer, a visible light branch, and a thermal imaging branch is constructed. A two-stage progressive, modal separation update strategy is adopted to sequentially generate dynamic residual updates of visible light modal-specific attributes and thermal imaging modal-specific attributes. A differentiable discrete selection strategy is adopted to generate modal participation masks based on Gaussian structure role routing variables; Constructing a joint loss function The model performs unified optimization on shared geometric parameters, modal-specific attributes, routing variables, embedding vectors, and two-stage dynamic deformation network parameters. Based on arbitrary input timestamps and arbitrary new viewpoint camera parameters, the model outputs new viewpoint rendering results for visible light and thermal imaging.
2. The dynamic temperature field reconstruction method based on multimodal embedding and modal routing as described in claim 1, characterized in that: The acquired visible light image sequence and thermal imaging image sequence are time-aligned by aligning start and end frames, unifying frame rate, or extracting frames to establish a one-to-one time index. Then, the images of the two modalities are cropped, downsampled, and matched for resolution to ensure that the visible light image and thermal imaging image are consistent in spatial size.
3. The dynamic temperature field reconstruction method based on multimodal embedding and modal routing as described in claim 1, characterized in that: Camera parameters and camera pose are estimated using visible light image sequences, and sparse point clouds are generated based on the estimated camera pose to initialize a 3D Gaussian set. The estimated camera parameters, camera pose, and sparse point clouds are used as geometric priors, and these geometric priors are used as a shared reference for joint modeling of visible light and thermal imaging modes.
4. The dynamic temperature field reconstruction method based on multimodal embedding and modal routing as described in claim 1, characterized in that: The spatial geometric parameters include Gaussian position, rotation, and scale; the visible light modal-specific attributes include visible light appearance features, visible light opacity, and visible light modal embedding; and the thermal imaging modal-specific attributes include thermal modal appearance features, thermal modal opacity, and thermal modal embedding.
5. The dynamic temperature field reconstruction method based on multimodal embedding and modal routing as described in claim 1, characterized in that: The structural role routing variables include shared structures, visible light mode-specific structures, and thermal imaging mode-specific structures; wherein, a shared structure represents a Gaussian structure that participates in both visible light mode and thermal imaging mode modeling; a visible light mode-specific structure represents a Gaussian structure that participates only in visible light mode modeling; and a thermal imaging mode-specific structure represents a Gaussian structure that participates only in thermal imaging mode modeling.
6. The dynamic temperature field reconstruction method based on multimodal embedding and modal routing as described in claim 1, characterized in that: The embedding layer of the dynamic deformation model outputs a temporal embedding that encodes the global time state and a modal embedding that encodes the state under a specified Gaussian mode. In the first stage of the update, the visible light branch predicts dynamic residual updates of geometric parameters, visible light appearance features, and visible light opacity, conditioned on temporal embedding and visible light modal embedding. Subsequently, the thermal imaging branch, based on the update results of the visible light branch, further predicts dynamic residual updates of geometric parameters, thermal modal appearance features, and thermal modal opacity, conditioned on temporal embedding and thermal modal embedding.
7. The dynamic temperature field reconstruction method based on multimodal embedding and modal routing as described in claim 1, characterized in that: The joint loss function for: ; in, Indicates the multimodal reconstruction loss. Indicates the prior loss of the route. This represents the modal embedding local smoothing regularization term. This indicates a time-embedded smoothing regularization term. , and These represent the weight coefficients of the corresponding loss terms; during the training phase, the gradient descent method is used to jointly optimize the above parameters.
8. The dynamic temperature field reconstruction method based on multimodal embedding and modal routing as described in claim 7, characterized in that: The multimodal reconstruction loss for: ; in, This represents the pixel-level reconstruction loss of the visible light mode. The structure similarity loss represents the visible light mode, and the pixel-level reconstruction loss represents the thermal imaging mode. and This represents the weighting coefficients used to balance the magnitudes of loss between different modes.
9. The dynamic temperature field reconstruction method based on multimodal embedding and modal routing as described in claim 8, characterized in that: A multimodal reconstruction regularization term is introduced, and weighting coefficients are set according to the visible light modal reconstruction loss and the thermal imaging modal reconstruction loss: ; 。 10. A computer-readable storage medium having a computer program stored thereon, which, when executed in a computer, causes the computer to perform the method of any one of claims 1 to 9.