A thermal infrared multi-view super-resolution three-dimensional reconstruction and new view synthesis method
Patent Information
- Application Number
- CN202611074453.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-20
- Publication Date
- 2026-08-18
AI Technical Summary
[0006]本发明的目的在于提供一种热红外多视角超分辨率三维重建与新视角合成方法,以解决现有技术中热红外场景三维重建容易受非均匀响应噪声、固定模式噪声以及热扩散模糊影响,从而导致几何恢复不稳定、细节表达不足以及高分辨率新视角生成质量不佳的问题
[0037] 1. This invention explicitly introduces thermal infrared imaging degradation modeling into the three-dimensional Gaussian sputtering training closed loop, which can decouple the real thermal radiation structure from the pseudo high-frequency texture caused by hardware, thereby reducing the adverse effects of pseudo high-frequency gradients on three-dimensional geometry optimization and reducing floating artifacts and geometry collapse phenomena.
Smart Images

Figure CN122597676A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of thermal infrared imaging, 3D visual reconstruction, and computer graphics, and particularly to a method for thermal infrared multi-view super-resolution 3D reconstruction and novel viewpoint synthesis. Specifically, this invention utilizes a 3D Gaussian sputtering scene representation and introduces a physically constrained infrared imaging calibration module, an intensity-condition appearance residual adaptation module, and a frequency-domain aware curriculum learning strategy during the training phase. This allows for the recovery of a 3D scene representation with high geometric and thermal radiation consistency from low-resolution thermal infrared multi-view observations, and the direct output of a high-resolution thermal infrared image of the target viewpoint during the inference phase. Background Technology
[0002] Thermal infrared imaging reflects the temperature distribution information in a scene by receiving long-wave infrared radiation emitted by the target object. It still has good imaging stability in environments such as night, backlight, smoke, and poor visibility, so it is widely used in scenarios such as night inspection, industrial inspection, security monitoring, disaster relief, emergency monitoring, and autonomous driving.
[0003] However, limited by the manufacturing cost, pixel size, diffraction limit, and imaging hardware performance of thermal infrared sensors, existing thermal infrared cameras typically only provide observations with low resolution, limited contrast, and scarce texture details. These characteristics make it significantly difficult to use thermal infrared images for 3D modeling and novel perspective synthesis, especially when recovering fine structural boundaries and stable geometries; low-resolution observations often fail to provide sufficient and effective supervision.
[0004] In recent years, 3D representation methods such as neural radiation fields and 3D Gaussian sputtering have achieved good results in visible light scene reconstruction and new perspective synthesis tasks. However, when these methods are directly transferred to thermal infrared scenes, problems such as increased floating artifacts, unstable geometric structures, and inconsistent results across different perspectives often occur. The fundamental reason is that thermal infrared imaging has a different degradation mechanism than visible light imaging: on the one hand, the real thermal radiation boundary usually exhibits low-frequency, smooth, and blurred characteristics under the effect of thermal diffusion; on the other hand, the non-uniform response noise and fixed-pattern noise in thermal infrared sensors often manifest as structured high-frequency textures. This phenomenon of "low-frequency real structure and high-frequency hardware artifacts" easily causes existing 3D reconstruction networks to misinterpret sensor noise as scene details, resulting in incorrect geometric augmentation and appearance fitting.
[0005] Furthermore, while existing 2D thermal infrared super-resolution methods can improve edge sharpness and local contrast in a single frame image, their lack of multi-view geometric consistency constraints means that directly using them as a preprocessing step for 3D reconstruction can easily introduce inconsistent pseudo-details across different viewpoints, thus affecting the stability of 3D structure restoration and the quality of new viewpoint synthesis. Therefore, how to balance physical imaging degradation modeling, multi-view geometric consistency constraints, and high-resolution detail restoration under low-resolution thermal infrared multi-view conditions remains a pressing technical problem to be solved in current technologies. Summary of the Invention
[0006] The purpose of this invention is to provide a method for thermal infrared multi-view super-resolution 3D reconstruction and new view synthesis, in order to solve the problems in the prior art where thermal infrared scene 3D reconstruction is easily affected by non-uniform response noise, fixed pattern noise and thermal diffusion blur, resulting in unstable geometric recovery, insufficient detail expression and poor quality of high-resolution new view generation.
[0007] To achieve the above objectives, the present invention adopts the following technical solution:
[0008] A method for thermal infrared multi-view super-resolution 3D reconstruction and novel view synthesis includes the following steps:
[0009] Acquire a sequence of low-resolution thermal infrared multi-view images of the scene and the corresponding camera parameters for each viewpoint. Construct a 3D Gaussian sputtering scene representation and render it at the target output resolution to obtain an ideal thermal infrared image.
[0010] The ideal thermal infrared image is mapped to a predicted image aligned with the observation domain through a differentiable imaging chain that includes a physical constraint infrared imaging calibration module and an intensity condition appearance residual adaptation module.
[0011] Joint optimization training is performed using the structural priors provided by the teacher network and the corresponding confidence masks.
[0012] After training, based on the camera parameters corresponding to the target viewpoint, the optimized 3D Gaussian sputtering scene representation is used to render at the target output resolution, outputting a high-resolution thermal infrared new perspective image from the target viewpoint.
[0013] In one possible implementation, the three-dimensional Gaussian sputtering scene representation is composed of multiple anisotropic three-dimensional Gaussian elements, each of which includes at least a three-dimensional center, a covariance matrix, an opacity, and a single-channel thermal radiation intensity parameter; the covariance matrix is represented in a parameterized form using rotation matrix and scale matrix decomposition.
[0014] In one possible implementation, when rendering the 3D Gaussian sputtering scene representation, each 3D Gaussian primitive is projected onto a 2D image plane, and pixel intensity is calculated using a depth-ordered alpha synthesis method. Specifically, Gaussian primitives contributing to the same pixel are arranged in depth order, and transmittance is accumulated during the alpha synthesis process. The pixel intensities of all sampled points form an ideal thermal infrared image corresponding to the target output resolution.
[0015] In one possible implementation, the physically constrained infrared imaging calibration module includes a learnable low-resolution spatial gain grid, a learnable low-resolution spatial bias grid, a parameterized radial vignetting function, and a viewpoint-dependent monotonic tone mapping function. The learnable low-resolution spatial gain grid and the learnable low-resolution spatial bias grid are upsampled to the resolution of the currently rendered image and then constrained by a hyperbolic tangent function to form spatial gain fields. and spatial bias field Furthermore, this is combined with the vignetting correction factor calculated from the parameterized radial vignetting function. A low-frequency calibration is performed on the ideal thermal infrared image, and a monotonic transformation is performed using a viewpoint-related monotonic tone mapping function to obtain a calibration prediction image aligned with the current viewpoint observation domain.
[0016] In one possible implementation, the intensity-conditional appearance residual adaptation module employs a pixel-wise multilayer perceptron network, including an input layer, two hidden layers, and an output layer, with the hidden layers connected by a ReLU activation function. For each spatial location in the input image, the rendering intensity at that location is used as input, and local multiplicative gain residuals and additive bias residuals are output, with range constraints applied to the residual amplitudes. Local affine residual modulation is performed on the input image based on the local multiplicative gain residuals and additive bias residuals to obtain the appearance adaptation prediction image.
[0017] In one possible implementation, the joint optimization training employs a frequency domain-aware curriculum learning strategy, comprising the following steps: generating a teacher super-resolution image using a pre-trained super-resolution teacher network, upsampling a low-resolution thermal infrared multi-view image, constructing a confidence mask based on the pixel difference and gradient difference between the two, and applying weighted constraints to the high-frequency components of the current prediction image and the teacher super-resolution image within a reliable region indicated by the confidence mask.
[0018] In one possible implementation, within the confidence region indicated by the confidence mask, a high-pass filter operator is used to extract the high-frequency components of the predicted image and the teacher super-resolution image, and a high-frequency residual loss is constructed to weight the high-frequency components.
[0019] In one possible implementation, the 3D Gaussian sputtering scene representation, the physically constrained infrared imaging calibration module, and the intensity-condition appearance residual adaptation module are jointly optimized based on a joint training objective, wherein the joint training objective is expressed as:
[0020] ,
[0021] Among them, pixel reconstruction loss Structural similarity loss is used to constrain pixel consistency between the predicted image and the observed image. Used to constrain the structural similarity between the predicted image and the observed image. For high-frequency residual loss; the regularization term Used to constrain the spatial gain field, spatial bias field, and learnable vignetting parameters in the radial vignetting function of the physical constraint infrared imaging calibration module. The modulation amplitude used for the constraint strength condition appearance residual adaptation module.
[0022] In one possible implementation, the regularization term of the physically constrained infrared imaging calibration module includes a smoothness constraint on the spatial gain field and the spatial bias field, as well as an amplitude constraint on the learnable vignetting parameter in the radial vignetting function, to suppress overfitting of the physically constrained infrared imaging calibration module to sensor non-uniform response, fixed pattern noise, or local abnormal intensity changes.
[0023] The regularization term of the intensity condition appearance residual adaptation module includes amplitude constraints on the local multiplicative gain residual and the local additive bias residual, which are used to limit the local affine residual modulation intensity of the intensity condition appearance residual adaptation module to avoid masking the three-dimensional geometric reconstruction error through excessive appearance modulation.
[0024] In one possible implementation, the weights of the structural similarity loss and the high-frequency residual loss are both scheduled using a cosine-incrementing method and dynamically adjusted according to the training iteration stage.
[0025] Secondly, embodiments of this application provide a thermal infrared multi-view super-resolution 3D reconstruction and new view synthesis device, comprising the following modules:
[0026] Data acquisition module: used to acquire low-resolution thermal infrared multi-view image sequences of the scene and the camera parameters corresponding to each view.
[0027] Rendering module: Constructs a three-dimensional Gaussian sputtering scene representation based on the low-resolution thermal infrared multi-view image sequence, and renders it at the target output resolution to obtain an ideal thermal infrared image.
[0028] Mapping module: Through a differentiable imaging chain including a physical constraint infrared imaging calibration module and an intensity condition appearance residual adaptation module, the ideal thermal infrared image is mapped into a predicted image aligned with the observation domain.
[0029] Training module: Using the structural priors and corresponding confidence masks provided by the teacher network, the three-dimensional Gaussian sputtering scene representation, the physical constraint infrared imaging calibration module, and the intensity condition appearance residual adaptation module are jointly optimized and trained.
[0030] Reconstruction module: After training, based on the camera parameters corresponding to the target viewpoint, the optimized 3D Gaussian sputtering scene representation is used to render at the target output resolution, outputting a high-resolution thermal infrared new viewpoint image under the target viewpoint.
[0031] Thirdly, embodiments of this application provide an electronic device, including a processor and a memory;
[0032] The memory is used to store computer programs.
[0033] When the processor executes the program stored in the memory, it implements any of the three-dimensional reconstruction and new perspective synthesis methods described in this application.
[0034] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program, which, when executed by a processor, implements any of the three-dimensional reconstruction and new perspective synthesis methods described in this application.
[0035] Fifthly, embodiments of this application provide a computer program product containing instructions that, when run on a computer, cause the computer to execute any of the three-dimensional reconstruction and new perspective synthesis methods described in this application.
[0036] Compared with the prior art, the present invention has at least the following beneficial effects:
[0037] 1. This invention explicitly introduces thermal infrared imaging degradation modeling into the three-dimensional Gaussian sputtering training closed loop, which can decouple the real thermal radiation structure from the pseudo high-frequency texture caused by hardware, thereby reducing the adverse effects of pseudo high-frequency gradients on three-dimensional geometry optimization and reducing floating artifacts and geometry collapse phenomena.
[0038] 2. This invention, through the strength condition appearance residual adaptation module, enhances the localized constrained intensity of the weak temperature difference region while maintaining the overall thermal radiation trend as basically consistent. This is beneficial to improving the effective monitoring intensity in the low-texture thermal infrared region and enhancing the detail recovery capability.
[0039] 3. This invention constructs a robust optimization path of "low-frequency structure first, then high-frequency details" through a frequency domain-aware course learning strategy, which can improve the convergence stability of the training process of low-resolution thermal infrared multi-view scenes and the boundary clarity of new view synthesis.
[0040] 4. This invention can directly recover high-resolution new perspective results from low-resolution thermal infrared multi-view input, avoiding the accumulation of errors caused by false details due to inconsistent perspectives in the traditional "two-dimensional super-resolution followed by three-dimensional reconstruction" process.
[0041] 5. This invention is applicable to tasks such as thermal infrared 3D reconstruction and new perspective synthesis in nighttime sensing, infrared inspection, industrial temperature measurement, disaster monitoring, target observation and emergency scenarios, and has good application value. Attached Figure Description
[0042] Figure 1 This is an overall flowchart of an embodiment of the present invention. Detailed Implementation
[0043] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments, but the scope of protection of the present invention is not limited to the following embodiments.
[0044] A method for thermal infrared multi-view super-resolution 3D reconstruction and novel view synthesis includes the following steps:
[0045] S1. Obtain a low-resolution thermal infrared multi-view image sequence of the scene and the camera parameters corresponding to each view.
[0046] S2. Construct a three-dimensional Gaussian sputtering scene representation based on the low-resolution thermal infrared multi-view image sequence, and render it at the target output resolution to obtain an ideal thermal infrared image.
[0047] S3. The ideal thermal infrared image is mapped into a predicted image aligned with the observation domain through a differentiable imaging chain that includes a physical constraint infrared imaging calibration module and an intensity condition appearance residual adaptation module.
[0048] S4. Using the structural priors and corresponding confidence masks provided by the teacher network, the three-dimensional Gaussian sputtering scene representation, the physical constraint infrared imaging calibration module, and the intensity condition appearance residual adaptation module are jointly optimized and trained.
[0049] S5. After training, based on the camera parameters corresponding to the target viewpoint, the optimized 3D Gaussian sputtering scene representation is used to render at the target output resolution, outputting a high-resolution thermal infrared new viewpoint image under the target viewpoint.
[0050] In one possible implementation, the three-dimensional Gaussian sputtering scene representation is composed of multiple anisotropic three-dimensional Gaussian elements, each of which includes at least a three-dimensional center, a covariance matrix, an opacity, and a single-channel thermal radiation intensity parameter; the covariance matrix is represented in a parameterized form by rotation matrix and scale matrix decomposition, such that the covariance matrix remains symmetric positive definite during the joint optimization training in step S4.
[0051] In one possible implementation, during the rendering of the 3D Gaussian sputtering scene representation in step S2, each 3D Gaussian primitive is projected onto a 2D image plane, and pixel intensity is calculated using a depth-ordered α-composition method. Specifically, Gaussian primitives contributing to the same pixel are arranged in depth order, and transmittance is accumulated during the α-composition process. On a high-density sampling grid corresponding to the target upsampling ratio, the 3D Gaussian sputtering scene representation is queried and rasterized. Based on the 2D projection footprint response, thermal radiation intensity, transparency, and depth order of the Gaussian primitives at each sampling point, the contribution of each Gaussian primitive to the pixel intensity of the corresponding sampling point is calculated. During the calculation, the effective contribution weight of the current Gaussian primitive is determined by accumulating the transmittance of Gaussian primitives preceding it, and the effective contributions of each Gaussian primitive are accumulated to determine the pixel intensity of each sampling point. The pixel intensities of all sampling points are arranged according to their spatial positions in the high-density sampling grid to form an ideal thermal infrared image corresponding to the target output resolution, rather than generating a low-resolution image first and then performing image spatial interpolation and upscaling.
[0052] In one possible implementation, the physically constrained infrared imaging calibration module includes a learnable low-resolution spatial gain grid, a learnable low-resolution spatial bias grid, a parameterized radial vignetting function, and a viewpoint-dependent monotonic tone mapping function. The learnable low-resolution spatial gain grid and the learnable low-resolution spatial bias grid are learnable calibration parameters in the physically constrained infrared imaging calibration module, used to generate a spatial calibration field consistent with the ideal thermal infrared image size based on the current rendered image resolution. Specifically, the spatial gain grid and the spatial bias grid are upsampled (using bilinear interpolation) to the current rendered image resolution, and then constrained by a hyperbolic tangent function to form spatial gain fields. and spatial bias field This is used to characterize the multiplicative and additive biases introduced by the non-uniform response and fixed-mode noise of thermal infrared sensors. First, based on the spatial gain field... Spatial bias field and the vignetting correction factor calculated from the parameterized radial vignetting function. A low-frequency calibration is performed on the ideal thermal infrared image to obtain a calibration intermediate image; then, the calibration intermediate image is monotonically transformed by the viewpoint-related monotonic tone mapping function to obtain a calibration prediction image aligned with the current viewpoint observation domain, which serves as the input to the intensity condition appearance residual adaptation module.
[0053] In one possible implementation, the radial halosing function V(r) is expressed as:
[0054] ,
[0055] in, This is the normalized radial distance from a pixel to the principal point of the image. and These are learnable vignetting parameters. Derived from the radial vignetting function. The two-dimensional vignetting correction field is obtained by taking values at each pixel location and used as the vignetting correction factor. .
[0056] In one possible implementation, the physically constrained infrared imaging calibration module first bases its calibration on the spatial gain field. Spatial bias field and the aforementioned vignetting correction factor For ideal thermal infrared images Perform low-frequency calibration to obtain intermediate calibration images. Then, the intermediate calibration image is processed. Numerical stabilization is performed, and the calibrated prediction image is obtained through monotonic intensity mapping corresponding to the viewpoint k. It satisfies:
[0057] ,
[0058] ,
[0059] in, This indicates the calibration of intermediate images. Indicates the value after numerical stabilization. , This represents the Sigmoid function. This represents the inverse function of the Sigmoid function. For an ideal thermal infrared image, This is the vignetting correction factor, used to characterize the radial photometric attenuation caused by the lens. and These are the mapping parameters corresponding to the viewpoint k. Furthermore, during the joint optimization training process, regularization constraints are set on the learnable calibration parameters of the physically constrained infrared imaging calibration module to suppress pixel-by-pixel overfitting.
[0060] In one possible implementation, the intensity condition appearance residual adaptation module employs a pixel-by-pixel multilayer perceptron network. The pixel-by-pixel multilayer perceptron network includes an input layer, two hidden layers, and an output layer, with each of the two hidden layers having 64 channels. The hidden layers are connected using the ReLU activation function. The input image for the intensity-conditional appearance residual adaptation module... Each spatial location in To the rendering intensity at that spatial location As input, the system outputs local multiplicative gain residuals and additive bias residuals, and imposes range constraints on the residual amplitudes to prevent geometric errors from being masked by excessive appearance degrees of freedom. The intensity-conditional appearance residual adaptation module adapts the input image based on the local multiplicative gain residuals and additive bias residuals. Local affine residual modulation is performed to obtain the appearance adaptation prediction image. .
[0061] In one possible implementation, the appearance adaptation prediction image output by the intensity condition appearance residual adaptation module satisfy:
[0062] ,
[0063] in, Indicates the predicted image at pixel location The intensity value at that location, This represents the input image input to the strength condition appearance residual adaptation module. The input image represents At pixel position The intensity value at the location; the input image The calibration prediction image output by the physically constrained infrared imaging calibration module ,Right now Correspondingly That is, the calibration prediction image. At pixel position The intensity value at that location; Represents the multiplicative gain residual. Indicates additive bias residuals, and A preset scaling factor is used to scale the multiplicative gain residual. and the additive bias residual The modulation amplitude is limited so that the intensity condition appearance residual adaptation module performs local intensity modulation on the input image within a limited range.
[0064] In one possible implementation, the joint optimization training employs a frequency-domain-aware curriculum learning strategy to leverage the structural priors provided by the teacher network, guiding the model to gradually transition from low-frequency geometric skeleton recovery to high-frequency detail recovery. The frequency-domain-aware curriculum learning strategy includes the following process: generating teacher super-resolution images using a pre-trained super-resolution teacher network. The pre-trained super-resolution teacher network is preferably a pre-trained SwinIR super-resolution network; the low-resolution thermal infrared multi-view image is upsampled to obtain... ,according to and The confidence mask is constructed from the pixel differences and gradient differences between the pixels. Weighted constraints are applied to the high-frequency components of the current predicted image and the teacher super-resolution image within the reliable region indicated by the confidence mask.
[0065] In one possible implementation, the confidence mask satisfy:
[0066] ,
[0067] in, and To control the hyperparameter of confidence decay sensitivity, This represents the spatial gradient operator.
[0068] In one possible implementation, within the confidence region indicated by the confidence mask, a high-pass filter operator is used to extract the high-frequency components of the predicted image and the teacher super-resolution image, and a high-frequency residual loss is constructed. To apply weighted constraints to the high-frequency components; the high-frequency residual loss satisfy:
[0069] ,
[0070] in, This represents a high-pass filter operator, which is implemented through the difference operation between the image and its average pooling low-pass result. That is, the input image is subtracted from the smoothed image obtained by average pooling to obtain the high-frequency response corresponding to edge details and local structural changes, so that the high-frequency residual loss focuses on constraining edge details and local structural changes. This represents the total number of pixels involved in the loss calculation.
[0071] In one possible implementation, the weight corresponding to the high-frequency residual loss A scheduling strategy that dynamically changes with the number of training iterations is adopted. A smaller value is taken in the early stage of training, and it is gradually increased to the preset maximum value in the middle and later stages of training, so that the model training process follows the course learning path of first restoring low-frequency structure and then restoring high-frequency details.
[0072] In one possible implementation, the joint training objective includes at least pixel reconstruction loss. Structural similarity loss High-frequency residual loss The regularization term of the physical constraint infrared imaging calibration module and the regularization term of the strength condition appearance residual adaptation module. The joint training objective is expressed as:
[0073] ,
[0074] in, Used to constrain pixel consistency between the predicted image and the observed image. The regularization term is used to constrain the structural similarity between the predicted image and the observed image. In the joint optimization training phase of step S4, the joint training objective is added to constrain the spatial gain field, spatial bias field, and learnable vignetting parameters in the radial vignetting function of the physically constrained infrared imaging calibration module, so as to improve the smoothness and stability of the physically constrained infrared imaging calibration module. The modulation amplitude used for the constraint strength condition appearance residual adaptation module.
[0075] Furthermore, the regularization term of the physically constrained infrared imaging calibration module satisfy:
[0076] ,
[0077] in, For spatial gain field, For spatial bias field, and These represent the spatial gain field and spatial bias field at the pixel location, respectively. The value at that location, and For the learnable vignetting parameters in the radial vignetting function, For physical calibration regularization weights.
[0078] in, Represents the total variation regularization term; for any two-dimensional calibration field The total variation regularization term Defined as:
[0079] ,
[0080] in, For spatial gain field or spatial bias field , This indicates the number of positions involved in the vertical adjacent difference calculation. This indicates the number of positions involved in the horizontal adjacent difference calculation.
[0081] The regularization term of the strength condition appearance residual adaptation module satisfy:
[0082] ,
[0083] For ease of explanation, the multiplicative modulation term in the aforementioned formula for calculating the appearance adaptation prediction image is denoted as... The additive bias term is denoted as ;Specifically, , .in, and These represent the pixel-by-pixel multilayer perceptron network at the pixel position. The local multiplicative gain residual and the local additive bias residual at the output, This indicates the number of pixel positions involved in the regularization calculation. This is the appearance regularization weight. The regularization term... By constraining the multiplicative modulation term Approaching 1, additive bias term The amplitude is set close to 0 to limit the modulation amplitude of the intensity condition appearance residual adaptation module, thus preventing it from masking geometric errors through excessive appearance freedom.
[0084] The weights of the structural similarity loss and the high-frequency residual loss are both scheduled using a cosine-incrementing method and dynamically adjusted according to the training iteration stages. In the early stage of training, pixel reconstruction and structure preservation are the main focus, while in the middle and later stages of training, the intensity of high-frequency detail supervision is gradually increased. This ensures that the model training process follows a frequency domain awareness learning path that gradually transitions from low-frequency geometric skeleton recovery to high-frequency detail recovery.
[0085] Since the three-dimensional Gaussian sputtering scene is represented as a resolution-independent continuous scene representation, it is possible to query and render on the sampling grid corresponding to the target output resolution, thereby generating a new thermal infrared perspective image with the corresponding resolution according to different target viewpoints.
[0086] Example 1
[0087] This embodiment provides a method for thermal infrared multi-view super-resolution 3D reconstruction and new viewpoint synthesis, the overall process of which is as follows: Figure 1 As shown, it includes the following steps:
[0088] Step 1: Data Acquisition and Preprocessing.
[0089] A low-resolution thermal infrared multi-view image sequence of the target scene is acquired, and camera parameters corresponding to each viewpoint are obtained. The camera parameters include at least a camera intrinsic matrix and a camera extrinsic matrix. In one embodiment, the thermal infrared image is a single-channel grayscale thermal radiation image. To facilitate training, radiometric normalization is performed on the original image, and the image size is standardized. To construct the low-resolution input required for training or testing, the original thermal infrared image can also be downsampled according to a preset magnification to obtain a low-resolution thermal infrared multi-view observation image sequence. The processed data is then divided into a training set and a test set.
[0090] Step 2: Construct a 3D Gaussian scene representation.
[0091] The scene is represented as N anisotropic 3D Gaussian elements. Each 3D Gaussian element includes at least a 3D center. Covariance matrix Opacity and single-channel thermal radiation intensity parameters To ensure that the covariance matrix is optimizable and remains positive definite, parameterization is performed using a rotation matrix and scale vector decomposition method, i.e.:
[0092] , ,
[0093] in, Represents the rotation matrix. Represents the scale matrix.
[0094] Step 3: Ideal thermal infrared image rendering.
[0095] For any training viewpoint v, based on the camera parameters corresponding to the viewpoint, a three-dimensional Gaussian primitive is projected onto the image plane, and an ideal thermal infrared image of that viewpoint is obtained through differentiable rasterization. In this invention, in order to directly learn high-resolution scene representations, the... Preferably, rendering is performed on the target high-resolution sampling grid, rather than rendering a low-resolution image first and then performing image space super-resolution processing. Since the 3D Gaussian scene representation itself is continuous, image generation for arbitrary scales can be achieved by increasing the sampling grid density.
[0096] Step 4: Physical constraint infrared imaging calibration.
[0097] Thermal infrared sensors often suffer from imaging biases such as non-uniform response, fixed-pattern noise, optical vignetting, and automatic gain control. To mitigate the adverse effects of these biases on geometric optimization, this embodiment incorporates a physically constrained infrared imaging calibration module to calibrate the ideal thermal infrared image. Perform forward calibration.
[0098] Specifically, firstly, a learnable low-resolution spatial gain grid is set up. and learnable low-resolution spatial bias mesh After being upsampled to the image resolution via bilinear interpolation, spatial gain fields are formed respectively. and spatial bias field Then construct the radial vignetting function. When viewpoint-dependent automatic gain control is present, a monotone tone mapping parameter corresponding to the viewpoint k is introduced. and Ultimately, we can obtain:
[0099] ,
[0100] ,
[0101] in, Indicates the value after numerical stabilization. . This represents the calibration prediction image output by the physically constrained infrared imaging calibration module, i.e., the ideal thermal infrared image. Passing through the space gain field in sequence Spatial bias field vignetting correction factor Low-frequency calibration and viewing angle The image obtained after corresponding monotonic intensity mapping; Aligned with the thermal infrared observation domain of the current viewpoint, it serves as the input image for subsequent intensity-condition appearance residual adaptation modules. Through this process, the ideal thermal radiation result obtained from 3D Gaussian field rendering can be mapped to an image domain that more closely approximates the actual sensor observation distribution, thereby suppressing the propagation of erroneous gradients caused by hardware artifacts.
[0102] Step 5: Modulation of appearance residual under intensity conditions.
[0103] Because thermal infrared images typically suffer from narrow dynamic range and weak local contrast, relying solely on raw observations for supervision can easily lead to a lack of effective geometric constraints in flat areas. Therefore, this embodiment includes an intensity-condition appearance residual adaptation module to adjust the calibration prediction image. Perform localized constrained reinforcement.
[0104] In one embodiment, the intensity condition appearance residual adaptation module employs a lightweight multilayer perceptron structure, with each pixel position... Input intensity As conditions, the local multiplicative gain residual and additive bias residual are output, and their variation range is limited by the hyperbolic tangent function and a preset scaling factor, ultimately yielding:
[0105] .
[0106] in, Indicates the predicted image at pixel location The intensity value at that location, Indicates the pixel position of the input image. The intensity value at the location; the input image The preferred calibration image is the output of the physically constrained infrared imaging calibration module. Right now Correspondingly That is, the calibration prediction image. At pixel position The intensity value at that location; Represents the multiplicative gain residual. Indicates additive bias residuals, and The preset scaling factor is used.
[0107] By employing the aforementioned local affine modulation, the discriminability of local temperature differences in weakly textured regions can be improved without significantly altering the overall thermal radiation distribution, thus providing more comprehensive supervisory information for subsequent geometric learning.
[0108] Step 6: Joint training in frequency domain awareness course.
[0109] To avoid disrupting the 3D geometric convergence process by directly introducing unstable high-frequency supervision in the early stages of training, this embodiment further incorporates a frequency-domain-aware curriculum learning strategy. This strategy preferably employs a pre-trained super-resolution teacher network, pre-generating teacher super-resolution images for each training perspective. Simultaneously, low-resolution thermal infrared multi-view images were upsampled to obtain... .
[0110] Furthermore, according to and The confidence mask is constructed from the pixel differences and gradient differences between the pixels. Specifically, it can be expressed as:
[0111] ,
[0112] in, and To control the hyperparameter of confidence decay sensitivity, This represents the spatial gradient operator. When the teacher's image is more consistent with the original observation in terms of low-frequency structure, the confidence level at the corresponding location is higher, indicating that the high-frequency details at that location are more reliable.
[0113] Based on this, the enhanced prediction image With teacher image Applying weighted residual constraints to the high-frequency components can be expressed as:
[0114] ,
[0115] in, This represents a high-pass filter operator, which is implemented through the difference operation between the image and its average pooling low-pass result. That is, the input image is subtracted from the smoothed image obtained by average pooling to obtain the high-frequency response corresponding to edge details and local structural changes, so that the high-frequency residual loss focuses on constraining edge details and local structural changes. This represents the total number of pixels involved in the loss calculation.
[0116] To ensure the training process follows a stable optimization path from coarse to fine, high-frequency loss weights are optimized. A cosine-incremental scheduling method is adopted, which prioritizes optimizing low-frequency structural consistency and overall radiation reconstruction in the preset initial iteration stage, and gradually increases the proportion of high-frequency detail constraints in subsequent iteration stages according to the preset scheduling strategy.
[0117] The overall joint training objective can be expressed as:
[0118] .
[0119] in, For pixel reconstruction loss, For structural similarity loss The weights are determined using a cosine-increasing scheduling method. For structural similarity loss, This is a regularization term for the physical constraint infrared imaging calibration module. This is a regularization term for the strength condition appearance residual adaptation module.
[0120] Step 7: Output a high-resolution new perspective from the target viewpoint.
[0121] After training, the camera parameters corresponding to the target viewpoint are input, and the trained 3D Gaussian sputtering scene representation is directly rendered at the target output resolution to obtain a high-resolution thermal infrared new viewpoint image from the target viewpoint. The physically constrained infrared imaging calibration module is mainly used for observation domain alignment during the training phase and can be bypassed during the inference output phase. The intensity condition appearance residual adaptation module can be selectively used to perform local intensity modulation on the rendered result according to output requirements. Because explicit absorption, decoupling, and adaptation of thermal infrared imaging degradation have been completed during the training phase, the obtained target viewpoint result has good geometric boundary integrity, thermal radiation consistency, and detail representation capability.
[0122] Example 2
[0123] In another embodiment, the spatial gain grid and spatial bias grid in the physically constrained infrared imaging calibration module can be set with different resolutions to adapt to thermal infrared sensors with different hardware specifications; the number of hidden layers and channels in the intensity condition appearance residual adaptation module can also be adjusted according to computing resources; the teacher network in the frequency domain awareness curriculum learning strategy can adopt other super-resolution networks that can provide relatively stable high-frequency priors. Those skilled in the art should understand that the above-mentioned parameter adjustments or equivalent module replacements based on the core ideas of physical degradation modeling, local constrained enhancement, and curriculum-based high-frequency supervision of this invention are all implementation scenarios contemplated and covered by this invention.
[0124] Example 3
[0125] To further verify the effectiveness of the method provided in this invention in low-resolution thermal infrared multi-view sequence super-resolution 3D reconstruction and new view synthesis tasks, this embodiment conducted a quantitative comparison experiment on the TI-NSD dataset. The experiment selected three typical thermal infrared scenarios: indoor, outdoor, and UAV (unmanned aerial vehicle) perspectives. The indoor scenario included 7 scenes, the outdoor scenario included 7 scenes, and the UAV perspective scenario included 6 scenes. Structural similarity (SSIM), peak signal-to-noise ratio (PSNR), and learned perceptual patch similarity (LPIPS) were used as evaluation metrics.
[0126] The comparison methods include: (1) three-dimensional Gaussian reconstruction methods based on two-dimensional super-resolution preprocessing (such as Bicubic+3DGS, Lanczos+3DGS, DifIISR+3DGS); (2) existing mainstream and thermal infrared three-dimensional Gaussian direct upsampling rendering methods (such as 3DGS, Thermal3DGS, Mip-Splatting, etc.); (3) existing super-resolution three-dimensional Gaussian methods (SRGS).
[0127] Table 1 Comparison of different methods on the TI-NSD dataset
[0128]
[0129] In summary, this invention organically combines three-dimensional Gaussian sputtering scene representation, a physically constrained infrared imaging calibration module, an intensity condition appearance residual adaptation module, and a frequency domain perception course learning strategy. It can achieve stable high-resolution three-dimensional reconstruction and new perspective synthesis under low-resolution thermal infrared multi-view conditions, and has good engineering application prospects.
[0130] This application also provides an electronic device, including a processor and a memory.
[0131] The memory is used to store computer programs.
[0132] When the processor executes a program stored in the memory, it implements any of the methods described in this application.
[0133] In one possible implementation, the electronic device of this application embodiment further includes a communication interface and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus.
[0134] The communication bus mentioned in the above electronic devices can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into address bus, data bus, control bus, etc.
[0135] The communication interface is used for communication between the aforementioned electronic devices and other devices.
[0136] The memory may include random access memory (RAM) or non-volatile memory (NVM), such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.
[0137] The processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.
[0138] In another embodiment provided in this application, a computer-readable storage medium is also provided, which stores a computer program that, when executed by a processor, implements any of the methods described in this application.
[0139] In another embodiment provided in this application, a computer program product containing instructions is also provided, which, when run on a computer, causes the computer to perform any of the methods described in this application.
[0140] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid state disk (SSD)).
[0141] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0142] The various embodiments in this specification are described in a related manner. Each embodiment focuses on the differences from other embodiments, and the same or similar parts between the various embodiments can be referred to each other.
[0143] The above description is merely a preferred embodiment of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application are included within the scope of protection of this application.
Claims
1. A method for thermal infrared multi-view super-resolution 3D reconstruction and novel viewpoint synthesis, characterized in that, Includes the following steps: Acquire a sequence of low-resolution thermal infrared multi-view images of the scene and the camera parameters corresponding to each view; construct a 3D Gaussian sputtering scene representation and render it at the target output resolution to obtain an ideal thermal infrared image; The ideal thermal infrared image is mapped to a predicted image aligned with the observation domain through a differentiable imaging chain that includes a physical constraint infrared imaging calibration module and an intensity condition appearance residual adaptation module. Joint optimization training is performed using the structural priors provided by the teacher network and the corresponding confidence masks. After training, based on the camera parameters corresponding to the target viewpoint, the optimized 3D Gaussian sputtering scene representation is used to render at the target output resolution, outputting a high-resolution thermal infrared new perspective image from the target viewpoint.
2. The method for thermal infrared multi-view super-resolution 3D reconstruction and new viewpoint synthesis according to claim 1, characterized in that, The three-dimensional Gaussian sputtering scene is represented by multiple anisotropic three-dimensional Gaussian elements. Each three-dimensional Gaussian element includes at least a three-dimensional center, a covariance matrix, an opacity, and a single-channel thermal radiation intensity parameter. The covariance matrix is represented in a parameterized form using rotation matrix and scale matrix decomposition.
3. The method for thermal infrared multi-view super-resolution 3D reconstruction and new viewpoint synthesis according to claim 1, characterized in that, When rendering a 3D Gaussian sputtering scene representation, each 3D Gaussian primitive is projected onto a 2D image plane, and pixel intensity is calculated using an α-synthesis method based on depth order. Gaussian primitives that contribute to the same pixel are arranged in depth order, and transmittance is accumulated during the α-synthesis process. The pixel intensity of all sampled points forms an ideal thermal infrared image corresponding to the target output resolution.
4. The method for thermal infrared multi-view super-resolution 3D reconstruction and new viewpoint synthesis according to claim 1, characterized in that, The physical constraint infrared imaging calibration module includes a learnable low-resolution spatial gain grid, a learnable low-resolution spatial bias grid, a parameterized radial vignetting function, and a viewpoint-dependent monotonic tone mapping function. The learnable low-resolution spatial gain grid and the learnable low-resolution spatial bias grid are upsampled to the resolution of the current rendered image, and then constrained in numerical range by the hyperbolic tangent function to form spatial gain fields respectively. and spatial bias field Furthermore, this is combined with the vignetting correction factor calculated from the parameterized radial vignetting function. After performing low-frequency calibration on the ideal thermal infrared image, a monotonic transformation is performed using the viewpoint-related monotonic tone mapping function to obtain a calibrated prediction image aligned with the current viewpoint observation domain.
5. The method for thermal infrared multi-view super-resolution 3D reconstruction and new viewpoint synthesis according to claim 1, characterized in that, The intensity-conditional appearance residual adaptation module employs a pixel-wise multilayer perceptron network, including an input layer, two hidden layers, and an output layer. The hidden layers are connected using the ReLU activation function. For each spatial location in the input image, the rendering intensity at that location is used as input, and local multiplicative gain residuals and additive bias residuals are output, with range constraints applied to the residual amplitudes. Based on the local multiplicative gain residuals and additive bias residuals, local affine residual modulation is performed on the input image to obtain the appearance adaptation prediction image.
6. The method for thermal infrared multi-view super-resolution 3D reconstruction and new viewpoint synthesis according to claim 1, characterized in that, The joint optimization training adopts a frequency domain-aware curriculum learning strategy, which includes the following process: generating a teacher super-resolution image using a pre-trained super-resolution teacher network, upsampling a low-resolution thermal infrared multi-view image, constructing a confidence mask based on the pixel difference and gradient difference between the two, and applying weighted constraints to the high-frequency components of the current prediction image and the teacher super-resolution image within the reliable region indicated by the confidence mask.
7. The method for thermal infrared multi-view super-resolution 3D reconstruction and new viewpoint synthesis according to claim 6, characterized in that, Within the confidence region indicated by the confidence mask, high-frequency components of the predicted image and the teacher super-resolution image are extracted using a high-pass filter operator, and a high-frequency residual loss is constructed to weight and constrain the high-frequency components.
8. The method for thermal infrared multi-view super-resolution 3D reconstruction and new viewpoint synthesis according to claim 7, characterized in that, The 3D Gaussian sputtering scene representation, the physically constrained infrared imaging calibration module, and the intensity condition appearance residual adaptation module are jointly optimized based on a joint training objective. The joint training objective includes: , Among them, pixel reconstruction loss Structural similarity loss is used to constrain pixel consistency between the predicted image and the observed image. Used to constrain the structural similarity between the predicted image and the observed image. For high-frequency residual loss; the regularization term Used to constrain the spatial gain field, spatial bias field, and learnable vignetting parameters in the radial vignetting function of the physical constraint infrared imaging calibration module. The modulation amplitude used for the constraint strength condition appearance residual adaptation module.
9. The method for thermal infrared multi-view super-resolution 3D reconstruction and new viewpoint synthesis according to claim 8, characterized in that, The regularization terms of the physical constraint infrared imaging calibration module include smoothness constraints on the spatial gain field and spatial bias field, as well as amplitude constraints on the learnable vignetting parameters in the radial vignetting function. The regularization term of the strength condition appearance residual adaptation module includes amplitude constraints on the local multiplicative gain residual and the local additive bias residual.
10. The method for thermal infrared multi-view super-resolution 3D reconstruction and new viewpoint synthesis according to claim 8, characterized in that, The weights of the structural similarity loss and the high-frequency residual loss are both scheduled using a cosine-incrementing method and are dynamically adjusted according to the training iteration stage.