Multi-dimensional complementary sharpness evaluation method for infrared images

By employing a multi-dimensional complementary sharpness evaluation method, the problems of scale variability and morphological non-ideality in sharpness measurement during autofocusing of infrared imaging systems are solved, achieving efficient and stable focusing of infrared images and improving the robustness and focusing accuracy of the system.

CN121190332BActive Publication Date: 2026-02-03NANJING ARTIFICIAL INTELLIGENCE CHIPS RES INST OF AUTOMATION CHINESE ACAD OF SCI
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511730690.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-24
Publication Date
2026-02-03
Estimated Expiration
2045-11-24

AI Technical Summary

Technical Problem

Existing autofocus methods in infrared imaging systems suffer from problems related to the variability of sharpness measurement scales and non-ideal morphology, which makes the search algorithm prone to getting trapped in local optima or oscillations, and unable to accurately determine the optimal focus.

Method used

A multi-dimensional complementary sharpness assessment method is adopted. By acquiring infrared image sequences, equipment parameters and environmental data, basic calibration, physical parameter estimation and normalization modeling are performed to construct physical consistent sharpness and thermal kernel multi-scale tensor sharpness indices. Physical gating fusion and unimodal enhancement are performed to generate a unimodal enhanced sharpness curve. Combined with focus search and control closed loop, the focusing process is optimized.

Benefits of technology

It improves the scale consistency and anti-interference stability of sharpness assessment, optimizes the shape of the sharpness curve, and ensures the accuracy and rapid convergence of the focusing process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121190332B_ABST
    Figure CN121190332B_ABST
Patent Text Reader

Abstract

The application discloses a kind of multi-dimension complementary definition evaluation methods for infrared image, comprising: obtaining infrared image sequence, equipment parameter set and environmental observation data;To infrared image sequence executes basic calibration, generates basic calibration image sequence;Based on basic calibration image sequence, equipment parameter set and environmental observation data, executes physical quantity estimation and normalization modeling, obtains normalized definition sub-index set and scene description vector;According to normalized definition sub-index set and scene description vector, multi-dimension complementary definition construction and unimodality enhancement are executed, and unimodal enhanced definition curve and real-time definition value are output.The application also includes using curve to execute focusing search, and can detect sidelobe risk in search, and dynamically write back parameters to suppress sidelobe.The application improves scale consistency and anti-interference stability, optimizes morphology, and improves convergence.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application pertains to sharpness assessment technology, and more particularly to a multi-dimensional complementary sharpness assessment method for infrared images. Background Technology

[0002] Infrared imaging technology is crucial in industrial and driver assistance fields due to its all-weather and interference-resistant advantages. Autofocus is the core function of an infrared system; its speed and robustness directly determine the system's ability to acquire clear thermal images, and are a key technological prerequisite for ensuring system effectiveness.

[0003] Currently, passive autofocus is the mainstream solution for infrared systems. This technology typically acquires a series of images at different focal lengths and calculates a predefined sharpness evaluation function for each frame. These functions mainly include operators based on image gradients (such as Tenengrad), operators based on statistical variance (such as Laplacian), or operators based on frequency domain analysis (such as Fourier high-frequency energy). The system employs search strategies such as hill climbing to find the focal position that maximizes the function value.

[0004] However, the aforementioned methods relying on empirical operators suffer from two major technical bottlenecks in practical applications: scale variability and morphological non-ideality in sharpness metrics. First, scale variability stems from the extreme sensitivity of the metric to off-focus factors (such as automatic gain control, AGC). When scene changes cause AGC to switch segments, the image contrast fluctuates drastically, leading to large, unrelated false peaks in gradient- or variance-based metrics. These false peaks often cause the search algorithm to lock onto incorrect locations. Morphological non-ideality arises because these empirical operators do not consider the physical characteristics of infrared imaging (such as thermal diffusion), resulting in sharpness curves often exhibiting strong sidelobes or broad, flat tops in low-texture or high-noise scenes. This non-ideal morphology makes traditional search algorithms prone to getting trapped in local optima (sidelobes) or oscillating at the peak, preventing convergence. Summary of the Invention

[0005] The purpose of this invention is to provide a multi-dimensional complementary sharpness evaluation method for infrared images to solve the aforementioned problems in the prior art.

[0006] According to one aspect of this application, a multi-dimensional complementary sharpness evaluation method for infrared images includes:

[0007] Acquire infrared image sequences, equipment parameter sets, and environmental observation data;

[0008] Perform basic calibration on the infrared image sequence to generate a basic calibration image sequence;

[0009] Based on the basic calibration image sequence, equipment parameter set and environmental observation data, physical parameter estimation and normalization modeling are performed to obtain a set of normalized sharpness sub-indices and scene description vectors;

[0010] Based on the normalized set of sharpness sub-indices and the scene description vector, multi-dimensional complementary sharpness construction and unimodal enhancement are performed to produce a unimodal enhanced sharpness curve and real-time sharpness value;

[0011] This includes performing multi-dimensional complementary clarity construction and unimodal enhancement:

[0012] Construct a physical consistency clarity metric;

[0013] Construct a clarity index for multiscale tensors of thermonuclear kernels;

[0014] Physical gating fusion is applied to the two types of indicators to generate fused clarity.

[0015] A single-peak enhancement transformation is applied to the fused sharpness and online self-calibration is performed, producing a single-peak enhancement sharpness curve and real-time sharpness values.

[0016] According to one aspect of this application, performing physical gating fusion includes:

[0017] Virtual defocus perturbation and noise injection are applied to the basic calibration image sequence to evaluate the unimodality prior and noise sensitivity of the two types of indicators, and to construct their respective free energy measures.

[0018] Extract the scene equivalent temperature from the scene description vector;

[0019] The physical gating weights are calculated using the Boltzmann distribution based on the free energy metric and the scene's equivalent temperature.

[0020] Physical gating weights are used to weight the two types of indicators to generate fusion clarity.

[0021] According to one aspect of this application, performing online self-calibration includes:

[0022] Perform a short-window micro-scan around the current focus to acquire a local single-peak enhanced sharpness sampling sequence;

[0023] The unimodality index of a local sampling sequence is calculated by determining the minimum value of the second derivative of the logarithmic curve of the sequence and combining it with the ratio of the secondary peak height to the primary peak height.

[0024] With the goal of maximizing the unimodality index, the parameters of the unimodal enhancement transform or the weights of the physical gated fusion are updated online.

[0025] Based on the updated parameters or weights, the final single-peak enhanced sharpness curve and real-time sharpness value are generated.

[0026] According to one aspect of this application, a physically consistent sharpness metric is constructed, including:

[0027] A frequency domain weighting function is constructed based on the synthesized modulation transfer function, which includes optical defocus variance and thermal diffusion equivalent time.

[0028] The amplitude spectrum is calculated for the basic calibration image sequence, and a frequency domain weighted integral is performed using a weighting function to obtain the original physical consistent sharpness;

[0029] Obtain a set of normalization transformers from the physical parameter estimation and normalization modeling steps, and call them to scale the original physical consistency sharpness.

[0030] According to one aspect of this application, a heat kernel multiscale tensor clarity index is constructed, including:

[0031] The set of heat core scales is determined based on the scene description vector;

[0032] A three-dimensional tensor is constructed by convolving the basic calibration image sequence with a heat kernel scale set.

[0033] Perform tensor decomposition on the three-dimensional tensor to extract directional energy singularities and scale modulus factor matrices;

[0034] Based on the directional energy singularity and scale modulus factor matrix, the directional energy concentration and scale energy concentration are calculated to form the clarity of the original thermonuclear multiscale tensor.

[0035] Obtain a set of normalization transformers from the physical parameter estimation and normalization modeling steps, and call them to scale the clarity of the original thermonuclear multiscale tensor.

[0036] According to one aspect of this application, the method further includes:

[0037] By utilizing a single-peak sharpness enhancement curve and real-time sharpness values, a focus search and control closed loop is executed to determine the optimal focus position;

[0038] The focus search and control closed loop includes performing three-point fitting, performing sidelobe detection, and generating a focus drive sequence;

[0039] The sidelobe detection process includes:

[0040] During the focus search process, the presence of sidelobe risk is determined based on the inconsistency between the ratio of secondary peaks to primary peaks or curvature.

[0041] In response to sidelobe risk, an increment of the sidelobe suppression parameter is generated;

[0042] The sidelobe suppression parameter increments are written back to the multi-dimensional complementary sharpness construction and unimodal enhancement steps to dynamically change the shape of the unimodal enhanced sharpness curve in subsequent sharpness calculations.

[0043] According to one aspect of this application, a three-point fitting is performed, including:

[0044] Collect real-time sharpness values ​​corresponding to the focus position within the search range;

[0045] A parabola is fitted between the focus position and the real-time sharpness value to obtain the candidate focus position and local curvature estimate;

[0046] Based on local curvature estimation, a step size suggestion for the next movement is set to achieve curvature adaptive step size.

[0047] According to one aspect of this application, a focus drive sequence is generated, comprising:

[0048] Obtain the lens mechanism model from the device parameter set;

[0049] Based on candidate focus positions or step size suggestions, trajectory planning with jerk constraints is used to generate an initial driving sequence;

[0050] By combining the lens mechanism model, reverse clearance compensation and feedforward friction compensation are performed on the initial drive sequence to generate the final focus drive sequence.

[0051] According to one aspect of this application, performing physical parameter estimation and normalization modeling includes:

[0052] Estimate the noise variance, automatic gain function and gain energy factor, and bit depth;

[0053] Estimate the scene equivalent temperature corresponding to the environmental observation data;

[0054] Construct a closed-loop normalized mapping to generate a set of normalized clarity sub-indices, and aggregate scene equivalent temperatures to generate scene description vectors.

[0055] According to one aspect of this application, performing data acquisition and basic calibration includes:

[0056] Adaptive scene modeling based on low-rank sparse decomposition is employed to perform response uniformization and black level correction in order to avoid reliance on mechanical shutter.

[0057] Dynamic bad pixels are detected through multi-scale consistency and time consistency.

[0058] Directional interpolation guided by structural tensor is used to repair dynamic bad pixels, thereby protecting edge structures and generating a basic calibration image sequence.

[0059] Beneficial effects: Through the above technical solutions, this invention solves the problems of inconsistent scales, susceptibility to interference, and poor curve shape of traditional indicators, improves scale consistency and anti-interference stability, optimizes the shape, and improves convergence. Attached Figure Description

[0060] Figure 1 This is a schematic diagram of the overall framework of a multi-dimensional complementary sharpness assessment method for infrared images.

[0061] Figure 2 This is a schematic diagram of the physical gating fusion process.

[0062] Figure 3 This is a schematic diagram of the online self-calibration process.

[0063] Figure 4 This is a schematic diagram of the process for constructing a physical consistency clarity index. Detailed Implementation

[0064] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0065] Example 1 describes the overall framework of a multi-dimensional complementary sharpness assessment method for infrared images, such as... Figure 1 As shown.

[0066] By combining the physical characteristics of infrared imaging (such as thermal diffusion, noise, and gain) with mathematical modeling, a sharpness evaluation index with physical interpretability, cross-device comparability, and single-peak searchability is constructed to solve the problems of poor sharpness curve shape, susceptibility to noise interference, and scene changes in traditional infrared focusing.

[0067] Step 101: Obtain infrared image sequences, equipment parameter sets, and environmental observation data.

[0068] In this embodiment, the infrared image sequence is the raw data stream to be evaluated for sharpness. The set of device parameters is data characterizing the inherent properties of the imaging system, such as sensor depth, nominal focal length, lens mechanism parameters, sensor response characteristics, and noise equivalent temperature difference (NETD). Environmental observation data includes data characterizing external conditions during imaging, such as scene temperature statistics, background temperature, humidity, visibility, or optional meteorological data.

[0069] Step 102: Perform basic calibration on the infrared image sequence to generate a basic calibration image sequence.

[0070] Basic calibration is a prerequisite for physical quantity estimation and sharpness calculation. The main task of this step is to eliminate or suppress non-ideal factors in the original infrared image sequence, such as fixed pattern noise, bad pixels, and scale drift caused by automatic gain variations. Specifically, this step may include performing black level correction, response homogenization, dynamic bad pixel repair, and preliminary estimation of automatic gain. The resulting basic calibration image sequence has a more consistent dynamic range and lower noise levels, providing a robust data foundation.

[0071] Step 103: Based on the basic calibration image sequence, equipment parameter set and environmental observation data, perform physical parameter estimation and normalization modeling to obtain a set of normalized sharpness sub-indices and scene description vectors.

[0072] This step extracts key physical and statistical features from the calibrated image and establishes a unified normalization framework. Physical parameter estimation includes: noise variance in statistically flat regions, estimating the precise shape of the automatic gain function, analyzing structural tensor features and power spectral slope, and estimating the scene's equivalent temperature. Normalization modeling aims to construct a closed mathematical model to eliminate the scale effects of noise, bit depth, and automatic gain on different sharpness metrics, ensuring comparability across different devices and scenes. Its output is the set of normalized sharpness sub-metrics. Simultaneously, the estimated physical features (such as anisotropy, temperature, and spectral slope) are aggregated into a scene description vector to guide subsequent fusion and enhancement.

[0073] Step 104: Based on the normalized sharpness sub-index set and scene description vector, perform multi-dimensional complementary sharpness construction and unimodal enhancement to produce a unimodal enhanced sharpness curve and real-time sharpness value.

[0074] This step involves constructing two or more complementary sharpness metrics across different dimensions in parallel. For example, a physically consistent sharpness metric based on a thermal diffusion and optical defocus synthesis model is constructed, along with a scale entropy sharpness metric based on thermonuclear multiscale tensor decomposition. Based on the scene description vector (especially the scene's equivalent temperature), a physical gating mechanism, such as one based on Boltzmann distribution, is used to adaptively fuse these two metrics, yielding an initial fused sharpness value. A single-peak enhancement transformation (e.g., through peak lifting and sidelobe suppression) is applied to this fused value, and a single-peak index is calculated online. The transformation parameters are adaptively adjusted using short-window microscanning to maximize this single-peak index, ultimately producing a well-formed, easily searchable single-peak enhanced sharpness curve and the real-time sharpness value of the current focus.

[0075] In this embodiment, the method may further include: performing a focus search and control closed loop. This step reads the single-peak enhanced sharpness curve and real-time sharpness value produced in step 104 as the target surface for focus search. The search strategy may employ methods such as interval bracketing or three-point parabolic fitting to estimate the peak position and local curvature, and adaptively adjust the search step size based on the curvature. When sidelobes or multi-peak risks are detected, this step can be linked in reverse to dynamically adjust the sidelobe suppression parameters of the single-peak enhanced transformation in step 104, actively reshaping the sharpness curve shape to facilitate the search, thus forming a closed loop of measurement and control.

[0076] In this embodiment, the method may further include: performing statistical evaluation and model self-calibration. This step is performed in the background or offline, summarizing the records of the focus search process (such as the number of convergence steps and sidelobe event counts) and evaluating the long-term performance of the overall focus quality and unimodality index. Based on the statistical evaluation results, the model parameters used in steps 103 and 104 can be updated through self-learning. For example, the free energy metric parameters of physical gated fusion (such as Boltzmann fusion) in step 104 can be updated, or the parameter boundaries of the unimodal enhancement transformation can be updated, forming the highest-level adaptive closed loop of the system and improving the robustness of the method to long-term operating condition changes.

[0077] Example 2 describes the physical parameter estimation and normalized modeling performed in step 103 of Example 1. Specifically:

[0078] Step 201: Perform physical parameter estimation. The aim is to extract the physical and statistical features required for subsequent normalization, fusion, and enhancement from the base calibration image sequence.

[0079] Specifically, this step may include:

[0080] Estimating noise variance: For example, by using a robust criterion that combines block-level variance with median absolute deviation, flat regions in an image can be identified, and after removing edges and textures, pixel fluctuations within the flat regions can be statistically analyzed to obtain a robust estimate of the sensor noise variance.

[0081] Estimating bit depth and dynamic range: Read the bit depth information (e.g., 12-bit or 14-bit) from the device parameter set and establish a bit depth consistency coefficient across devices, for example, normalize the grayscale to a uniform dynamic range (e.g., 0-65535) through linear mapping.

[0082] Estimating structural and spectral features: For example, the structure tensor is calculated for each frame of the image, and the eigenvalues ​​are solved to obtain the anisotropy intensity and principal direction; simultaneously, a log-linear fit is performed on the radial average power spectrum to obtain the power spectrum slope. These features reflect the scene content (e.g., whether it is sparse edges or dense texture).

[0083] Estimating scene equivalent temperature: For example, using the brightness temperature coarse mapping obtained in the basic calibration steps (such as in Example 6) or the device calibration relationship, the temperature distribution of the current frame is estimated, and the number of peaks, peak width and overall temperature level are calculated to obtain the scene equivalent temperature.

[0084] Step 202: Generate the scene description vector. The multiple physical parameters estimated in step 201 are aggregated to form a structured data vector, namely the scene description vector.

[0085] Preferably, the scene description vector can aggregate, but is not limited to, anisotropic intensity, principal direction, power spectral slope, scene equivalent temperature, and temperature distribution characteristics. This vector constitutes a compact description of the physical properties of the current imaging scene and will serve as the basis for physical gating fusion and prior setting of single-peak enhancement parameters in subsequent steps (such as in Example 4).

[0086] Step 203: Estimate the automatic gain function and gain energy factor. Automatic gain control (AGC) is one of the main sources of scale drift in sharpness metrics in infrared imaging. Accurate estimation of its morphology is necessary for normalization.

[0087] Preferably, this step incorporates preliminary estimation results from the basic calibration (such as in Example 6).

[0088] In the basic calibration phase, local fitting of the grayscale mapping can be used to identify the segments and switching points of the automatic gain. Based on this, this step performs a more refined fitting for each identified segment, for example, using monotonic constraint splines to fit the grayscale histogram within the segment to construct the automatic gain function.

[0089] Furthermore, the derivative of the automatic gain function and the weighted norm of the current gray-level distribution are calculated to obtain the gain energy factor G, which quantitatively describes the amplification (or compression) effect of the automatic gain on the overall image contrast.

[0090] Step 204: Construct a closed-loop normalized mapping to eliminate the scaling effects of estimated noise, bit depth, and automatic gain on the sharpness sub-index.

[0091] In this embodiment, for each type of sharpness sub-index S to be calculated subsequently... i (For example, the physical consistency sharpness S in Example 3) phys or thermonuclear multiscale tensor clarity S tensor For each of these, a corresponding normalized mapping is established. Preferably, this closed-form normalized mapping (also known as NAA-Norm-IR) adopts the following unified normalization formula: Stilde i =(S i -μ i (σn 2 )) / σ i (σ n 2 )*2 -b *G -1 Among them, Stilde i It is the normalized sharpness sub-index; S i It is the original calculated sharpness sub-index; σ n 2 is the noise variance; b is the bit depth; G is the gain energy factor. i (σ n 2 ) and σ i (σ n 2 These are the respective sub-indicators of sharpness, S. i With only noise variance σ n 2 The theoretical expectation and theoretical standard deviation under the given conditions.

[0092] In a preferred implementation, in order to determine μ i and σ i It needs to be based on S i The mathematical definition (e.g., frequency domain integral or tensor decomposition type) determines the appropriate index noise response model. Within the flat region, μ can be obtained through analytical calculation based on this model or robust correction using sample quantile statistics. i and σ i The specific numerical value. Using this formula, the Stilde value can be calculated for different devices, different gains, and different noise levels. i All of them were unified to the same scale, forming a set of normalized sharpness sub-indices, which provided comparable inputs at the same scale for subsequent physical gating fusion (such as in Example 4).

[0093] Step 205: Summarize all the outputs of this embodiment to form a set, which will serve as the input for subsequent steps. This set preferably includes: a set of normalized transformers instantiated for various sub-indices (i.e., the formula and its parameters from step 204), the scene description vector generated in step 202, and the noise variance estimated in step 201, etc.

[0094] Example 3 describes the specific implementation of the multi-dimensional complementary sharpness construction in step 104 of Example 1. In this example, two types of core sharpness indicators with physical meaning and structural insight will be constructed in parallel. Together, they constitute the input source of the normalized sharpness sub-indicator set mentioned in Example 2.

[0095] Step 301, construct a physical consistency resolution metric, such as Figure 4 As shown.

[0096] This step aims to construct a sharpness index S based on an imaging physics model. phys This model unifies optical defocusing blur with the thermal diffusion blur effect unique to infrared.

[0097] Specifically, this embodiment introduces a synthetic fuzzy variance model, the expression of which can be: σ eff 2 =σ opt 2 +2*κ*t. Where, σ eff 2 It is the total equivalent fuzzy variance; σ opt 2 The blur variance caused by lens optical defocus is a function of the focal length z, and it reaches its minimum at the optimal focus; κ is the thermal diffusivity; and t is the thermal diffusivity equivalent time. This model attributes the blurring in infrared imaging to the combined effect of the optical system and thermal conduction (or their mathematical equivalents).

[0098] Based on this model, the physical consistency sharpness index S phys It is constructed as a frequency-domain weighted integral:

[0099] S phys =∫(W(f)*|I hat (f)| 2 df), the integration interval is from f=0 to infinity. Where, I hat (f) is the amplitude spectrum of the basic calibration image sequence; f is the spatial frequency.

[0100] The frequency domain weighting function W(f) is designed to have a form related to the synthesized modulation transfer function (MTF). Preferably, an inverse weighting design is used, such as W(f) = f 2 *exp(2*π 2 *σ ref 2 *f 2 ), where σ ref It is the nominally best imaging (i.e., σ) opt The fuzzy variance corresponding to the minimum value.

[0101] W(f) assigns higher weights to high-frequency components that attenuate due to ambiguity, in σ opt 2 When deviating from the optimal point (i.e., when the blur increases), S phys The value will decrease monotonically, ensuring the unimodality of this indicator.

[0102] In a preferred implementation, to improve S phys Robustness of computation, in calculating amplitude spectrum I hat (f) Prior to this, a preprocessing step may be performed on the basic calibration image.

[0103] For example, to reduce spectral leakage, a two-dimensional cosine window and boundary mirror expansion can be used. Furthermore, to avoid the energy dominating the integration results from low-frequency background (such as large areas of sky or walls), a flat region mask can be used to suppress or downweight the energy corresponding to extremely low frequencies and fringe noise directions in the frequency domain. The calculated original S... phys The value is fed into the normalization transformer constructed in Example 2 (step 204) to produce a normalized physical consistent sharpness Stilde. phys .

[0104] Step 302: Construct the clarity index of the thermonuclear multiscale tensor.

[0105] This step aims to construct a supplementary sharpness index S. tensor Starting from the scale-space structure and energy concentration of the image, it provides information related to S. phys Different dimensions of information.

[0106] Specifically, this step first requires constructing a three-dimensional tensor T(x, y, s). Unlike traditional Gaussian pyramids, the scale axis s in this embodiment is driven by a physical heat kernel (i.e., the Green's function of the heat conduction equation). For example, a discrete set of heat kernel scales s is constructed. k =sqrt(2*κ*t k ), where κ is the thermal diffusivity, t k For a logarithmically uniformly distributed equivalent time.

[0107] Using thermonuclear scale s k Convolve the base calibration image sequence to generate a set of scale images, and stack these scale images along the s-axis to form a three-dimensional tensor T(x, y, s).

[0108] Based on this, tensor decomposition is performed on the three-dimensional tensor T(x, y, s), such as Higher-Order Singular Value Decomposition (HOSVD) or Tucker Decomposition. The decomposition operation breaks down the tensor into a core tensor and factor matrices for each modulus (space, scale).

[0109] This step extracts key features from the decomposition results to construct a sharpness index. For example, it extracts the first two singular values ​​λ1 and λ2 representing the directional energy, i.e., the directional energy singular values ​​λ1 and λ2, and the factor matrix U3 representing the scale modulus.

[0110] Thermonuclear multiscale tensor clarity index S tensor It can be constructed from the following formula:

[0111] S tensor=(λ1 / λ2)*exp(-H(U3)); where (λ1 / λ2) represents the energy concentration in a given direction. When the image is clear, the energy is concentrated in a few principal directions, making λ1 much larger than λ2. When the image is blurry, the energy is evenly distributed in all directions, and the ratio is close to 1.

[0112] H(U3) is the information entropy of the scale modulus factor matrix U3. The term exp(-H(U3)) correspondingly characterizes the scale energy concentration. When the image is sharp, its energy is concentrated on a few key scales, H(U3) is small, and this term is large. When the image is blurry, the energy is diffused to all scales, H(U3) is large, and this term is small.

[0113] In summary, S tensor It is a multiplicative metric that takes the maximum value when the energy of the image exhibits strong concentration in both direction and scale (i.e., the image is sharp).

[0114] The calculated original S tensor The value is also fed into the normalization transformer constructed in Example 2 (step 204) to produce a normalized thermonuclear multiscale tensor sharpness Stilde. tensor Simultaneously, the tensor decomposition process also outputs tensor decomposition auxiliary statistics (e.g., scale energy spectrum), which are used for sidelobe risk assessment in Example 4.

[0115] According to one aspect of this application, in Embodiment 3, the physical model parameters κ and thermal diffusion equivalent time t involved in step 301 (constructing a physically consistent sharpness index) and step 302 (constructing a thermonuclear multiscale tensor sharpness index) are further explained as follows:

[0116] The synthetic fuzzy variance σ in step 301 eff 2 =σ opt 2 +2 / κt and the thermonuclear scale s=sqrt{2 / κt} in step 302 represent a physics-inspired mathematical model. Specifically, this model borrows the Green's kernel mathematical form of the heat conduction equation, using the 2 / κt term to parametrically model the non-optical blurring component unique to infrared images (this blurring characteristic is mathematically similar to thermal diffusion). This is a mathematical analogy, rather than a direct measurement of the physical thermal diffusion process.

[0117] In this model, κ is a model coefficient, and its value is not the physical thermal diffusivity, but preferably it can be used as a system-level calibration parameter. For example, κ can be an empirical constant obtained by fitting a set of infrared images with known different levels of fuzziness to the model. This constant characterizes the rate at which the fuzziness characteristics of the infrared system evolve with equivalent time t.

[0118] Similarly, t (thermal diffusion equivalent time) is also a model parameter used to characterize the analytical scale. Its value (or t...) k Discrete sets can be heuristically set based on prior knowledge of the typical size of the target to be analyzed. For example, to analyze the blurriness of fine textures or small targets (e.g., 3x3 pixels) in an image, a small set of t values ​​can be set, such as t set ={0.1, 0.2, 0.5} (the unit can be any time unit, used only as a model parameter). Accordingly, to analyze the structural or fuzzy characteristics of medium-sized targets (e.g., 20x20 pixels), a set of t values ​​within a medium range can be set, such as t set ={0.8, 1.5, 3.0}, which allows the model to match the fuzzy evolution characteristics at different scales.

[0119] According to one aspect of this application, the construction of the physical consistency sharpness index and scale alignment can also be as follows:

[0120] Read the basic calibration image sequence and flat region mask, select a two-dimensional cosine window and boundary mirror expansion for each frame to reduce spectral leakage and edge effects, calculate the amplitude spectrum and frequency sampling specifications, and map the flat region mask to the frequency domain for subsequent frequency band suppression to obtain the output amplitude spectrum and frequency sampling specifications.

[0121] Based on the amplitude spectrum, frequency sampling specifications, and flat region mask, a spectrum suppression mask is constructed with the low-frequency energy of the flat region as a reference. The weights of extremely low frequencies and stripe noise directions are reduced to avoid background energy dominating the integral, thus obtaining the spectrum suppression mask.

[0122] Based on the scene description vector and the nominal optimal imaging variance in the device parameter set, a weighting function is set according to the synthesis concept of thermal diffusion and optical defocus. The weighting function is obtained by fine-tuning the weighting slope and cutoff buffer band at the high-frequency end based on the power spectrum slope and anisotropy intensity in the scene description vector.

[0123] Based on the amplitude spectrum, spectral suppression mask, and weighting function, numerical integration is performed in the frequency domain using frequency-by-frequency sampling. Low-frequency blanking and fringe suppression are then combined with the spectral suppression mask, and anti-aliasing protection is applied to the high-frequency tail (automatically converged to the upper limit based on the frequency sampling specifications). This yields physically consistent sharpness.

[0124] Read the set of physical consistent sharpness and normalization transformers, call the closed normalization mapping of NAA-Norm-IR, and obtain the normalized physical consistent sharpness.

[0125] Example 4 illustrates the fusion, enhancement, and self-calibration process in the multi-dimensional complementary sharpness construction and unimodal enhancement of step 104 in Example 1. This example uses the normalized sharpness sub-index set (e.g., Stilde) produced in Example 3 above. physand Stilde tensor The scene description vector generated in Example 2 is used as input.

[0126] Specifically, it includes the following steps:

[0127] Step 401: Perform physical gating fusion on the two types of indicators. That is, based on the scene characteristics and the real-time reliability of each indicator, the Stilde... phys and Stilde tensor Adaptively blended into a single blended clarity S fuse .like Figure 2 As shown.

[0128] In this embodiment, the fusion weights are determined based on a virtual free energy metric E. i For Stilde phys and Stilde tensor The i-th index in the equation has a free energy E. i The aim is to quantify its unreliability or tendency to cause search failures.

[0129] The virtual free energy measure E i This can be evaluated by applying a virtual defocus perturbation and noise injection to the current image. Specifically:

[0130] By applying Gaussian smoothing with a very small standard deviation to the image (simulating micro-defocus), the metric Stilde was observed. i The descent rate and curvature are used to obtain the unimodal prior score M. i .

[0131] By injecting noise with a variance consistent with that estimated in Example 2 into the flat region, the statistical Stilde was obtained. i The variance of the noise is used to obtain the noise sensitivity score. i .

[0132] By analyzing the tensor decomposition auxiliary statistics (such as scale energy spectrum) output in Example 3, the multimodal risk is estimated, and the sidelobe tendency score is obtained. i .

[0133] Preferably, the free energy measure E i It can be constructed as a linear combination of the above scores:

[0134] E i =a1*(1-M i )+a2*Noise i +a3*Sidelobe i Where a1, a2, and a3 are calibration coefficients, which can be updated through offline self-learning. Virtual free energy measurement E iThe higher the value, the better the Steld indicator. i The worse the unimodality, the more susceptible it is to noise or the more likely it is to produce sidelobes.

[0135] Based on this, the free energy measure E i The scene equivalent temperature T extracted from the scene description vector in Example 2 eff The physical gating weight w is calculated using the Boltzmann distribution. i :

[0136] w i =exp(-E i / (k*T eff )) / (∑ j (exp(-E j / (k*T eff )))); where k is the Boltzmann constant (or set to 1). The physical meaning of this physical gating is: free energy E i The lower the value of the indicator (the more reliable the indicator), the higher its weight. i The scene temperature T eff The higher the value, the smoother the weight distribution (softmax effect).

[0137] Finally, the combined sharpness S fuse We obtain S by weighted summation. fuse =∑ i (w i *Stilde i ).

[0138] Step 402: Apply a single-peak enhancement transform to the fused sharpness and perform online self-calibration.

[0139] Its goal is to achieve enhanced fusion clarity S fuse (There may still be flat peaks or shoulder lobes) Reconstructed to a final sharpness S with a sharp single peak that is easy to search. star .

[0140] First, regarding the fusion clarity S fuse A single-peak enhancement transformation is applied. This transformation preferably takes the following form:

[0141] S star =log(1+exp(α*S fuse ))-γ*LSE k (S fuse *d k The first term, log(1+exp(...)), is a Softplus function whose parameter α controls the peak elevation and sharpening. The second term is the sidelobe suppression term, where γ is the suppression intensity parameter, LSE. k It is the logarithmic and exponential smoothing maxima operator, dk It is a local kernel.

[0142] In a preferred implementation, the initial values ​​or ranges of α and γ, and the kernel d k The form of γ is not fixed, but is pre-defined by the Scene Description Vector (SDV) in Example 2. For example, if the anisotropy intensity in the SDV is high, the initial range of γ can be increased, and the kernel d k It can be molded into a shape that matches the anisotropic direction.

[0143] Perform online self-calibration to maximize the unimodality of the sharpness curve, such as... Figure 3 As shown.

[0144] The core of this process is defining a computable unimodality index, M. This M must quantify the peak sharpness and sidelobe suppression of the curve.

[0145] In this embodiment, to calculate the unimodality index M, the system performs a short-window micro-scan around the current focal point, for example, performing a very small step-size focal length micro-motion of ±2 to 3 steps near the current focal point z, and acquiring the corresponding local S. star Sampling sequence.

[0146] Based on this local sampling sequence, the unimodality index M is calculated, and its definition can be as follows:

[0147] M=min z (d 2 / dz 2 [smooth(log(S star (z)))])-η*(Secondary peak height / Main peak height);

[0148] Where, d 2 / dz 2 [smooth(log(S star [z] is the second derivative (i.e., curvature) of the logarithmic sharpness curve. min(...) represents taking its minimum value (representing the sharpness of the peak, the more negative the better); the second term is the ratio of the height of the secondary peak to the height of the primary peak (side lobe penalty term); η is the tradeoff coefficient.

[0149] With the objective of maximizing the unimodality index M (i.e., minimizing M or maximizing one of its positive transforms), the parameters (α, γ) of the unimodal enhancement transform or the weights (w) of the physically gated fusion are updated online. i ).

[0150] In a preferred implementation, the update can employ a limited-amplitude mirror gradient step optimization algorithm to perform small updates on α and γ. If an excessively high secondary peak ratio is detected, this step will also invoke kernel d. kThe adaptive adjustment is used to specifically suppress sidelobes.

[0151] After 1-2 rounds of iterative updates, the system generates the final single-peak enhanced sharpness curve (i.e., S) based on the final parameters. star (the function form of z) and the real-time sharpness value (i.e., the S corresponding to the current focus z). star value).

[0152] According to one aspect of this application, in Embodiment 4, the scene equivalent temperature T used in step 401 (physical gating fusion) is... eff Its application in the Boltzmann distribution is further explained below:

[0153] In this step, the Boltzmann distribution is used to calculate the physical gating weight w. i It is a mathematical borrowing. Its function is equivalent to the Softmax weight assigner commonly used in machine learning, used to weight two (or more) sharpness sub-indices (S... phys and S tensor ) Perform soft weighting. E i It is a virtual free energy or cost function composed of measurable indicators (such as unimodality, noise sensitivity).

[0154] In this model, T eff It is not a precise measurement of physical temperature, but rather a temperature parameter used by the Softmax allocator to control the sharpness or smoothness of weight allocation. Specifically, when a lower T is set... eff When the value is equal to 1, the weight distribution tends to be sharp (i.e., winner-takes-all), causing the free energy E to decrease. i The significantly lower metric receives the majority of the weight (e.g., w1=0.9, w2=0.1). When a higher T is set... eff When the values ​​are equal, the weight distribution tends to be smoother (e.g., w1=0.55, w2=0.45), which is suitable for two indicators E. i Approaching or in situations with high uncertainty.

[0155] T eff The specific values ​​can be adaptively set based on the scene temperature distribution characteristics (such as peak width and number of peaks) estimated in Example 2 (step 201). This is a heuristic mapping. For example, if the temperature distribution statistics show a single, sharp peak (indicating a simple scene with concentrated contrast), the system can determine that the scene has high determinism and set a lower T value. eff Value (e.g., T) eff The normalization parameter is set to 0.5 to allow for decisive weight allocation. Conversely, if the temperature distribution exhibits multiple peaks or a broad flat top (indicating a complex scene with interwoven target and background temperatures), the system can determine that the scene has high uncertainty and set a higher T value.eff Value (e.g., T) eff The normalization parameter value is 2.0 to perform more conservative and smoother weight fusion.

[0156] According to one aspect of this application, the construction of the free energy metric and the fusion with Boltzmann gating can also be:

[0157] Based on the baseline calibration image sequence, normalized physical consistency sharpness, and normalized thermonuclear multiscale tensor sharpness, a Gaussian smoothing with minimal standard deviation is constructed for the current frame to simulate the slight increase in the variance of the synthetic blur. The two types of sharpness are recalculated separately, and the relative reduction and local second-order curvature are obtained to form a unimodal prior score (one for each type of index). The unimodal prior score is then output.

[0158] Based on the basic calibration image sequence and noise variance, Gaussian noise with the same variance as the noise is injected into the flat area. The two types of sharpness are repeatedly calculated, and the variance and relative sensitivity are statistically analyzed to obtain the noise sensitivity score (one for each type of index).

[0159] The tensor decomposition auxiliary statistics and normalized thermonuclear multiscale tensor clarity are read, the energy ratio of secondary peaks to primary peaks is extracted from the scale energy spectrum, and potential multi-peaks or scale sidelobes are quantified to obtain sidelobe tendency scores.

[0160] Based on the unimodality prior score, noise sensitivity score, and sidelobe tendency score, the free energy measure of each index is constructed by linear combination, resulting in a set of free energy measures.

[0161] Read the set of free energy measurements and the scene equivalent temperature, calculate the gating weights using the Boltzmann distribution, and obtain the physical gating weights.

[0162] The fused sharpness is calculated based on the normalized physical consistency sharpness, the normalized thermonuclear multiscale tensor sharpness, and the physical gating weights. A stability check is performed on temporal abrupt changes (if the weight abrupt changes exceed the threshold, a smooth permutation is triggered) to obtain the initial value of the fused sharpness.

[0163] According to one aspect of this application, the single-peak enhancement transform and sidelobe suppression can also be:

[0164] By reading the physical gating weights and the anisotropy intensity and power spectral slope in the scene description vector, and considering strong anisotropy and steep high spectrum as priors for sidelobe risk, the ranges of the initial values ​​for the single-peak enhancement parameter and the sidelobe suppression parameter are set.

[0165] Based on the initial values ​​of fused sharpness, single-peak enhancement parameters, and sidelobe suppression parameters, execute:

[0166] S star =log(1+exp(α•S fuse))−γ•LSE k (S fuse *d k ), where the single-peak enhancement parameter α raises the peak, the sidelobe suppression parameter γ weakens the sidelobes, and the convolution kernel d k The scale and shape are guided by the anisotropic intensity. Numerical stability is considered for large S... fuse Soft saturation protection is applied to obtain the initial value of single-peak enhanced sharpness.

[0167] Based on the scene description vector and physical gating weights, the convolution kernel d k The directional and scale weights are fine-tuned to make the suppression stronger in high-risk directions and less suppressive in the main information direction, thus generating the convolution kernel weight adjustment amount.

[0168] The initial value of the single-peak sharpness enhancement is read, and the boundary points of the curve are protected to avoid misjudgment caused by edge distortion due to convolution in the short-window micro-scan. The initial value of the single-peak sharpness enhancement after boundary protection is output.

[0169] Example 5: This example details the implementation of the focus search and control closed loop in Example 1. This example uses the single-peak enhanced sharpness curve (target plane) and real-time sharpness value (current value) generated in Example 4 as input.

[0170] Step 501: Perform focus search.

[0171] In a preferred embodiment, the search does not begin with a blind global scan.

[0172] Preferably, the first step of the search is initial interval selection. This step utilizes historical priors (e.g., the optimal focus position and local curvature of the last focus, which can be provided by the statistical module in Example 8) and combines the local difference (first derivative estimation) of the current focus to intelligently construct an initial search interval that includes the main peak. For example, if the scene has not changed significantly (which can be determined by histogram or SDV changes), the search interval can be scaled and centered at the previous peak position, reducing the invalid search range.

[0173] Step 502: Perform three-point fitting and curvature adaptive step size.

[0174] Within the bracket range determined in step 501, the system performs a search.

[0175] Specifically, the system collects real-time sharpness values ​​S corresponding to at least three focal positions z1, z2, and z3 within the interval. star (z1), S star (z2), S star (z3).

[0176] In a preferred implementation, the selection of sampling points is constrained. The system will actively avoid the automatic gain switching adjacent bands identified in Embodiment 2 (or Embodiment 6). This is because automatic gain switching will affect the sharpness value S. star A sudden change unrelated to focus occurs, and if the sampling point falls into this region, it will cause false peaks in subsequent fitting.

[0177] After acquiring the sampling points, the system analyzes the three sets of data (z... i S star (z i By fitting a parabola, candidate focal positions (i.e., the vertex of the parabola) and local curvature estimates (i.e., the second-order coefficients of the parabola) are obtained.

[0178] This local curvature estimation is used to achieve an adaptive curvature step size. Specifically, based on the local curvature estimation, a suggested step size for the next move is set. For example, if the curvature estimate is large (the peak is very sharp), the step size for the next move should be reduced accordingly for a finer search; if the curvature estimate is small (the peak is very flat), the step size can be increased appropriately for a faster approximation.

[0179] Step 503: Perform sidelobe detection and reverse linkage. During the search process, the system continuously detects sidelobe risks.

[0180] The criteria for determining sidelobe risk may include: during the search process, based on (the collected S...) star The determination can be made by the ratio of the secondary peak to the primary peak in the sequence, or by the inconsistency of curvature (e.g., excessively large fitting residuals or abnormal curvature signs) in the parabolic fit.

[0181] In a preferred implementation, such as step 502, if the candidate focal position is adjacent to the automatic gain switching band, it will also be judged as a sidelobe risk (false peak risk).

[0182] In response to sidelobe risk, the system generates an increment of the sidelobe suppression parameter.

[0183] This increment is written back to the multi-dimensional complementary clarity construction and unimodality enhancement steps (i.e., Example 4).

[0184] Specific write-back operations may include:

[0185] Increase the value of the sidelobe suppression parameter γ in the single-peak enhancement transform of Example 4.

[0186] Preferably, the convolution kernel d in Example 4 is adjusted. k The weights are adjusted to apply stronger suppression at specific scales or directions of detection.

[0187] Optionally, the equivalent temperature T of the Boltzmann-gated scenario in Example 4 can be temporarily increased. effThe lower limit is used to smooth the fusion weights and reduce the volatility of the curve.

[0188] Through this reverse linkage, the sharpness curve S star The shape of (z) is dynamically changed in the calculation of the next frame (becoming smoother or having lower sidelobes) to actively adapt to the search algorithm, making it easier to find the true peak.

[0189] Step 504: Generate the focus drive sequence.

[0190] This step converts the candidate focus position or step size suggestions generated in step 502 into precise motor control commands.

[0191] Obtain the lens mechanism model from the set of device parameters. This model preferably includes the lens backlash, static and dynamic friction models, and constraints on maximum speed, acceleration, and jerk.

[0192] Based on candidate focal positions or step size suggestions, an initial drive sequence is generated using jerk-constrained trajectory planning (e.g., S-curve trajectory planning). This ensures smooth motor start-up and shutdown, avoiding inaccurate sharpness measurements due to vibration.

[0193] By combining the lens mechanism model, reverse clearance compensation (e.g., adding an additional compensation displacement when the direction of motion changes) and feedforward friction compensation (e.g., applying an anti-friction torque in advance according to the velocity curve) are performed on the initial drive sequence to generate the final focus drive sequence.

[0194] Step 505: Determine the optimal focus position.

[0195] The search process needs a robust termination criterion.

[0196] In a preferred implementation, the termination criterion is not based on the sharpness curve S. star Instead of a single threshold value, a multi-criteria termination criterion is adopted.

[0197] This criterion can be combined with the following indicators:

[0198] Position convergence: For example, the change in the candidate focus position obtained by two consecutive three-point fittings is less than the minimum step distance of the motor.

[0199] Curvature convergence: The local curvature estimate has stabilized within a reasonable range (e.g., greater than a certain threshold, indicating that it has reached the peak; and the rate of change is small enough).

[0200] Uncertainty assessment: The focus position uncertainty estimated based on curvature and noise level has converged to an acceptable range.

[0201] When the multi-criteria termination criterion is met, the search stops, and the last candidate focus position is output as the optimal focus position.

[0202] Example 6: A preferred and detailed implementation of the basic calibration in step 102 of Example 1. This example focuses on providing a high-quality, highly stable basic calibration image sequence for all subsequent steps through correction and anchoring methods in the early stages of data acquisition.

[0203] Step 601: Perform shutter-independent response uniformization and black level correction.

[0204] This step supports the constraints regarding response uniformity and black level correction. Traditional infrared correction, also known as non-uniformity correction (NUC), heavily relies on a mechanical shutter (or baffle) closing in front of the lens to obtain a uniform radiation field for calibrating pixel responses. However, the action of the mechanical shutter inevitably introduces minute vibrations and can cause a momentary disruption of the thermal equilibrium of the sensor's focal plane. During precision focusing (as in Example 5), disturbances caused by non-focusing factors can lead to abnormal jumps in the sharpness curve, severely interfering with the convergence of the search algorithm.

[0205] To address this issue, this embodiment preferably employs a scene-adaptive modeling approach to achieve shutter-independent correction. For example, a low-rank plus sparse decomposition (LRSD) model can be used. This model assumes that within a short time window, an infrared image sequence can be modeled as a superposition of three parts: a real scene that changes slowly over time (corresponding to the mathematically low-rank part), a spatially fixed pattern noise (such as stripe noise, corresponding to the mathematically sparse part), and a random noise term.

[0206] Specifically, by performing robust principal component analysis (RPCA) or a similar low-rank matrix recovery algorithm on an image sequence within a time window (e.g., 100 frames), fixed pattern noise can be robustly separated from the sparse parts, and the black level (fixed bias) and the gain coefficients (i.e., non-uniformity correction coefficients) of each pixel can be estimated using robust regression. This method does not rely on a mechanical shutter and can continuously and adaptively update the correction coefficients during normal system imaging, ensuring real-time performance and accuracy of the correction while avoiding interference from vibration and thermal disturbances.

[0207] Step 602: Perform dynamic dead pixel self-updating and structure-fidelity interpolation. This step supports limitations regarding dynamic dead pixels and directional interpolation. Traditional dead pixel correction relies on a static dead pixel table calibrated at the factory. However, after long-term operation or drastic temperature changes, new or intermittent dead pixels may appear on the sensor, which the static dead pixel table cannot cover. These dead pixels typically appear as constant bright or dark spots in the image. If left untreated, they will be mistakenly identified as extremely strong high-frequency details (edges) by subsequent sharpness algorithms (such as in Example 3), resulting in a fixed false peak on the sharpness curve that is unrelated to the true focus.

[0208] In this embodiment, a dynamic bad pixel self-updating strategy is adopted. For example, by using multi-scale consistency test (comparing whether the statistical characteristics of a pixel and its neighborhood are consistent at different scales) and temporal consistency test (detecting whether the gray value of a pixel undergoes a non-physical jump in the temporal domain), newly emerging bad pixels and intermittent bad pixels are detected and located in real time, generating a dynamically updated bad pixel mask.

[0209] When repairing detected bad pixels, directional interpolation guided by the structure tensor is preferably used to protect the true edge structure of the image. Specifically, traditional mean or median interpolation is isotropic; when a bad pixel falls exactly on an edge, this interpolation averages the gray levels on both sides of the edge, resulting in blurred edges. In this embodiment, the structure tensor of the bad pixel's neighborhood (e.g., a 5x5 window) is first calculated, and its eigenvectors are solved to determine the dominant direction of the neighborhood (i.e., the direction of the edge). The interpolation process (e.g., one-dimensional linear or spline interpolation) is performed along this dominant direction (i.e., parallel to the edge), rather than perpendicular to the edge. Directional interpolation avoids introducing additional blur when repairing edge bad pixels, maximizing the protection of high-frequency information required for subsequent sharpness calculations.

[0210] Step 603: Perform environmental assimilation and coarse brightness temperature mapping anchoring. This step is used to initially correlate the unitless grayscale values ​​output by the sensor (e.g., 0-4095) with brightness temperatures that have clear physical meaning. This is for estimating the scene equivalent temperature (T) in Example 2. eff The key physical parameters provide a physical basis.

[0211] Specifically, this step utilizes known sensor response curves (if any) and factory calibration constants from the device parameter set, combined with environmental observation data (e.g., the average background temperature of the scene or sky temperature measured by external sensors) as one or more radiation anchor points. For example, it can be assumed that the true brightness temperature of a large background area in the image (such as the sky or distant mountains) should be close to the background temperature value in the environmental observation data. Using this anchor point (i.e., a set of [grayscale value, brightness temperature value] corresponding points), the system can construct a simplified (e.g., piecewise linear or polynomial) coarse mapping from grayscale to brightness temperature. This step incorporates environmental physical information (rather than pure image statistics) into the early calibration, reducing the need for subsequent T... eff The deviation in the estimate.

[0212] Step 604: Perform coarse estimation and segmentation of the automatic gain function. This step is a prerequisite for the precise estimation of the automatic gain function (step 203) in Example 2.

[0213] Automatic gain control (AGC) in infrared imaging systems is typically piecewise linear (or nonlinear), with different gain slopes (i.e. contrast stretching) in different grayscale input ranges.

[0214] In this embodiment, monotonic constraint splines are preferably used to locally fit the frame-by-frame grayscale mapping (i.e., the relationship between input and output grayscale), and combined with the detection of grayscale change rate and saturation pixel ratio, to explicitly identify the segmentation points and switching points of automatic gain. The output gain segment division information has a dual purpose: first, it is used in Embodiment 2 (step 203) to perform individual refined fitting on each segment, improving the estimation accuracy of the gain energy factor G; second, it is used in Embodiment 5 (steps 502, 503) to treat these switching bands as high-risk areas and actively avoid them during focus search, so as to prevent the focus algorithm from misjudging them as false peaks due to sudden gain changes.

[0215] Step 605: Generate data quality tags and consistency reports.

[0216] This step preferably involves real-time evaluation of the quality of the generated base calibration image sequence. For example, its signal-to-noise ratio (which can be estimated using the flat area and noise variance from Example 2), the proportion of saturated pixels (overexposed or underexposed), and scene stability (e.g., through inter-frame difference estimation) can be calculated.

[0217] The evaluation results are used to form a quality label. This label is used as a confidence prior in subsequent steps to improve the robustness of the algorithm. For example, if the quality label indicates that the current signal-to-noise ratio is low, Embodiment 2 (step 201) can automatically increase the threshold for identifying flat regions when estimating the noise variance; and Embodiment 4 (step 401) can automatically increase the threshold for calculating the free energy E. i When this is the case, a noise sensitivity score can be given. iHigher weights (i.e., increasing the corresponding coefficient a2) make the fusion result more inclined towards the sub-index with stronger noise resistance.

[0218] Example 7: Calculation process of physical gating fusion and unimodality index. This example illustrates the calculation process of free energy metric, Boltzmann gating weight, and unimodality index.

[0219] Step 701, calculation of physical gating fusion.

[0220] This step exemplifies the computation flow of step 401 (physical gating fusion) in embodiment four.

[0221] Suppose that at a certain moment, the system has completed the calculations of Example 3 and obtained two normalized sharpness sub-indices (i.e., a set of normalized sharpness sub-indices):

[0222] Physical consistency sharpness Stilde phys =1.50; Thermonuclear Multiscale Tensor Resolve Stilde tensor =1.20;

[0223] Meanwhile, the system has completed the calculations for Example 2, obtaining the scene equivalent temperature T from the scene description vector. eff To simplify the calculation, we assume k*T. eff After calibration, the equivalent normalized energy of the product of Boltzmann constant and temperature is 0.8.

[0224] The system executes the evaluation process of Example 4 (step 401), and obtains the free energy measure E of each of the two indicators through virtual defocus perturbation, noise injection, and sidelobe risk assessment. i Assume the evaluation results are as follows:

[0225] For Stilde phys Its unimodal prior score M i =0.9 (Good), Noise Sensitivity i =0.2 (low), sidelobe tendency i =0.1 (low).

[0226] For Stilde tensor Its unimodal prior score M i =0.7 (Medium), Noise Sensitivity i =0.6 (slightly high), sidelobe tendency i =0.5 (Medium).

[0227] Assume the free energy formula defined in Example 4 is E i =a1*(1-M i)+a2Noise i +a3Sidelobe i The calibration coefficients (which can be updated by Example 8) are currently a1=1.0, a2=1.0, and a3=1.0.

[0228] Then E phys =1.0*(1-0.9)+1.00.2+1.00.1=0.4. Therefore, E tensor =1.0*(1-0.7)+1.00.6+1.00.5=1.4.

[0229] E phys Much smaller than E tensor This indicates that the system determines Steld. phys The current frame is a more reliable indicator.

[0230] Based on this, the system calculates the physical gating weight w using the Boltzmann distribution. i :w phys =exp(-E phys / (kT eff )) / (exp(-E phys / (kT eff ))+exp(-E tensor / (k*T eff )));

[0231] w phys =exp(-0.4 / 0.8) / (exp(-0.4 / 0.8)+exp(-1.4 / 0.8));

[0232] w phys =exp(-0.5) / (exp(-0.5)+exp(-1.75));

[0233] w phys =0.6065 / (0.6065+0.1738)=0.777 (i.e. 77.7%).

[0234] Accordingly, w tensor =1-w phys =0.223 (i.e., 22.3%).

[0235] Ultimately, the fusion of sharpness S fuse We obtain the result through weighted summation:

[0236] S fuse =w phys *Stilde phys +w tensor *Stilde tensor S fuse=0.777*1.50+0.223*1.20=1.1655+0.2676=1.4331.

[0237] This calculation process shows that the system automatically uses S with lower free energy (i.e., more reliable performance). phys A higher fusion weight was assigned.

[0238] Step 702, Calculation of the unimodality index M.

[0239] This step exemplifies the calculation process of the unimodality index M in step 402 (online self-calibration) of Embodiment 4.

[0240] Assuming the system is in the short-window microscan phase, local single-peak sharpness enhancement sampling sequences S are acquired around the current focus z=0 at three positions: z=-1, 0, and +1 (unit: focal length step). star (z). To simplify the calculation, assume that the sampled S star The value has undergone logarithmic transformation and light smoothing (corresponding to smooth(log(S))). star (z)))), its sequence is:

[0241] log(S star (-1))=1.80log(S star (0))=2.00(current main peak)log(S star (+1))=1.70.

[0242] The system first calculates the minimum of the second derivative (curvature) of the logarithmic curve of the sequence. A central difference approximation is used here (assuming a step size of s). s For 1):

[0243] d 2 / dz 2 [log(S star (0))]≈(log(S star (+1))-2log(S star (0))+log(S star (-1))) / (s s 2 )d 2 / dz 2 [log(S star (0))]≈(1.70-22.00+1.80) / (1 2 = -0.5. This value of -0.5 quantifies the sharpness of the peak (the more negative, the sharper).

[0244] Meanwhile, the system detects secondary peaks within the sampling window (or a wider window). For example, a secondary peak is detected at z=+5, and its height is 0.1 times the height of the primary peak (corresponding to the ratio of the secondary peak height to the primary peak height).

[0245] Assume the unimodality index M is defined as: M = min(second derivative) - η*(ratio of secondary peak to primary peak) and the tradeoff coefficient η (calibrated by the system) = 1.0. Then M = -0.5 - 1.0*(0.1) = -0.6.

[0246] The M value of -0.6 will be used as the target for this round of online self-calibration. In the next step, the system will make small updates to the α and γ parameters by (for example) mirroring the gradient step and recalculate the M value to make the M value smaller (e.g., -0.7), that is, the peak apex is sharper (curvature is more negative) or the secondary peaks are suppressed less.

[0247] Example 8 describes a third layer, namely an offline or background self-learning closed-loop mechanism, in addition to real-time focusing (as in Examples 4 and 5).

[0248] Step 801: Summarize the focusing process records and real-time sharpness values.

[0249] In this embodiment, the system continuously collects and summarizes the focusing process record output by Embodiment 5 (e.g., step 505) in the background (e.g., when the CPU / GPU is idle) upon completion of each focusing task. The focusing process record is a structured log. Preferably, this record includes not only the final optimal focus position, but also:

[0250] Performance data: Total number of steps and total time required to reach convergence.

[0251] Trajectory data: The complete search trajectory (i.e., S corresponding to each z) star (Value), local curvature estimation of each three-point fit.

[0252] Event data: count of sidelobe events that occurred, number of times and magnitude of triggering sidelobe suppression (such as gamma boost).

[0253] Quality data: final unimodality index M and focus location uncertainty.

[0254] Scene snapshot: Scene description vector at the moment of successful focus (from Example 2).

[0255] Step 802: Perform a long-term statistical evaluation.

[0256] This step performs statistical analysis on the focus records collected in step 801 (e.g., the most recent 1000 times) to uncover potential, systemic performance bottlenecks. For example, the system may perform the following analysis:

[0257] Performance metrics: Calculate the average number of convergence steps, focus success rate (e.g., the proportion of failures not due to side lobes), and timeout failure rate.

[0258] Quality indicators: Calculate the mean and variance of the final unimodality index M, and the mean of the sidelobe ratio.

[0259] Scene correlation analysis: This analyzes the correlation between specific scene description vectors and focus failures (e.g., high sidelobe ratio). For example, through cluster analysis, the system may identify scenes with high anisotropy and low Tf. eff In this scenario, the average sidelobe event count is significantly higher than in other scenarios.

[0260] Step 803: Update model parameters and policy boundaries.

[0261] This step is the core self-learning process in this embodiment. Based on the statistical evaluation results of step 802, the system adaptively updates the hyperparameters or model coefficients used in Embodiments 4 and 5. Specifically, the update may include:

[0262] 1. Update the free energy metric parameters: If step 802 finds that, in a highly anisotropic scenario, S tensor Sidelobe tendency score (from Example 3) i It is systematically underestimated (i.e., S) tensor This leads to a greater number of sidelobe events. The system can update the free energy formula E in Example 4 (step 401) offline. i =a1*(1-M i )+a2Noise i +a3Sidelobe I Among them, S tensor The corresponding coefficient a3 (side lobe penalty) was increased from 1.0 to 1.2.

[0263] 2. Update the boundary of the single-peak enhancement parameter: If step 802 finds that the current average number of convergence steps is too high, and the final peak curvature (the first term of M) is generally too small (the peak is too flat), the system can update the prior value range of the α (peak elevation) parameter in Example 4 (step 402) offline, for example, by setting its lower limit α. min The value was increased from 0.5 to 0.8, which raises the starting point for real-time self-calibration (step 402).

[0264] 3. Update the search strategy boundary: If step 802 finds that false peaks caused by switching the adjacent band of automatic gain control (from Example 6, step 604) account for 30% of all focus failure cases, the system can update the safety width of the avoidance band in Example 5 (steps 502, 503) offline, expanding it from 3 steps to 5 steps.

[0265] Step 804: Generate and apply the strategy to update the results.

[0266] The parameters updated in step 803 (e.g., a3=1.2, α) min =0.8, avoidance width=5) are used to form the strategy update result. This result is stored and read and applied during the next system initialization or calibration cycle by the parameter initialization process of Examples 2, 4, 5, and 6.

[0267] In summary, this embodiment, through a long-term, background statistical and self-learning closed loop, enables the method of the present invention to continuously optimize its internal model parameters, automatically adapt to equipment aging (such as changes in sensor noise characteristics) or unforeseen specific difficult working conditions, and further improve the robustness and long-term performance of the overall focusing system.

[0268] According to one aspect of this application, constructing a closed normalized mapping includes:

[0269] For the sharpness sub-index, a normalization formula S for scale alignment is established. ~ i =(S i -μ i (σ n 2 )) / σ i (σ n 2 )*2 -b *G -1 ; where μ i With σ i For the sharpness sub-index S i Corresponding to the noise variance σ n 2 The expected value and standard deviation, b is the bit depth, and G is the gain energy factor;

[0270] The normalization formula is used to quantitatively eliminate the scale effects of noise, bit depth, and automatic gain on the sharpness sub-indices, so as to obtain a normalized set of sharpness sub-indices.

[0271] According to one aspect of this application, estimating the scene equivalent temperature includes:

[0272] In the physical parameter estimation and normalization modeling steps, the equivalent temperature of the scene is estimated;

[0273] Scene equivalent temperature is used to calculate physical gating weights through Boltzmann distribution, establishing a physical scene adaptation foundation for physical gating fusion.

[0274] According to one aspect of this application, performing data acquisition and basic calibration includes:

[0275] Adaptive scene modeling based on low-rank sparse decomposition is employed to perform response uniformization and black level correction in order to avoid reliance on mechanical shutter.

[0276] Dynamic bad pixels are detected through multi-scale consistency and time consistency.

[0277] Directional interpolation guided by structural tensor is used to repair dynamic bad pixels, thereby protecting edge structures and generating a basic calibration image sequence.

[0278] According to one aspect of this application, performing data acquisition and basic calibration also includes:

[0279] Using environmental observation data as radiation anchor points, a coarse brightness-temperature mapping is constructed;

[0280] Brightness-temperature coarse mapping is used in the physical parameter estimation and normalization modeling steps as a physical basis for estimating the equivalent temperature of the scene.

[0281] According to one aspect of this application, estimating the automatic gain function and gain energy factor includes:

[0282] By locally fitting the grayscale mapping, the segmentation and switching points of the automatic gain are identified;

[0283] Refine the fit for each segment to construct the automatic gain function;

[0284] The derivative of the automatic gain function and the weighted norm of the gray-level distribution are calculated to obtain the gain energy factor.

[0285] According to one aspect of this application, the construction of thermonuclear multiscale tensors and the calculation of scale entropy clarity can also be:

[0286] Based on the scene description vector and the set of device parameters, a discrete scale set is constructed according to the thermal diffusivity κ and the typical target size. The upper limit is adjusted by combining the scene's equivalent temperature to avoid excessively wide smoothing. This yields the thermal core scale set.

[0287] Read the base calibration image sequence and the set of heat core scales. For each scale s k Scale images are generated by hot kernel convolution, and a stack of scale images is obtained by normalizing with mirror boundaries and energy order preservation.

[0288] Based on scale image stacking and flat region masking, local energy normalization is first performed at each scale to weaken cross-scale intensity bias. Then, the images are stacked along the spatial and scale axes to form a three-dimensional tensor T(x,y,s), and the effective voxel mask is recorded to obtain the tensor data structure.

[0289] Read the tensor data structure. Use higher-order singular value decomposition or Tucker decomposition, select the rank based on the cumulative energy threshold and stability criterion, and output the modulus factor matrices and core tensors to obtain the tensor decomposition results.

[0290] Read the tensor decomposition results. Calculate the probability distribution p of the first two singular values ​​λ1 and λ2 of the directional energy and the scale modulus factor matrix U3. k This constitutes the clarity of the thermonuclear multiscale tensor and the scale energy spectrum.

[0291] Read the clarity of the thermonuclear multiscale tensor, the set of normalized transformers, and the normalized boundary configuration. Call NAA-Norm-IR to complete the scale alignment and obtain the normalized clarity of the thermonuclear multiscale tensor; at the same time, output the tensor decomposition auxiliary statistics (directional concentration curve, scale energy spectrum peak position and peak width).

[0292] In summary, this method includes: acquiring infrared image sequences, equipment parameter sets, and environmental observation data; performing basic calibration to generate a basic calibration image sequence; and based on the basic calibration image sequence, equipment parameter sets, and environmental observation data, performing physical parameter estimation and normalized modeling to obtain a set of normalized sharpness sub-indices and a scene description vector.

[0293] The method further includes: constructing a physically consistent sharpness index based on the thermal diffusion-optical defocus synthesis modulation transfer function, and constructing a thermal kernel multiscale tensor sharpness index; based on the free energy metric obtained by evaluating the unimodality prior and noise sensitivity of each index, and combined with the scene equivalent temperature, performing physical gating fusion of the two types of indices through Boltzmann distribution to generate fused sharpness; applying a unimodal enhancement transformation to the fused sharpness, and performing short-window microscanning to calculate the unimodality index, and maximizing this index to self-calibrate the transformation parameters online, finally producing a unimodal enhanced sharpness curve.

[0294] This invention can also perform focus search through three-point fitting and curvature adaptive step size, and dynamically write back the sidelobe suppression parameters to reshape the curve shape when sidelobe risk is detected, thus solving the problems of inconsistent index scales, susceptibility to gain interference and poor curve shape of traditional indicators.

[0295] First, to address the scale variability and spurious peak issues caused by automatic gain control (AGC) switching, this solution mathematically and quantitatively eliminates the scale influence of gain, bit depth, and noise variance on sharpness metrics by constructing a closed-loop normalized mapping. Simultaneously, through the linkage between automatic gain segmentation identification and subsequent search strategies, the search algorithm can proactively avoid gain switching bands, thus mitigating the interference of spurious peaks.

[0296] To address the morphological non-ideals (i.e., sidelobes and wide flat tops) caused by traditional empirical operators, this solution provides a multi-layered approach:

[0297] The traditional empirical operators were replaced by physical consistency clarity based on a physical model (thermal-optical coupling model) and thermonuclear multiscale tensor clarity based on energy concentration, which ensured a better unimodal trend in the construction.

[0298] Through physical gating fusion, adaptive weighting is performed based on the free energy of each indicator (including sidelobe tendency and noise sensitivity assessment) to dynamically select and combine the indicators with the most ideal morphology in the current scenario.

[0299] By employing a single-peak enhancement transform and online self-calibration, with the goal of maximizing a calculable single-peak performance index, the fused curve is actively sharpened and sidelobe suppressed, thus achieving a positive reshaping of the curve shape.

[0300] By employing sidelobe detection and a reverse linkage mechanism, a collaborative closed loop between measurement and search is established. When the search algorithm detects sidelobe risk, it can dynamically write back the suppression parameters to the measurement module, causing the sharpness curve to actively deform to facilitate the search, thus solving the problem of the separation between measurement and control in traditional schemes.

[0301] The preferred embodiments of the present invention have been described in detail above. However, the present invention is not limited to the specific details in the above embodiments. Within the scope of the technical concept of the present invention, various equivalent transformations can be made to the technical solutions of the present invention, and these equivalent transformations all fall within the protection scope of the present invention.

Claims

1. A multi-dimensional complementary sharpness evaluation method for infrared images, characterized in that, include: Acquire infrared image sequences, equipment parameter sets, and environmental observation data; Perform basic calibration on the infrared image sequence to generate a basic calibration image sequence; Based on the basic calibration image sequence, equipment parameter set and environmental observation data, physical parameter estimation and normalization modeling are performed to obtain a set of normalized sharpness sub-indices and scene description vectors; Based on the normalized set of sharpness sub-indices and the scene description vector, multi-dimensional complementary sharpness construction and unimodal enhancement are performed to produce a unimodal enhanced sharpness curve and real-time sharpness value; This includes performing multi-dimensional complementary clarity construction and unimodal enhancement: Construct a physical consistency clarity metric; Construct a clarity index for multiscale tensors of thermonuclear kernels; Physical gating fusion is applied to the two types of indicators to generate fused clarity. A single-peak enhancement transformation is applied to the fused sharpness and online self-calibration is performed, producing a single-peak enhancement sharpness curve and real-time sharpness values; Among them, physical gating fusion includes: Virtual defocus perturbation and noise injection are applied to the basic calibration image sequence to evaluate the unimodality prior and noise sensitivity of the two types of indicators, and to construct their respective free energy measures. Extract the scene equivalent temperature from the scene description vector; The physical gating weights are calculated using the Boltzmann distribution based on the free energy metric and the scene's equivalent temperature. Physical gating weights are used to weight the two types of indicators to generate fusion clarity; Construct a physically consistent clarity metric, including: A frequency domain weighting function is constructed based on the synthesized modulation transfer function, which includes optical defocus variance and thermal diffusion equivalent time. The amplitude spectrum is calculated for the basic calibration image sequence, and a frequency domain weighted integral is performed using a weighting function to obtain the original physical consistent sharpness; Obtain a set of normalization transformers from the physical parameter estimation and normalization modeling steps, and call them to scale the original physical consistency sharpness.

2. The method according to claim 1, characterized in that, Perform online self-calibration, including: Perform a short-window micro-scan around the current focus to acquire a local single-peak enhanced sharpness sampling sequence; The unimodality index of a local sampling sequence is calculated by determining the minimum value of the second derivative of the logarithmic curve of the sequence and combining it with the ratio of the secondary peak height to the primary peak height. With the goal of maximizing the unimodality index, the parameters of the unimodal enhancement transform or the weights of the physical gated fusion are updated online. Based on the updated parameters or weights, the final single-peak enhanced sharpness curve and real-time sharpness value are generated.

3. The method according to claim 1, characterized in that, Construct a multi-scale tensor clarity index for heat kernels, including: The set of heat core scales is determined based on the scene description vector; A three-dimensional tensor is constructed by convolving the basic calibration image sequence with a heat kernel scale set. Perform tensor decomposition on the three-dimensional tensor to extract directional energy singularities and scale modulus factor matrices; Based on the directional energy singularity and scale modulus factor matrix, the directional energy concentration and scale energy concentration are calculated to form the clarity of the original thermonuclear multiscale tensor. Obtain a set of normalization transformers from the physical parameter estimation and normalization modeling steps, and call them to scale the clarity of the original thermonuclear multiscale tensor.

4. The method according to claim 1, characterized in that, The method also includes: By utilizing a single-peak sharpness enhancement curve and real-time sharpness values, a focus search and control closed loop is executed to determine the optimal focus position; The focus search and control closed loop includes performing three-point fitting, performing sidelobe detection, and generating a focus drive sequence; The sidelobe detection process includes: During the focus search process, the presence of sidelobe risk is determined based on the inconsistency between the ratio of secondary peaks to primary peaks or curvature. In response to sidelobe risk, an increment of the sidelobe suppression parameter is generated; The sidelobe suppression parameter increments are written back to the multi-dimensional complementary sharpness construction and unimodal enhancement steps to dynamically change the shape of the unimodal enhanced sharpness curve in subsequent sharpness calculations.

5. The method according to claim 4, characterized in that, Perform a three-point fit, including: Collect real-time sharpness values ​​corresponding to the focus position within the search range; A parabola is fitted between the focus position and the real-time sharpness value to obtain the candidate focus position and local curvature estimate; Based on local curvature estimation, a step size suggestion for the next movement is set to achieve curvature adaptive step size.

6. The method according to claim 5, characterized in that, Generate the focus drive sequence, including: Obtain the lens mechanism model from the device parameter set; Based on candidate focus positions or step size suggestions, trajectory planning with jerk constraints is used to generate an initial driving sequence; By combining the lens mechanism model, reverse clearance compensation and feedforward friction compensation are performed on the initial drive sequence to generate the final focus drive sequence.

7. The method according to claim 1, characterized in that, Performing physical parameter estimation and normalization modeling includes: Estimate the noise variance, automatic gain function and gain energy factor, and bit depth; Estimate the scene equivalent temperature corresponding to the environmental observation data; Construct a closed-loop normalized mapping to generate a set of normalized clarity sub-indices, and aggregate scene equivalent temperatures to generate scene description vectors.

8. The method according to claim 1, characterized in that, Perform data acquisition and basic calibration, including: Adaptive scene modeling based on low-rank sparse decomposition is employed to perform response uniformization and black level correction in order to avoid reliance on mechanical shutter. Dynamic bad pixels are detected through multi-scale consistency and time consistency. Directional interpolation guided by structural tensor is used to repair dynamic bad pixels, thereby protecting edge structures and generating a basic calibration image sequence.

Citation Information

Patent Citations

  • Focusing method and device, electronic equipment and storage medium

    CN111432125A