Image sharpness evaluation method

By generating a multi-scale baseline sharpness response set and an information-physical sharpness kernel, and combining failure posterior probability and gating weights, the problem of image sharpness assessment being susceptible to interference in complex scenarios in existing technologies is solved, and a more robust and accurate sharpness assessment is achieved.

CN121147229BActive Publication Date: 2026-02-03NANJING ARTIFICIAL INTELLIGENCE CHIPS RES INST OF AUTOMATION CHINESE ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511697771.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-19
Publication Date
2026-02-03
Estimated Expiration
2045-11-19

AI Technical Summary

Technical Problem

Existing image sharpness assessment methods have a superficial fusion mechanism in complex scenes, making them susceptible to interference. They lack a dynamic quantification and discrimination mechanism, which makes the assessment results easily affected by false peaks, and they cannot know their own reliability. Furthermore, the complementarity of the algorithms is not fully utilized.

Method used

By acquiring the original image and optical modulation transfer function, a multi-scale baseline sharpness response set, information-physical sharpness kernel, uncertainty measure and evidence vector are generated. Based on the failure posterior probability and gating weight, failure posterior coupling processing is performed to generate the final sharpness assessment result.

Benefits of technology

It improves the robustness of evaluation in complex scenarios and provides more robust, accurate and interpretable clarity evaluation results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121147229B_ABST
    Figure CN121147229B_ABST
Patent Text Reader

Abstract

The application discloses an image sharpness evaluation method, comprising: obtaining an original image and an optical modulation transfer function; generating a multi-scale baseline sharpness response set, an information-physical sharpness kernel, an uncertainty measure and an evidence vector based on the original image and the optical modulation transfer function; estimating a failure posterior probability based on the evidence vector; generating a gating weight based on the failure posterior probability and the uncertainty measure; applying the failure posterior probability and the gating weight to perform a failure posterior coupling process on the multi-scale baseline sharpness response set and the information-physical sharpness kernel to generate a final sharpness evaluation result. The coupling process preferably includes applying an analytical bias correction and performing a cross-scale band selection. The application improves the evaluation robustness in a complex scene through a failure posterior coupling mechanism.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of computational imaging and machine vision, and in particular to a method for evaluating image sharpness. Background Technology

[0002] Image sharpness assessment is a key technology in computational imaging and machine vision. It plays an important role in applications such as autofocus, video surveillance diagnostics, and remote sensing imaging, and is fundamental to ensuring the performance of downstream tasks and improving system intelligence.

[0003] Currently, sharpness assessment methods mainly revolve around spatial gradients (such as Tenengrad and SMD), frequency domain energy (such as Laplacian and Vollath F4), and statistical features. To improve adaptability, existing research also employs a fusion strategy that applies fixed weights or simple linear weights to multiple algorithms.

[0004] Current technologies still face challenges in dealing with complex imaging scenarios, including superficial fusion mechanisms and susceptibility to interference. Specifically, fusion strategies often rely on fixed weights and lack dynamic quantitative discrimination mechanisms for the failure states of individual algorithms. This makes the fusion results susceptible to spurious peaks (such as noise-induced peaks) from individual algorithms. Such empirical indicators cannot distinguish between identifiable structures and high-frequency noise at the physical level, nor can they generate evaluation uncertainty, leaving the system unable to recognize the reliability of its evaluation. Furthermore, algorithm complementarity often remains at the level of weight reduction, lacking mechanisms to use scene evidence (such as noise intensity) to correct analytical biases in failed algorithms. Summary of the Invention

[0005] The purpose of this invention is to provide an image sharpness evaluation method to solve the aforementioned problems existing in the prior art.

[0006] According to one aspect of this application, a method for evaluating image sharpness includes:

[0007] Obtain the original image and optical modulation transfer function;

[0008] Based on the original image and optical modulation transfer function, a multi-scale baseline sharpness response set, information-physical sharpness kernel, uncertainty measure and evidence vector are generated.

[0009] Based on the evidence vector, estimate the posterior probability of failure;

[0010] Gating weights are generated based on the posterior probability of failure and uncertainty measure;

[0011] By applying failure posterior probability and gating weights, failure posterior coupling processing is performed on the multi-scale baseline sharpness response set and the information-physical sharpness kernel to generate the final sharpness assessment result.

[0012] Optionally, based on the evidence vector, the posterior probability of failure is estimated, including:

[0013] For the evidence vectors of the evaluation target and the pre-stored reference target, perform robust standardization and outlier pruning.

[0014] By applying self-supervised learning for ranking, the estimation model is optimized by minimizing the ranking loss function, which is used to maximize the margin between the non-failure probability of the reference target and the non-failure probability of the target to be evaluated.

[0015] In the estimation model training, domain adaptive regularization is combined to calculate and output the failure posterior probability.

[0016] Optionally, the evidence vector includes:

[0017] Texture density, noise intensity, spectral slope, motion indication, and two-domain consistent generalized likelihood ratio statistics.

[0018] Optionally, gating weights are generated based on the posterior probability of failure and uncertainty measure, including:

[0019] A non-failure metric for calculating the posterior probability of failure;

[0020] Calculate the confidence level of the exponential decay of uncertainty measure under temperature parameter control;

[0021] The non-failure metric is multiplied by the exponential decay confidence level and then normalized to produce the gate control weight.

[0022] Optionally, by applying the posterior probability of failure and gating weights, posterior failure coupling processing is performed on the multi-scale baseline sharpness response set and the information-physical sharpness kernel, including:

[0023] Monitor the posterior probability of failure;

[0024] When the posterior probability of failure indicates a high risk of failure, analytical bias correction is applied to the corresponding response in the multi-scale baseline sharpness response set to produce a corrected baseline sharpness response.

[0025] By applying gating weights, the baseline sharpness response after correction and the information-physical sharpness kernel are fused to generate the final sharpness assessment result.

[0026] Optionally, a resolution bias correction is applied to produce a corrected baseline sharpness response, including:

[0027] When the corresponding response is a Laplacian sharpness response, noise intensity is extracted from the evidence vector;

[0028] By applying noise intensity, the noise-induced overestimation term is subtracted from the Laplacian sharpness response to produce the corrected baseline sharpness response.

[0029] Optionally, applying resolution bias correction also includes:

[0030] Extract texture density from evidence vectors;

[0031] By applying texture density to the Laplacian sharpness response after noise-induced overestimation has been removed, false high responses caused by overly dense textures are suppressed, resulting in a corrected baseline sharpness response.

[0032] Optionally, generate information-physical sharpness kernel and uncertainty measure, including:

[0033] Fisher information is calculated based on the optical modulation transfer function, combined with the power spectrum of the observable signal and the power spectrum of noise derived from the original image.

[0034] Based on Fisher information, the information-physical clarity kernel and the Cramer-Rao lower bound as an uncertainty measure are derived respectively.

[0035] Optionally, based on the optical modulation transfer function, and combining the power spectrum of the observable signal derived from the original image with the power spectrum of the noise, the Fisher information is calculated, including:

[0036] Determine the sensitivity of the optical modulation transfer function to the defocus parameter;

[0037] Fisher information is obtained by integrating or summing the square of the sensitivity with the ratio of the power spectrum of the observable signal and the power spectrum of the noise in the frequency domain.

[0038] Optionally, the method further includes:

[0039] Extract the bandpass energy in the frequency domain and the phase consistency in the spatial domain from the original image;

[0040] Construct joint likelihood functions for the structural and noise assumptions, respectively;

[0041] The ratio of joint likelihood functions is calculated to generate a two-domain consistent generalized likelihood ratio statistic, which is then used as one of the components of the evidence vector to estimate the posterior probability of failure.

[0042] Beneficial effect: Through the above technical solutions, the present invention improves the evaluation robustness in complex scenarios. Attached Figure Description

[0043] Figure 1 This is a schematic diagram of the overall framework of an image sharpness assessment method.

[0044] Figure 2This is a flowchart illustrating the process of generating gating weights based on the posterior probability of failure and uncertainty measurement.

[0045] Figure 3 This is a schematic diagram of the calculation process for Fisher's information content I.

[0046] Figure 4 This is a schematic diagram of the statistical calculation process for the two-domain consistent generalized likelihood ratio. Detailed Implementation

[0047] For ease of understanding, some symbols in this specification are defined as follows. Unless otherwise stated, the following symbols have the following meanings:

[0048] S _L (x,y) is the local sharpness response based on the brightness gradient or Laplacian operator, used to characterize the high-frequency texture intensity at pixel (x,y).

[0049] S _v (x,y) is the sharpness response based on local variance or energy statistics, used to characterize local contrast.

[0050] S _s (x,y) represents the overall sharpness response at scale s, which can be expressed by S _L S _v Other scale features are obtained by combining them according to preset weights.

[0051] S _info (x,y) represents the sharpness value given by the information-physical sharpness kernel, which is derived from the image power spectrum and optical modulation transfer function based on Fisher information and Cramer-Rao lower bound.

[0052] MTF(f;θ) represents the optical modulation transfer function, which indicates the modulation attenuation characteristics at spatial frequency f and defocus parameter θ.

[0053] I _sig (f), I _noise (f) represents the observable signal power spectrum and noise power spectrum at frequency f, respectively.

[0054] J _s The Fisher information content corresponding to scale s can be obtained by integrating the optical modulation transfer function sensitivity and signal-to-noise ratio in the frequency domain.

[0055] CRLB _s Let J represent the Cramer-Rao lower bound corresponding to scale s, and be the Fisher information quantity. _s The function is used as a measure of uncertainty.

[0056] φ、φ _p φ_q This represents an evidence vector composed of features such as multi-scale sharpness response, information-physical sharpness kernel, generalized likelihood ratio statistics, and scene parameters; where φ _p φ _q This represents a pair of evidence vectors that have a ranking relationship.

[0057] F _post,i,s This represents the posterior probability of failure output by the estimation model under scenario i and scale s, used to characterize the possibility of response failure or distortion at the corresponding scale.

[0058] C _fused (x,y) is the fusion confidence map, representing the overall confidence of the multi-scale sharpness response and the information-physical kernel after failure a posteriori weighted fusion at pixel (x,y).

[0059] R _conf_s The fusion confidence level corresponding to scale s can be understood as C. _fused The confidence level of the components in the scale dimension or the aggregate confidence level of scale s.

[0060] S _hat_t,s (x,y) represents the prior sharpness estimate at time t and scale s, which can be obtained by scaling the final sharpness map or information-physical sharpness kernel at the previous time step and is used as a reference target in spatiotemporal regularization.

[0061] S _t,s (x,y) represents the observation sharpness response calculated from the current frame image at time t and scale s.

[0062] IFT(x,y) represents the information-gradient tensor, which is used to comprehensively characterize the coupling between local information content and gradient direction.

[0063] C represents the scene complexity index, which can be obtained by combining information such as texture density, motion intensity, brightness dynamic range, and failure posterior statistics.

[0064] τ represents the temperature parameter, used to adjust the steepness of the failure posterior or scene complexity mapping.

[0065] γ represents the time smoothing coefficient, which is used to control the first-order or higher-order smoothing intensity of the sharpness map over time.

[0066] ρ _min ρ _max This represents the lower and upper bounds of the normalization interval for scene complexity, used to compress the complexity of different scenes into a unified range.

[0067] S _final(x,y) represents the final sharpness map after multi-scale fusion, failure correction and spatiotemporal regularization, which can be used as the basis for autofocus control or image quality assessment.

[0068] Example 1 provides an image sharpness evaluation method based on multi-algorithm weighted fusion, which can serve as the background or foundation for subsequent embodiments of the present invention. This method is used to improve the robustness of evaluation for specific scenes (such as infrared or low-light images) by fusing multiple complementary algorithms.

[0069] The image preprocessing steps involve acquiring real-time image frames from the camera, and if the acquired image is color, converting it to grayscale. To preserve image details, especially edge information in infrared images, while removing noise interference, a light denoising process can be performed on the image, such as using a Gaussian filter (e.g., 3×3 kernel, σ=0.3).

[0070] Next, a multi-algorithm sharpness calculation step is performed. Three complementary sharpness evaluation algorithms are applied to the preprocessed image:

[0071] (1) Laplacian Variance Algorithm: This algorithm reflects the severity of edge changes. Specifically, the second derivative L of the image is calculated using a Laplacian operator (e.g., a 3×3 kernel), and the variance Var(L) of the Laplacian transform result is calculated. The final Laplacian variance sharpness S _Laplacian It can be represented as S _Laplacian =Var(L) / k1, where k1 is a scaling factor (e.g., a suggested value of 100) used to normalize the results to a reasonable range.

[0072] (2) Vollath F4 autocorrelation algorithm: This algorithm evaluates the texture sufficiency of an image by measuring the autocorrelation difference between adjacent pixels and pixels in between, and is particularly suitable for sharpness evaluation of infrared images. The specific calculation method is as follows: calculate the autocorrelation difference F4 in the horizontal direction. _h =Σ[I(x,y)×I(x+1,y)]-Σ[I(x,y)×I(x+2,y)]; Calculate the autocorrelation difference F4 in the vertical direction. _v =Σ[I(x, y)×I(x, y+1)]-Σ[I(x, y)×I(x, y+2)]; x and y are pixel coordinates.

[0073] Final Vollath F4 resolution: S _vollath =(F4 _h +F4 _v ) / (2×width×height), where width is the width of the image and height is the height of the image.

[0074] (3) SMD Gray-Level Difference Algorithm: This algorithm reflects the sufficiency of image details by calculating the sum of gray-level differences between adjacent pixels. The specific calculation method is as follows: calculate the sum of horizontal differences SMD. _h =Σ|I(x+1, y)-I(x, y)| and the sum of the differences in the vertical direction SMD _v =Σ|I(x,y+1)-I(x,y)|;The final SMD resolution S _smd =(SMD _h +SMD _v ) / k2, where k2 is the scaling factor, which can be adjusted according to the image resolution.

[0075] Perform a weighted fusion calculation to determine the overall sharpness. The results of the three algorithms described above are weighted and fused to obtain the overall sharpness score S. The fusion process can be represented as: S = w1 × S _Laplacian +w2×S _vollath +w3×S _smd The weighting coefficients w1, w2, and w3 are determined based on the algorithm characteristics and the actual scenario. For example, in a preferred configuration, w1 = 0.3 (Laplacian variance, good at edge detection), w2 = 0.5 (Vollath F4, highly sensitive to infrared images), and w3 = 0.2 (SMD grayscale difference, reflecting pixel-level details). The principle of this weighting configuration is to ensure a relatively balanced weighting of the three indicators, avoiding the dominance of a single indicator, while giving slightly higher weights to algorithms such as Volath F4 according to the needs of specific scenarios (such as infrared or low contrast).

[0076] Furthermore, to improve the stability of sharpness assessment, a time-series smoothing optimization step can be performed to effectively suppress single-frame noise interference and make the sharpness curve smoother. One specific implementation is to introduce a sliding window averaging algorithm: the currently calculated sharpness score S(t) is added to a historical record queue; then, a fixed-length historical window is maintained (e.g., the most recent 5-10 frames); finally, the average of all sharpness values ​​within this window is calculated as the smoothed sharpness S. _smooth =ΣS(ti) / N, where N is the window size and i ranges from 0 to N-1; when the history exceeds the window size, remove the oldest record to keep the queue length constant.

[0077] This embodiment also includes a sharpness feedback and autofocus assistance step. This step is used to determine image quality based on the smoothed sharpness score and trigger autofocus. Specifically, it may include: recording the historical best sharpness S. _best and the corresponding focal length position F _best ; when the smoothed clarity S _smooth Significantly lower than the best historical resolution (S) _bestWhen the sharpness decreases by more than 20%, the autofocus adjustment is triggered. The focus direction is determined based on the sharpness change trend ΔS=S(t)-S(t-1). For example, if ΔS>0, the current direction is continued, and if ΔS<0, the focus is reversed. The focus length is gradually adjusted using optimization strategies such as hill climbing algorithm, and the sharpness value after each adjustment is recorded. When the sharpness reaches a local maximum or there is no significant improvement after multiple consecutive adjustments, the adjustment is stopped, and the optimal focus length position is saved.

[0078] Example 2: A detailed description of the overall framework of the image sharpness evaluation method proposed in this invention, such as... Figure 1 As shown, this method is designed to overcome the failure problem of traditional methods in complex scenarios (such as noise, weak texture, fixed patterns, etc.). By introducing a physical model, failure posterior estimation and adaptive coupling mechanism, it provides more robust, accurate and interpretable sharpness assessment results.

[0079] Step 201: Obtain the original image and optical modulation transfer function.

[0080] In this embodiment, the original image refers to real-time or offline image data acquired by an image acquisition device (such as a visible light camera, an infrared thermal imaging camera, etc.). The optical modulation transfer function (MTF) is a key physical parameter describing the imaging performance of an optical system, characterizing the system's ability to transfer different spatial frequencies (i.e., details). There are several ways to obtain the MTF: for example, it can be obtained by acquiring camera and lens metadata and parsing it; or, in some implementations, it can be obtained from pre-acquired image data through blind estimation or calibration methods. Obtaining the MTF provides crucial physical priors for subsequent calculations of the information-physical sharpness kernel.

[0081] Step 202: Based on the original image and optical modulation transfer function, generate a multi-scale baseline sharpness response set, an information-physical sharpness kernel, an uncertainty measure, and an evidence vector.

[0082] This step is a composite step used to extract important information from the input data needed for subsequent coupled decision-making.

[0083] Generating a multi-scale baseline sharpness response set involves decomposing the original image at multiple scales, such as generating a multi-scale image pyramid. At multiple scales (or resolutions) of the pyramid, the response values ​​of various traditional, complementary sharpness algorithms are calculated. These baseline algorithms preferably include, but are not limited to: Laplacian sharpness response (e.g., Laplacian variance), Vollath F4 sharpness response (autocorrelation), and SMD sharpness response (grayscale difference). Through multi-scale computation, sharpness information at different spatial frequencies can be captured to address textures at different scales.

[0084] The generation of the information-physical sharpness kernel refers to the introduction of a sharpness definition based on physical optics models and information theory. In this invention, sharpness is defined as the discernibility of defocus parameters. Specifically, the generation of this kernel (detailed in subsequent embodiments) is based on the optical modulation transfer function (MTF), combined with the power spectrum of the observable signal derived from the original image and the power spectrum of the noise, to calculate the Fisher information content. For example, the information-physical sharpness kernel S... _info It can be expressed as a function of Fisher information I, such as S _info =log(1+I) represents the amount of information that the current image state (determined by MTF, signal, and noise) can theoretically provide to distinguish the degree of defocus.

[0085] Generating an uncertainty metric refers to simultaneously producing a metric for quantifying the reliability of the assessment results during the calculation of the information-physical sharpness kernel. In this embodiment, the uncertainty metric is preferably the Cramer-Rao lower bound (CRLB) corresponding to the Fisher information quantity I. CRLB is the theoretical lower bound for estimating the variance of a parameter (here, the defocus parameter). The smaller the CRLB, the higher the discernibility and the lower the uncertainty; conversely, the larger the CRLB, the higher the uncertainty and the lower the reliability. For example, CRLB can be calculated as CRLB = 1 / (I + ε), where ε is a small constant to prevent the denominator from being zero.

[0086] Generating an evidence vector refers to extracting features that characterize the physical and statistical properties of the current image scene. These features serve as the basis for subsequent determinations of validity. As detailed in subsequent embodiments, the evidence vector is preferably extracted from the multi-scale representation and initial analysis of the original image. Specific content may include (but is not limited to): texture density TD, uniformity U, and noise intensity σ. n Or noise estimation, spectral slope β, motion indication, and discriminative statistics obtained in the process of generating the above metrics, such as the two-domain consistent generalized likelihood ratio statistic GLR, Fisher information I, and Cramer-Rao lower bound CRLB itself.

[0087] Step 203: Estimate the posterior probability of failure based on the evidence vector.

[0088] This step is one of the mechanisms by which this invention achieves failure coupling. Traditional baseline sharpness algorithms (such as Laplacian, Vollath, SMD) fail in specific scenarios (such as high noise, uniform regions, and overly dense textures), i.e., producing erroneous (pseudo-high or pseudo-low) sharpness responses. This step is used to explicitly model a failure latent variable for each baseline algorithm at each scale, and calculate the posterior probability of the algorithm being in a failed state at that scale based on the evidence vector Z extracted in step 202, denoted as P(F)._post =1|Z).

[0089] In other words, this step replaces the traditional approach of arbitrary weighting or fixed rules of experience. Instead, it uses a learnable, physico-statistically evidence-driven model to quantitatively estimate the unreliability of each algorithm in the current scenario. For example, when the evidence vector indicates high noise intensity and low texture density, the model will automatically (through learned parameters) increase the posterior probability of failure for the Laplacian algorithm. This estimation process will be detailed in subsequent embodiments.

[0090] Step 204: Generate gating weights based on the posterior probability of failure and uncertainty measure.

[0091] This step is used to generate the weights needed for fusion. The generation of these weights takes into account information from two dimensions:

[0092] Unreliability of the algorithm: the posterior probability of failure P(F) estimated in step 203 _post (Provided by) The higher the posterior probability of algorithm failure, the lower its corresponding weight should be.

[0093] Physical uncertainty: Provided by the uncertainty measure (preferably Cramer-Rao lower bound CRLB) generated in step 202. The higher the CRLB of a scale or algorithm (the higher the uncertainty), the lower its corresponding weight should be.

[0094] In this embodiment, the gating weight w _i,s (Corresponding to the i-th algorithm, s-th scale) Preferably, it is related to the non-failure probability (1-P(F)). _post The confidence level is directly proportional to the uncertainty-based confidence level. For example, the confidence level can be expressed as an exponential decay form of CRLB under the control of the temperature parameter τ, such as softmax(-CRLB). _s / τ) or exp(-CRLB / τ). The final gating weight is the product of the two and then normalized, i.e., w _i,s ∝(1-P(F _post ,i,s))×exp(-CRLB _s / τ). This weighting method ensures that only responses that are both error-free and have low physical uncertainty can dominate in the final fusion.

[0095] Step 205: Apply the failure posterior probability and gating weights to perform failure posterior coupling processing on the multi-scale baseline sharpness response set and the information-physical sharpness kernel to generate the final sharpness assessment result.

[0096] This step utilizes the decision information generated in steps 203 and 204 to intelligently couple and fuse the raw sharpness response generated in step 202. The coupling process preferably includes the following mechanisms:

[0097] Analytical bias correction (correction): This mechanism is triggered by the posterior probability of failure. When the posterior probability of failure in step 203 is P(F... _post When a baseline response (such as the Laplacian) is indicated to be at high failure risk, the system does not simply reduce its weight (achieved through gating weights), but instead applies an analytical bias correction. For example, if the failure is caused by noise-induced overestimation, the system extracts the noise intensity σn from the evidence vector and from the Laplacian response S. _L Excluding overestimated items, such as S _L =max(0, S) _L -κ·σn 2 ); where κ is the correction coefficient. The correction mechanism (which will be detailed in subsequent embodiments) embodies the complementarity between algorithms, that is, a strong algorithm (or evidence) helps a weak algorithm that has failed, rather than simply replacing the head.

[0098] Cross-scale band selection: This mechanism utilizes all information to dynamically select the optimal subset of scales for fusion. It prioritizes scales with high optical sensitivity (e.g., large optical modulation transfer function (MTF) derivative), low uncertainty (small Cramer-Rao lower bound (CRLB)), and low posterior probability of failure (P(F)). _post (Small) scale. If all scales fail, bandwidth migration or redundancy fallback strategies may be triggered. This mechanism ensures that fusion always occurs on the most reliable set of scales.

[0099] Adaptive re-fusion: After (optionally) performing bias correction and band selection, the system applies the gating weights w generated in step 204 to the biased baseline sharpness response set and the information-physical sharpness kernel S. _info Perform weighted fusion. The preferred fusion method here is adaptive, for example, the information-physical clarity kernel S. _info The fusion ratio ρ* can be dynamically determined based on the scene complexity C, such as ρ*=clamp(ρ _min +(ρ _max -ρ _min )·C γ ); where ρ _min and ρ _max These represent the upper and lower limits of the proportion, and γ is the sensitivity coefficient. The final fusion response S _fused_fused =ρ*·S _info +(1-ρ*)·Σw _i ·S _i_corr Among them, w _i It is the gating weight, S_i_corr It is the baseline sharpness response after correction.

[0100] Temporal robustness: To obtain smooth and robust results, the fused response S _fused_fused It will also undergo time-series robustness processing, including time-axis smoothing based on fusion confidence (e.g., freezing the update of the final sharpness map S if the motion indication is high). _final (t)=γ·S _final (t-1)+(1-γ)·S _fused_fused (t) and rollback protection for abnormal jumps.

[0101] This coupled processing is a systematic mechanism stack that organically combines physical kernel, uncertainty, baseline response and scene evidence through a post-failure a posteriori to generate a final sharpness assessment result including a final sharpness map, a global sharpness score, confidence intervals and (preferably) diagnostic evidence.

[0102] Example 3 provides the process for generating the information-physical clarity kernel, uncertainty, and evidence vector, and describes in detail how to generate all the input signals necessary for the subsequent failure coupling mechanism based on the original image and optical modulation transfer function.

[0103] Step 301: Extract initial scene features.

[0104] This step serves as the scene prior input for all subsequent calculations and is preferably performed immediately after acquiring the original image. The initial scene features are a vector describing the physical and statistical properties of the image. In a preferred embodiment, this feature vector specifically includes:

[0105] Noise estimation or noise intensity (e.g., σn): can be obtained by analyzing flat regions in an image or by using a blind noise estimator, and is used to quantify the noise level of an image.

[0106] Phase consistency (PC): A spatial domain metric used to characterize the saliency of structural information (such as edges and corners) in an image, and is insensitive to changes in illumination.

[0107] Texture density TD: Used to quantify the sufficiency of texture in local regions of an image, for example, by calculating local gradient variance or entropy.

[0108] Spectral slope β: By analyzing the attenuation characteristics of the Fourier spectrum of the image (e.g., fitting 1 / f) β This is obtained to distinguish natural textures (β close to 1) from noise or artificial structures.

[0109] Motion indication (MI): For image sequences, it can be estimated using inter-frame differencing or optical flow methods to indicate whether motion exists in the scene.

[0110] The initial scene features mentioned above are output for use in subsequent steps (such as noise power spectrum estimation, evidence vector construction, scene adaptive parameter generation, etc.).

[0111] Step 302: Define the frequency band and prepare the bandpass energy.

[0112] This step is performed in the frequency domain. A multi-scale image pyramid is obtained. For each scale of the pyramid (i.e., an image at a specific resolution), one or more frequency bands of interest, B, are defined. The bandpass energy E(f) of the image at that scale in each frequency band B is calculated, where f represents the spatial frequency. Simultaneously, the phase coherence PC(f) of the corresponding frequency band is calculated. The above frequency domain metrics (frequency band set, E(f), PC(f)) are output for subsequent physical kernel calculations.

[0113] Step 303: Estimate the noise power spectrum.

[0114] Based on the initial scene features extracted in step 301 (especially noise estimation), and combined with the frequency band set and bandpass energy produced in step 302, the noise power spectrum S of each frequency band is calculated. _N (f) and the global noise intensity σn are estimated. _N (f) is the key denominator term for subsequent calculations of the signal-to-noise ratio.

[0115] Step 304: Calculate the information-physical clarity kernel, Fisher information, and Cramer-Rao lower bound.

[0116] This step is used to improve clarity from empirical measures (such as gradients) to physical discernibility.

[0117] Fisher information is calculated based on the optical modulation transfer function, combined with the power spectrum of the observable signal derived from the original image and the power spectrum of the noise.

[0118] In this invention, Fisher information (I) is used to quantitatively describe the total amount of identifiable information that the current image (defined by its signal and noise characteristics) can provide for an unknown parameter (i.e., the defocus parameter b). The Cramer-Rao lower bound (CRLB) is the reciprocal of Fisher information (CRLB = 1 / (I + ε)), representing the theoretical minimum of the estimation variance that can be achieved when estimating the defocus parameter b. Therefore, CRLB is used as an uncertainty measure in this invention: the smaller the CRLB, the lower the uncertainty and the higher the confidence level.

[0119] like Figure 3 As shown, the calculation process of Fisher information I preferably includes:

[0120] Determining the sensitivity of the optical modulation transfer function (MTF) to the defocus parameter is a crucial process connecting physical optics and information content. This sensitivity can be specifically expressed as the value of the partial derivative of the MTF(H) with respect to the defocus parameter b at b=0 (i.e., at the focal point). For example, its form can be derived as (dH... _b / db)| b=0 =-4*π 2 *f 2 *MTF(f).

[0121] Calculate the power spectrum S of the observable signal _X (f). Due to the real S _X (f) Unknown, the present invention preferably uses the bandpass energy E(f) and phase coherence PC(f) from step 302 to jointly approximate, for example S _X (f) = a*E(f)*(1+b*PC(f)), where a and b are coefficients.

[0122] Based on the square of the sensitivity, and the power spectrum of the observable signal S _X (f) and noise power spectrum S _N The ratio (f) (i.e., signal-to-noise ratio) is integrated or summed in the frequency domain to obtain the Fisher information content. To improve robustness, a frequency weight W is preferably introduced. _f Weight W _f =normalize((PC(f)*E(f)) / (S _N (f)+ε)) is used to enhance the frequency with both high energy and phase consistency.

[0123] Fisher's information content I = Σ _f [((dH _b / db)| b=0 ) 2 *(S _X (f) / (S _N (f)+ε))*W _f ].

[0124] Based on Fisher information I, the information-physical resolution kernel S is derived respectively. _info Compared to the Cramer-Rao lower bound (CRLB), which is a measure of uncertainty. _info Preferably defined as a monotonically increasing function of I, such as S _info =log(1+I) to balance its dynamic range. CRLB is calculated as before, as CRLB=1 / (I+ε).

[0125] In some alternative implementations, when other evidence (such as the GLR in step 306) indicates that the current region is of low confidence (e.g., low GLR and spectral slope β close to 1, indicating uniform noise), a suppression gate can be applied to the calculated Fisher information I and an upper limit truncation can be set on CRLB to prevent false high identifiability or low uncertainty from being generated in pure noise regions.

[0126] In step 304, the preferred implementation of the description regarding the calculation of the physical sharpness kernel, Fisher information content, and Cramer-Rao lower bound is as follows:

[0127] Before calculating the Fisher information content I, construct the frequency weights W. _f Weight W _f When integrating (summing) Fisher information in the frequency domain, priority should be given to frequencies that possess both high energy and high phase consistency and are contributed by the real structure, while suppressing the contribution of noise or artifact frequencies. Therefore, the frequency weight W... _f Preferably, it is constructed as follows: W _f =normalize((PC(f)*E(f)) / (S _N (f)+ε)); where PC(f) and E(f) are the phase coherence and bandpass energy at frequency point f (from step 302), respectively, and S _N (f) is the noise power spectrum at that frequency (from step 303), and ε is a small constant to prevent the denominator from being zero.

[0128] Another important step in calculating Fisher's information content I is to approximate the power spectrum S of the observable signal. _X (f). Due to the true S _X (f) Unknown. Preferably, this invention does not directly use the bandpass energy E(f) (which may contain noise or non-structural energy), but instead employs a joint approximation using E(f) and phase coherence PC(f). This approximation is superior to using E(f) alone because it utilizes PC(f) as a priori information indicating the structure. The preferred formula for this joint approximation is: S _X (f)=a*E(f)*(1+b*PC(f)); where the coefficients a and b can be calibrated and determined by methods such as energy conservation constraints or quantile fitting.

[0129] Step 305: Calculate the multi-scale response of the baseline algorithm.

[0130] These steps are performed in parallel. A multi-scale image pyramid is acquired, and at each scale, the response values ​​of several traditional, mechanistically complementary sharpness assessment algorithms are calculated. These algorithms preferably constitute a multi-scale baseline sharpness response set. For example:

[0131] Laplacian Clarity Response S _L For example, calculating the variance of the Laplacian operator response is sensitive to edges.

[0132] Vollath F4 Clarity Response S _v Based on autocorrelation algorithms, it is sensitive to texture in low-contrast images such as infrared.

[0133] SMD Clarity Response S _s The algorithm, based on grayscale difference, reflects pixel-level details. After calculation, these responses undergo uniform boundary processing and scale normalization for subsequent fusion.

[0134] Step 306: Calculate the two-domain consistent generalized likelihood ratio statistics.

[0135] This step is used to provide structural evidence, distinguishing real textures from noise / fixed patterns.

[0136] Dual-domain consistency refers to the statistical consistency between the spatial domain (represented by phase consistency PC) and the frequency domain (represented by bandpass energy E). Real structures (such as edges) typically possess both high energy and high phase consistency, while noise or fixed patterns do not.

[0137] like Figure 4 As shown, the calculation process preferably includes:

[0138] Extract the bandpass energy E in the frequency domain and the phase coherence PC in the spatial domain from the original image (the result of step 302 can be reused).

[0139] Construct joint likelihood functions for the structural and noise hypotheses, respectively. Specifically, construct two hypotheses: the structural hypothesis H1, i.e., E and PC are highly correlated; and the noise hypothesis H0, i.e., E and PC are uncorrelated. Fit the likelihood functions L(·|H1) and L(·|H0) for the joint distribution of E and PC under these two hypotheses, respectively.

[0140] The ratio of joint likelihood functions is calculated to generate the two-domain consistent generalized likelihood ratio statistic GLR = L(·|H1) / L(·|H0). A higher GLR value indicates a greater probability that the current region represents a true structure. A lower GLR value suggests noise, a uniform region, or a fixed pattern / stripes. Therefore, GLR is a good indicator that can be used for subsequent failure assessment.

[0141] In some alternative implementations, the repression factor g=σ(α) can also be derived based on GLR. _g (logGLR-τ _g It identifies scales with fixed pattern and stripe risks and outputs this information (GLR, g, risk markers) together.

[0142] In step 306, the preferred implementation of the description regarding the calculation of the two-domain consistent generalized likelihood ratio (GLR), which includes details of the underlying model, is as follows:

[0143] The step of constructing the joint likelihood function specifically refers to establishing statistical models for the structural hypothesis H1 and the noise hypothesis H0. The physical meaning of the structural hypothesis H1 is that the observed bandpass energy E and phase consistency PC both originate from real, salient image structures, and E and PC should have a high statistical correlation. The physical meaning of the noise hypothesis H0 is that the observed E and PC are generated only by random noise, a uniform background, or a fixed pattern, and E and PC should not have a correlation, or exhibit distribution characteristics significantly different from H1. In a preferred implementation of this invention, a Gamma-Beta joint distribution model is used to accurately characterize the two hypotheses. Specifically, it is assumed that the energy E follows a Gamma distribution, and the phase consistency PC follows a Beta distribution. A coupling parameter is used to describe the correlation under the H1 hypothesis, while this coupling parameter is zero under the H0 hypothesis. By fitting the joint distribution model, L(·|H1) and L(·|H0) can be calculated, yielding a GLR statistic with strong discriminative power.

[0144] Step 307: Execute the standard unchanged and normalized.

[0145] This step is a preferred engineering step of the present invention to improve the cross-device robustness of the solution. Obtain the normalized information – Physical Sharpness Kernel S – produced in steps 304 and 305. _info and three types of baseline sharpness response (S _L S _v S _s Then, in conjunction with the initial scene features (such as resolution, total energy, etc.) from step 301, cross-resolution normalization is performed.

[0146] Normalization is used to eliminate differences in evaluation scales introduced by variations in camera resolution, lens focal length, or sensor type. For example, all sharpness responses can be normalized to a standard resolution benchmark or a standard energy range. This step ensures that sharpness scores for the same scene are comparable when shot on different devices.

[0147] Step 308: Summarize the output and construct the evidence vector.

[0148] This step summarizes all the outputs of this embodiment. It combines the information from step 304—the physical sharpness kernel S… _info Fisher information content I, Cramer-Rao lower bound CRLB, Laplacian sharpness response S in step 305 _LVollath F4 resolution response S _v SMD Sharpness Response _s The GLR, inhibitory factor g, and risk markers from step 306 are summarized.

[0149] Furthermore, the evidence vector φ is formally constructed. The evidence vector φ is a high-dimensional vector, constructed by assembling it according to scale. For each scale s, the corresponding evidence vector φ... _s Preferably, it includes: the initial scene features (texture density TD, noise intensity σn, spectral slope β, motion indication MI) from step 301; and the key statistics calculated in this embodiment (generalized likelihood ratio statistic GLR, Fisher information I, Cramer-Rao lower bound CRLB). Simultaneously, the baseline sharpness response (S) is included. _L S _v S _s ) is the component to be evaluated in this vector.

[0150] Evidence vector φ along with information-physical clarity kernel S _info and S _L S _v S _s They are output together as the input for performing failure posterior estimation and coupling processing.

[0151] Example 4 provides a specific implementation method for the fusion of Failure Post-hoc Coupled FPC and cross-algorithm correction, and describes in detail the process of using evidence vectors and clarity kernels to perform failure judgment, weight allocation, bias correction and scale selection.

[0152] Step 401: Estimate the posterior probability of failure based on the evidence vector.

[0153] This step is used to quantitatively estimate each baseline algorithm (such as S) based on the evidence vector φ. _L S _v S _s The posterior probability P(F) of failure at each scale s _post =1|φ).

[0154] The estimation process preferably employs an unlabeled ranking self-supervised learning method to avoid large-scale manual annotation. Specific implementation methods include:

[0155] Obtain the evidence vector φ. This vector contains information such as texture, noise, spectral slope, motion, GLR, and CRLB.

[0156] The target to be evaluated (current frame t) _q ) and a pre-stored reference target (e.g., the known sharpest frame t obtained by focal length scanning) _pThe evidence vectors are used to perform robust standardization and outlier pruning to improve the stability of the model.

[0157] Self-supervised learning for ranking is applied, and the estimation model is optimized by minimizing the ranking loss function. The ranking loss function maximizes the margin between the non-failure probability of the reference target and the non-failure probability of the target to be evaluated. The ranking loss function L... _rank Used to maximize the reference target t _p Non-failure probability and the target to be evaluated t _q The margin between the non-failure probabilities. For example, the loss function can be expressed as L _rank =Σmax{0,1-[μ(1-F _post (t _p ))-μ(1-F _post (t _q ))]}。 Among them, F _post (t) represents the failure probability output by the model for target t, and μ is the scoring function. This loss function forces the model to learn a meaningful failure discrimination boundary by distancing the non-failure probabilities of known good and bad samples.

[0158] In the estimation model training, domain adaptive regularization is incorporated to calculate and output the posterior probability of failure. To enable the model to generalize to different types of cameras (e.g., infrared vs. visible light), domain adaptive regularization (e.g., gradient inversion layer or maximum mean difference (MMD) loss) is added during training to ensure that the failure features learned by the model are cross-domain general.

[0159] The parameters α of the model (e.g., a small neural network or logistic regression classifier) ​​are optimized in the self-supervised process described above. For any input evidence vector φ, the model outputs the failure posterior probability, e.g., P(F... _post =1|φ)=σ(α T *φ / T _cal ), where σ is the sigmoid function, T _cal These are the calibration temperature parameters.

[0160] In one specific implementation, the training samples are derived from a combination of infrared autofocus sequences and visible light autofocus sequences, totaling no fewer than 5,000 sets of focus sequences. Each set of sequences contains 20-50 frames of images at different focal lengths, with resolutions of 640×480 or 1280×720. The estimation model is preferably a fully connected neural network with two hidden layers, each containing 32 and 16 hidden units respectively. The activation function is ReLU, and the output layer uses a sigmoid function to output the posterior probability of failure. The Adam optimizer is used during training, with an initial learning rate of 1×10⁻⁶. -3The batch size is 64, and the training rounds are no less than 30. The training process includes: pairing evidence vectors (φ... _p ,φ _q Using this as input, forward propagation yields the corresponding posterior failure probability, and the ranking loss L is calculated. _rank The domain adaptive regularization term is backpropagated and the model parameters α are updated until convergence on the validation set.

[0161] Step 402: Generate gate weights based on the posterior probability of failure and uncertainty measure.

[0162] This step is used to generate an intelligent fusion weight that must also consider whether the algorithm has subjectively failed (F). _post And whether it is objectively credible in terms of physics (Cramer-Rao lower bound CRLB).

[0163] like Figure 2 As shown, its generation process preferably includes:

[0164] The non-failure metric for calculating the posterior probability of failure is (1-P(F)). _post,i )), where P(F _post,i ) is the failure probability estimated by the i-th algorithm in step 401.

[0165] Calculate the exponential decay confidence level of the uncertainty measure (i.e., the Cramer-Rao lower bound CRLB) under the control of the temperature parameter τ. This is expressed as exp(-CRLB / τ). The temperature parameter τ controls the degree to which the uncertainty (CRLB) suppresses the confidence level: the smaller τ is, the stronger the suppression.

[0166] Multiply the non-failure metric by the exponential decay confidence level, then normalize to produce the gate control weight w. _i w _i ∝(1-P(F _post,i ))*exp(-CRLB / τ). Weight w _i A high value will only be obtained if algorithm i is not invalid and its physical basis is reliable.

[0167] Step 403: Apply cross-algorithm parsing bias correction.

[0168] The failure posterior probability P(F) generated by system monitoring step 401 _post When the posterior probability of failure is P(F) _post When indicating a high risk of failure (e.g., P(F) _post (> a certain threshold, such as 0.8), the system will not only pass the gating weight w in step 402 _iTo reduce its fusion ratio, it also actively applies analytical bias correction to the corresponding responses in the multi-scale baseline sharpness response set, in order to produce a corrected baseline sharpness response S. _i_corr .

[0169] Here are some examples of corrective measures for preferred choices:

[0170] Case 1: Targeting Laplacian Sharpness Response S _L Noise-induced overestimation failure.

[0171] Triggering condition: When the corresponding response is a Laplacian sharpness response, P(F) _post,L The noise intensity σn is very high, and the evidence vector φ indicates that the noise intensity σn is very high. 2 Very high.

[0172] Correction process: Extracting noise intensity σn from the evidence vector 2 .

[0173] Applying noise intensity, from the Laplacian sharpness response S _L The noise-induced overestimation term from the resolution is subtracted to produce the corrected baseline sharpness response. For example, the corrected S... _L =max(0, S) _L -κ1*K _L *σn 2 ), where κ1 and K _L It is a correction coefficient that can be obtained by minimizing inconsistency optimization.

[0174] Case 2: Further suppress the false high response caused by overly dense textures in the Laplacian response that has been implemented in Case 1 (with noise subtracted).

[0175] Triggering condition: P(F) _post,L The value is very high, and the evidence vector φ indicates that the texture density TD exceeds a certain threshold τ. _TD .

[0176] Correction process: Extract texture density (TD) from the evidence vector.

[0177] Applying texture density, for the Laplacian sharpness response S after subtracting noise-induced overestimation. _L (From Case 1) Further subtracting the over-dense texture suppression term suppresses the false high response caused by over-dense texture, producing a corrected baseline sharpness response. For example, S _L ''=S _L '-κ3*max(0,TD-τ _TD ), where κ3 is the correction coefficient, which controls the suppression strength and is minimized by the information-physical clarity kernel S. _infoIt is estimated by the degree of local inconsistency.

[0178] Case 3: Targeting Vollath Sharpness Response (S) _v Conservative failure (i.e., low response) in uniform regions.

[0179] Triggering condition: P(F) _post,V The uniformity of the scene is very high, and the evidence vector φ indicates that the scene uniformity U is very high (or the texture density is very low).

[0180] Correction process: Obtain the phase consistency metric (PC) from the evidence vector.

[0181] Applying the phase consistency metric PC to the Vollath sharpness response S _v Compensation is performed for low response in uniform regions. For example, S after correction. _v '=S _v +κ2*(1-U)*PC. Although the phase coherence PC value is also low in the uniform region, its signal-to-noise ratio as a structural indicator is much higher than that of the Vollath algorithm, and it can be used as a reliable compensation source. This compensation is preferably also gated by GLR, that is, it is only performed when GLR indicates a weak structure.

[0182] Step 404: Perform cross-scale band selection.

[0183] Not all scales are equally important or equally reliable; risk must be mitigated by selective fusion, which should only be performed on the optimal subset of scales.

[0184] The band selection process preferably includes:

[0185] The scale sensitivity G is calculated at each scale based on the optical modulation transfer function (MTF). The sensitivity G is preferably calculated from the sensitivity term (dH) in the Fisher information calculation. _b The integral (or summation) over the frequency band B at that scale is obtained by / db, for example, G=Σ _f ∈B(4*π 2 *f 2 *MTF(f)) 2 , where f is the spatial frequency. G quantifies the physical sensitivity of this scale to defocusing.

[0186] Combining scale sensitivity G and failure posterior probability P(F) _post The uncertainty measure Cramer-Rao lower bound (CRLB) is used to generate the set of activated scales.

[0187] The combination rule is: prioritize selecting items that simultaneously satisfy high G (physical sensitivity) and low P (F). _post (Algorithm not invalid) and low CRLB (physically reliable) scale.

[0188] In some alternative implementations, if the set is empty, or contains only scales contaminated by fixed pattern and stripe markings, the system will trigger a bandwidth migration (e.g., recalculating using a wider or narrower band B) or a redundancy fallback (e.g., falling back to a conservative, predefined safety scale).

[0189] Step 405: Perform pre-fusion and credibility generation.

[0190] This step applies gated weights within the activated scale set and fuses the corrected baseline sharpness response with the information-physical sharpness kernel.

[0191] This step produces the pre-fused sharpness response S. _pre_fused_fused And fusion credibility C _fused_fused S _pre_fused_fused These are multi-scale response maps or vectors, which will be further adaptively aggregated into a final sharpness map in Example 5. _fused_fused It is for S _pre_fused_fused Quantification of overall credibility (e.g., by aggregating the Cramer-Rao lower bound CRLB and the failure posterior probability P(F) at the selected scale). _post The result obtained will be used for timing robustness processing in Example 5.

[0192] Steps 401 to 405 constitute a complete and systematic failure post-hoc coupling mechanism. This mechanism works in concert to ensure the robustness and intelligence of the fusion decision.

[0193] According to one aspect of this application, prior to step 401, the explicit modeling premise of the failure post-hoc coupling mechanism of this invention is supplemented as follows:

[0194] Unlike traditional empirical weighting methods, this invention assigns weights to each baseline algorithm A. _i (e.g. A) _i For each scale s, a binary failure latent variable F is explicitly defined. _i,s ∈{0,1}. Where, F _i,s =1 explicitly indicates algorithm A _i It is in a failed state at scale s (e.g., dominated by noise, in a uniform region, or with an overly dense texture); F _i,s =0 indicates that it is in a normal state. Explicit modeling transforms failure from a vague empirical judgment into a probabilistic event that can be quantitatively estimated. Therefore, the task of step 401 is to estimate the posterior probability of the latent variable, P(F), based on the physical-statistical evidence set Z (i.e., the evidence vector φ), using a model (such as ranking self-supervised learning): _i,s =1|Z).

[0195] According to one aspect of this application, in step 402, the preferred implementation of the description regarding the generation of gating weights, which includes a dual-evidence gating collaborative mechanism, is as follows:

[0196] Gating weight w _i,s The preferred calculation formula is: w _i,s ∝(1-P(F _post,i,s ))×exp(-CRLB _s / τ); where P(F) _post,i,s This represents the estimated posterior probability of failure, CRLB. _s It is the uncertainty measure of scale s (Cramer-Rao lower bound), and τ is the temperature parameter.

[0197] This invention employs a dual evidence gating mechanism for the following reasons:

[0198] (1-P(F _post,i,s This represents subjective failure judgment: judging algorithm A based on sufficient scene evidence (such as noise, texture, GLR, etc.). _i Does it perform well in the current scenario?

[0199] exp(-CRLB _s / τ) represents the objective physical credibility: based on the optical modulation transfer function (MTF) and signal characteristics (S _X S _N The two types of evidence are orthogonal and complementary, used to determine whether scale *s* itself possesses high identifiability (i.e., low uncertainty) at the physical level. Single-weighting mechanisms (e.g., using only P or only Cramer-Rao lower bounds (CRLB)) have significant drawbacks: for example, an algorithm may not fail (low P), but its scale may be physically unreliable (high CRLB); or, a scale may be physically reliable (low CRLB), but baseline algorithms running on that scale (such as Laplacian) may fail due to high noise (high P). The dual-gating mechanism of this invention, by multiplying the two, ensures that only responses that are both physically reliable and not failed receive high weights in the final fusion, achieving robustness exceeding that of single-weighting mechanisms.

[0200] Example 5 provides a preferred implementation of applying failure posterior probability and gating weights to perform failure posterior coupling processing on the multi-scale baseline sharpness response set and information-physical sharpness kernel to generate the final sharpness evaluation result. It can be connected with the output of Example 3 and Example 4. It describes in detail how to perform adaptive refusion and temporal robustness on the pre-fused sharpness response to generate the final evaluation result.

[0201] Within the set of activated scales, let R be the fusion confidence of each scale s._conf_s The fusion confidence level C produced by step 405 _fused The components at this scale are given; R _conf Represents {R _conf_s A set of}.

[0202] Step 501: Generating scene adaptive parameters.

[0203] In this embodiment, a scene adaptive parameter generator is preferably provided.

[0204] Assemble the adaptively generated input vector. This vector preferably includes: pre-fused sharpness response, fusion confidence, failure posterior probability; and initial scene features (such as noise, phase consistency, texture density, motion indicators, etc.), gating weights, and the set of activated scales.

[0205] Next, an adaptive parameter generator is used to process the input vector. This generator preferably employs a rule-learning hybrid mechanism. For example, some parameters (such as time windows based on noise intensity) can be generated from an expert rule base, while other parameters (such as temperature parameters related to the posterior distribution of failure) can be generated through a small learning model (such as a shallow neural network or a regressor).

[0206] The generator produces a series of scene-adaptive parameters to dynamically adjust the behavior of subsequent steps. These include: a temperature parameter τ for recalculating gating weights, a time window N for time-series robustness, a refined scale set for refusion, and a threshold set for anomaly detection.

[0207] Step 502: Perform adaptive refusion and spatial aggregation.

[0208] This step defines scene complexity based on the distribution of the posterior probability of failure and the statistics in the evidence vector; dynamically determines the fusion ratio of the information-physical sharpness kernel according to the scene complexity; and performs weighted fusion of the corrected baseline sharpness response and the information-physical sharpness kernel to generate the fused scale response.

[0209] A preferred implementation is as follows:

[0210] Within the refined scale set generated in step 501, the gating weights are recalculated. This calculation method is similar to that used to generate gating weights based on the posterior probability of failure and uncertainty metric, but it uses the adaptive temperature parameter τ newly generated in step 501. For example, wi ∝ (1-F _post,i )*exp(-CRLB / τ).

[0211] Define the scenario complexity C. The scenario complexity C is preferably determined by the divergence of the failure posterior probability (i.e., the F-values ​​between different algorithms and different scales)._post The criteria for determining the consistency of the generalized likelihood ratio (GLR) and the fusion credibility are jointly defined. A high scenario complexity C value indicates a complex scenario and difficulty in discrimination.

[0212] Based on the scene complexity C, dynamically determine the information-physical clarity kernel S. _info The fusion ratio ρ*. For example, using the clamp function: ρ*=clamp(ρ _min +(ρ _max -ρ _min )*C γ ). Where ρ _min and ρ _max These represent the upper and lower limits of the proportion, and γ is the sensitivity coefficient. In simple scenarios (low C), more reliance can be placed on the baseline response; in complex scenarios (high C), it is necessary to increase the S value, which has a stronger physical meaning and is more robust. _info The percentage.

[0213] Perform weighted fusion. The corrected baseline sharpness response S _i_corr With Information-Physical Clarity Kernel S _info Perform weighted fusion to generate the fused scale response S _fused_fused For example, S _fused_fused =ρ**S _info +(1-ρ*)*Σ _ i w _i *S _i_corr .

[0214] Perform spatial aggregation. Merge the multi-scale response S... _fused_fused The precursor S of the final sharpness map is formed through spatial aggregation (e.g., weighted upsampling or interpolation based on fusion confidence). _final,pre .

[0215] Step 503: Perform timing robustness and global confidence output.

[0216] This step aggregates and generates a fusion credibility score; based on the fusion credibility score, and combined with the motion indications in the evidence vector, the fused scale response (i.e., the S output in step 502) is evaluated. _final,pre Perform time-series robustness processing; time-series robustness processing includes timeline smoothing and abnormal jump rollback protection to produce the final sharpness assessment results.

[0217] A preferred implementation is as follows:

[0218] Perform time-axis smoothing. Based on the fused scale response S produced in step 502. _final,pre (t) and the final output S from the previous time step _final (t-1), perform time smoothing: S_final (t)=γ*S _final (t-1)+(1-γ)*S _final,pre (t).

[0219] Generate adaptive smoothing coefficients. The smoothing coefficient γ is dynamically generated based on the fusion confidence C. _fused_fused To determine this, for example, a higher confidence level results in a smaller γ, indicating greater trust in the current frame; a lower confidence level results in a larger γ, indicating greater reliance on historical values.

[0220] Motion gating is performed. This smoothing also incorporates the motion indicator (MI) from the evidence vector. A high MI indicates rapid scene motion, and continuing smoothing in this case would result in motion blur. Preferably, updates are frozen, for example, by forcing γ=1.0, i.e., S... _final (t)=S _final (t-1), until the motion stops.

[0221] Execute abnormal transition rollback protection. Monitor the scale response S after fusion. _final,pre (t) relative to S _final The change is (t-1). If an abnormal transition occurs that exceeds the adaptive threshold (from step 501), rollback protection is triggered, for example, the current frame S is discarded. _final,pre (t) and (optionally) adjust the temperature parameters and time window for the next frame to deal with sudden interference.

[0222] Step 504: Generate the final result package.

[0223] This step encapsulates the final sharpness assessment results produced in step 503.

[0224] Calculate the global score and confidence interval for the final sharpness map S. _final (t) Perform spatial aggregation to obtain the global sharpness score S _global Meanwhile, its variance Var is approximately calculated based on uncertainty propagation (using Cramer-Rao lower bound CRLB and fusion confidence). _global And output the confidence interval CI=[S _global ±z _α *sqrt(Var _global )).

[0225] Optionally, evidence consistency constraints are enforced. As a preferred implementation, in calculating S... _global Previously, evidence consistency regularization could be added. That is, if the peak shape indicated by GLR or PC corresponds to the information-physical clarity kernel S... _info If the indicated peak shape is contradictory, the contribution at that moment or scale is suppressed to further reduce misjudgments of fixed patterns or false clarity.

[0226] Preferably, the final sharpness assessment result is a result package containing diagnostic evidence. This diagnostic evidence summarizes the key intermediate states during the operation of this method, and may specifically include: the dominant scale of the current frame, the main sources of failure determined (e.g., high noise or weak texture), the suppressed fixed pattern and stripe scales, and the key adaptive parameters generated in step 501 (e.g., temperature τ, time window N).

[0227] According to one aspect of this application, in step 501, the preferred implementation of the description of the Scene Adaptive Parameter Generator (SPG) is as follows:

[0228] The specific implementation of the Scene Adaptive Parameter Generator (SPG) is as follows: Construct the input vector X of the SPG. _SPG This vector is a high-level summary of the scene state in the current frame, and preferably includes: X _SPG =[β,σn,TD,U,MI,summary(C _fused ), summary(F _post Among them, the spectral slope β, noise intensity σn, texture density TD, uniformity U, and motion indication MI are all derived from initial scene features or evidence; summary(C _fused ) and summary(F _post These are statistical summaries (e.g., mean, variance) of fusion confidence and failure posterior probability in scale or algorithm dimension.

[0229] Next, SPG uses a rule-learning hybrid mechanism to process X. _SPG To generate an adaptive parameter set Π _SPG ={τ,N,Σ _refine , threshold set}, where Σ _refine This is for a refined set of scales. The rules section of this mechanism provides interpretable, fast-response baseline adjustments, such as:

[0230] When X _SPG Indicates high noise and low texture (e.g., β < 1.2 and σ). _N When the temperature is high, the rule is triggered by increasing the temperature parameter τ (reducing the sensitivity to the Cramer-Rao lower bound CRLB) and increasing the time window N (strengthening timing smoothing).

[0231] When X _SPG When indicating excessively dense texture pseudostructure (e.g., high TD but low GLR), the rule is triggered: in the refined scale set Σ _refine High-frequency scales are removed from the middle.

[0232] When X _SPG When a high-speed motion is indicated (e.g., MI high), the rule is triggered: freeze timing updates in step 503.

[0233] According to one aspect of this application, in step 502, the preferred specific implementation of the description regarding the definition of scene complexity C is as follows:

[0234] Scene complexity C is a key metric used to drive the information-physical clarity kernel fusion ratio ρ*. A high C value indicates difficulty in scene discrimination, and the information-physical clarity kernel S should be increased. _info The proportion of C. The optimal formula for calculating C is: C = mean _s [var _i (F _post (i, s))]×(1-mean _s (GLR _s ))×(1-mean _s (C _fused_s ));

[0235] The three components of the formula have clear physical meanings:

[0236] Post-failure divergence mean _s [var _i (F _post [(i, s))]: Calculate the posterior probability F of failure among different algorithms i across all scales. _post The variance. If the variance is high (e.g., Laplacian fails, Vollath works), it indicates disagreement among algorithms, complex scenarios, and an increased C-value.

[0237] GLR inverse contribution (1-mean) _s (GLR _s A high GLR indicates a clear structure. Therefore, a high (1-GLR) indicates an unclear structure or noise interference, a complex scene, and an increased C value.

[0238] Fusion credibility reverse contribution (1-mean) _s (C _fused_s )): C _fused High indicates high credibility. Therefore, (1-C _fused A high C value indicates low credibility and a complex scenario.

[0239] The three terms multiplied together constitute a robust assessment of the scenario complexity.

[0240] According to one aspect of this application, in step 503, the preferred implementation of the description regarding the execution of timing robustness processing is as follows:

[0241] The timing robustness processing includes not only time-axis smoothing based on motion indicator (MI) and smoothing coefficient (γ), but also preferably a complete abnormal jump rollback protection mechanism. This mechanism is specifically as follows:

[0242] In executing S _final (t)=γ·S _final (t-1)+(1-γ)·S _final,pre Before (t), the system detects the fused scale response S. _final,pre (t) relative to S _final Does the transition at (t-1) exceed the exception threshold generated by the Scene Adaptive Parameter Generator (SPG) (step 501)? If it does not exceed the threshold, normal time smoothing is performed. If it exceeds the threshold (i.e., an exception transition occurs), rollback protection is triggered: the system discards the current SPG. _final,pre (t) (i.e., S) _final (t)=S _final (t-1)), and (optionally) feed this abnormal event back to the Scene Adaptive Parameter Generator (SPG). In the next frame, the SPG will (based on rules) dynamically adjust the output parameters, such as increasing the temperature parameter τ and the time window N, to enter a more conservative evaluation mode and prevent the system from crashing under sudden interference.

[0243] Example 6 provides an evaluation method based on information-gradient tensor consistency (IFT), which can be used as another optional evaluation mechanism or can complement Examples 4 and 5.

[0244] Step 601: Construct the information-gradient tensor and analyze the eigenvalues.

[0245] This step calculates the spatial gradient of the information-physical sharpness kernel and at least one response in the multi-scale baseline sharpness response set; based on the spatial gradient, an information-gradient tensor is constructed; and by analyzing the eigenvalues ​​of the information gradient tensor, the consistency and divergence of the algorithm are quantified.

[0246] A preferred implementation is as follows:

[0247] Calculate the spatial gradient. For the information-physical clarity kernel S... _info Laplacian Sharpness Response S _L Vollath F4 resolution response S _v and SMD Sharpness Response S _s Calculate the spatial gradients (▽x, ▽y) in the x and y directions respectively.

[0248] Construct the information - gradient tensor T. Stack all gradient vectors into a tensor T. For example, T = [▽xS _L , ▽yS _L ;▽xS _v , ▽yS _v ;▽xS _s , ▽yS _s ;▽xS_info , ▽yS _info ].

[0249] Analyze the eigenvalues. Calculate the covariance matrix (T) of this tensor. T *T) and perform eigenvalue decomposition on it to obtain eigenvalues ​​λ1≥λ2≥λ3≥....

[0250] Quantifying consistency and divergence. The above eigenvalues ​​have clear physical meanings:

[0251] λ1 (maximum eigenvalue): represents the main consistent direction of all algorithms (baseline + physical kernel), that is, the direction in which the sharpness increases the fastest.

[0252] λ3 / λ1 (disagreement): If λ3 is much larger than λ1, it indicates that there is a significant disagreement between the algorithms (i.e., they are contradicting each other) and the consistency is low.

[0253] λ2 / λ1 (anisotropy): used to distinguish between defocus (isotropy, λ2≈λ1) and motion blur (anisotropy, λ2<<λ1).

[0254] Step 602: Perform adaptive scale selection.

[0255] This step performs adaptive scale selection based on algorithm consistency and divergence; adaptive scale selection is used to prioritize scales with high algorithm consistency and low divergence.

[0256] A preferred implementation is as follows:

[0257] Calculate the quality score for each scale. For each scale s, use the eigenvalues ​​λ1(s) and λ3(s) produced in step 601 to calculate a consensus-discordance joint score.

[0258] Selection rule: For example, s*=arg max _s [λ1(s)*(1-κ*(λ3(s) / λ1(s)))].

[0259] This rule is used to achieve direction consistency and divergence suppression, and prioritizes the selection of scales s* with strong main consistency direction signals (large λ1) and small algorithm divergence (small λ3 / λ1).

[0260] This method can be used to uniformly decouple spurious peaks caused by sparse textures, noise, or motion.

[0261] In some alternative implementations, the selected scale s* can be fed back to the scene parameter generator SPG to adaptively adjust the bandpass, window, or step size of subsequent processing.

[0262] Step 603, Spatiotemporal robustness and parameter generation mapping.

[0263] This step is an optional supplementary implementation in this embodiment (Information-Gradient Tensor Consistency IFT Scheme).

[0264] Spatiotemporal variational optimization with Cramer-Rao lower bound CRLB weighting is performed. As an alternative to temporal robustness and global confidence output, this embodiment can employ spatiotemporal variational optimization for robustness. For example, minimizing the cost function E=Σ _t Σ _s CRLB _s -1 *||S _t,s -S _hat_t,s || 2 +λ _t *||d _t S||1+λ _X *||▽ 2 S||1. Wherein, CRLB _s -1 As confidence weights, the last two terms are time and L1 regularization terms for smoothing, respectively. This optimization can also be gated (frozen update) by motion indication MI. _hat_t,s As a priori estimate of the ideal sharpness distribution, it is used to guide spatiotemporal regularization, so that the optimization result closely approximates the observed response S. _t,s It also maintains smoothness in both time and space.

[0265] The response spectrum is mapped to Info-MTF parameters. As a specific implementation of SPG in step 602, a continuous mapping can be used. For example, the spectral slope β, phase coherence PC, and noise intensity σ can be mapped. n Using these parameters as input, the information-physical clarity kernel S is continuously mapped and generated through a Gaussian process or kernel regression model. _info The required bandwidth parameter or the required temperature parameter τ.

[0266] Example 7 provides an application of autofocus closed-loop control based on sharpness assessment, describing how the assessment method of the present invention is implemented in a specific application scenario.

[0267] Step 701: Generate focus control command.

[0268] This step involves converting the final sharpness assessment results into control commands for the focusing motor.

[0269] In the basic scheme of Example 1, a hill-climbing algorithm is usually used, which compares the magnitudes of S(t) and S(t-1) to decide whether to continue or reverse. This algorithm is prone to crossing peaks or jittering at noisy or false peaks.

[0270] In a preferred embodiment of the present invention, the generation of control commands is based on more comprehensive multidimensional information:

[0271] Based on the final sharpness map S _final The gradient of the clarity map provides spatial directionality, which is more advantageous than a single global score S. _global It provides more comprehensive control information, which can guide the focusing motor to move in the direction of the greatest gradient.

[0272] Based on fusion credibility C _fused_fused Fusion reliability is used as the brake or throttle for focus control.

[0273] When the fusion credibility C _fused_fused High: This indicates that the current evaluation results are reliable, and the system can use a larger step size to quickly approach the focus.

[0274] When the fusion credibility C _fused_fused Low values ​​(e.g., in cases of weak texture, strong noise, or algorithm failure): indicate that the current evaluation result is unreliable. In this case, the system should not blindly climb the hill, but should instead implement conservative strategies, such as reducing the focus step size, stopping movement and waiting for the scene to stabilize, or triggering a full-focal-length scan to rebuild credibility.

[0275] The confidence-gated control logic can effectively avoid the peak crossing and back-and-forth jitter problems of traditional hill-climbing algorithms in the failure domain (such as noise false peaks and uniform regions).

[0276] Step 702: Export the evaluation result package and closed-loop interface.

[0277] The output evaluation result package preferably includes: the final sharpness map S _final Global Sharpness Score _global Confidence interval CI.

[0278] The results package also preferably includes diagnostic evidence, such as: dominant scale, main failure source, suppressed fixed pattern scale, and key adaptive parameters.

[0279] This method also provides an interface to output the above result package and the focus control command (optional) generated in step 701. This allows the present invention not only to function as an independent evaluation module, but also to be deeply embedded in a closed-loop control system, providing the system with decision-making basis (control commands) and decision interpretation (diagnostic evidence).

[0280] In a preferred embodiment, the parameters of the present invention can be selected from the following representative value ranges. The temperature parameter τ can range from 0.1 to 5. When τ is small, the weight change corresponding to the posterior probability of failure is steeper; when τ is large, the weight change is smoother. The time smoothing coefficient γ can range from 0.6 to 0.95. When γ is small, the system is more sensitive to changes in the sharpness of the current frame; when γ is large, the sharpness map is more stable over time. The lower limit of the scene complexity normalization interval ρ _min It can be taken as 0.2~0.5, with an upper limit ρ. _max A value of 0.5 to 0.9 can be used to adapt to scenes with different texture densities and noise levels.

[0281] Furthermore, the scene complexity index C can be calculated not only based on the statistics of the evidence vector, but also supplementary indices such as the variance of the focus curve slope and the sparsity of the sharpness peak distribution. The time smoothing strategy can employ first-order IIR filtering, fixed-length sliding window averaging, Savitzky-Golay smoothing, or recursive estimation based on Kalman filtering. These different values ​​and implementation methods can be selected according to the specific system's requirements for response speed and stability, without affecting the basic principles of the method of this invention.

[0282] In a typical application scenario, the method of this invention can be deployed in an autofocus closed-loop system. Specifically, the image acquisition module acquires the original image of the current frame at each focus position and outputs it to the sharpness evaluation module. The sharpness evaluation module calculates the multi-scale sharpness response, the information-physical sharpness kernel, the failure posterior probability, and the final sharpness map S according to the method of this invention. _final And the fusion credibility graph C _fused The autofocus control module is based on S... _final The global peak position, the slope on both sides of the peak, and C _fused The spatial distribution of the image generates the next focus step command, driving the focus actuator to move along the direction of increasing sharpness to a new focus position. Once the new focus position is achieved, the image acquisition module acquires a new frame image and sends it to the sharpness evaluation module. This process iterates repeatedly until the sharpness curve reaches point C. _fused Once a stable optimal or suboptimal peak value is reached within the displayed high-confidence region, the autofocus process ends. This establishes a closed-loop control chain of "image acquisition -> sharpness assessment -> failure identification and correction -> focus control -> image acquisition," enabling the method of this invention to be directly embedded into existing autofocus systems for engineering applications.

[0283] This invention introduces failure posterior probability estimation based on evidence vectors, providing a real-time, quantitative assessment of the unreliability of each baseline algorithm in the current scenario (such as high noise and weak texture), replacing fixed weights, and combining uncertainty to generate gated weights, ensuring that only non-failed and physically reliable responses can dominate the fusion result, effectively suppressing the interference of pseudo-peaks, and solving the problem of being biased by pseudo-peaks due to the lack of a dynamic quantification discrimination mechanism.

[0284] This invention introduces an information-physical sharpness kernel and Fisher information based on the optical modulation transfer function, elevating the definition of sharpness from empirical gradients to the level of physical discernibility, thus naturally enabling it to distinguish between imageable structures and high-frequency noise. Simultaneously, the computation process of this physical model (such as the Cramer-Rao lower bound) naturally and synchronously produces an evaluation uncertainty metric, giving the system important self-awareness and solving the problems of being unable to distinguish between structure and noise at the physical level and being unable to generate uncertainty.

[0285] This invention proposes an analytical bias correction mechanism. When an algorithm (such as Laplacian) is detected to have failed for some reason (such as noise), the system does not simply reduce its weight, but instead uses scene evidence (such as estimated noise intensity) to perform analytical subtraction or compensation on its response (e.g., subtracting noise-induced overestimation or compensating for low responses in uniform regions). This correction mechanism achieves deep complementarity between algorithms, rather than a simple head-swapping, thus solving the problem that complementarity remains at the weight reduction level.

[0286] The preferred embodiments of the present invention have been described in detail above. However, the present invention is not limited to the specific details in the above embodiments. Within the scope of the technical concept of the present invention, various equivalent transformations can be made to the technical solutions of the present invention, and these equivalent transformations all fall within the protection scope of the present invention.

Claims

1. A method for evaluating image sharpness, characterized in that, include: Obtain the original image and optical modulation transfer function; Based on the original image and optical modulation transfer function, a multi-scale baseline sharpness response set, information-physical sharpness kernel, uncertainty measure and evidence vector are generated. Based on the evidence vector, estimate the posterior probability of failure; Gating weights are generated based on the posterior probability of failure and uncertainty measure; By applying failure posterior probability and gating weights, failure posterior coupling processing is performed on the multi-scale baseline sharpness response set and the information-physical sharpness kernel to generate the final sharpness assessment result; Among these, estimating the posterior probability of failure based on the evidence vector includes: For the evidence vectors of the evaluation target and the pre-stored reference target, perform robust standardization and outlier pruning. By applying self-supervised learning for ranking, the estimation model is optimized by minimizing the ranking loss function, which is used to maximize the margin between the non-failure probability of the reference target and the non-failure probability of the target to be evaluated. In the estimation model training, domain adaptive regularization is combined to calculate and output the failure posterior probability; By applying failure posterior probability and gating weights, failure posterior coupling processing is performed on the multi-scale baseline sharpness response set and the information-physical sharpness kernel, including: Monitor the posterior probability of failure; When the posterior probability of failure indicates a high risk of failure, analytical bias correction is applied to the corresponding response in the multi-scale baseline sharpness response set to produce a corrected baseline sharpness response. By applying gating weights, the baseline sharpness response after correction and the information-physical sharpness kernel are fused to generate the final sharpness assessment result.

2. The method according to claim 1, characterized in that, The evidence vector includes: Texture density, noise intensity, spectral slope, motion indication, and two-domain consistent generalized likelihood ratio statistics.

3. The method according to claim 1, characterized in that, Based on the posterior probability of failure and uncertainty measure, gating weights are generated, including: A non-failure metric for calculating the posterior probability of failure; Calculate the confidence level of the exponential decay of uncertainty measure under temperature parameter control; The non-failure metric is multiplied by the exponential decay confidence level and then normalized to produce the gate control weight.

4. The method according to claim 1, characterized in that, Applying resolution bias correction produces a corrected baseline sharpness response, including: When the corresponding response is a Laplacian sharpness response, noise intensity is extracted from the evidence vector; By applying noise intensity, the noise-induced overestimation term is subtracted from the Laplacian sharpness response to produce the corrected baseline sharpness response.

5. The method according to claim 4, characterized in that, Applying analytical bias correction also includes: Extract texture density from evidence vectors; By applying texture density to the Laplacian sharpness response after noise-induced overestimation has been removed, false high responses caused by overly dense textures are suppressed, resulting in a corrected baseline sharpness response.

6. The method according to claim 1, characterized in that, Generate information - physical sharpness kernel and uncertainty measure, including: Fisher information is calculated based on the optical modulation transfer function, combined with the power spectrum of the observable signal and the power spectrum of noise derived from the original image. Based on Fisher information, the information-physical clarity kernel and the Cramer-Rao lower bound as an uncertainty measure are derived respectively.

7. The method according to claim 6, characterized in that, Based on the optical modulation transfer function, and combining the power spectrum of the observable signal and the power spectrum of noise derived from the original image, Fisher information is calculated, including: Determine the sensitivity of the optical modulation transfer function to the defocus parameter; Fisher information is obtained by integrating or summing the square of the sensitivity with the ratio of the power spectrum of the observable signal and the power spectrum of the noise in the frequency domain.

8. The method according to claim 1 or 2, characterized in that, The method also includes: Extract the bandpass energy in the frequency domain and the phase consistency in the spatial domain from the original image; Construct joint likelihood functions for the structural and noise assumptions, respectively; The ratio of joint likelihood functions is calculated to generate a two-domain consistent generalized likelihood ratio statistic, which is then used as one of the components of the evidence vector to estimate the posterior probability of failure.

Citation Information

Patent Citations

  • Coal mine low-illumination blurred image restoration method

    CN110415193A

  • Equipment imaging definition evaluation method, system and device

    CN116843683A