A multi-focus image fusion method and system based on semi-smooth Newton method

CN122415352BActive Publication Date: 2026-09-22HUAQIAO UNIVERSITY +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610825352.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-09
Publication Date
2026-09-22
Estimated Expiration
2046-06-09

AI Technical Summary

Technical Problem

但该类深度学习方法普遍存在模型可解释性差、推理运算耗时久的问题,难以满足显微成像等实际场景对成像速度、运行稳定性的严苛要求

Benefits of technology

[0035]通过采用上述方案,本发明具有如下的优点和有益效果:本发明首次将二阶优化算法应用于多聚焦图像融合,借助深度展开技术把半光滑牛顿算法迭代过程映射为可学习网络结构,兼顾严谨物理可解释性与优异特征学习能力,突破了传统融合方法可解释性弱、泛化能力不足的局限。通过增设自适应特征迭代重构单元、聚焦梯度感知模块与多尺度频率增强单元,可同步在空间域、频率域及边缘域完成多维度特征提取与融合,大幅强化融合图像全局聚焦性能与细节还原能力。引入可学习边缘卷积,以 Laplacian、Sobel 经典边缘算子初始化卷积核,在保留边缘检测物理机理的同时赋予网络自适应学习优化能力。搭配乘子更新模块与矩阵更新模块,实现散焦矩阵和拉格朗日乘子动态自适应调整,显著加快算法收敛速度,提升融合精度。相较于现有技术,本发明在融合视觉质量、边缘细节保留及整体计算效率上均实现明显提升,实用性与工程应用价值突出。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122415352B_ABST
    Figure CN122415352B_ABST
Patent Text Reader

Abstract

The application discloses a multi-focus image fusion method and system based on a semi-smooth Newton method, and belongs to the technical field of image processing. The semi-smooth Newton algorithm is used for solving, a proximal operator and a generalized Jacobian matrix are introduced, and the optimization solving process is expanded into a deep network. An adaptive feature iterative reconstruction unit is designed, auxiliary variable updating, a learnable edge convolution, focus gradient perception and a multi-scale frequency enhancement structure are integrated, and iterative optimization is realized in combination with a feature reconstruction block, a multiplier updating module and a matrix updating module. The second-order optimization algorithm is introduced into multi-focus image fusion for the first time, has physical interpretability and strong feature learning capability, and has better fusion effect than existing methods, and has important application value in high-precision imaging fields such as microscopy and remote sensing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing technology, and in particular to a multi-focus image fusion method and system based on the semi-smooth Newton method. Background Technology

[0002] Due to the inherent limitations of optical lens depth of field, a single imaging device can only achieve a clear image of the area near the focal plane. Scenes outside the focal plane are prone to defocusing and blurring, making it impossible to obtain a clear image across the entire field, which greatly reduces the effective information content and image quality. Multi-focus image fusion technology can fuse multiple source images with different local focuses to reconstruct a clear, fully focused image across the entire field. This can effectively improve image clarity and information integrity, and has extremely high application value in high-precision imaging fields such as microscopic imaging and remote sensing monitoring.

[0003] Existing multi-focus image fusion algorithms are mainly divided into two categories: traditional algorithms and deep learning algorithms. Traditional algorithms include spatial domain and transform domain methods, which suffer from drawbacks such as weak feature extraction capabilities and insufficient fusion accuracy. In recent years, deep learning-based fusion algorithms have been widely used due to their superior fusion results, and are mainly divided into decision graph-based methods and end-to-end methods. However, these deep learning methods generally suffer from poor model interpretability and long inference computation times, making it difficult to meet the stringent requirements of imaging speed and operational stability in practical scenarios such as microscopic imaging.

[0004] Current multi-focus image fusion technology still faces three major challenges: insufficient extraction of local focus features, difficulty in maintaining global consistency, and susceptibility to artifacts at depth boundaries. To address these technical shortcomings, existing technologies urgently need a multi-focus image fusion scheme that balances interpretability, computational efficiency, and fusion accuracy to adapt to various high-precision, high-real-time imaging applications. Summary of the Invention

[0005] The purpose of this invention is to provide a multi-focus image fusion method and system based on the semi-smooth Newton method. The multi-focus image fusion problem is modeled as a variational optimization problem with non-smooth regularization constraints. The optimization algorithm is mapped to a learnable deep network through deep expansion technology, which has both physical interpretability and strong learning ability.

[0006] To achieve the above objectives, the present invention provides the following technical solution: a multi-focus image fusion method based on the semi-smooth Newton method, comprising the following steps: S1. Problem Modeling: Construct a mathematical model for multi-focus image fusion, including dataset construction, degradation model establishment, Kronecker integral solution, and variational optimization problem establishment; S2. Problem Solving and Optimization: The variational optimization problem is solved using the semi-smooth Newton algorithm. A Newton linear system is established by introducing Lagrange dual variables and the generalized Jacobian matrix. S3. Network Unfolding and Image Reconstruction: The optimization algorithm is unfolded into a deep network, and the feature reconstruction blocks are iteratively optimized to finally reconstruct the fully focused image.

[0007] Preferably, the specific process of problem modeling in step S1 includes: S11: Dataset construction. The DUTS-TR dataset is used as the source data. This dataset contains 10,553 pairs of high-quality images and corresponding binary masks. Multi-focus image pairs are generated by Gaussian blur. The Gaussian blur parameters are adaptively calculated according to the image resolution. The original image is retained as the ground truth of full focus. The training set and test set are divided in a 9:1 ratio. S12: Degradation model establishment, obtaining two source images with different focal regions generated in step S11. ,in Representing the real space of the image, For the number of channels, The spatial resolution of the image height. To determine the spatial resolution of the image width, the source image... Considered as a potential ideal all-focused image After the defocusing operator matrix Degradation observations after action, and the establishment of a degradation model. ,in The defocus operator matrix is ​​represented as an identity matrix in the focused region and as a low-pass filter in the defocused region. Represents the real space of the matrix. For full-focus image Vectorization operator The resulting column vector after expansion The source image is vectorized. This refers to sensor noise; among which The product of the number of channels, height, and width of the image, and the total dimension after vectorization. Represent the real space of vectors; S13: Kronecker integral solution, to avoid defocusing operator matrix The high computational overhead introduced by this problem is addressed by employing the Kronecker integral technique. Based on the Kronecker product and vectorization equivalence theorem, for any matrix... , , The following relationship holds:

[0008] in This represents the Kronecker product operation. and Let be the transformation matrix. Let the matrix be to be transformed; the one-dimensional degradation process is equivalently reduced back to the two-dimensional tensor space, and the degradation process is modeled as:

[0009] in, For the matrix form of the first Source image, corresponding to the aforementioned vectorized form Vertical defocus matrix and horizontal defocus matrix Characterize the low-pass degradation properties in the row and column directions respectively, satisfying , The noise is in the reconstructed tensor form; To avoid explicitly extracting the mask, the two observation images are jointly stitched together along the channel dimension to construct a unified high-dimensional observation model.

[0010] in, This indicates a splicing operation along the channel dimension. This is the stitched image of the joint observation. and These are two source images with different focus regions; in the focus region, the vertical and horizontal defocus matrices degenerate into identity matrices. In the defocused region, the operator behaves as a low-pass blur operator, which significantly reduces computational complexity and memory usage through Kronecker integral solutions. S14: The variational optimization problem is established. Based on the maximum a posteriori probability (MAP) estimation, the reconstruction of the full-focus image F from the degraded observations is equivalent to minimizing the energy functional.

[0011] in For data fidelity, the Frobenius norm is used. Measure the difference between the reconstructed image and the observed image. This is a non-smoothing regularization term used to apply high-frequency structural priors to natural images. For balancing parameters.

[0012] Preferably, the specific process of problem solving and optimization in step S2 includes: S21: Based on the cyclic properties of the Frobenius norm and the trace, and using matrix differential theory, define the residual function:

[0013] in, Let be the residual function, and when the residual is zero, it corresponds to the optimal solution of the optimization problem. Indicates the first The estimated value of the fully focused image during each iteration; It is a proximal operator used to handle implicit updates of non-smooth regularization terms; This is the iteration step size parameter; S22: By explicitly introducing Lagrange dual variables As a multiplier, it is responsible for extracting and accumulating the non-smooth high-frequency prior information of the image during iteration, thereby transforming the residual function into a fixed-point root solving problem. :

[0014] in This indicates a feature fusion operation, used to combine dual variables. The encoded high-frequency prior information is injected into the input of the proximal operator; S23: To introduce second-order curvature information, the generalized Jacobian matrix is ​​used. A semi-smooth Newton equation is established, and the generalized chain rule is applied to expand it, deriving the first... Newtonian linear systems of steps:

[0015] in, It is the identity matrix. Let be the generalized Jacobian derivative matrix of the proximal operator at the current iteration point. For the first The update increment of the step; the high-precision iterative update step size is obtained by solving the linear system.

[0016] Preferably, the specific process of network unrolling and image reconstruction in step S3 includes: S31: The iterative process of the optimization algorithm is unfolded into a deep network. The network contains M = 6 feature reconstruction blocks. Each feature reconstruction block corresponds to one iterative optimization of the semi-smooth Newton algorithm. The algorithm is mapped into a learnable network structure through deep unfolding technology, which has both physical interpretability and strong learning ability. S32: For the first There are feature reconstruction blocks, among which It consists of four modules: multiplier update module, matrix update module, auxiliary variable update module, and adaptive feature iterative reconstruction unit; S33: The multiplier update module updates according to the current image state. And auxiliary variables, update the Lagrange multipliers through a convolutional neural network. The Lagrange multipliers It is responsible for extracting and accumulating non-smooth high-frequency prior information of the image during iteration; S34: The matrix update module updates the first image based on the current image state and observation data through a normalization operation. Vertical defocus matrix corresponding to each feature reconstruction block and horizontal defocus matrix The matrix update module first calculates the gradient information of the current image, and then extracts the first derivative of the image using a learnable derivative operator. The convolution is implemented by initializing the convolution kernel as a central difference operator, and then normalizing the gradient magnitude to between 0 and 1. The normalization formula is as follows: , where G is the gradient magnitude; S35: The auxiliary variable update module calculates the local focusing response and updates the Lagrange dual variable. This provides structural prior guidance for the next iteration stage; S36: The adaptive feature iteration reconstruction unit includes a multi-scale feature extraction unit, a focused gradient perception module, a multi-scale frequency enhancement module, and a learnable edge convolution, constructing a cross-domain feature representation to implicitly fit the second-order Newton update direction, and using a learnable step size. The formula for updating and merging the image is as follows: ; S37: For the last feature reconstruction block, i.e. It contains only adaptive feature iteration reconstruction units, and finally performs image reconstruction and outputs the final fusion result.

[0017] Preferably, the multi-scale feature extraction unit in step S36 adopts a parallel three-branch architecture; the first branch extracts local focusing features through channel attention and cascaded convolution, and calculates the spatial maximum amplitude as a physical prior constraint, as shown in the following formula:

[0018]

[0019] in, As input features, This indicates the channel attention module. Indicates the kernel size as Convolution operation, To modify the activation function of the linear unit, The output features of the first branch, Features The spatial maximum amplitude map obtained by taking the maximum value along the channel dimension is used for subsequent adaptive modulation of the frequency domain output. For channel dimensions; The second branch achieves frequency sensing of the global receptive field through Fast Fourier Transform, performs feature mapping in the complex domain, and finally utilizes... The constraint reduces ringing artifacts; the calculation formula is as follows:

[0020]

[0021] in, and These are the forward and inverse Fast Fourier Transforms for real numbers, respectively. Represents a complex linear layer. For batch normalization operations, These are intermediate features after frequency domain transformation. The output features of the second branch, The hyperbolic tangent activation function is used to normalize the response to... interval, Represents the element-wise Hadamard product operation; The third branch consists of a focus gradient perception module and a multi-scale frequency enhancement module, which utilizes learnable edge convolution to extract accurate focus boundary features.

[0022] Preferably, the focused gradient sensing module includes three parallel paths, and its output is normalized by a complex layer to obtain intermediate features. The paths are represented as follows: Standard convolution path: ; Channel compression - extended path: ; Isotropic learnable edge paths: ; in For learnable edge convolution functions, The kernel size is The convolution operation can learn edge convolutions, and the depthwise convolution kernel is initialized with a Laplacian operator:

[0023] The output formula of the focused gradient sensing module is: , ( ) is a complex linear normalization function.

[0024] Preferably, the multi-scale frequency enhancement module includes three parallel paths for extracting directional edge features: Standard convolution path: ; Multi-scale path: ; Directional learnable edge paths: ; in, This indicates a splicing operation along the channel dimension. Learnable edge convolutions in the horizontal direction Learnable edge convolutions in the vertical direction. and Initialize them as Sobel-x and Sobel-y operators respectively: Sobel-x operator:

[0025] Sobel-y operator:

[0026] The final output of the multi-scale frequency enhancement module for:

[0027] The final output is a fusion of the three modules: .

[0028] Preferably, the image reconstruction process in step S37 includes: At the network input end, the source image and spliced ​​together The input features are downsampled to reduce spatial resolution and increase the number of channels to obtain the basic input features. Enter the iterative network; At the network output, the output of the feature reconstruction block in the last iteration stage is combined with the initial concatenated features. The global residuals are summed, then the image is restored to its original resolution using an upsampling module, and finally reconstructed using a lightweight output module to produce the final fully focused fused image. .

[0029] Preferably, a hybrid loss function is used for training the network from the input to the output:

[0030] The Charbonnier loss, which measures pixel-level differences in the spatial domain, is expressed as:

[0031] The frequency domain loss constraint for global structural consistency is expressed as:

[0032] The edge loss enhancement effect on edge preservation is expressed as:

[0033] in, For the total loss function, For Charbonnier loss, and These are the weighting coefficients for the frequency domain loss and the edge loss, respectively. For frequency domain loss, For edge loss, This represents the number of image samples in the batch. For sample index, For the network prediction of the first A fused image, For the first A fully focused ground truth image, The square of the Frobenius norm. To prevent gradient vanishing, a small positive constant, For Fourier transform operators, The Laplacian edge detection operator is used to calculate the second derivative of an image to highlight edge and detail information. for Norm.

[0034] A multi-focus image fusion system based on the semi-smooth Newton's method, employing any of the image fusion methods described above, includes the following modules: Problem modeling module: Constructing a mathematical model for multi-focus image fusion, including dataset construction, degradation model establishment, Kronecker integral solution, and variational optimization problem establishment; Problem Solving and Optimization Module: The semi-smooth Newton algorithm is used to solve variational optimization problems. By introducing Lagrange dual variables and the generalized Jacobian matrix, a Newtonian linear system is established. Network Unfolding and Image Reconstruction Module: The optimization algorithm is unfolded into a deep network, and through iterative optimization of feature reconstruction blocks, the fully focused image is finally reconstructed.

[0035] By adopting the above scheme, this invention has the following advantages and beneficial effects: This invention is the first to apply a second-order optimization algorithm to multi-focus image fusion. It uses deep unrolling technology to map the iterative process of the semi-smooth Newton algorithm into a learnable network structure, balancing rigorous physical interpretability with excellent feature learning capabilities, overcoming the limitations of traditional fusion methods in terms of weak interpretability and insufficient generalization ability. By adding an adaptive feature iteration reconstruction unit, a focus gradient perception module, and a multi-scale frequency enhancement unit, multi-dimensional feature extraction and fusion can be completed simultaneously in the spatial, frequency, and edge domains, significantly enhancing the global focusing performance and detail restoration capabilities of the fused image. The introduction of learnable edge convolution, using classic Laplacian and Sobel edge operators to initialize the convolution kernels, preserves the physical mechanism of edge detection while endowing the network with adaptive learning and optimization capabilities. Combined with a multiplier update module and a matrix update module, dynamic adaptive adjustment of the defocus matrix and Lagrange multipliers is achieved, significantly accelerating the algorithm's convergence speed and improving fusion accuracy. Compared with existing technologies, this invention achieves significant improvements in fused visual quality, edge detail preservation, and overall computational efficiency, demonstrating outstanding practicality and engineering application value. Attached Figure Description

[0036] Figure 1 This is a detailed flowchart of the present invention; Figure 2 This is a diagram of the overall network architecture of the present invention; Figure 3 This is a detailed framework diagram of the main modules of the present invention. Detailed Implementation

[0037] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention. Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to represent selected embodiments of the invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0038] Reference manual attached Figure 1 This invention provides a multi-focus image fusion method based on the semi-smooth Newton method, comprising the following steps: S1: Problem modeling, constructing a mathematical model for multi-focus image fusion, including dataset construction, degradation model establishment, Kronecker integral solution and variational optimization problem establishment; S2: Problem Solving and Optimization. An improved semi-smooth Newton algorithm is used to solve variational optimization problems. By introducing Lagrange dual variables and generalized Jacobian matrices, a Newtonian linear system is established. S3: Network Unfolding and Image Reconstruction. The optimization algorithm is unfolded into a deep network, and through iterative optimization of feature reconstruction blocks, the fully focused image is finally reconstructed.

[0039] Furthermore, step S1, problem modeling, specifically includes the following steps: S11: Input two source images with different focus areas. ( For the number of channels, (for spatial resolution); instead of using any mask heuristics, we use the source image... To be considered a potentially ideal all-focused image Defocusing operator matrix after specific spatial variation Degradation observation after action, i.e. ,in Here is the defocus operator matrix. For vectorized full-focus images, The source image is vectorized. This is sensor noise; S12: To avoid the computational overhead of high-dimensional matrices, the Kronecker integral solution is introduced, and based on the Kronecker product and vectorization equivalence theorem, the one-dimensional degradation process is equivalently reduced to the two-dimensional tensor space. Therefore, the degradation process is modeled as formula (1). At the same time, to avoid explicitly extracting the mask, the two observation images are jointly stitched together in the channel dimension to construct a unified high-dimensional observation model formula (2).

[0040]

[0041] Among them, the vertical defocus matrix and horizontal defocus matrix These characterize the low-pass degradation properties in the row and column directions, respectively. For the reconstructed tensor form noise, in the focused region, the vertical and horizontal defocus matrices degenerate into the identity matrix E, and in the defocus region, the operator behaves as a low-pass blur operator. S13: Reconstructing full-focus images from degraded observations based on maximum a posteriori probability (MAP) estimation. This is equivalent to minimizing the following energy functional:

[0042] The first item is the data fidelity item, which requires the generated data to be of high quality. Conforms to the joint observation pattern; second item λ is a non-smoothing regularization term used to apply high-frequency structural priors (such as gradient sparsity) to the natural image; λ is a balance parameter.

[0043] Furthermore, step S2 specifically includes the following steps: S21: Based on the cyclic properties of the Frobenius norm and trace, the data fidelity term in formula (3) is related to... The physical gradient can be precisely defined as:

[0044] According to Fermat's theorem, the optimal solution satisfies the first-order extremum condition as follows:

[0045] By introducing step size parameters and proximal operators This is transformed into a forward-backward splitting iteration format:

[0046] At the optimal point, the state of the network will no longer change, that is, the fixed-point condition is satisfied: Substituting this convergence condition into formula (6) and rearranging the operators on the right side of the equation, we define a residual function. :

[0047] Substituting formula (4) into formula (7a) and expanding, we get:

[0048] S22: Since the proximal operator is non-smooth, the Jacobian matrix of the classical Newton's method does not exist. This is addressed by explicitly introducing Lagrange dual variables. As a multiplier, it is responsible for extracting and accumulating the non-smooth high-frequency priors of the image during iteration, thereby transforming equation (7) into a problem of solving fixed-point roots. :

[0049] To accelerate optimization and provide data-driven prior guidance, initial dual variables... The input joint observation I is extracted via a convolution kernel with residuals, and its mathematical expression is:

[0050] S23: To introduce second-order curvature information, the generalized Jacobian matrix is ​​used. Establish the semi-smooth Newton equation:

[0051] Expanding using the generalized chain rule The Newtonian linear system at step t is rigorously derived as follows:

[0052] in, It is the identity matrix. To update the increment, The generalized Jacobian derivative of the proximal operator at the current point is used to obtain a high-precision iterative update step size by solving the linear system. Compared with the first-order method, the improved semi-smooth Newton method utilizes second-order information, resulting in faster convergence speed and higher solution accuracy.

[0053] As per the instruction manual Figure 2 As shown, the network adopts a deep unfolded architecture, which unfolds the iterative process of the optimization algorithm into... Each feature reconstruction block (FRBlock) corresponds to one iteration of the semi-smooth Newton algorithm; in the first stage, a 4x downsampling is used to reduce the spatial resolution from... Reduce to The number of channels is from Expand to Then, it goes through 6 consecutive FRBlock modules.

[0054] As per the instruction manual Figure 3 As shown, to avoid the catastrophic computational redundancy caused by high-dimensional matrix inversion in classical Newtonian linear systems, an Adaptive Feature Iteration Reconstruction Unit (AFIRU) is added as a nonlinear surrogate solver in each FRBlock. The core of AFIRU, the Multi-Scale Feature Extraction Unit (MSFEU), abandons traditional serial convolution and adopts a three-domain decoupled parallel architecture. By simultaneously performing orthogonal feature extraction in the spatial, frequency, and edge domains, it improves the full-focus fidelity of the fused image from the underlying physical dimension. The specific branching mechanism is as follows: Local Spatial Domain Branch: This branch utilizes the Channel Attention (CA) mechanism and cascaded convolutions to extract locally focused regions. The CA module compresses the spatial dimension into channel descriptors through global average pooling, and then... The bottleneck layer (fully connected network) and sigmoid gating function adaptively assign high weights to channels with high-frequency focusing responses. This is followed by two layers with ReLU activation. Convolutional layers introduce non-linear representations. This branch performs aggregation operations along the channel dimension to extract the spatial maximum amplitude. This serves as the physical upper limit for subsequent frequency domain transformations, preventing abnormal responses induced by spectral manipulation.

[0055] Global Frequency Domain Branch: This branch utilizes the real-number Fast Fourier Transform (rFFT) to achieve lossless global frequency sensing, effectively overcoming the limitations of the local receptive field in spatial convolution. Input features, after being stably distributed through batch normalization (BN), are mapped to the complex frequency domain via rFFT. Within the frequency domain, a learnable complex weight matrix is ​​used... and Adaptive filtering and reconstruction are performed on the high and low frequency components of the spectrum. After reconstruction back to the spatial domain via inverse FFT, a cross-domain modulation mechanism is introduced: utilizing the hyperbolic tangent function. Normalize the frequency domain response to The interval, and the aforementioned spatial priors Element-wise Hadamard Product is performed. This mechanism ensures that the high-frequency gain of the global spectrum is absolutely limited to the maximum true extremum in the local space, thereby reducing ringing artifacts.

[0056] Edge Enhancement Branch: Precise segmentation of the focus and defocus boundaries is crucial for eliminating halos and artifacts. This branch includes a Focus Gradient Awareness Module (FBPM) and a Multi-Scale Frequency Enhancement Module (MSFEM). Internally, these modules use Learnable Edge Convolution (LEconv) to extract high-frequency features. FBPM provides isotropic features and comprises three parallel paths: standard... Convolution extracts local basic features ; and Cascaded channel compression-expansion path extraction of multi-scale receptive field features LEconv extracts second-order derivative features Its LEconv module employs depthwise separable convolutions, with its depthwise convolution kernels initialized using the classic Laplacian operator. The LEconv kernels in the MSFEM module are initialized with the Sobel operator to extract directional first-order gradients. All operator parameters are adaptively fine-tuned during training.

[0057] Furthermore, step S3 specifically includes the following steps: S31: The iterative solution process of the non-smooth optimization algorithm is expanded into a deep network, the main body of which consists of... It consists of cascaded feature reconstruction blocks (FRBlocks), each FRBlock strictly corresponding to one complete iteration of the semi-smooth Newton algorithm.

[0058] For the t-th FRBlock ( It contains the following modules: Multiplier Update Module (MPU): Updates Lagrange multipliers ; Matrix Update Module (MQU): Updates the defocus matrix ; Auxiliary Variable Update Module (ACPU): Calculates the local focusing response and updates the Lagrange dual variable. This provides structural prior guidance for the next iteration stage.

[0059] The Adaptive Feature Iterative Reconstruction Unit (AFIRU) comprises a Multi-Scale Feature Extraction Unit (MSFEU), a Focused Gradient Aware Module (FBPM), a Multi-Scale Frequency Enhancement Module (MSFEM), and a Learnable Edge Convolution (LEconv). It constructs cross-domain feature representations to implicitly fit the second-order Newton update direction. And by learning step size Update the merged image:

[0060] The last FRBlock ( It only contains the AFIRU module and no longer updates multipliers and matrices.

[0061] Furthermore, the overall network architecture consists of an initialization downsampling module, It consists of cascaded feature reconstruction blocks (FRBlocks) and an upsampling output module. In each iteration phase... middle, Internally, it contains, in sequence, an Adaptive Feature Iteration Reconstruction Unit (AFIRU), a Matrix Multiplication Unit (MPU), a Matrix Update Unit (MQU), and an Auxiliary Variable Update Module (ACPU). Among them, the core multi-domain feature extraction is performed by the Multi-Scale Feature Extraction Unit (MSFEU) embedded in each unit.

[0062] The Multi-Scale Feature Extraction Unit (MSFEU) employs a parallel three-branch architecture to explicitly decouple local spatial features, global frequency features, and edge high-frequency features. Its input features are denoted as... .

[0063] The first branch of MSFEU (spatial amplitude branch) extracts local focusing features through channel attention (CA) and cascaded convolutions, and calculates the spatial maximum amplitude as a physical prior constraint. The calculation formula is as follows:

[0064]

[0065] in, This indicates the channel attention module. This represents the maximum spatial amplitude of the feature in the channel dimension, which is used for subsequent adaptive modulation of the frequency domain output.

[0066] The second branch of MSFEU (global frequency branch) achieves frequency sensing of the global receptive field through Fast Fourier Transform (rFFT), performs feature mapping in the complex domain, and finally utilizes... The constraint reduces ringing artifacts; the calculation formula is as follows:

[0067]

[0068] Where rFFT and irFFT are the forward and inverse fast Fourier transforms of real numbers, respectively. This represents a complex linear layer. This branch utilizes spatial magnitude priors. Adaptive amplitude modulation is achieved by nonlinearly scaling the frequency domain characteristics after inverse transformation.

[0069] The third branch of MSFEU (edge ​​enhancement branch) consists of a focus gradient-aware module (FBPM) and a multi-scale frequency enhancement module (MSFEM), which uses learnable edge convolution (LEconv) to extract accurate focus boundary features.

[0070] Furthermore, the Focused Gradient Awareness Module (FBPM) comprises three parallel paths, and its output is normalized by a Complex LayerNorm to obtain intermediate features. The paths are represented as follows: Standard convolution path: ; Channel compression - extended path: ; Isotropic learnable edge paths: ; in The depthwise convolution kernel is initialized as a Laplacian operator:

[0071] The output formula for FBPM is: .

[0072] The Multi-Scale Frequency Enhancement Module (MSFEM) also contains three parallel paths for extracting directional edge features: Standard convolution path: ; Multi-scale path: ; Directional learnable edge paths:

[0073] in, and Initialize them as Sobel-x and Sobel-y operators respectively: Sobel-x operator:

[0074] Sobel-y operator:

[0075] The convolution kernel parameters of all the classical physical operators mentioned above are fully learnable during training. The final output of MSFEM (i.e. )for:

[0076] The final output is a fusion of the three modules:

[0077] Furthermore, the image reconstruction process includes: At the network input end, the source image and spliced ​​together The input features are downsampled to reduce spatial resolution and increase the number of channels to obtain the basic input features. Enter the iterative network.

[0078] At the network output, the last iteration stage Output With initial splicing features The global residuals are summed, then upsampled to restore the original resolution, and finally reconstructed using a lightweight output module (LGOutput) to produce the final fully focused fused image. .

[0079] Furthermore, a hybrid loss function is used for end-to-end training:

[0080] The Charbonnier loss, which measures pixel-level differences in the spatial domain, can be expressed as:

[0081] The frequency domain loss constraint on global structural consistency can be expressed as:

[0082] Edge loss enhances the edge preservation capability and can be expressed as:

[0083] Edge information of the image is extracted using the Laplacian operator to ensure a clear transition at the boundaries of the focused region in the fused image; whereby... This represents the Laplacian operator, which calculates the second derivative of an image to highlight edges and details. During training, the AdamW optimizer is used to minimize the mixture loss function, and the initial learning rate is set to... The weight decay is set to The final learning rate was set to The learning rate scheduler adopts a linear decay strategy; the total number of training epochs is 500, the batch size is 8, and each batch contains 8 image patches with a resolution of 128×128; the gradient of the loss function with respect to the network parameters is calculated through the backpropagation algorithm, and the network parameters are updated through the AdamW optimizer. The network learns the mapping relationship from the source image to the full-focus image, while maintaining consistency in the spatial domain, frequency domain and edge domain.

[0084] The above description represents the preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications are also considered to be within the scope of protection of the present invention.

Claims

1. A multi-focus image fusion method based on the semi-smooth Newton's method, characterized in that, Includes the following steps: S1. Problem Modeling: Construct a mathematical model for multi-focus image fusion, including dataset construction, degradation model establishment, Kronecker integral solution, and variational optimization problem establishment; S2. Problem Solving and Optimization: The variational optimization problem is solved using the semi-smooth Newton algorithm. A Newton linear system is established by introducing Lagrange dual variables and the generalized Jacobian matrix. S3. Network Unfolding and Image Reconstruction: The optimization algorithm is unfolded into a deep network, and iterative optimization is performed through feature reconstruction blocks to finally reconstruct the fully focused image; wherein, the specific process of network unfolding and image reconstruction in step S3 includes: S31: The iterative process of the optimization algorithm is unfolded into a deep network. The deep network contains M = 6 feature reconstruction blocks. Each feature reconstruction block corresponds to one iterative optimization of the semi-smooth Newton algorithm. The algorithm is mapped into a learnable network structure through deep unfolding technology. S32: For the first There are feature reconstruction blocks, among which It consists of four modules: multiplier update module, matrix update module, auxiliary variable update module, and adaptive feature iterative reconstruction unit; S33: The multiplier update module updates according to the current image state. And auxiliary variables, update the Lagrange multipliers through a convolutional neural network. The Lagrange multipliers It is responsible for extracting and accumulating non-smooth high-frequency prior information of the image during iteration; S34: The matrix update module updates the first image based on the current image state and observation data through a normalization operation. Vertical defocus matrix corresponding to each feature reconstruction block and horizontal defocus matrix The matrix update module first calculates the gradient information of the current image, and then extracts the first derivative of the image using a learnable derivative operator. The convolution is implemented by initializing the convolution kernel as a central difference operator, and then normalizing the gradient magnitude to between 0 and 1. The normalization formula is as follows: , where G is the gradient magnitude; S35: The auxiliary variable update module calculates the local focusing response and updates the Lagrange dual variable. This provides structural prior guidance for the next iteration stage; S36: The adaptive feature iteration reconstruction unit includes a multi-scale feature extraction unit, a focused gradient perception module, a multi-scale frequency enhancement module, and a learnable edge convolution, constructing a cross-domain feature representation to implicitly fit the second-order Newton update direction, and using a learnable step size. The formula for updating and merging the image is as follows: ; S37: For the last feature reconstruction block, i.e. It contains only adaptive feature iteration reconstruction units, and finally performs image reconstruction and outputs the final fusion result.

2. The multi-focus image fusion method based on the semi-smooth Newton method according to claim 1, characterized in that, The specific process of problem modeling in step S1 includes: S11: Dataset construction. The DUTS-TR dataset is used as the source data. This dataset contains 10,553 pairs of high-quality images and corresponding binary masks. Multi-focus image pairs are generated by Gaussian blur. The Gaussian blur parameters are adaptively calculated according to the image resolution. The original image is retained as the ground truth of full focus. The training set and test set are divided in a 9:1 ratio. S12: Degradation model establishment, obtaining two source images with different focal regions generated in step S11. ,in Representing the real space of the image, For the number of channels, The spatial resolution of the image height. To determine the spatial resolution of the image width, the source image... Considered as a potential ideal all-focused image After the defocusing operator matrix Degradation observations after action, and the establishment of a degradation model. ,in The defocus operator matrix is ​​represented as an identity matrix in the focused region and as a low-pass filter in the defocused region. Represents the real space of the matrix. For full-focus image Vectorization operator The resulting column vector after expansion The source image is vectorized. This refers to sensor noise; among which The product of the number of channels, height, and width of the image, and the total dimension after vectorization. Represent the real space of vectors; S13: Kronecker integral solution, to avoid defocusing operator matrix The high computational overhead introduced by this problem is addressed by employing the Kronecker integral technique. Based on the Kronecker product and vectorization equivalence theorem, for any matrix... , , The following relationship holds: in This represents the Kronecker product operation. and Let be the transformation matrix. Let the matrix be to be transformed; the one-dimensional degradation process is equivalently reduced back to the two-dimensional tensor space, and the degradation process is modeled as: in, For the matrix form of the first Source image, corresponding to the aforementioned vectorized form Vertical defocus matrix and horizontal defocus matrix Characterize the low-pass degradation properties in the row and column directions respectively, satisfying , The noise is in the reconstructed tensor form; To avoid explicitly extracting the mask, the two observation images are jointly stitched together along the channel dimension to construct a unified high-dimensional observation model. in, This indicates a splicing operation along the channel dimension. This is the stitched image of the joint observation. and These are two source images with different focus regions; in the focus region, the vertical and horizontal defocus matrices degenerate into identity matrices. In the defocused region, the operator behaves as a low-pass blur operator, which significantly reduces computational complexity and memory usage through Kronecker integral solutions. S14: The variational optimization problem is established. Based on the maximum a posteriori probability (MAP) estimation, the reconstruction of the full-focus image F from the degraded observations is equivalent to minimizing the energy functional. in For data fidelity, the Frobenius norm is used. Measure the difference between the reconstructed image and the observed image. This is a non-smoothing regularization term used to apply high-frequency structural priors to natural images. For balancing parameters.

3. The multi-focus image fusion method based on the semi-smooth Newton method according to claim 2, characterized in that, The specific process of solving and optimizing the problem in step S2 includes: S21: Based on the cyclic properties of the Frobenius norm and the trace, and using matrix differential theory, define the residual function: in, Let be the residual function, and when the residual is zero, it corresponds to the optimal solution of the optimization problem. Indicates the first The estimated value of the fully focused image during each iteration; It is a proximal operator used to handle implicit updates of non-smooth regularization terms; This is the iteration step size parameter; S22: By explicitly introducing Lagrange dual variables As a multiplier, it is responsible for extracting and accumulating the non-smooth high-frequency prior information of the image during iteration, thereby transforming the residual function into a fixed-point root solving problem. : in This indicates a feature fusion operation, used to combine dual variables. The encoded high-frequency prior information is injected into the input of the proximal operator; S23: To introduce second-order curvature information, the generalized Jacobian matrix is ​​used. A semi-smooth Newton equation is established, and the generalized chain rule is applied to expand it, deriving the first... Newtonian linear systems of steps: in, It is the identity matrix. Let be the generalized Jacobian derivative matrix of the proximal operator at the current iteration point. For the first The update increment of the step; the high-precision iterative update step size is obtained by solving the linear system.

4. The multi-focus image fusion method based on the semi-smooth Newton method according to claim 1, characterized in that, In step S36, the multi-scale feature extraction unit adopts a parallel three-branch architecture; the first branch extracts local focusing features through channel attention and cascaded convolution, and calculates the spatial maximum amplitude as a physical prior constraint, as shown in the following formula: in, As input features, This indicates the channel attention module. Indicates the kernel size as Convolution operation, To modify the activation function of the linear unit, The output features of the first branch, Features The spatial maximum amplitude map, obtained by taking the maximum value along the channel dimension, is used for subsequent adaptive modulation of the frequency domain output. For channel dimensions; The second branch achieves frequency sensing of the global receptive field through Fast Fourier Transform, performs feature mapping in the complex domain, and finally utilizes... The constraint reduces ringing artifacts; the calculation formula is as follows: in, and These are the forward and inverse Fast Fourier Transforms for real numbers, respectively. Represents a complex linear layer. For batch normalization operations, These are intermediate features after frequency domain transformation. The output features of the second branch, The hyperbolic tangent activation function is used to normalize the response to... interval, Represents the element-wise Hadamard product operation; The third branch consists of a focus gradient sensing module and a multi-scale frequency enhancement module, which utilizes learnable edge convolution to extract accurate focus boundary features.

5. The multi-focus image fusion method based on the semi-smooth Newton method according to claim 4, characterized in that, The focused gradient sensing module includes three parallel paths, whose outputs are normalized by complex layers to obtain intermediate features. The paths are represented as follows: Standard convolution path: ; Channel compression - extended path: ; Isotropic learnable edge paths: ; in For learnable edge convolution functions, The kernel size is The convolution operation can learn edge convolutions and initialize the depthwise convolution kernel with a Laplacian operator: The output formula of the focused gradient sensing module is: , ( ) is a complex linear normalization function.

6. The multi-focus image fusion method based on the semi-smooth Newton method according to claim 5, characterized in that, The multi-scale frequency enhancement module contains three parallel paths for extracting directional edge features: Standard convolution path: ; Multi-scale path: ; Directional learnable edge paths: ; in, This indicates a splicing operation along the channel dimension. Learnable edge convolutions in the horizontal direction Learnable edge convolutions in the vertical direction. and Initialize them as Sobel-x and Sobel-y operators respectively: Sobel-x operator: Sobel-y operator: The final output of the multi-scale frequency enhancement module for: The final output is a fusion of the three modules: .

7. The multi-focus image fusion method based on the semi-smooth Newton method according to claim 6, characterized in that, The image reconstruction process in step S37 includes: At the network input end, the source image and spliced ​​together The input features are downsampled to reduce spatial resolution and increase the number of channels to obtain the basic input features. Enter the iterative network; At the network output, the output of the feature reconstruction block in the last iteration stage is combined with the initial concatenated features. The global residuals are summed, then the image is restored to its original resolution using an upsampling module, and finally reconstructed using a lightweight output module to produce the final fully focused fused image. .

8. The multi-focus image fusion method based on the semi-smooth Newton method according to claim 7, characterized in that, A hybrid loss function is used for training the network from its input to its output: The Charbonnier loss, which measures pixel-level differences in the spatial domain, is expressed as: The frequency domain loss constraint for global structural consistency is expressed as: The edge loss enhancement effect on edge preservation is expressed as: in, For the total loss function, For Charbonnier loss, and These are the weighting coefficients for the frequency domain loss and the edge loss, respectively. For frequency domain loss, For edge loss, This represents the number of image samples in the batch. For sample index, For the network prediction of the first A fused image, For the first A fully focused ground truth image, The square of the Frobenius norm. To prevent gradient vanishing, a small positive constant, For Fourier transform operators, The Laplacian edge detection operator is used to calculate the second derivative of an image to highlight edge and detail information. for Norm.

9. A multi-focus image fusion system based on the semi-smooth Newton's method, characterized in that, Image fusion is performed using the multi-focus image fusion method based on the semi-smooth Newton method as described in any one of claims 1-8, and includes the following modules: Problem modeling module: Constructing a mathematical model for multi-focus image fusion, including dataset construction, degradation model establishment, Kronecker integral solution, and variational optimization problem establishment; Problem Solving and Optimization Module: The semi-smooth Newton algorithm is used to solve variational optimization problems. By introducing Lagrange dual variables and the generalized Jacobian matrix, a Newtonian linear system is established. Network Unfolding and Image Reconstruction Module: The optimization algorithm is unfolded into a deep network, and through iterative optimization of feature reconstruction blocks, the fully focused image is finally reconstructed.

Citation Information

Patent Citations

  • Multi-focus image fusion method and system based on multi-scale image fusion network

    CN121414596A

  • Explanatable compressed sensing image reconstruction method and system based on value domain-null space decomposition

    CN121544678A