A method for enhancing robustness of an image fusion model based on a sliding mode variable structure control
By adaptively controlling the fusion of high- and low-frequency sub-images of infrared and visible light images through the sliding mode variable structure control algorithm, the robustness and adaptability problems of image fusion models in complex environments in existing technologies are solved, and higher quality fused images are generated.
Patent Information
- Application Number
- CN202511129562.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-13
- Publication Date
- 2025-10-24
- Estimated Expiration
- 2045-08-13
AI Technical Summary
Existing infrared and visible light image fusion methods lack robustness in complex environments, have poor adaptability to dynamic scenes, and the fusion results are easily distorted.
An image fusion model based on sliding mode variable structure control is adopted. The image is decomposed into high-frequency and low-frequency layers through co-occurrence analysis and shearlet transform. The high- and low-frequency sliding mode surfaces and control laws are designed to adaptively adjust the fusion weights and realize the adaptive fusion of infrared and visible light images.
The robustness and adaptability of the image fusion model have been improved, which can effectively reduce the impact of complex environments and changes in lighting conditions on the fusion effect, and generate more comprehensive and reliable fused images.
Smart Images

Figure CN120634883B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image fusion, and in particular to a method for enhancing the robustness of an image fusion model (infrared and visible light image fusion model) based on sliding mode variable structure control. Background Art
[0002] Infrared and visible light image fusion is a key research area in computer vision and image processing. It aims to combine the thermal radiation information of infrared images with the texture details of visible light images to produce a more comprehensive and reliable fused image. This technology has widespread applications in military reconnaissance, security monitoring, medical diagnosis, and autonomous driving.
[0003] Common image fusion methods can be roughly divided into traditional methods and deep learning methods. Traditional methods can be categorized into methods such as multi-scale transformation, sparse representation, and optimization models; deep learning methods include Convolutional Neural Network (CNN) architecture design, Generative Adversarial Network (GAN) variants, attention mechanisms, and Transformer applications. However, fusion models are sensitive to noise and interference, lack robustness in complex environments, and have poor adaptability to dynamic scenes. In recent years, Wang proposed a method for infrared and visible light image fusion based on the Proportional-Integral-Derivative (PID) control algorithm. While this method improves its adaptability to dynamic scenes, its ability to suppress noise and interference is weak, and the fusion rules may not match the current imaging conditions, causing the fusion algorithm to fail.
[0004] Currently, the main fusion algorithms rely on manually designed fusion parameters or rules based on engineering experience or large amounts of data samples to achieve the fusion of infrared and visible light images. Because the adaptability of manually designed fusion rules is extremely limited, control algorithms are often required to improve the universality of the fusion model. PID control, a typical control algorithm, is the most widely used. It combines the proportional (P), integral (I), and differential (D) steps to adjust the system output, allowing it to quickly and stably reach the target value. However, the PID control algorithm has poor robustness and cannot adapt to complex and changing weather and lighting conditions. Summary of the Invention
[0005] The present invention aims to solve the technical problems in the prior art of poor robustness, low universality and easy distortion of fusion results of fusion methods, and provides a method for enhancing the robustness of image fusion models based on sliding mode variable structure control.
[0006] To solve the above technical problems, the technical scheme of the present application is as follows:
[0007] A method for enhancing the robustness of an image fusion model based on a sliding mode variable structure control, wherein the image fusion model is an infrared and visible light image fusion model, comprising the following steps:
[0008] Step 1: obtaining original infrared images and original visible light images from the same scene and having completed image registration;
[0009] Step 2: decomposition:
[0010] Based on the co-occurrence analysis shearlet transform, the original infrared images and the original visible light images obtained in Step 1 are decomposed into high-frequency layers and low-frequency layers, respectively;
[0011] Step 3: high-frequency subgraph fusion:
[0012] Designing a high-frequency sliding mode surface and a control law to guide the fusion of the visible light high-frequency subgraph and the infrared high-frequency subgraph obtained in Step 2;
[0013] Step 4: low-frequency subgraph fusion:
[0014] Designing a low-frequency sliding mode surface and a control law to guide the fusion of the visible light low-frequency subgraph and the infrared low-frequency subgraph obtained in Step 2;
[0015] Step 5: inverse transform:
[0016] Performing co-occurrence analysis shearlet inverse transform on the fused high-frequency images and low-frequency images obtained in Steps 3 and 4, respectively, to reconstruct the fusion image.
[0017] In the above technical scheme, Step 2 specifically comprises the following steps:
[0018] Step 21: multi-scale decomposition;
[0019] ;
[0020] Wherein, represents the low-frequency subgraph under different scales; is the scale number, and the maximum value is ; When , it represents the infrared image, and when , it represents the visible light image; represents the high-frequency subgraph under different scales; is a co-occurrence filter operator; represents the low-frequency subgraph under a scale;
[0021] The co-occurrence filter is defined as follows:
[0022] ;
[0023] ;
[0024] wherein, and are the output and input pixel values, respectively; and are the index signs of the pixels; is the normalization factor, a weight assignment term, whose value represents the contribution of the input pixel to the output pixel ; is the normalized co-occurrence matrix, represents the number of pixels ; is a Gaussian filter, as shown in the following equation,
[0025] ;
[0026] wherein the change of the filter scale parameter controls the filtering of the edge texture in the image;
[0027] is calculated by the co-occurrence matrix, as shown in the following equation;
[0028] ;
[0029] wherein, is the co-occurrence information provided by the co-occurrence matrix; and correspond to the frequencies of the pixels and in the images, which are the histograms of the pixel values, and are defined as follows:
[0030] ;
[0031] ;
[0032] wherein, is the Euclidean distance between the pixels and ; the co-occurrence matrix parameter is set to ; the operator defines a Boolean relationship: the Boolean value in the square brackets is false, and the expression takes the value of 0, and vice versa; represents the pixel value of the pixel ; represents the frequency of the pixel in the corresponding image; refers to the pixel or ;
[0033] Step22, multi-directional decomposition;
[0034] Step221, generating a Meyer window function with a Meyer wavelet function , defined as follows: ;
[0035] ;
[0036] wherein, represents a wavelet function, represents a wavelet function of a variable , is a variable greater than 0, is an integer variable from 1 to 15, is a variable whose value is determined by ;
[0037] Step222, resampling the Meyer window function into a pseudo-polarized coordinate grid under a filter window, and mapping the sampling result to a Cartesian coordinate system;
[0038] Step223, transforming the shear wave filter set in the frequency domain to the time domain through inverse Fourier transform to obtain the final filter set, denoted as , the value of the parameter is only 0 or 1; is the number of decomposition directions, whose value is ; denotes the number of scales;
[0039] Step 224, performing convolution operation on the detail layer component and the filter set in the time domain to obtain the final multi-directional detail sub-band image, i.e., high-frequency sub-image ;
[0040] .
[0041] In the above technical solution, the third step is specifically:
[0042] Step31, defining a high-frequency coefficient fusion sliding mode surface function :
[0043] ;
[0044] wherein, and represent infrared local energy and visible light local energy, respectively; and respectively represent the infrared gradient component and the visible light gradient component; and are high-frequency coefficient fusion sliding mode surface parameters;
[0045] Step 32, design a high-frequency coefficient fusion control law :
[0046] ;
[0047] wherein, is a sign function;
[0048] Step 33, the high-frequency coefficient selection is determined by :
[0049] ;
[0050] wherein, is a high-frequency fusion coefficient, represents an infrared high-frequency subgraph, represents a visible light high-frequency subgraph.
[0051] In the above technical solution, the fourth step is specifically:
[0052] Step 41, define a low-frequency coefficient fusion sliding mode surface function :
[0053] ;
[0054] wherein, and respectively represent the infrared low-frequency component average value and the visible light low-frequency component average value; and respectively represent the infrared low-frequency component variance and the visible light low-frequency component variance; is a low-frequency coefficient fusion sliding mode surface parameter;
[0055] Step 42, design a low-frequency coefficient fusion control law :
[0056] ;
[0057] wherein, is a control gain, is a saturation function, satisfying:
[0058] ;
[0059] wherein, δ is the boundary layer thickness, is a sign function;
[0060] Step 43, low-frequency coefficient fusion weight guides low-frequency coefficient fusion:
[0061] ;
[0062] wherein, is a low-frequency fusion coefficient, represents an infrared low-frequency subgraph, represents a visible light low-frequency subgraph.
[0063] The present application has the following beneficial effects:
[0064] The image fusion model (infrared and visible light image fusion model) robustness enhancement method based on sliding mode variable structure control of the present application adopts sliding mode variable structure control (Sliding Mode Control, SMC) as the fusion rule regulation means, and is introduced into the image fusion field due to its insensitivity to system parameter changes and external disturbances, and its core advantages include:
[0065] Strong robustness: strong inhibition ability to system uncertainty and external disturbance;
[0066] Fast response: can realize finite time convergence;
[0067] Design flexibility: different control objectives can be achieved through sliding mode surface design.
[0068] The image fusion model (infrared and visible light image fusion model) robustness enhancement method based on sliding mode variable structure control of the present application, under the framework of co-occurrence analysis shear wave transform, designs respective sliding mode surface functions and control laws for infrared and visible light high and low frequency subgraphs, adaptively fuses high and low frequency subgraphs, enhances the robustness of the fusion model, and can reduce the influence of different imaging environments (such as rain, fog, etc.) and sudden changes in light conditions on the fusion effect. BRIEF DESCRIPTION OF DRAWINGS
[0069] The present application will be further described in detail below in combination with the drawings and specific embodiments.
[0070] Figure 1 It is a flowchart of the image fusion model robustness enhancement method based on sliding mode variable structure control of the present application.
[0071] Figure 2 It is a schematic diagram of the overall framework principle of the image fusion model robustness enhancement method based on sliding mode variable structure control of the present application.
[0072] Figure 3 It is a schematic diagram of the pseudo-polarization coordinate grid after two-direction merging with a window size of 15x15.
[0073] Figure 4Schematic diagram of the saturation function principle.
[0074] Figure 5 This is a schematic diagram showing the comparison results of the comparative test between the method of the present invention and three existing image fusion methods. DETAILED DESCRIPTION
[0075] The inventive concept of the present invention is:
[0076] The image fusion model robustness enhancement method based on sliding mode variable structure control of the present invention is a method for adaptive fusion of infrared images and visible light images to solve the shortcomings of existing fusion models such as poor robustness, low universality and easy distortion of fusion results.
[0077] Currently, most fusion models rely on extensive sample training or the manual design of fusion rules based on the engineering design experience of researchers to achieve effective fusion of infrared and visible light images. Because the adaptability of a single fusion rule is limited, robust control algorithms are being introduced to improve the adaptability of fusion models. This invention leverages the robustness of sliding mode variable structure control and its insensitivity to system parameter changes and external disturbances to effectively enhance the adaptability of infrared and visible light image fusion models.
[0078] This invention uses a sliding mode variable structure control algorithm to regulate the fusion rules, replacing the single fusion rule used in traditional algorithms. When fusing infrared and visible light images, the fusion weights are adjusted in real time based on different imaging environments (such as rain and fog) and lighting conditions. This design significantly improves the model's adaptability. By directly utilizing a discontinuous control law, the infrared and visible light fusion system's state moves along a preset sliding mode surface, resulting in strong robustness.
[0079] The present invention will be described in detail below with reference to the accompanying drawings.
[0080] like Figure 1 As shown ( Figure 1 Only the method outline of the present invention is shown in the figure). The present invention provides a method for enhancing the robustness of an image fusion model based on sliding mode variable structure control, wherein the image fusion model is an infrared and visible light image fusion model. The present invention implements dynamic weight adjustment through sliding mode variable structure control. The principle diagram thereof is shown in Figure 2, and includes the following steps:
[0081] The first step is to obtain the original infrared image and the original visible light image from the same scene and complete the image registration.
[0082] Step 2: Decomposition:
[0083] Based on the co-occurrence analysis shear wave transform, the original infrared image and the original visible light image obtained in the first step are decomposed into high-frequency layers and low-frequency layers respectively; that is, the original infrared image and the original visible light image are processed by the co-occurrence analysis shear wave transform respectively; specifically including the following steps:
[0084] The co-occurrence filter is used for multi-scale decomposition of the image, and the high-frequency layer subgraph is decomposed in multiple directions by the discrete shear wave transform, so that different texture information can be effectively targeted, the edges of large-scale contours and texture regions are separated into different levels and saved respectively, and the infrared low-frequency subgraph and a series of infrared high-frequency subgraphs and visible light low-frequency subgraphs and a series of visible light high-frequency subgraphs are obtained.
[0085] Step 21, co-occurrence filter - multi-scale decomposition:
[0086] ;
[0087] Among them, represents the low-frequency subgraph under different scales; is the scale number (the maximum value is ); When , it represents the infrared image, and when , it represents the visible light image; represents the high-frequency subgraph under different scales; is the co-occurrence filter operator; represents the low-frequency subgraph under the scale.
[0088] The co-occurrence filter is a kind of local linear filter, and its definition is as follows:
[0089] ;
[0090] ;
[0091] In the formula, and are the pixel values of the output and the input; and both represent the index mark of the pixel; and is a normalization factor, which can be regarded as a weight allocation item, and its value represents the contribution of the input pixel to the output pixel ; is a normalized co-occurrence matrix, represents the number of pixels , is a Gaussian filter, and its specific content is shown in the following formula,
[0092] ;
[0093] where the change of filter scale parameter controls the filtering of edge texture in the image, which cannot avoid the problem of edge blurring in the smoothing process and needs to rely on the guidance of image co-occurrence information distribution weight.
[0094] It can be obtained by co-occurrence matrix calculation, as shown in the following formula:
[0095] ;
[0096] where, is the co-occurrence information provided by the co-occurrence matrix; and correspond to the frequency of the pixel and in the image, which is the histogram of the pixel value, and is defined as follows:
[0097] ;
[0098] ;
[0099] In the formula, is the Euclidean distance between the pixels and ; the co-occurrence matrix parameter is set to ; the operator defines a Boolean relationship: the Boolean value in the square brackets is false, and the expression takes the value 0, and vice versa. represents the pixel value of the pixel ; represents the frequency of the pixel in the corresponding image; refers to the pixel or .
[0100] Step 22, discrete shear wave transform-multiple direction decomposition:
[0101] The specific steps of direction localization realized by adaptive shear wave filter are as follows:
[0102] Step 221, generate Meyer window function with Meyer wavelet function , which is defined as follows: ;
[0103] ;
[0104] In the formula, representative wavelet function, representative variable representative wavelet function, is a variable greater than 0, is an integer variable from 1 to 15, is a variable whose value is determined by ;
[0105] Step 222, the Meyer window function is put into the pseudo-polarization coordinate grid under the filter window (window size is 15x15) (as shown in Figure 3 ) for resampling, and the sampling results are mapped into the Cartesian coordinate system;
[0106] Step 223, the shear wave filter bank in the frequency domain is transformed into the time domain by inverse Fourier transform to obtain the final filter bank, denoted as , the value of parameter is only 0 or 1; is the number of decomposition directions, whose value is ; denotes the number of scales;
[0107] Step 224, convolution operation is performed on the detail layer component and the filter bank in the time domain to obtain the final multi-directional detail sub-band image, i.e. high-frequency subgraph .
[0108] ;
[0109] Third step, high-frequency subgraph fusion:
[0110] The high-frequency sliding mode surface and control law are designed to guide the fusion of the visible light high-frequency subgraph and the infrared high-frequency subgraph obtained in the second step. A sliding mode fusion strategy is designed for the high-frequency coefficients based on local energy and gradient, which includes the following steps
[0111] Step 31, define the high-frequency coefficient fusion sliding mode surface function :
[0112] ;
[0113] wherein, and represent the infrared local energy and the visible light local energy, respectively; and represent the infrared gradient component and the visible light gradient component, respectively; and are both high-frequency coefficient fusion sliding mode surface parameters.
[0114] Step 32, design the high-frequency coefficient fusion control law :
[0115] ;
[0116] wherein, is a sign function.
[0117] Step 33, the high-frequency coefficient selection is determined by :
[0118] ;
[0119] wherein, is a high-frequency fusion coefficient, represents an infrared high-frequency subgraph, represents a visible light high-frequency subgraph.
[0120] Fourth step, low-frequency subgraph fusion:
[0121] The low-frequency sliding mode surface and control law are designed to guide the fusion of the visible light low-frequency subgraph and the infrared low-frequency subgraph obtained in the second step. The sliding mode fusion strategy is designed for the low-frequency coefficient, which includes the following steps:
[0122] Step 41, design a low-frequency coefficient fusion sliding mode surface function (i.e. a low-frequency adaptive weighted sliding mode surface):
[0123] ;
[0124] wherein, and represent the average value of the infrared low-frequency component and the average value of the visible light low-frequency component, respectively; and and represent the variance of the infrared low-frequency component and the variance of the visible light low-frequency component, respectively; is a low-frequency coefficient fusion sliding mode surface parameter.
[0125] Step 42, design a low-frequency coefficient fusion control law:
[0126] ;
[0127] wherein, is a control gain, is a saturation function (as shown in Figure 4 ), which satisfies:
[0128] ;
[0129] wherein, δ is the boundary layer thickness, is a sign function.
[0130] In order to suppress chattering during the control process, the following Figure 4 middle( x, y ) coordinate system and ( s, u The saturation function shown in the ) coordinate system essentially replaces the relay characteristics with saturation characteristics to mitigate the discontinuity of control switching. It is used to demonstrate a linear relationship within the boundary layer and a fixed value outside the boundary layer. The saturation function has a three-segment structure and two switching planes: and , and are the upper and lower bounds of the boundary layer, respectively; is a constant, which is the boundary layer thickness; in addition is the boundary layer width.
[0131] Step 43, using the low-frequency coefficient fusion weight (the low-frequency coefficient fusion weight refers to the low-frequency coefficient fusion control law ) guides the fusion of low-frequency coefficients:
[0132] ;
[0133] in, is the low-frequency fusion coefficient, represents the infrared low-frequency sub-image, Represents the visible light low-frequency sub-image.
[0134] Step 5, inverse transform:
[0135] The co-occurrence analysis and shearlet inverse transform are performed on the fused high-frequency image and low-frequency image obtained in the third and fourth steps respectively to reconstruct the fused image.
[0136] The final fusion result is obtained by co-occurrence analysis and shearlet inverse transformation, such as Figure 5 As shown in the figure, the present invention is compared with three existing image fusion methods, among which (a) is an infrared image (Infrared), (b) is a visible light image (Visible), (c) is an anisotropic diffusion filter (ADF) method, (d) is a unified unsupervised fusion (U2Fusion) method, (e) is a generative adversarial network fusion (FusionGAN) method, and (f) is the present invention (Ours) method. Through the above experiments, it can be seen that the image fused by the method of the present invention has higher infrared target brightness, richer scene details, and higher contrast. Therefore, based on the strong robust performance of the sliding mode variable structure control algorithm, the present invention guides the adaptive asymptotic fusion of high-frequency sub-images and low-frequency sub-images, which can effectively improve the universality of the fusion model to imaging environments and conditions.
[0137] Obviously, the above embodiments are merely example for clearly illustrating but not limitation to the embodiments. Based on the above description, other different forms of changes or variations can be made by those skilled in the art. Here, all the embodiments need not and can not be enumerated. The obvious changes or variations derived from the above description are still within the protection scope of the present application.
Claims
1. A method for enhancing robustness of an image fusion model based on sliding mode variable structure control, characterized in that, Wherein the image fusion model is an infrared and visible light image fusion model, comprising the following steps: Step 1, obtaining original infrared images and original visible light images from the same scene and having completed image registration; Step 2, decomposition: Based on the co-occurrence analysis shearlet transform, the original infrared images and the original visible light images obtained in step 1 are decomposed into high-frequency layers and low-frequency layers respectively; Step 3, high-frequency subgraph fusion: A high-frequency sliding surface and a control law are designed to guide the fusion of the visible light high-frequency subgraph and the infrared high-frequency subgraph obtained in step 2; Step 4, low-frequency subgraph fusion: A low-frequency sliding surface and a control law are designed to guide the fusion of the visible light low-frequency subgraph and the infrared low-frequency subgraph obtained in step 2; Step 5, inverse transform: The fused high-frequency images and low-frequency images obtained in steps 3 and 4 are respectively subjected to co-occurrence analysis shearlet inverse transform to reconstruct the fused images; Step 3 is specifically: Step 31, define high frequency coefficient fusion sliding mode surface function : ; wherein, and respectively represent the infrared local energy and the visible light local energy; and respectively represent the infrared gradient component and the visible light gradient component; and are both high-frequency coefficient fusion sliding mode surface parameters; Step 32, design high frequency coefficient fusion control law : ; wherein is a sign function; Step 33, the high frequency coefficient selection is determined by Step 33, the high frequency coefficient selection is determined by ; wherein, is a high frequency fusion coefficient, represents an infrared high frequency subgraph, represents a visible light high frequency subgraph; Step 4 is specifically: Step 41, define the low frequency coefficient fusion sliding mode surface function : ; wherein, and respectively represent the average of the infrared low-frequency component and the average of the visible light low-frequency component; and respectively represent the variance of the infrared low-frequency component and the variance of the visible light low-frequency component; is the low-frequency coefficient fusion sliding mode surface parameter; Step 42, design low frequency coefficient fusion control law : ; wherein is a control gain, is a saturation function satisfying: ; wherein δ is the boundary layer thickness, is the sign function; Step 43, low-frequency coefficient fusion guided by low-frequency coefficient fusion weight: ; wherein, is a low frequency fusion coefficient, represents an infrared low frequency subgraph, represents a visible light low frequency subgraph.
2. The method of claim 1, wherein the image fusion model robustness enhancement based on a sliding mode variable structure control is characterized by, Step 2 is specifically Comprising the following steps: Step 21, multi-scale decomposition; ; wherein, represent low frequency subgraphs at different scales; is the number of scales, and the maximum value is ; is taken as 1, and is taken as 0, then the infrared image is represented, and when is taken as 1, and is taken as 0, then the visible light image is represented; represent high frequency subgraphs at different scales; is a co-occurrence filtering operator; low frequency subgraphs at different scales; The co-occurrence filter is defined as follows: ; ; wherein, and are the output and input pixel values, respectively; and are the index of the pixel; is the normalization factor, and is the weight assignment term, whose value represents the contribution of the input pixel to the output pixel ; and is the normalized co-occurrence matrix, represents the number of pixels , is a Gaussian filter, as shown in the following equation, ; wherein the filter scale parameter controls the filtering of edge texture in the image; The co-occurrence matrix calculation is obtained as follows; ; wherein, is co-occurrence information, provided by a co-occurrence matrix; and correspond to pixels in the image and are frequencies, a histogram of pixel values, defined as follows: ; ; wherein is a pixel and is the Euclidean distance between the pixels is set to ; the operator defines a Boolean relation: the Boolean value inside the square brackets is false, the expression takes the value 0, and vice versa, the value 1; represents the pixel value of the pixel ; represents the frequency of the pixel in the corresponding image; refers to the pixel or ; Step 22, multi-directional decomposition; Step 221, generating a Meyer window function with a Meyer wavelet function , defined as follows: ; ; wherein represents a wavelet function, represents a variable represents a wavelet function, is a variable greater than 0, is an integer variable from 1 to 15, is a variable whose value is determined by ; Step 222, the Meyer window function is put into the pseudo-polarized coordinate grid under the filter window for resampling, and the sampling results are mapped into the Cartesian coordinate system; Step 223, transform the shear wave filter set in the frequency domain to the time domain by inverse Fourier transform to obtain the final filter set, denoted as , the value of the parameter is only 0 or 1; is the number of decomposition directions, and the value is ; indicates the number of scales. Step 224, with the filter bank in the time domain and the detail layer component performing a convolution operation to obtain a final multi-directional detail subband image, i.e., a high-frequency subgraph ; 。
Citation Information
Patent Citations
Infrared image and visible image fusion method
CN108389158A
Infrared and low-illumination image fusion method based on non-subsampled shear transformation
CN117994146A