Image fusion model robustness enhancement method based on sliding mode variable structure control
The high- and low-frequency decomposition of infrared and visible light images is performed through the sliding mode variable structure control algorithm, and the sliding surface control law is designed to solve the robustness and adaptability problems of the existing model in complex environments and achieve a more stable image fusion effect.
Patent Information
- Application Number
- CN202511129562.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-13
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2045-08-13
AI Technical Summary
The existing infrared and visible light image fusion model is not robust enough in complex environments, has poor adaptability to dynamic scenes, and the fusion results are easily distorted.
A sliding mode variable structure control algorithm is adopted to decompose infrared and visible light images into high-frequency and low-frequency layers through co-occurrence analysis and shearlet transform. The corresponding sliding surface and control law are designed to fuse the high- and low-frequency sub-images respectively, and the fusion weight can be dynamically adjusted.
The robustness of the image fusion model has been enhanced, which can effectively resist changes in system parameters and external disturbances, respond quickly, and improve the adaptability in different imaging environments and lighting conditions.
Smart Images

Figure CN120634883A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image fusion, and in particular to a method for enhancing the robustness of an image fusion model (infrared and visible light image fusion model) based on sliding mode variable structure control. Background Art
[0002] Infrared and visible light image fusion is a key research area in computer vision and image processing. It aims to combine the thermal radiation information of infrared images with the texture details of visible light images to produce a more comprehensive and reliable fused image. This technology has widespread applications in military reconnaissance, security monitoring, medical diagnosis, and autonomous driving.
[0003] Common image fusion methods can be roughly divided into traditional methods and deep learning methods. Traditional methods can be categorized into methods such as multi-scale transformation, sparse representation, and optimization models; deep learning methods include Convolutional Neural Network (CNN) architecture design, Generative Adversarial Network (GAN) variants, attention mechanisms, and Transformer applications. However, fusion models are sensitive to noise and interference, lack robustness in complex environments, and have poor adaptability to dynamic scenes. In recent years, Wang proposed a method for infrared and visible light image fusion based on the Proportional-Integral-Derivative (PID) control algorithm. While this method improves its adaptability to dynamic scenes, its ability to suppress noise and interference is weak, and the fusion rules may not match the current imaging conditions, causing the fusion algorithm to fail.
[0004] Currently, the main fusion algorithms rely on manually designed fusion parameters or rules based on engineering experience or large amounts of data samples to achieve the fusion of infrared and visible light images. Because the adaptability of manually designed fusion rules is extremely limited, control algorithms are often required to improve the universality of the fusion model. PID control, a typical control algorithm, is the most widely used. It combines the proportional (P), integral (I), and differential (D) steps to adjust the system output, allowing it to quickly and stably reach the target value. However, the PID control algorithm has poor robustness and cannot adapt to complex and changing weather and lighting conditions. Summary of the Invention
[0005] The present invention aims to solve the technical problems in the prior art of poor robustness, low universality and easy distortion of fusion results of fusion methods, and provides a method for enhancing the robustness of image fusion models based on sliding mode variable structure control.
[0006] In order to solve the above technical problems, the technical solutions of the present invention are as follows:
[0007] A method for enhancing the robustness of an image fusion model based on sliding mode variable structure control, wherein the image fusion model is an infrared and visible light image fusion model, comprises the following steps:
[0008] The first step is to obtain the original infrared image and the original visible light image from the same scene and complete the image registration;
[0009] Step 2: Decomposition:
[0010] Based on the co-occurrence analysis, the shearlet transform is used to decompose the original infrared image and the original visible light image obtained in the first step into high-frequency layer and low-frequency layer respectively;
[0011] The third step is to fuse high-frequency sub-graphs:
[0012] Design a high-frequency sliding surface and control law to guide the fusion of the visible light high-frequency sub-image and the infrared high-frequency sub-image obtained in the second step;
[0013] Step 4: low-frequency sub-image fusion:
[0014] Design a low-frequency sliding surface and control law to guide the fusion of the visible light low-frequency sub-image and the infrared low-frequency sub-image obtained in the second step;
[0015] Step 5, inverse transform:
[0016] The co-occurrence analysis and shearlet inverse transform are performed on the fused high-frequency image and low-frequency image obtained in the third and fourth steps respectively to reconstruct the fused image.
[0017] In the above technical solution, the second step specifically includes the following steps:
[0018] Step 21, multi-scale decomposition;
[0019] ;
[0020] in, Represents low-frequency subgraphs at different scales; is the scale number, the maximum value is ; Pick When represents the infrared image, take When represents the visible light image; Represents high-frequency subgraphs at different scales; is the co-occurrence filtering operator; represent Low-frequency subgraphs at scale;
[0021] The co-occurrence filter is defined as follows:
[0022] ;
[0023] ;
[0024] Where, and are the output and input pixel values respectively; and Both represent the index flag of the pixel; Is the normalization factor, which is the weight distribution term, and its value represents the input pixel To output pixel contribution; is the normalized co-occurrence matrix, Representative pixels the number of is a Gaussian filter, as shown below:
[0025] ;
[0026] Among them, the filter scale parameter The change of controls the filtering of edge texture in the image;
[0027] It is obtained by calculating the co-occurrence matrix, as shown below;
[0028] ;
[0029] in, is the co-occurrence information, provided by the co-occurrence matrix; and Corresponding pixels in the image and The frequency of is the histogram of pixel values, which is defined as follows:
[0030] ;
[0031] ;
[0032] Where, It's a pixel and Euclidean distance between; co-occurrence matrix parameters is set to ; operator A Boolean relationship is defined: if the Boolean value in the square brackets is false, the expression evaluates to 0, otherwise it evaluates to 1; Representative pixels Pixel value of Represents the pixel in the corresponding image frequency; Refers to pixels or ;
[0033] Step 22, multi-directional decomposition;
[0034] Step 221, generate Meyer window function using Meyer wavelet function , defined as follows: ;
[0035] ;
[0036] Where, represents the wavelet function, Representative variables The wavelet function of is a variable greater than 0, is an integer variable ranging from 1 to 15. is a variable whose value is determined by Decide;
[0037] Step 222, placing the Meyer window function into the pseudo-polarization coordinate grid under the filter window for resampling, and mapping the sampling result to the Cartesian coordinate system;
[0038] Step 223, transform the shearlet filter bank in the frequency domain to the time domain through inverse Fourier transform to obtain the final filter bank, which is recorded as ,parameter The value of is only 0 or 1; is the number of decomposition directions, and its value is ; Indicates the scale number;
[0039] Step 224, the filter bank in the detail layer component and time domain Perform convolution operation to obtain the final multi-directional detail sub-band image, that is, the high-frequency sub-image ;
[0040] .
[0041] In the above technical solution, the third step is specifically as follows:
[0042] Step 31: Define the high-frequency coefficient fusion sliding surface function :
[0043] ;
[0044] in, and represent infrared local energy and visible light local energy respectively; and Represent the infrared gradient component and the visible light gradient component respectively; and All are high frequency coefficients fused with sliding surface parameters;
[0045] Step 32: Design high-frequency coefficient fusion control law :
[0046] ;
[0047] in, is a symbolic function;
[0048] Step 33, high frequency coefficient selection is done by Decide:
[0049] ;
[0050] in, is the high frequency fusion coefficient, represents the infrared high-frequency sub-image, Represents the visible light high-frequency sub-graph.
[0051] In the above technical solution, the fourth step is specifically as follows:
[0052] Step 41: Define the low-frequency coefficient fusion sliding surface function :
[0053] ;
[0054] in, and Represent the average value of infrared low-frequency component and visible light low-frequency component respectively; and represent the variance of infrared low-frequency component and visible light low-frequency component respectively; is the low-frequency coefficient fusion sliding surface parameter;
[0055] Step 42: Design the low-frequency coefficient fusion control law :
[0056] ;
[0057] in, To control the gain, is a saturation function that satisfies:
[0058] ;
[0059] Where δ is the boundary layer thickness, is a symbolic function;
[0060] Step 43: Use the low-frequency coefficient fusion weight to guide the low-frequency coefficient fusion:
[0061] ;
[0062] in, is the low-frequency fusion coefficient, represents the infrared low-frequency sub-image, Represents the visible light low-frequency sub-image.
[0063] The present invention has the following beneficial effects:
[0064] The robustness enhancement method for the image fusion model (infrared and visible light image fusion model) based on sliding mode variable structure control (SMC) in this paper uses sliding mode variable structure control (SMC) as the fusion rule control method. SMC has been introduced into the field of image fusion due to its insensitivity to system parameter changes and external interference. Its core advantages include:
[0065] Strong robustness: strong ability to suppress system uncertainty and external disturbances;
[0066] Fast response: able to achieve finite time convergence;
[0067] Design flexibility: Different control objectives can be achieved through sliding surface design.
[0068] The robustness enhancement method of the image fusion model (infrared and visible light image fusion model) based on sliding mode variable structure control of the present invention designs respective sliding mode surface functions and control laws for high and low frequency sub-images of infrared and visible light under the framework of co-occurrence analysis shearlet transform, adaptively fuses high and low frequency sub-images, enhances the robustness of the fusion model, and can reduce the impact of different imaging environments (such as rain, fog, etc.) and sudden changes in lighting conditions on the fusion effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0069] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0070] Figure 1 Schematic diagram of the flow of the method for enhancing the robustness of the image fusion model based on sliding mode variable structure control of the present invention.
[0071] Figure 2 Schematic diagram of the overall framework principle of the image fusion model robustness enhancement method based on sliding mode variable structure control of the present invention.
[0072] Figure 3 Schematic diagram of the pseudo-polarization coordinate grid after merging two directions with a window size of 15×15.
[0073] Figure 4Schematic diagram of the saturation function principle.
[0074] Figure 5 This is a schematic diagram showing the comparison results of the comparative test between the method of the present invention and three existing image fusion methods. DETAILED DESCRIPTION
[0075] The inventive concept of the present invention is:
[0076] The image fusion model robustness enhancement method based on sliding mode variable structure control of the present invention is a method for adaptive fusion of infrared images and visible light images to solve the shortcomings of existing fusion models such as poor robustness, low universality and easy distortion of fusion results.
[0077] Currently, most fusion models rely on extensive sample training or the manual design of fusion rules based on the engineering design experience of researchers to achieve effective fusion of infrared and visible light images. Because the adaptability of a single fusion rule is limited, robust control algorithms are being introduced to improve the adaptability of fusion models. This invention leverages the robustness of sliding mode variable structure control and its insensitivity to system parameter changes and external disturbances to effectively enhance the adaptability of infrared and visible light image fusion models.
[0078] This invention uses a sliding mode variable structure control algorithm to regulate the fusion rules, replacing the single fusion rule used in traditional algorithms. When fusing infrared and visible light images, the fusion weights are adjusted in real time based on different imaging environments (such as rain and fog) and lighting conditions. This design significantly improves the model's adaptability. By directly utilizing a discontinuous control law, the infrared and visible light fusion system's state moves along a preset sliding mode surface, resulting in strong robustness.
[0079] The present invention will be described in detail below with reference to the accompanying drawings.
[0080] like Figure 1 As shown ( Figure 1 Only the method outline of the present invention is shown in the figure). The present invention provides a method for enhancing the robustness of an image fusion model based on sliding mode variable structure control, wherein the image fusion model is an infrared and visible light image fusion model. The present invention implements dynamic weight adjustment through sliding mode variable structure control. The principle diagram thereof is shown in Figure 2, and includes the following steps:
[0081] The first step is to obtain the original infrared image and the original visible light image from the same scene and complete the image registration.
[0082] Step 2: Decomposition:
[0083] Based on the co-occurrence analysis shearlet transform, the original infrared image and the original visible light image obtained in the first step are decomposed into a high-frequency layer and a low-frequency layer respectively; that is, the co-occurrence analysis shearlet transform is used to process the original infrared image and the original visible light image respectively; specifically, the following steps are included:
[0084] The co-occurrence filter is used to perform multi-scale decomposition of the image, and then the discrete shearlet transform is used to perform multi-directional decomposition of the high-frequency layer sub-image. This can more effectively separate the large-scale contours and texture area edges into different levels and save them separately according to the difference in texture information, thereby obtaining infrared low-frequency sub-images and a series of infrared high-frequency sub-images and visible light low-frequency sub-images and a series of visible light high-frequency sub-images.
[0085] Step 21, co-occurrence filter - multi-scale decomposition:
[0086] ;
[0087] in, Represents low-frequency subgraphs at different scales; is the scale number (the maximum value is ); Pick When represents the infrared image, take When represents the visible light image; Represents high-frequency subgraphs at different scales; is the co-occurrence filtering operator; represent Low-frequency subgraph at scale.
[0088] The co-occurrence filter is a type of local linear filter defined as follows:
[0089] ;
[0090] ;
[0091] Where, and are the output and input pixel values; and Both represent the index mark of the pixel; is a normalization factor, which can be regarded as a weight distribution term, and its value represents the input pixel To output pixel contribution; is the normalized co-occurrence matrix, Representative pixels the number of is a Gaussian filter, and its specific content is shown in the following formula:
[0092] ;
[0093] Among them, the filter scale parameter The change of can control the filtering of edge texture in the image. This parameter cannot avoid the problem of blurred edges during the smoothing process, and filtering needs to be guided by the weight allocation of image co-occurrence information.
[0094] It can be obtained by calculating the co-occurrence matrix, as shown below;
[0095] ;
[0096] in, is the co-occurrence information, provided by the co-occurrence matrix; and Corresponding pixels in the image and The frequency of is the histogram of the pixel value, which is defined as follows:
[0097] ;
[0098] ;
[0099] Where, It's a pixel and Euclidean distance between; co-occurrence matrix parameters is set to ; operator A Boolean relationship is defined: if the Boolean value in the square brackets is false, the expression evaluates to 0, otherwise it evaluates to 1; Representative pixels Pixel value of Represents the pixel in the corresponding image frequency; Refers to pixels or .
[0100] Step 22, Discrete Shearlet Transform - Multi-directional Decomposition:
[0101] The specific steps of directional localization achieved by adaptive shearlet filter are as follows:
[0102] Step 221, generate Meyer window function using Meyer wavelet function , defined as follows: ;
[0103] ;
[0104] Where, represents the wavelet function, Representative variables The wavelet function of is a variable greater than 0, is an integer variable ranging from 1 to 15. is a variable whose value is determined by Decide;
[0105] Step 222, place the Meyer window function into the pseudo-polarization coordinate grid under the filter window (window size is 15×15) (e.g. Figure 3 Resample the result to the Cartesian coordinate system;
[0106] Step 223, transform the shearlet filter bank in the frequency domain to the time domain through inverse Fourier transform to obtain the final filter bank, which is recorded as ,parameter The value of is only 0 or 1; is the number of decomposition directions, and its value is ; Indicates the scale number;
[0107] Step 224, the filter bank in the detail layer component and time domain Perform convolution operation to obtain the final multi-directional detail sub-band image, that is, the high-frequency sub-image .
[0108] ;
[0109] The third step is to fuse high-frequency sub-graphs:
[0110] Design high-frequency sliding surface and control law to guide the fusion of visible light high-frequency sub-graph and infrared high-frequency sub-graph obtained in the second step. Use sliding surface based on local energy and gradient to design sliding mode fusion strategy for high-frequency coefficients, which includes the following steps:
[0111] Step 31: Define the high-frequency coefficient fusion sliding surface function :
[0112] ;
[0113] in, and represent infrared local energy and visible light local energy respectively; and Represent the infrared gradient component and the visible light gradient component respectively; and Both are high frequency coefficients fused with sliding surface parameters.
[0114] Step 32: Design high-frequency coefficient fusion control law :
[0115] ;
[0116] in, is a symbolic function.
[0117] Step 33, high frequency coefficient selection is done by Decide:
[0118] ;
[0119] in, is the high frequency fusion coefficient, represents the infrared high-frequency sub-image, Represents the visible light high-frequency sub-graph.
[0120] Step 4: low-frequency sub-image fusion:
[0121] Design a low-frequency sliding mode surface and control law to guide the fusion of the visible light low-frequency sub-image and the infrared low-frequency sub-image obtained in the second step. Design a sliding mode fusion strategy for the low-frequency coefficients, which specifically includes the following steps:
[0122] Step 41: Design the low-frequency coefficient fusion sliding surface function (i.e., low-frequency adaptive weighted sliding surface) :
[0123] ;
[0124] in, and represent the average value of infrared low-frequency component and the average value of visible light low-frequency component respectively; and and represent the variance of infrared low-frequency component and visible light low-frequency component respectively; is the low-frequency coefficient fusion sliding surface parameter.
[0125] Step 42: Design the low-frequency coefficient fusion control law :
[0126] ;
[0127] in, To control the gain, is a saturation function (e.g. Figure 4 shown), satisfying:
[0128] ;
[0129] Where δ is the boundary layer thickness, is a symbolic function.
[0130] In order to suppress chattering during the control process, the following Figure 4 middle( x,y ) coordinate system and ( s,u The saturation function shown in the ) coordinate system essentially replaces the relay characteristics with saturation characteristics to mitigate the discontinuity of control switching. It is used to demonstrate a linear relationship within the boundary layer and a fixed value outside the boundary layer. The saturation function has a three-segment structure and two switching planes: and , and are the upper and lower bounds of the boundary layer, respectively; is a constant, which is the boundary layer thickness; in addition is the boundary layer width.
[0131] Step 43, using the low-frequency coefficient fusion weight (the low-frequency coefficient fusion weight refers to the low-frequency coefficient fusion control law ) guides the fusion of low-frequency coefficients:
[0132] ;
[0133] in, is the low-frequency fusion coefficient, represents the infrared low-frequency sub-image, Represents the visible light low-frequency sub-image.
[0134] Step 5, inverse transform:
[0135] The co-occurrence analysis and shearlet inverse transform are performed on the fused high-frequency image and low-frequency image obtained in the third and fourth steps respectively to reconstruct the fused image.
[0136] The final fusion result is obtained by co-occurrence analysis and shearlet inverse transformation, such as Figure 5 As shown in the figure, the present invention is compared with three existing image fusion methods, among which (a) is an infrared image (Infrared), (b) is a visible light image (Visible), (c) is an anisotropic diffusion filter (ADF) method, (d) is a unified unsupervised fusion (U2Fusion) method, (e) is a generative adversarial network fusion (FusionGAN) method, and (f) is the present invention (Ours) method. Through the above experiments, it can be seen that the image fused by the method of the present invention has higher infrared target brightness, richer scene details, and higher contrast. Therefore, based on the strong robust performance of the sliding mode variable structure control algorithm, the present invention guides the adaptive asymptotic fusion of high-frequency sub-images and low-frequency sub-images, which can effectively improve the universality of the fusion model to imaging environments and conditions.
[0137] Obviously, the above embodiments are merely examples for clarity of explanation and are not intended to limit the implementation methods. Those skilled in the art will readily appreciate that other variations or modifications based on the above descriptions are possible. It is not necessary and impossible to enumerate all implementation methods here. Obvious variations or modifications arising therefrom remain within the scope of protection of the present invention.
Claims
1. A method for enhancing the robustness of an image fusion model based on sliding mode variable structure control, characterized in that: The image fusion model is an infrared and visible light image fusion model, which includes the following steps: The first step is to obtain the original infrared image and the original visible light image from the same scene and complete the image registration; Step 2: Decomposition: Based on the co-occurrence analysis, the shearlet transform is used to decompose the original infrared image and the original visible light image obtained in the first step into high-frequency layer and low-frequency layer respectively; The third step is to fuse high-frequency sub-graphs: Design a high-frequency sliding surface and control law to guide the fusion of the visible light high-frequency sub-image and the infrared high-frequency sub-image obtained in the second step; Step 4: low-frequency sub-image fusion: Design a low-frequency sliding surface and control law to guide the fusion of the visible light low-frequency sub-image and the infrared low-frequency sub-image obtained in the second step; Step 5, inverse transform: The co-occurrence analysis and shearlet inverse transform are performed on the fused high-frequency image and low-frequency image obtained in the third and fourth steps respectively to reconstruct the fused image.
2. The method for enhancing the robustness of an image fusion model based on sliding mode variable structure control according to claim 1, characterized in that: The second step is specific The following steps are involved: Step 21, multi-scale decomposition; ; in, Represents low-frequency subgraphs at different scales; is the scale number, the maximum value is ; Pick When represents the infrared image, take When represents the visible light image; Represents high-frequency subgraphs at different scales; is the co-occurrence filtering operator; represent Low-frequency subgraphs at scale; The co-occurrence filter is defined as follows: ; ; Where, and are the output and input pixel values respectively; and Both represent the index flag of the pixel; Is the normalization factor, which is the weight distribution term, and its value represents the input pixel To output pixel contribution; is the normalized co-occurrence matrix, Representative pixels the number of is a Gaussian filter, as shown below: ; Among them, the filter scale parameter The change of controls the filtering of edge texture in the image; It is obtained by calculating the co-occurrence matrix, as shown below; ; in, is the co-occurrence information, provided by the co-occurrence matrix; and Corresponding pixels in the image and The frequency of is the histogram of pixel values, which is defined as follows: ; ; Where, It's a pixel and Euclidean distance between; co-occurrence matrix parameters is set to ; operator A Boolean relationship is defined: if the Boolean value in the square brackets is false, the expression evaluates to 0, otherwise it evaluates to 1; Representative pixels Pixel value of Represents the pixel in the corresponding image frequency; Refers to pixels or ; Step 22, multi-directional decomposition; Step 221, generate Meyer window function using Meyer wavelet function , defined as follows: ; ; Where, represents the wavelet function, Representative variables The wavelet function of is a variable greater than 0, is an integer variable ranging from 1 to 15. is a variable whose value is determined by Decide; Step 222, placing the Meyer window function into the pseudo-polarization coordinate grid under the filter window for resampling, and mapping the sampling result to the Cartesian coordinate system; Step 223, transform the shearlet filter bank in the frequency domain to the time domain through inverse Fourier transform to obtain the final filter bank, which is recorded as ,parameter The value of is only 0 or 1; is the number of decomposition directions, and its value is ; Indicates the scale number; Step 224, the filter group of detail layer components and time domain Perform convolution operation to obtain the final multi-directional detail sub-band image, that is, the high-frequency sub-image ; 。 3. The method for enhancing the robustness of an image fusion model based on sliding mode variable structure control according to claim 1, characterized in that: The third step is as follows: Step 31: Define the high-frequency coefficient fusion sliding surface function : ; in, and represent infrared local energy and visible light local energy respectively; and Represent the infrared gradient component and the visible light gradient component respectively; and All are high frequency coefficients fused with sliding surface parameters; Step 32: Design high-frequency coefficient fusion control law : ; in, is a symbolic function; Step 33, high frequency coefficient selection is done by Decide: ; in, is the high frequency fusion coefficient, represents the infrared high-frequency sub-image, Represents the visible light high-frequency sub-graph.
4. The method for enhancing the robustness of an image fusion model based on sliding mode variable structure control according to claim 1, characterized in that: The fourth step is as follows: Step 41: Define the low-frequency coefficient fusion sliding surface function : ; in, and Represent the average value of infrared low-frequency component and visible light low-frequency component respectively; and represent the variance of infrared low-frequency component and visible light low-frequency component respectively; is the low-frequency coefficient fusion sliding surface parameter; Step 42: Design the low-frequency coefficient fusion control law : ; in, To control the gain, is a saturation function that satisfies: ; Where δ is the boundary layer thickness, is a symbolic function; Step 43: Use the low-frequency coefficient fusion weight to guide the low-frequency coefficient fusion: ; in, is the low-frequency fusion coefficient, represents the infrared low-frequency sub-image, Represents the visible light low-frequency sub-image.
Citation Information
Patent Citations
Infrared image and visible image fusion method
CN108389158A
Infrared and visible light image fusion method based on NSCT and DWT
CN110110786A
Infrared and low-illumination image fusion method based on non-subsampled shear transformation
CN117994146A
Robust iterative learning control method for series inverted pendulums in finite frequency range
WO2022012155A1