Multi-focus image fusion method and system based on structural similarity and region segmentation

By employing a multi-focus image fusion method based on structural similarity and region segmentation, and utilizing visual saliency detection, multi-scale detail enhancement, and structural similarity segmentation, the problems of pixel value distortion and boundary blurring in existing technologies are solved, generating high-quality multi-focus images.

CN115829895BActive Publication Date: 2026-06-02ZHUJIANG SWITCH FACTORY GUANGDONG PROV

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
ZHUJIANG SWITCH FACTORY GUANGDONG PROV
Filing Date
2022-12-08
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing multi-focus image fusion methods are prone to pixel value distortion, high computational complexity, blurred boundaries or artificial artifacts in the fusion results, and it is difficult to accurately distinguish and integrate pixels with different focusing attributes.

Method used

A multi-focus image fusion method based on structural similarity and region segmentation is adopted. Through visual saliency detection, Gaussian filtering and multi-scale detail enhancement based on logarithmic energy difference, structural similarity exponential segmentation and pixel selection rules, the method can accurately distinguish between fully focused, fully defocused and uncertain regions and generate fully focused images.

Benefits of technology

It achieves the ability to accurately distinguish and integrate pixels with different focusing attributes while maintaining image sharpness and boundary clarity, thus generating high-quality all-focus images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115829895B_ABST
    Figure CN115829895B_ABST
Patent Text Reader

Abstract

The application discloses a multi-focus image fusion method and system based on structural similarity and region segmentation, and the method comprises the following steps: performing weighted processing on multi-focus source images based on a fusion strategy of visual saliency detection, performing enhancement processing on pre-fusion images through a multi-scale detail enhancement strategy of logarithmic energy difference, and obtaining detail-enhanced pre-fusion images; performing segmentation processing on the detail-enhanced pre-fusion images based on a structural similarity index, and obtaining a region segmentation decision graph; and performing fusion processing on the region segmentation decision graph according to a pixel selection rule, and obtaining a final fusion image. The system comprises a pre-fusion module, an enhancement module, a segmentation module and a fusion module. Through the application, pixels with different focusing properties can be accurately distinguished and integrated to generate a full-focus image. The application can be widely applied to the field of image fusion technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image fusion technology, and in particular to a multi-focus image fusion method and system based on structural similarity and region segmentation. Background Technology

[0002] Due to the limitations of optical lenses in ordinary digital cameras, it is difficult to obtain a clear image of the entire scene in focus using a single camera. Clear textures and details are generally distributed within the focused area, revealing more features about the scene or objects. Therefore, to obtain complete information and ensure all objects in the scene are in focus and clear, specific multi-focus image fusion methods are typically used to extract focused pixel information from different source images of the same scene. Finally, this pixel information is integrated to achieve scene fusion. For example, in deep learning algorithms, a common approach is to use convolutional neural networks (CNNs) to fuse different images. CNNs can extract features from input data through the superposition of convolutional and pooling layers, and finally connect fully connected layers for classification. Transform domain-based multi-focus image fusion algorithms first decompose the source image into different sub-bands, then design different fusion methods based on the features of different sub-bands to fuse them, and finally reconstruct the fused result. However, due to the multi-layered and multi-directional decomposition, transform domain-based methods can effectively extract images at different scales. The pixel information in the source image has a strong ability to preserve scene information. However, since the pixel information of the fusion result cannot be completely derived from the focus area of ​​a single source image, a certain degree of pixel value distortion may occur. In addition, multi-layer and multi-directional decomposition also brings high computational complexity. The spatial domain-based algorithm can directly process the pixels on the source image and determine the focus decision map by extracting the salient information of the focus area on the source image to achieve the fusion of multi-focus images. However, due to the complexity of the image, sometimes the focus measurement value is the largest in the out-of-focus area. Therefore, the initial decision map still has many shortcomings. This situation often occurs around the boundary between the focus and out-of-focus areas. This also leads to the fusion results obtained by many algorithms having blur or artificial artifacts at the boundary, and even retaining a large amount of pixel information from the out-of-focus area. Furthermore, many methods require post-processing algorithms to make the decision map achieve the desired effect, such as consistency verification and small area removal. Therefore, most of the small area of ​​focus will inevitably be deleted, resulting in an incomplete focus area and a suboptimal fusion result. Summary of the Invention

[0003] To address the aforementioned technical problems, the present invention aims to provide a multi-focus image fusion method and system based on structural similarity and region segmentation, which can accurately distinguish pixels with different focusing attributes and integrate them together to generate a full-focus image.

[0004] The first technical solution adopted in this invention is a multi-focus image fusion method based on structural similarity and region segmentation, comprising the following steps:

[0005] A fusion strategy based on visual saliency detection is used to weight images from multiple aggregation sources to obtain a pre-fused image;

[0006] The pre-fused image is enhanced by a multi-scale detail enhancement strategy using Gaussian filtering and logarithmic energy difference, resulting in a pre-fused image with enhanced details.

[0007] The pre-fused image after detail enhancement is segmented based on the structural similarity index to obtain a region segmentation decision map;

[0008] The region segmentation decision map is fused according to the pixel selection rule to obtain the final fused image.

[0009] Furthermore, the step of weighting the multi-source images to obtain the pre-fused image using the fusion strategy based on visual saliency detection specifically includes:

[0010] Acquire the first multi-aggregation source image and the second multi-aggregation source image;

[0011] Visual saliency detection is performed on the first and second multi-cluster source images to obtain the corresponding saliency values;

[0012] The significance values ​​are normalized to obtain normalized significance values;

[0013] Based on the normalized significance value, a weighted average is performed on the first and second multi-source images to obtain the pre-fused image.

[0014] Furthermore, the step of enhancing the pre-fused image using a multi-scale detail enhancement strategy based on Gaussian filtering and logarithmic energy difference to obtain a detail-enhanced pre-fused image specifically includes:

[0015] The pre-fused image is convolved using a Gaussian function to obtain detailed information about the pre-fused image.

[0016] Based on the detail information of the pre-fused image, the high-frequency components to be fused in the pre-fused image are obtained;

[0017] The energy of the high-frequency component to be fused is obtained by using a log-energy high-frequency component fusion algorithm and then the logarithm is taken to obtain the log-energy value of the high-frequency component to be fused.

[0018] The difference in log-energy between different high-frequency components to be fused is calculated and a preset threshold is introduced for judgment.

[0019] Based on the judgment result, the corresponding high-frequency component fusion rule is selected to perform fusion processing on the high-frequency component to be fused, and the fused high-frequency component is obtained.

[0020] High-frequency components are added to the pre-fused image to obtain a pre-fused image with enhanced details.

[0021] Furthermore, the expression for the high-frequency component fusion rule is as follows:

[0022]

[0023] In the above formula, FH(x, y) represents the high-frequency components to be fused, HM(x, y) represents the decision graph of the high-frequency components, λ represents the preset threshold, and DV represents the difference in log-energy between different high-frequency components to be fused. This represents different high-frequency components to be fused. This represents the weighting coefficient.

[0024] Furthermore, the step of segmenting the pre-fused image after detail enhancement based on the structural similarity index to obtain a region segmentation decision map specifically includes:

[0025] Obtain multi-source images, and based on the focal regions of the images, calculate the structural similarity between the multi-source images and the pre-fused images after detail enhancement to obtain a similarity score map;

[0026] Store the similarity score graph into a matrix to generate a rating graph;

[0027] Based on the multi-channel signal of the image, the score map is obtained by iterative filtering of the score map through a recursive filter.

[0028] The initial two-region segmentation decision map is obtained by performing decision processing on the fraction map using preset decision rules;

[0029] A consistency detection and judgment is performed on the center pixel and surrounding pixels of the initial two-region segmentation decision map to obtain the two-region segmentation decision map, which includes a fully focused region and a fully defocused region.

[0030] The fractional map is analyzed based on the pixel difference method of RF to obtain a three-region segmentation decision map, which includes a fully focused region, a fully defocused region, and an uncertain region.

[0031] By integrating the two-region segmentation decision map and the three-region segmentation decision map, a region segmentation decision map is obtained.

[0032] Furthermore, the expression for the preset decision rule is as follows:

[0033]

[0034] In the above formula, IMP(x, y) represents the initial two-region splitting decision graph, B1(x, y), B2(x, y), ..., B T (x, y) represents the corresponding fractional graph.

[0035] Furthermore, the expression for the pixel selection rule in the RF-based pixel difference method is as follows:

[0036]

[0037] In the above formula, RDM(x, y) represents the three-region segmentation decision map, and β represents the parameter used to control the accuracy.

[0038] Furthermore, the step of fusing the region segmentation decision map according to the pixel selection rule to obtain the final fused image specifically includes:

[0039] The two-region segmentation decision map and the three-region segmentation decision map are comprehensively processed by the preset pixel selection rules to obtain the final decision map;

[0040] To generate a fully focused and visually appealing image, the final decision map and the pre-fused image are fused together to obtain the final fused image.

[0041] Furthermore, the expression for the preset pixel selection rule is as follows:

[0042]

[0043] In the above formula, FDM(x,y) represents the final decision graph, OMP(x,y) represents the two-region split decision graph, and RDM(x,y) represents the three-region split decision graph.

[0044] The second technical solution adopted in this invention is: a multi-focus image fusion system based on structural similarity and region segmentation, comprising:

[0045] The pre-fusion module performs weighted processing on multi-aggregate source images based on a fusion strategy of visual saliency detection to obtain a pre-fused image;

[0046] The enhancement module is used to enhance the pre-fused image through a multi-scale detail enhancement strategy using Gaussian filtering and logarithmic energy difference, resulting in a pre-fused image with enhanced details.

[0047] The segmentation module performs segmentation processing on the pre-fused image after detail enhancement based on the structural similarity index to obtain a region segmentation decision map;

[0048] The fusion module is used to fuse the region segmentation decision map according to the pixel selection rules to obtain the final fused image.

[0049] The beneficial effects of the method and system of this invention are as follows: This invention performs weighted fusion processing on multi-aggregate source images through a fusion strategy of visual saliency detection to obtain a pre-fused image. The source image is regarded as a combination of fully focused regions, fully defocused regions, and uncertain regions. Furthermore, the pre-fused image is enhanced by a high-frequency component fusion rule based on log-energy. By calculating the log-energy of different high-frequency components to be fused, the sharpness of the source image can be effectively evaluated, thereby reflecting the salient information on the source image. In addition, the source image is segmented by structural similarity index through non-boundary region decision maps and boundary region decision maps. The generated two-region segmentation decision maps and three-region segmentation decision maps are used to accurately distinguish these three regions, which can effectively realize the fusion between images with different focus, accurately distinguish pixels with different focus attributes, and integrate them together to generate a fully focused image. Attached Figure Description

[0050] Figure 1 This is a flowchart of the steps of the multi-focus image fusion method based on structural similarity and region segmentation of the present invention;

[0051] Figure 2 This is a structural block diagram of the multi-focus image fusion system based on structural similarity and region segmentation of the present invention;

[0052] Figure 3 These are schematic diagrams of two multi-aggregation source images used in the simulation experiment of this invention;

[0053] Figure 4 It is a difference image between the fusion result obtained by the existing method and the method of the present invention and the same source image;

[0054] Figure 5 This is a structural block diagram of the multi-aggregation image fusion method based on structural similarity and region segmentation of the present invention. Detailed Implementation

[0055] The present invention will now be described in further detail with reference to the accompanying drawings and specific embodiments. The step numbers in the following embodiments are only for ease of explanation and do not limit the order of the steps. The execution order of each step in the embodiments can be adapted according to the understanding of those skilled in the art.

[0056] Reference Figure 1 and Figure 5 This invention provides a multi-focus image fusion method based on structural similarity and region segmentation, which includes the following steps:

[0057] S1. A fusion strategy based on visual saliency detection is used to weight the multi-aggregate source images to obtain a pre-fused image;

[0058] Specifically, the corresponding multi-aggregate source images I1 and I2 are obtained. This invention utilizes a fusion strategy based on visual saliency detection to obtain the pre-fused image. Visual saliency detection is a method that can effectively detect salient structures in an image. This fusion strategy can effectively avoid the loss of contrast and provide a good visual effect for the fused image. The algorithm defines pixel saliency based on the comparison of a pixel with all other pixels, and I... p Let S(p) represent the intensity value of pixel p in image I. The salience value of pixel p, S(p), can be defined as follows:

[0059] S(p)=|I p -I 1 |+|I p -I 2 |+…+|I p -I N |

[0060] In the above formula, N represents the total number of pixels in I. If two pixels have the same intensity value, their salience is equal. Therefore, the above formula can be rewritten as follows:

[0061]

[0062] In the above formula, j represents pixel intensity, and M j Let S(p) represent the number of pixels with intensity equal to j, and L represent the number of gray levels (256 in this invention). Then, S(p) is normalized to [0, 1].

[0063] Let S t Representative source image I t The visual saliency detection value is denoted by t = (1, 2, ..., T), where T represents the number of source images, which is taken as 2 in this invention. Using S... t The pre-fused image power factor (PF) can be obtained using the following weighted average rule, as shown in the formula below:

[0064] PF=W F ×I1+(1-W F )×I2

[0065] Where the weight value W F It can be defined as follows,

[0066]

[0067] To further enhance the contrast, brightness, and structural information of the pre-fusion results and facilitate the extraction of salient features, this invention proposes a multi-scale detail enhancement strategy based on Gaussian filtering and logarithmic energy difference.

[0068] S2. The pre-fused image is enhanced by a multi-scale detail enhancement strategy using Gaussian filtering and logarithmic energy difference to obtain a pre-fused image with enhanced details.

[0069] Specifically, the first step is to acquire detailed information at different scales.

[0070] D m,t (x, y) = I t (x, y) - I t (x, y)*G m,σ

[0071] In the above formula, D m,t (x, y) represents the source image I. t High-frequency components of (x, y) at scale m, * represents the convolution operator, G m,σ This represents a Gaussian function with a window size of m×m and a standard deviation of σ, where m∈M is the scale, and M is set to 5 in this invention;

[0072] Subsequently, the high-frequency components to be fused are obtained according to the following rules.

[0073]

[0074] However, due to the diversity and complexity of images, a single fusion strategy is often insufficient to effectively process various types of images. Therefore, to enable the algorithm of this invention to combine the advantages of different fusion rules to achieve optimal fusion performance, this invention introduces a high-frequency component fusion rule based on log-energy. By analyzing the log-energy differences between different high-frequency components to be fused, a suitable fusion strategy is selected. Calculating the log-energy of different high-frequency components to be fused can effectively evaluate the sharpness of the source image, thereby reflecting the significant information in the source image. The algorithm can be roughly divided into two steps: the first is to calculate the energy (i.e., the square of the pixel value) of the high-frequency components to be fused, and the second is to calculate the logarithm of the energy. The calculation method is defined as follows:

[0075]

[0076] The difference in log-energy (DV) between different high-frequency components to be fused is calculated using the following formula:

[0077] DV = |E1 - E2|

[0078] Furthermore, a suitable fusion rule is selected by judging the relationship between DV and the threshold λ. The expression of the fusion rule is as follows:

[0079]

[0080] In the above formula, HM(x, y) represents the decision graph of the high-frequency components, and FH(x, y) represents the fused high-frequency components. and They are respectively and Weighting coefficients;

[0081] in Furthermore, the threshold λ is set to [value missing] in this invention. Furthermore, HM(x, y) can be obtained by the following formula, the specific expression of which is shown below:

[0082]

[0083] Finally, the present invention adds the obtained fused high-frequency components to the pre-fused image to enhance the detail information, resulting in a detail-enhanced pre-fused image EPF, the expression of which is as follows:

[0084] EPF(x,y)=IF(x,y)+FH(x,y)

[0085] S3. Based on the structural similarity index, the pre-fused image after detail enhancement is segmented to obtain a region segmentation decision map;

[0086] Specifically, in the multi-focus fusion task, the source image can be roughly divided into three regions: the fully focused region, the fully defocused region, and the uncertain region. The uncertain region contains pixels in both the focused and defocused areas of the image, and usually exists around the focus boundary of the image. Therefore, in order to achieve a good fusion effect, it is necessary to classify each pixel in detail. In the proposed algorithm, this invention uses two decision maps to segment the source image into three regions, namely the non-boundary region decision map and the boundary region decision map.

[0087] Structural similarity (SSIM), as a structural similarity index for measuring image quality, is widely used in image processing to evaluate the performance of fusion. It separates the similarity measurement task into three comparisons: brightness, contrast, and structure. Combining these three components yields an overall similarity measure. In the algorithm of this invention, the SSIM values ​​between the source image and the pre-fused image after detail enhancement in different focused regions are calculated to evaluate structural similarity. Finally, the results are aggregated into a matrix to obtain a score map that reflects the brightness, contrast, and structural information of the source image.

[0088] S31. Obtain the two-region segmentation decision map;

[0089] S311. Obtain the similarity score image;

[0090] First, obtain the fractional graph (SCM), whose expression is shown below:

[0091] SCM t =ssim(I t EPF)

[0092] In the above formula, ssim(·) is the SSIM operator, and SCM t For source image I t Similarity score map with EPF. SSIM between each image is calculated as follows:

[0093]

[0094] In the above formula, (i, j) represents the image pair (multi-aggregate source image and pre-fused image after detail enhancement processing), μ i σ represents the average pixel value of source image i. i σ represents the standard deviation of the source image i. ij Let c1 and c2 represent the covariance of the image pair (i, j), and c1 and c2 represent constants to prevent the denominator from being zero.

[0095] The SSIM range is between [0, 1]. A higher SSIM indicates that the structures between the two images are more similar. Finally, the calculation results are stored in a matrix to generate a rating map.

[0096] Then, the fractional image obtained by processing with a recursive filter is expressed as follows:

[0097] B t =RF(SCM) t , σ s , σ r (N)

[0098] In the above formula, σ s σ r In this invention, N is set to 3, 0.25, and 3, respectively.

[0099] S312. Obtain the corresponding score chart;

[0100] Recursive filters are edge-preserving filters that can reduce the impact of noise without eliminating significant information such as edges and textures in an image. In recursive filters, the distance between points in a high-dimensional signal can be preserved in a low-dimensional space, so high-dimensional signals can be processed by a low-dimensional filter kernel.

[0101] For a multi-channel signal I, the one-dimensional domain transform is achieved by the following expression, which is shown below:

[0102]

[0103] In the above formula, u∈Ω represents a point of the signal in the original domain Ω, ct(u) preserves the geodesic distance from the origin to u in the transform domain, and σ s σ is the standard deviation of the filter over the spatial range. r denoted as the standard deviation of the filter within the signal range, and c represents the number of signal channels.

[0104] For a discrete one-dimensional signal I[n], the recursive filter is represented as follows:

[0105] J[n]=(1-a d )I[n]+a d J[n-1]

[0106] In the above formula, J[n] represents the filtered signal at position n, and a∈[0,1] represents the feedback coefficient, which can be expressed by the following formula:

[0107]

[0108] In the above formula, σ represents the standard deviation of the filter used in the nth iteration, and N represents the total number of iterations. This invention allows us to see σ in each iteration. i d represents the adjacent sample x in the transform domain n and x n-1 The distance between them can be expressed by the following formula, which is:

[0109]

[0110] Two-dimensional filtering of an image can be decomposed into one-dimensional filtering performed in each dimension. Therefore, since each row and each column can be regarded as a discrete signal, recursive filtering must be performed, and multiple iterations are used to achieve this. By combining different parameters, this invention can write the recursive filter operation as RF(IN, σ s , σ r , N), where IN is the input image.

[0111] S313, Two-Region Segmentation Decision Diagram;

[0112] Finally, the initial two-region segmentation decision graph (IMP) is generated according to the following rules.

[0113]

[0114] The initial decision map obtained at this time cannot accurately determine the focus attribute of pixels. In the out-of-focus region, there are still some small areas composed of erroneous pixels. To address this, the present invention introduces a consistency verification technique to optimize the obtained initial decision map. By analyzing the consistency between the center pixel and the surrounding pixels within a fixed window, it is determined whether a pixel is in the focus or out-of-focus region, as shown below:

[0115]

[0116] In the above formula, OMP(x, y) represents the final two-region segmentation decision map, and δ represents the square neighborhood window centered at (x, y), which is set to 23 in this invention;

[0117] The two-region segmentation decision map OMP obtained above only roughly divides the source image into two regions (fully in focus and fully out of focus regions).

[0118] S32. Obtain the three-region segmentation decision map;

[0119] To further segment the source image into three regions, we propose a pixel difference analysis scheme based on radiometrics (RF). Specifically, generating the boundary region segmentation decision map requires the following four steps;

[0120] First, the difference map (DM) between the fractional images obtained from different source images is calculated, which can be mathematically represented as:

[0121] DM(x,y)=|SCM1(x,y)-SC2(x,y)|

[0122] Then, the difference map is filtered using RF to obtain the differenced blur map (DBM), the specific expression of which is shown below:

[0123] DBM(x,y)=RF(DM(x,y),σ s , σ r (N)

[0124] In the above formula, DBM represents the output result after filtering, and σ s σ r In this invention, N is set to 3, 0.25, and 3, respectively;

[0125] Next, we will use B t The expression for obtaining the blurred difference map (BDM) is shown below:

[0126] BDM(x,y)=|B1(x,y)-B2(x,y)|

[0127] Finally, BDM(x, y) is compared with DBM(x, y) to generate a three-region segmentation decision map. Furthermore, BDM(x, y) can only approach DBM(x, y) at a corresponding location when the pixel is in full focus or full defocus; otherwise, it will be smaller than DBM(x, y). Based on this characteristic, this invention generates a three-region segmentation decision map (RDM) using the following pixel selection rule, the specific expression of which is shown below:

[0128]

[0129] In the above formula, β represents the parameter used to control the accuracy, which is set to 0.5 in this invention;

[0130] For source image I1, the pixel is considered a focused pixel when RDM(x,y)=1, a defocused pixel when RDM(x,y)=2, and an uncertain pixel when RDM(x,y)=0.5. For source image I2, except for the uncertain pixels which are the same as those in I1, the remaining cases are complementary. In addition, uncertain pixels are generally located on or around the boundary between the focused and defocused areas.

[0131] S4. The region segmentation decision map is fused according to the pixel selection rules to obtain the final fused image.

[0132] Specifically, in order to obtain a final decision map with accurate boundaries and a complete focused region, this invention proposes the following pixel selection rule to combine the advantages of two-region segmentation focused decision map (OMP) and three-region segmentation decision map (RDM). The specific algorithm is as follows:

[0133]

[0134] For obtaining the final decision map FDM, only pixels whose pixel focus attributes are consistent in both the OMP and RDM decision maps are taken as the determined pixels. Otherwise, uncertain pixels are set to 0.5 in FDM.

[0135] Using the final decision map FDM and the pre-fused image PF, a fully focused and visually appealing and natural-looking fused image F can be generated, as shown in the following expression:

[0136]

[0137] In FDM, pixels identified as being in focus or out of focus are directly copied from the corresponding source image to the fusion result F. For uncertain pixels, this invention will obtain them from the pre-fused image PF, thereby achieving a smooth transition from one source image to another.

[0138] Reference Figure 2 A multi-focus image fusion system based on structural similarity and region segmentation includes:

[0139] The pre-fusion module performs weighted processing on multi-aggregate source images based on a fusion strategy of visual saliency detection to obtain a pre-fused image;

[0140] The enhancement module is used to enhance the pre-fused image through a multi-scale detail enhancement strategy using Gaussian filtering and logarithmic energy difference, resulting in a pre-fused image with enhanced details.

[0141] The segmentation module performs segmentation processing on the pre-fused image after detail enhancement based on the structural similarity index to obtain a region segmentation decision map;

[0142] The fusion module is used to fuse the region segmentation decision map according to the pixel selection rules to obtain the final fused image.

[0143] The simulation experiment of this invention is as follows:

[0144] To further demonstrate the advantages and effectiveness of this invention, a comparative experiment was conducted with five state-of-the-art existing image fusion algorithms. The performance of each algorithm was analyzed based on subjective visual evaluation. Figure 3 (a) and Figure 3 (b) shows two multi-focus source images. Figure 4 (a) to (f) show the difference maps of the following algorithms: a multi-focus image boundary finding algorithm based on multi-scale morphological focus measurement (BF), a multi-focus image fusion algorithm based on multi-scale focus measurement and generalized random walk (GRW), a multi-aggregate image fusion algorithm based on adaptive and gradient joint constraints in an unsupervised generative adversarial network (MFF-GAN), a unified unsupervised image fusion network (U2Fusion), a multi-functional squeeze decomposition network for real-time image fusion (SDNet), and the fusion algorithm of this scheme (SSRS). Figure 4 The difference graph clearly shows Figure 4 (c) through (e) show a large amount of residual information in the background region, proving that this method cannot directly obtain pixel information from the source image, which reduces the clarity of the fusion result. Figure 4 (a) shows that the fusion result of this method loses a large amount of focused pixel information and cannot effectively determine the focused attributes of different pixels. Figure 4 (b) Residual artifacts appear at the boundaries, which can lead to blurred boundaries in the fusion result. Figure 4 (f) were not affected by the above-mentioned problems, thus proving that the method proposed in this invention is superior to other existing comparative methods in maintaining the clarity of the fused image and the accuracy of detecting the focused pixels.

[0145] The content of the above method embodiments is applicable to this system embodiment. The specific functions implemented in this system embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments.

[0146] The above is a detailed description of the preferred embodiments of the present invention. However, the present invention is not limited to the embodiments described. Those skilled in the art can make various equivalent modifications or substitutions without departing from the spirit of the present invention. All such equivalent modifications or substitutions are included within the scope defined by the claims of this application.

Claims

1. A multi-focus image fusion method based on structural similarity and region segmentation, characterized in that, Includes the following steps: A fusion strategy based on visual saliency detection is used to weighted process multi-source images to obtain a pre-fused image. Specifically, this includes: acquiring a first multi-source image and a second multi-source image; performing visual saliency detection on the first and second multi-source images to obtain corresponding saliency values; normalizing the saliency values ​​to obtain normalized saliency values; and performing a weighted average on the first and second multi-source images based on the normalized saliency values ​​to obtain the pre-fused image. A multi-scale detail enhancement strategy based on Gaussian filtering and logarithmic energy difference is used to enhance the pre-fused image, resulting in a detail-enhanced pre-fused image. Specifically, this involves: convolving the pre-fused image with a Gaussian function to obtain its detail information; identifying the high-frequency components to be fused based on the detail information; obtaining the energy of the high-frequency components using a log-energy high-frequency component fusion algorithm and taking the logarithm to obtain their log-energy values; calculating the difference between the log-energy values ​​of different high-frequency components and introducing a preset threshold for judgment; selecting the corresponding high-frequency component fusion rule based on the judgment result to fuse the high-frequency components, obtaining the fused high-frequency components; and adding the fused high-frequency components to the pre-fused image to obtain the detail-enhanced pre-fused image. The pre-fused image after detail enhancement is segmented based on structural similarity index to obtain a region segmentation decision map. Specifically, this process includes: acquiring multi-source images; calculating structural similarity between the multi-source images and the pre-fused image after detail enhancement based on the image's focal region to obtain a similarity score map; storing the similarity score map in a matrix to generate a scoring map; iteratively filtering the scoring map using a recursive filter based on the image's multi-channel signal to obtain a score map; performing decision processing on the score map using preset decision rules to obtain an initial two-region segmentation decision map; performing consistency detection on the center pixel and surrounding pixels of the initial two-region segmentation decision map to obtain a second-region segmentation decision map, which includes a fully focused region and a fully defocused region; analyzing the score map using the pixel difference method based on RF to obtain a third-region segmentation decision map, which includes a fully focused region, a fully defocused region, and an uncertain region; and integrating the second-region segmentation decision map and the third-region segmentation decision map to obtain a final region segmentation decision map. The region segmentation decision maps are fused according to pixel selection rules to obtain the final fused image. Specifically, this includes: comprehensively processing the two-region segmentation decision maps and the three-region segmentation decision maps according to preset pixel selection rules to obtain the final decision map; and considering the effect of generating full focus and visual quality, fusing the final decision map and the pre-fused image to obtain the final fused image.

2. The multi-focus image fusion method based on structural similarity and region segmentation according to claim 1, characterized in that, The expression for the high-frequency component fusion rule is as follows: In the above formula, Indicates the high-frequency components of fusion. Decision graphs representing high-frequency components. Indicates the preset threshold. This represents the difference in log-energy between different high-frequency components to be fused. , This represents different high-frequency components to be fused. , This represents the weighting coefficient.

3. The multi-focus image fusion method based on structural similarity and region segmentation according to claim 1, characterized in that, The expression for the preset decision rule is as follows: In the above formula, This represents the initial two-region partitioning decision graph. This represents the corresponding fractional graph.

4. The multi-focus image fusion method based on structural similarity and region segmentation according to claim 1, characterized in that, The expression for the pixel selection rule in the RF-based pixel difference method is as follows: In the above formula, This represents a decision diagram for three-region partitioning. This refers to the parameters used to control precision.

5. The multi-focus image fusion method based on structural similarity and region segmentation according to claim 1, characterized in that, The expression for the preset pixel selection rule is as follows: In the above formula, This represents the final decision diagram. This represents a two-region partitioning decision graph. This represents a decision graph for three-region partitioning.

6. A multi-focus image fusion system based on structural similarity and region segmentation, characterized in that, Includes the following modules: The pre-fusion module performs weighted processing on multi-source images based on a visual saliency detection fusion strategy to obtain a pre-fused image. Specifically, it includes: acquiring a first multi-source image and a second multi-source image; performing visual saliency detection on the first and second multi-source images to obtain corresponding saliency values; normalizing the saliency values ​​to obtain normalized saliency values; and performing a weighted average processing on the first and second multi-source images based on the normalized saliency values ​​to obtain the pre-fused image. The enhancement module is used to enhance the pre-fused image using a multi-scale detail enhancement strategy based on Gaussian filtering and logarithmic energy difference, resulting in a detail-enhanced pre-fused image. Specifically, this includes: performing convolution operations on the pre-fused image using a Gaussian function to obtain detail information; obtaining the high-frequency components to be fused from the pre-fused image based on the detail information; obtaining the energy of the high-frequency components to be fused using a log-energy high-frequency component fusion algorithm and taking the logarithm to obtain the log-energy value of the high-frequency components to be fused; calculating the difference between the log-energy values ​​of different high-frequency components to be fused and introducing a preset threshold for judgment; selecting the corresponding high-frequency component fusion rule according to the judgment result to fuse the high-frequency components to be fused, obtaining the fused high-frequency components; and adding the fused high-frequency components to the pre-fused image to obtain the detail-enhanced pre-fused image. The segmentation module performs segmentation processing on the pre-fused image after detail enhancement based on the structural similarity index to obtain a region segmentation decision map. Specifically, this includes: acquiring a multi-source image; calculating the structural similarity between the multi-source image and the pre-fused image after detail enhancement based on the image's focal region to obtain a similarity score map; storing the similarity score map in a matrix to generate a scoring map; iteratively filtering the scoring map using a recursive filter based on the image's multi-channel signal to obtain a score map; performing decision processing on the score map using preset decision rules to obtain an initial two-region segmentation decision map; performing consistency detection and judgment on the center pixel and surrounding pixels of the initial two-region segmentation decision map to obtain a two-region segmentation decision map, which includes a fully focused region and a fully defocused region; analyzing the score map using the pixel difference method based on RF to obtain a three-region segmentation decision map, which includes a fully focused region, a fully defocused region, and an uncertain region; and integrating the two-region segmentation decision map and the three-region segmentation decision map to obtain a region segmentation decision map. The fusion module is used to fuse the region segmentation decision maps according to pixel selection rules to obtain the final fused image. Specifically, it includes: comprehensively processing the two-region segmentation decision maps and the three-region segmentation decision maps according to preset pixel selection rules to obtain the final decision map; and considering the effect of generating full focus and visual effects, fusing the final decision map and the pre-fused image to obtain the final fused image.