A Multimodal Image Fusion Method Based on Hesitant Fuzzy Variable Granularity Dictionary Learning

By adopting a hesitant fuzzy variable granularity dictionary learning method, the problems of insufficient image fuzziness and adaptive feature extraction capability in existing technologies are solved, achieving more efficient multimodal image fusion and improving the robustness and visual quality of image fusion.

CN121190320BActive Publication Date: 2026-06-02SICHUAN UNIVERSITY OF SCIENCE AND ENGINEERING

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SICHUAN UNIVERSITY OF SCIENCE AND ENGINEERING
Filing Date
2025-09-11
Publication Date
2026-06-02

Smart Images

  • Figure CN121190320B_ABST
    Figure CN121190320B_ABST
Patent Text Reader

Abstract

This invention relates to the field of image fusion technology, specifically to a multimodal image fusion method based on hesitant fuzzy variable-granularity dictionary learning. First, the source image is divided into blocks by adaptively selecting the granularity based on image quality. Then, features of the image blocks are extracted and their hesitant fuzzy membership degrees are calculated to quantify the uncertainty information in the image. Next, a joint overcomplete dictionary and sparse coefficients are obtained through dictionary learning, and hesitant fuzzy entropy and granularity coefficients are fused to construct adaptive weights. Finally, these weights are used to fuse the sparse coefficients and reconstruct the fused image. This method effectively solves the problems of insufficient image fuzziness processing, multi-scale structural representation, and adaptive feature extraction capabilities.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image fusion technology, and in particular to a multimodal image fusion method based on hesitant fuzzy variable granularity dictionary learning. Background Technology

[0002] Image fusion is a technique that integrates image information from multiple sensors into a single image, aiming to preserve key features of the source images to improve their recognizability, clarity, and information integrity. In applications such as multimodal medical imaging, remote sensing mapping, and infrared-visible light imaging, different modalities often reflect complementary information about the target object, such as structure and function, thermal features, and texture details. Therefore, how to effectively achieve complementary information fusion has become an important research topic in the fields of computer vision and intelligent perception.

[0003] Traditional image fusion methods largely rely on image decomposition and reconstruction strategies, such as multi-scale transformations and filter design. However, these methods are often limited by their representational capabilities when dealing with blurry, low-contrast, or weakly featured images. In recent years, sparse representation, dictionary learning, and deep learning methods have been gradually introduced into image fusion tasks, driving the development from pixel-level to decision-level fusion. In particular, dictionary learning methods, whose core idea is to sparsely reconstruct input data by linearly combining a small number of atoms in a dictionary, can achieve consistent modeling and selective fusion of cross-modal information in a sparse feature domain, effectively improving the robustness and interpretability of image fusion.

[0004] However, existing dictionary learning methods still have shortcomings when dealing with image ambiguity, multi-scale structural representation, and adaptive feature extraction. They also face challenges such as high computational complexity, limited adaptive dictionary size, and sensitivity to noise. Therefore, to improve the efficient application of dictionary learning in image fusion, this invention constructs a multimodal image fusion method based on hesitant fuzzy sets, leveraging their flexible representation advantages and uncertainty information representation capabilities. Summary of the Invention

[0005] The purpose of this invention is to provide a multimodal image fusion method based on hesitant fuzzy variable granularity dictionary learning, which solves the problem that existing methods are insufficient in handling image fuzziness, multi-scale structural representation and adaptive feature extraction.

[0006] To achieve the above objectives, this invention provides a multimodal image fusion method based on hesitant fuzzy variable-granularity dictionary learning, comprising the following steps:

[0007] Multimodal source images are acquired, and the quality of each source image is evaluated using an image quality assessment method. Based on the relationship between the evaluation results and a preset threshold, the granularity of image block division is adaptively selected. Based on the selected granularity, each source image is divided into non-overlapping image blocks.

[0008] For each corresponding image block pair, extract its multiple features and perform normalization processing. Map the normalized features to the fuzzy membership space, and calculate the hesitant fuzzy membership degree of each image block accordingly.

[0009] All image patches are converted into column vectors to form training samples. A dictionary learning algorithm is used to learn a joint overcomplete dictionary from the training samples. Using this dictionary, sparse coding is performed on each pair of image patch column vectors to obtain their sparsity coefficients.

[0010] Hesitant fuzzy entropy of each image block is calculated based on hesitant fuzzy membership degree. Adaptive fusion weights are constructed for the sparse coefficients of each pair of image blocks in combination with the current granularity coefficients, and the weights are normalized.

[0011] Using the normalized weights, the sparse coefficients of each pair of image blocks are weighted and fused to obtain the fused coefficients. The fused image blocks are reconstructed using the joint overcomplete dictionary and the fused coefficients. The windowing overlap method is used to reconstruct all the fused image blocks into the final fused image.

[0012] The process involves acquiring multimodal source images, evaluating the quality of each source image using an image quality assessment method, and adaptively selecting the granularity of image block division based on the relationship between the evaluation results and a preset threshold. Based on the selected granularity, each source image is divided into non-overlapping image blocks.

[0013] If the image quality evaluation value of any source image is higher than the division threshold, then a fine 4×4 granularity is used for block division; otherwise, a larger 8×8 granularity is used for block division.

[0014] For each corresponding image patch pair, multiple features are extracted and normalized. The normalized features are then mapped to a fuzzy membership space, and the hesitant fuzzy membership degree of each image patch is calculated accordingly. The specific steps include:

[0015] For each non-overlapping image patch Extract multiple features: f1|peak signal-to-noise ratio, f2|local gradient, ..., f k |Edge information}, and perform feature normalization to map it to a fuzzy membership space.

[0016]

[0017] in, Represents image blocks The j-th feature, Represents image blocks The j-th feature; then, based on this, construct respectively hesitant fuzzy membership degree

[0018]

[0019] in, Represents image blocks Membership degree under the first feature Represents image blocks Membership degree under the second feature Represents image blocks Membership degree under the k-th feature; Represents image blocks Membership degree under the first feature Represents image blocks Membership degree under the second feature Represents image blocks Membership degree under the k-th feature Represents the i-th image patch Image hesitant blur membership, Represents the i-th image patch The degree of hesitation and fuzzy membership.

[0020] In this process, all image patches are converted into column vectors to form training samples. A dictionary learning algorithm is used to learn a joint overcomplete dictionary from the training samples. Using this dictionary, sparse coding is performed on each pair of image patch column vectors to obtain their sparsity coefficients.

[0021] The K-SVD algorithm is used to obtain the joint dictionary D∈R. d×K and image blocks sparse coefficient matrix

[0022]

[0023] Where T1 represents the sparsity threshold.

[0024] Specifically, the hesitant fuzzy entropy of each image patch is calculated based on the hesitant fuzzy membership degree. Combined with the current granularity coefficient, an adaptive fusion weight is constructed for the sparse coefficients of each pair of image patches, and the weight is normalized.

[0025] For the corresponding image patch Construct sparse coefficient matrices respectively Fusion weights

[0026] in,

[0027]

[0028] in, and Representing image blocks The hesitant fuzzy entropy, Let represent the variable granularity coefficient, k represent the total number of features, g represent the granularity size, and λ1 and λ2 be the weight values, where λ1 + λ2 = 1. Represents the i-th image patch Image hesitant blur membership, Represents the i-th image patch The degree of hesitation and fuzzy membership.

[0029] The normalization process yields the final weight information.

[0030]

[0031] in, Represents the sparse coefficient matrix The weight, Represents the sparse coefficient matrix The weight, Represents the sparse coefficient matrix The normalized weights, Represents the sparse coefficient matrix The normalized weights.

[0032] In this process, the sparse coefficients of each pair of image blocks are weighted and fused using normalized weights to obtain fused coefficients. The fused image blocks are then reconstructed using a joint overcomplete dictionary and the fused coefficients. Finally, a windowed overlap method is used to reconstruct all fused image blocks into the final fused image.

[0033] Use the obtained final weight information For sparse coefficient matrix To merge:

[0034]

[0035] in, This represents the sparse matrix after fusion.

[0036] In this process, the sparse coefficients of each pair of image blocks are weighted and fused using normalized weights to obtain fused coefficients. The fused image blocks are then reconstructed using a joint overcomplete dictionary and the fused coefficients. Finally, a windowed overlap method is used to reconstruct all fused image blocks into the final fused image.

[0037] Reconstruct the complete image from all image patches using the Hann window:

[0038]

[0039] Among them, I F The fused image patch obtained through sparse matrix and dictionary Let represent the i-th fused image patch, and D represent the joint dictionary. Let represent the sparse coefficients after fusion, and Hann(g) represent the Hanning window function.

[0040] This invention presents a multimodal image fusion method based on hesitant fuzzy variable-granularity dictionary learning. Traditional image processing and feature extraction methods often struggle to simultaneously address data uncertainty and multi-granularity structural information. Hesitant fuzzy sets, by introducing "hesitation degree" to model information uncertainty, better represent ambiguous or vague regions in an image; variable-granularity theory provides a framework for understanding and processing image information from multiple scales. Introducing both into the dictionary learning process allows the constructed sparse representation dictionary to adapt to both data uncertainty and information needs at different granularity levels, improving the model's generalization and expressive power. This method is particularly valuable in image fusion tasks. The core of image fusion is how to extract and retain key information from the original image. Hesitant fuzzy variable-granularity dictionary learning can more sensitively identify edges, textures, and salient regions in the image during feature extraction, providing more accurate and stable support for the fusion of different modalities. Simultaneously, through multi-granularity processing, a better balance can be achieved between local details and global structure, resulting in better visual quality and semantic integrity in the fusion results. It effectively solves the problems of insufficient image blur processing, multi-scale structural representation, and adaptive feature extraction capabilities. Attached Figure Description

[0041] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below.

[0042] Figure 1 This is a flowchart of the steps of the multimodal image fusion method based on hesitant fuzzy variable granularity dictionary learning according to the first embodiment of the present invention. Detailed Implementation

[0043] The embodiments of the present invention are described in detail below. Examples of the embodiments are shown in the accompanying drawings. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain the present invention, but should not be construed as limiting the present invention.

[0044] The first embodiment of this application is as follows:

[0045] Please see Figure 1 ,in, Figure 1 This is a flowchart of the steps of the multimodal image fusion method based on hesitant fuzzy variable granularity dictionary learning according to the first embodiment of the present invention.

[0046] This invention provides a multimodal image fusion method based on hesitant fuzzy variable granularity dictionary learning, comprising the following steps:

[0047] S101: Acquire multimodal source images, evaluate the quality of each source image using an image quality assessment method, and adaptively select the granularity of image block division based on the relationship between the evaluation results and a preset threshold. Based on the selected granularity, divide each source image into non-overlapping image blocks.

[0048] S102: For each corresponding image block pair, extract its multiple features and perform normalization processing, map the normalized features to the fuzzy membership space, and calculate the hesitant fuzzy membership degree of each image block accordingly.

[0049] S103: Convert all image patches into column vectors to form training samples. Use a dictionary learning algorithm to learn a joint overcomplete dictionary from the training samples. Use this dictionary to sparsely encode each pair of image patch column vectors to obtain their sparsity coefficients.

[0050] S104: Calculate the hesitant fuzzy entropy of each image block based on the hesitant fuzzy membership degree, combine it with the current granularity coefficient, construct an adaptive fusion weight for the sparse coefficients of each pair of image blocks, and normalize the weight.

[0051] S105: Using the normalized weights, the sparse coefficients of each pair of image blocks are weighted and fused to obtain the fused coefficients. The fused image blocks are reconstructed using the joint overcomplete dictionary and the fused coefficients. The windowing overlap method is used to reconstruct all the fused image blocks into the final fused image.

[0052] Specifically, assume the input image is divided into two modalities: input: two modalities: I1∈R M×N ,I2∈R M ×N ;

[0053] Output: Fused image I F ;

[0054] Step 1: Evaluate the input images I1 and I2 using the National Image Quality Evaluation (NIQE) method, where NIQE(·) is the Natural Image Quality Evaluation algorithm.

[0055] q1 = NIQE(I1), q2 = NIQE(I2),

[0056] Step 2: Set the segmentation threshold T0 and select the granularity based on the image quality assessment results:

[0057] If (q1+q2) / 2≤T0, then g=4;

[0058] If (q1+q2) / 2>T0, then g=8;

[0059] Step 3: Divide the image into g×g non-overlapping blocks according to the granularity of the division:

[0060]

[0061] in, This represents the non-overlapping block of the input image I1. This represents the non-overlapping block of the input image I2;

[0062] Step 4: For each non-overlapping block Extract multiple features: f1|peak signal-to-noise ratio, f2|local gradient, ..., f k |Edge information}, and perform feature normalization to map it to a fuzzy membership space.

[0063]

[0064] in, Represents image blocks The j-th feature, Represents image blocks The j-th feature;

[0065] Then, based on this, construct separately hesitant fuzzy membership degree

[0066]

[0067] Step 5: Learning fuzzy-to-granular dictionary concepts:

[0068] Step 5.1: Represent the image patches as column vectors:

[0069]

[0070] Among them, X 1 X represents the input image in column vector form I1. 2 This represents the input image in I2 column vector form;

[0071] Step 5.2: Obtain the joint dictionary D∈R using the K-SVD algorithm. d×K and image blocks sparse coefficient matrix

[0072]

[0073] Where T1 represents the sparsity threshold.

[0074] Step 5.3: For the corresponding image patch Construct sparse coefficient matrices respectively Fusion weights

[0075]

[0076] in,

[0077]

[0078] in, and Representing image blocks The hesitant fuzzy entropy, Let represent the variable granularity coefficient, k represent the total number of features, g represent the granularity size, and λ1 and λ2 be the weight values, where λ1 + λ2 = 1. Represents the i-th image patch Image hesitant blur membership, Represents the i-th image patch The degree of hesitation and fuzzy membership.

[0079] Then, normalization is performed to obtain the final weight information.

[0080]

[0081] in, Represents the sparse coefficient matrix The weight, Represents the sparse coefficient matrix The weight, Represents the sparse coefficient matrix The normalized weights, Represents the sparse coefficient matrix The normalized weights.

[0082] Step 6: Coefficient matrix fusion and reconstruction:

[0083] Step 6.1: Use the obtained final weight information For sparse coefficient matrix To merge:

[0084]

[0085] in, This represents the sparse matrix after fusion.

[0086] Step 6.2: Use the Hann window to reconstruct all image patches into a complete image:

[0087]

[0088] Among them, I F The fused image patch obtained through sparse matrix and dictionary Let represent the i-th fused image patch, and D represent the joint dictionary. Let represent the sparse coefficients after fusion, and Hann(g) represent the Hanning window function.

[0089] Traditional image processing and feature extraction methods often struggle to simultaneously address data uncertainty and multi-granularity structural information. Hesitant fuzzy sets, by introducing "hesitation degree" to model information uncertainty, can better represent ambiguous or vague regions in images. Variable granularity theory provides a framework for understanding and processing image information from multiple scales. Introducing both into the dictionary learning process allows the constructed sparse representation dictionary to adapt to both data uncertainty and information needs at different granular levels, improving the model's generalization ability and expressive power.

[0090] This method is particularly valuable in image fusion tasks. The core of image fusion is how to extract and retain key information from the original image. Hesitant fuzzy variable-granularity dictionary learning can more accurately identify edges, textures, and salient regions in the image when extracting features, providing more precise and stable support for the fusion of different modalities. Simultaneously, through multi-granularity processing, a better balance can be achieved between local details and global structure, resulting in better performance in both visual quality and semantic integrity of the fusion result.

[0091] The above-disclosed embodiments are merely one or more preferred embodiments of this application and should not be construed as limiting the scope of this application. Those skilled in the art can understand that all or part of the processes for implementing the above embodiments and equivalent changes made in accordance with the claims of this application still fall within the scope of this application.

Claims

1. A multi-modal image fusion method based on hesitant fuzzy variable granularity dictionary learning, characterized in that, Includes the following steps: Multimodal source images are acquired, and the quality of each source image is evaluated using an image quality assessment method. Based on the relationship between the evaluation results and a preset threshold, the granularity of image block division is adaptively selected. Based on the selected granularity, each source image is divided into non-overlapping image blocks. For each corresponding image patch pair, its various features are extracted and normalized. The normalized features are then mapped to a fuzzy membership space, and the hesitant fuzzy membership degree of each image patch is calculated accordingly. Specifically, for each non-overlapping image patch... Extracting multiple features And perform feature normalization to map to a fuzzy membership space. , : in, Represents image blocks The j-th feature, Represents image blocks The j-th feature; Then, based on this, construct separately , hesitant fuzzy membership degree , , in, Represents image blocks Membership degree under the first feature Represents image blocks Membership degree under the second feature Represents image blocks Membership degree under the k-th feature; Represents image blocks Membership degree under the first feature Represents image blocks Membership degree under the second feature Represents image blocks Membership degree under the k-th feature Represents the i-th image patch Image hesitant blur membership, Represents the i-th image patch The degree of hesitant and ambiguous membership; All image patches are converted into column vectors to form training samples. A dictionary learning algorithm is used to learn a joint overcomplete dictionary from the training samples. Using this dictionary, sparse coding is performed on each pair of image patch column vectors to obtain their sparsity coefficients. Hesitant fuzzy entropy is calculated for each image patch based on hesitant fuzzy membership degree. Adaptive fusion weights are then constructed for the sparse coefficients of each pair of image patches, combined with the current granularity coefficients. These weights are then normalized. Specifically, for the corresponding image patch... Construct sparse coefficient matrices respectively Fusion weights in, in, and Representing image blocks The hesitant fuzzy entropy, , denoted by , where k represents the total number of features and g represents the granularity. , The weight value, and , Represents the i-th image patch Image hesitant blur membership, Represents the i-th image patch The degree of hesitant and ambiguous membership; Using the normalized weights, the sparse coefficients of each pair of image blocks are weighted and fused to obtain the fused coefficients. The fused image blocks are reconstructed using the joint overcomplete dictionary and the fused coefficients. The windowing overlap method is used to reconstruct all the fused image blocks into the final fused image.

2. The multimodal image fusion method based on hesitant fuzzy variable-granularity dictionary learning as described in claim 1, characterized in that, Multimodal source images are acquired, and the quality of each source image is evaluated using an image quality assessment method. Based on the relationship between the evaluation results and a preset threshold, the granularity of image patch division is adaptively selected. Based on the selected granularity, each source image is divided into non-overlapping image patches. If the image quality evaluation value of any source image is higher than the division threshold, then a fine 4×4 granularity is used for block division; otherwise, a larger 8×8 granularity is used for block division.

3. The multimodal image fusion method based on hesitant fuzzy variable-granularity dictionary learning as described in claim 1, characterized in that, All image patches are converted into column vectors to form training samples. A dictionary learning algorithm is used to learn a joint overcomplete dictionary from the training samples. Using this dictionary, sparse encoding is performed on each pair of image patch column vectors to obtain their sparsity coefficients. Use the K-SVD algorithm to obtain the joint dictionary and image blocks sparse coefficient matrix in, This represents the sparsity threshold.

4. The multimodal image fusion method based on hesitant fuzzy variable-granularity dictionary learning as described in claim 1, characterized in that, The final weight information is obtained through normalization. , : in, Represents the sparse coefficient matrix The weight, Represents the sparse coefficient matrix The weight, Represents the sparse coefficient matrix The normalized weights, Represents the sparse coefficient matrix The normalized weights.

5. The multimodal image fusion method based on hesitant fuzzy variable-granularity dictionary learning as described in claim 4, characterized in that, Using normalized weights, the sparse coefficients of each pair of image patches are weighted and fused to obtain fused coefficients. The fused image patches are then reconstructed using a joint overcomplete dictionary and the fused coefficients. A windowed overlap method is then used to reconstruct all fused image patches into the final fused image. Use the obtained final weight information , For sparse coefficient matrix To merge: in, This represents the sparse matrix after fusion.

6. The multimodal image fusion method based on hesitant fuzzy variable-granularity dictionary learning as described in claim 5, characterized in that, Using normalized weights, the sparse coefficients of each pair of image patches are weighted and fused to obtain fused coefficients. The fused image patches are then reconstructed using a joint overcomplete dictionary and the fused coefficients. A windowed overlap method is then used to reconstruct all fused image patches into the final fused image. Reconstruct the complete image from all image patches using the Hann window: in, The fused image patch obtained through sparse matrix and dictionary Let represent the i-th fused image patch, and D represent the joint dictionary. Represents the sparse coefficients after the i-th fusion. This represents the Hanning window function.