Image fusion method combining convolution features and gradient histogram weighting
By combining convolutional features and gradient histogram weighting, the problem that existing image fusion methods are difficult to preserve image edges and details is solved, and a higher quality image fusion effect is achieved, with good noise immunity and scalability.
Patent Information
- Application Number
- CN202411936839.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-26
- Publication Date
- 2025-05-16
AI Technical Summary
Existing image fusion methods are difficult to effectively retain and enhance the edge, detail and contrast information of the image, especially under noise interference, resulting in blurring of details and degradation of the fused image.
The image fusion method combining convolutional features and gradient histogram weighting is adopted to generate weight information by calculating the local gradient histogram, and extracting features with convolutional neural network, weighted fusion and reconstruction are performed to finally generate high-quality fusion images.
This method can better preserve edge details information in the image, enhance the contrast and detail clarity of the image, and has good noise immunity and application scenario scalability.
Smart Images

Figure CN120013772A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image fusion, and in particular to an image fusion method combining convolution features with gradient histogram weighting. Background Art
[0002] With the rapid development of the information age, multi-sensor image fusion technology has become a widely used research direction. Especially in the fields of airborne remote sensing, medical image processing, security monitoring, etc., different sensors (such as visible light, infrared, X-ray, ultrasound, etc.) can provide rich and diverse information. Through the fusion processing of multi-source images, various types of information can be effectively combined to improve the accuracy and clarity of scene understanding. However, how to retain and enhance the edge, details and contrast information of the image during the fusion process has become a technical problem.
[0003] Traditional image fusion methods are usually based on multi-scale transforms (such as wavelet transform, Laplacian pyramid, etc.) or weighted strategies based on statistical features (such as mean, variance). Although these methods work well in some applications, they generally have several defects. First, traditional methods usually rely on regular feature selection and are difficult to adaptively capture the details and edge features of the image; second, noise interference can significantly affect the fusion quality, leading to problems such as blurred details and reduced contrast. In particular, although multi-scale transform methods such as Laplacian pyramid and wavelet transform capture the frequency components of the image at different scales, due to the lack of effective description of the edges and details of the image, the fused image often fails to reflect the ideal visual effect.
[0004] In recent years, deep learning has made remarkable achievements in the field of computer vision, especially convolutional neural networks (CNNs), which have shown strong feature extraction capabilities in tasks such as image classification, detection, segmentation, fusion, and decision support. Adaptive multi-source image fusion can be achieved by learning the complex feature relationships of multi-source sensor images through deep neural networks and combining the correlation of various parts of the image. Compared with traditional methods, deep learning methods have better performance in noise suppression, detail enhancement, and edge preservation. However, existing image fusion methods based on deep learning still have certain shortcomings when dealing with edge details. Specifically, the convolution operation of the convolutional neural network is usually not sensitive enough to edge and detail features when extracting local features, resulting in blurred edges and unclear edges of some fused images. Therefore, how to better retain detail information based on the convolutional neural network has become a key issue to further improve the fusion effect. Summary of the invention
[0005] The purpose of the present invention is to design an image fusion method combining convolution features and gradient histogram weighting. The method first calculates the gradients of the input image in the horizontal and vertical directions within a 5x5 neighborhood, and uses the gradient values in the two directions to calculate the final gradient amplitude map, thereby generating a local gradient histogram; then, weight information is generated according to the gradient histogram, and normalization is performed, and at the same time, the input image is subjected to feature extraction using a convolutional neural network; finally, the extracted feature map is subjected to weighted fusion and fusion feature map reconstruction to generate the final fused image, and the overall method has strong robustness and image fusion quality is better than general algorithms.
[0006] The invention objective of the present invention is mainly achieved by the following technical solutions:
[0007] Provided is an image fusion method combining convolution features and gradient histogram weighting, comprising:
[0008] Obtaining detail feature distribution and intensity distribution of the original input image at the multi-source sensor end;
[0009] For each original input image, different weights are assigned based on detail feature distribution and intensity distribution;
[0010] Adaptively extract the features of the original input image through a convolutional neural network;
[0011] The weights are used to guide the feature fusion process, so that the edge detail information in the image is better preserved in the fused image.
[0012] Furthermore, the detailed feature distribution and intensity distribution of the original input image at the multi-source sensor end are obtained, including:
[0013] For any original image, calculate the horizontal gradient G x With the vertical gradient G y , thereby calculating the gradient amplitude of the original image;
[0014] For each pixel position (x, y), the local gradient histogram is calculated based on the gradient amplitude in its 5x5 neighborhood N(x, y);
[0015] The local gradient histogram is used to calculate the gradient weighted value of each pixel in the original image, where areas with higher gradient intensity obtain higher weights; the gradient weighted value represents the distribution of detail features and intensity distribution.
[0016] Furthermore, the convolutional neural network is used to adaptively extract the features of the original input image, including:
[0017] The original feature map is extracted through a series of convolutional layers and pooling layers.
[0018] Furthermore, weights are used to guide the feature fusion process, including:
[0019] The gradient weighted value of each pixel in multiple original images corresponds to each pixel of the feature map, and the fused weighted feature map is calculated;
[0020] The deconvolution operation corresponding to the convolution is performed on the fused weighted feature map to reconstruct the final fused image.
[0021] Furthermore, the fused weighted feature map F (x,y) :
[0022] F (x,y) =F 1(x,y) ·W 1(x,y) +F 2(x,y) ·W 2(x,y) +…+F n(x,y) ·W n(x,y) ;
[0023] Among them, F (x,y) It is formed by fusing n weighted feature maps, F 1(x,y) is the feature map of the first original image, W 1(x,y) is the gradient weighted value of the first original image, F 2(x,y) is the feature map of the second original image, W 2(x,y) is the gradient weighted value of the second original image, F n(x,y) is the feature map of the last original image, W n(x,y) is the gradient weighted value of the last original image, and (x, y) is any coordinate of the fused weighted feature map.
[0024] Furthermore, each layer of deconvolution operation is expressed as:
[0025]
[0026] Among them, F (l) represents the feature map of layer l, T (l) is the deconvolution kernel of the lth layer, * represents the convolution operation, Represents the feature map of the l+1th layer obtained by deconvolution.
[0027] Furthermore, the local gradient histogram H k(x,y) It is expressed as:
[0028]
[0029] Among them, H h (i, x, y) is the i-th bin value of the local gradient histogram at position (x, y), which represents the gradient distribution of the pixel neighborhood. k(x′, y′) is the gradient magnitude of N(x, y).
[0030] Furthermore, the gradient weighted value W (x,y) The calculation formula is:
[0031]
[0032] N is the number of bins in the histogram.
[0033] Beneficial effects:
[0034] The present invention proposes to use a convolutional neural network to extract image features with strong adaptability. The convolutional neural network has a strong feature learning ability and can adaptively perform weighted learning based on the gradient features and histogram distribution of multi-source images. In the feature fusion process, the network can adaptively perform weighted processing on the importance of each image area to ensure that the fused image can optimally retain detail information. This adaptive characteristic enables the method to cope with multi-source image fusion tasks in different scenarios.
[0035] The present invention combines the gradient histogram to enhance the edge and image detail information after fusion. The gradient histogram can effectively reflect the distribution and intensity information of the edge and detail areas in the image. Therefore, by weighting the gradient histogram, the fusion method can highlight and retain the edge information and detail features in the input image, and the generated fused image has a clearer contour and a more delicate visual effect.
[0036] The method proposed in the present invention has good noise resistance. The gradient histogram weighting method can have a stronger suppression effect on the noise in the image. When calculating the gradient information, the gradient histogram statistics can effectively smooth the high-frequency noise in the image and prevent the noise from being overly prominent in the histogram. By combining the deep feature extraction capability of the convolutional neural network, the network can automatically suppress noise interference, thereby improving the clarity of the fused image.
[0037] The method proposed in the present invention has good scalability in application scenarios. The gradient histogram weighting method proposed in the article can be extended to a variety of image fusion scenarios, including infrared and visible light image fusion, CT and MRI medical image fusion, remote sensing image multispectral fusion, etc. Especially in application scenarios with high requirements for detail preservation, this method can significantly improve the detail clarity and visual effect of the image. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] To further illustrate the technical content of the present invention, the present invention is described in detail below with reference to the accompanying drawings and implementation examples, wherein:
[0039] Figure 1 It is an overall block diagram of an image fusion method combining convolution features and gradient histogram weighting in the present invention.
[0040] Figure 2(a) is the original infrared image.
[0041] Figure 2(b) is the original visible light image.
[0042] Figure 3 This is a conventional fusion effect diagram using the Laplace transform algorithm in an embodiment of the present invention.
[0043] Figure 4 This is an image fusion effect diagram of an embodiment of the present invention using a convolution feature combined with a gradient histogram weighted algorithm. DETAILED DESCRIPTION
[0044] In order to solve these problems, this patent proposes an image fusion method based on gradient histogram weighting. The gradient histogram is a statistical feature description of the image gradient information, which is used to measure the edge detail distribution of the image. By performing histogram statistics on the gradient amplitude of the image, the detail feature distribution and intensity distribution of the image can be obtained. In the gradient histogram weighting method, the gradient information of each source image is quantized and calculated in a histogram manner, and different weights are assigned according to the gradient histogram distribution to highlight the detail-rich parts of each image, so that the details and edges can be effectively enhanced during the fusion process. In addition, the method combines a convolutional neural network to further optimize the image fusion effect through the nonlinear features learned by the network. By weighting the gradient histogram of the input image, the convolutional neural network can be guided to adaptively adjust the feature extraction and feature fusion process, thereby better retaining the edge detail information in the image and enhancing the image contrast and detail clarity.
[0045] The present invention provides an image fusion method combining convolution features and gradient histogram weighting, such as Figure 1 As shown, the following steps are included: Step 1: Obtain the original input image for image fusion from the multi-source sensor, such as an infrared image and visible light images , the resolution of both images is ; Step 2: Input the infrared image in step 1 and visible light images Each pixel in Calculate the horizontal gradient With vertical gradient , thereby calculating the gradient amplitude of the two images and ; Step 3: For the position of each pixel of the two images in step 2 According to its 5x5 neighborhood Gradient magnitude statistics in local gradient histogram and .
[0049] Step 4: Use the local gradient histogram H in step 3 b(x,y) With H v(x,y) , calculate the gradient weighted value W of each pixel in the two images respectively b(x,y ) and W v(x,y) , where regions with higher gradient strength receive higher weights;
[0050] Step 5: Input the visible light image and infrared image in step 1 into the convolutional neural network for feature extraction, and extract the infrared feature map F through a series of convolutional layers and pooling layers. b(x,y) With visible light characteristic map F v(x,y) , combined with the weight W in step 4 b(x,y) With W v(x,y) , weight the feature maps respectively to obtain the final fusion feature map F (x,y) ;
[0051] Step 6: Weighted fusion feature map F in step 5 (x,y) Deconvolution is used to upsample layer by layer, so as to reconstruct the fused feature map into a fused image I with visual effects. Fused , and the resolution is restored to the original resolution H×W.
[0052] The images in step 1 are two images to be fused, namely, visible light images I and v With infrared image I b ;
[0053] As a preferred embodiment of the present invention, the specific processing process of step 2 is: first, I v The image is calculated using the horizontal gradient kernel With the vertical gradient calculation kernel Perform convolution operations by multiplying and adding corresponding elements to obtain the gradient values G in two directions x With G y , then for G x With G y The final gradient value G is obtained by using the average root mean square algorithm. v(x,y) , and similarly we get I b The gradient value G of the image b(x,y) .
[0054] As a preferred solution of the present invention, the specific processing process of step 3 is to select a 5x5 neighborhood with each pixel as the center, and convert the gradient amplitude G v(x,y)The value of is counted in the interval [0, 255] to generate a local gradient histogram H v(x,y) , and similarly we can get I b The local gradient histogram H of the image b(x,y) .
[0055] As a preferred embodiment of the present invention, the specific processing process of step 4 is as follows: v The local gradient histogram H of each pixel position (x, y) in the image v(x,y) , the weighted value W of the pixel position is determined according to the average gradient amplitude of the area v(x,y) , and similarly we can get I b The weighted value W at each pixel position in the image b(x,y) .
[0056] As a preferred embodiment of the present invention, the specific processing process of step 5 is as follows: first, v and I b The images are respectively subjected to eigenvalue F using the residual convolutional neural network structure ResNet18 v(x,y) and F b(x,y) Extraction, using W v(x,y) and W b(x,y) The convolution features of each image are weighted and added pixel by pixel to obtain the weighted feature map P. (x,y) .
[0057] As a preferred solution of the present invention, the specific processing process of step 6 is to use 5 layers of deconvolution to perform upsampling layer by layer, so that the feature map is gradually restored to a higher resolution, thereby reconstructing a fused image.
[0058] The present invention proposes an image fusion method combining convolution features with gradient histogram weighting, including input of infrared images and visible light images, calculation of the gradient amplitude of each image, generation of local gradient histograms, determination of local weighted values of pixels, extraction of image convolution features and calculation of weighted fusion features, and reconstruction of fused images.
[0059] First, as shown in Figure 2(a) and Figure 2(b), the infrared image I directly output by the infrared sensor is input. b and the visible light image I output by the visible light sensor v , the width of the two images is W, the height is H, and the grayscale value of the image is within [0, 255]. The gradient calculation in the horizontal and vertical directions of each image is performed as follows:
[0060]
[0061] Among them, G x is the horizontal gradient of the image, G yis the vertical gradient of the image, I x is the input infrared or visible light image, and its value is I b and I v .
[0062] Furthermore, the gradient magnitude of each pixel is calculated by the horizontal gradient value and the vertical gradient value:
[0063]
[0064] Among them, G k(x,y) represents the gradient amplitude at the corresponding pixel position, and k represents the infrared image source or the visible light image source.
[0065] Furthermore, we take a 5x5 neighborhood N(x, y) centered at each pixel and calculate the gradient magnitude G of its neighborhood. k The value of (x′, y′) generates the local gradient histogram H k(x,y) , the value of each bin is H k (i, x, y) represents the number of pixels in the neighborhood within the corresponding gradient amplitude range:
[0066]
[0067] Among them, H k (i, x, y) is the i-th bin value of the local gradient histogram at position (x, y), which represents the gradient distribution of the pixel neighborhood.
[0068] Furthermore, after obtaining the local gradient histogram of each pixel position, the weighted value W of each pixel is calculated. k(x,y) , where regions with higher gradient strength receive higher weights. The weight calculation formula is as follows:
[0069]
[0070] Where N is the number of bins in the histogram, and W is k(x,y) Can be the infrared image source weight W b(x,y) Or the visible light image source weight W v(x,y) , so that W k(x,y) The weight values are normalized to between [0, 1].
[0071] Furthermore, the infrared image I b and visible light image I v Input into the residual convolutional neural network RenNet18, and extract the feature map F through a series of convolutional layers and pooling layers v(x,y) and F b(x,y) , as shown below:
[0072] F v(x,y)=ResNet18(I v )
[0073] F b(x,y) =ResNet18(I b )
[0074] Then, the local weighted value W at each pixel position is used k(x,y) , weight the convolution features of each image pixel by pixel to obtain the fused weighted feature map F (x,y) :
[0075] F (x,y) =F v(x,y) ·W v(x,y) +F b(x,y) ·W b(x,y)
[0076] Furthermore, through 5 layers of deconvolution, the weighted fusion feature map F (x,y) Reconstruct the final fused image I Fused , each layer operation of deconvolution can be expressed as:
[0077]
[0078] Among them, F (l) represents the feature map of layer l, T (l) is the deconvolution kernel of the lth layer, and * represents the convolution operation. After multiple layers of transposed convolution, the output feature map F (l) Convert it to the original resolution, and then use a convolution layer Conv to adjust the number of channels to generate the final fused image:
[0079] I Fused =Conv(F (l) )
[0080] In order to verify the effectiveness of the image fusion method combining convolution features and gradient histogram weighting provided by the present invention, this example conducts experiments on two infrared images and visible light images with a resolution of 640×480. Figure 3 In order to fuse the two original images using conventional Laplace transform, the result image contains rich detail information, but its color is distorted, the image edge details are blurred, and the overall visual effect is not good.
[0081] according to Figure 4 It can be seen that the image fused using the method provided by the present invention can not only significantly improve the content and detail information of the original image, but also restore the color of the original image itself, presenting a better visual effect in enhancing the image quality.
[0082] The specific embodiments described above further illustrate the objectives, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above description is only a specific embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the protection scope of the present invention.
Claims
1. An image fusion method combining convolution features and gradient histogram weighting, characterized in that: include: Obtaining detail feature distribution and intensity distribution of the original input image at the multi-source sensor end; For each original input image, different weights are assigned based on detail feature distribution and intensity distribution; Adaptively extract the features of the original input image through a convolutional neural network; The weights are used to guide the feature fusion process, so that the edge detail information in the image is better preserved in the fused image.
2. The method according to claim 1, characterized in that: Obtain the detail feature distribution and intensity distribution of the original input image from the multi-source sensor, including: For any original image, calculate the horizontal gradient G x With the vertical gradient G y , thereby calculating the gradient amplitude of the original image; For each pixel position (x, y), the local gradient histogram is calculated based on the gradient amplitude in its 5x5 neighborhood N(x, y); The local gradient histogram is used to calculate the gradient weighted value of each pixel in the original image, where areas with higher gradient intensity obtain higher weights; the gradient weighted value represents the distribution of detail features and intensity distribution.
3. The method according to claim 2, characterized in that The convolutional neural network is used to adaptively extract the features of the original input image, including: The original feature map is extracted through a series of convolutional layers and pooling layers.
4. The method according to claim 3, characterized in that Use weights to guide the feature fusion process, including: The gradient weighted value of each pixel in multiple original images corresponds to each pixel of the feature map, and the fused weighted feature map is calculated; The deconvolution operation corresponding to the convolution is performed on the fused weighted feature map to reconstruct the final fused image.
5. The method according to claim 4, characterized in that The fused weighted feature map F (x,y) : F (x,y) =F 1(x,y) ·W 1(x,y) +F 2(x,y) ·W 2(x,y) +…+F n(x,y) ·W n(x,y) ; Among them, F (x,y) It is formed by fusing n weighted feature maps, F 1(x,y) is the feature map of the first original image, W 1(x,y) is the gradient weighted value of the first original image, F 2(x,y) is the feature map of the second original image, W 2(x,y) is the gradient weighted value of the second original image, F n(x,y) is the feature map of the last original image, W n(x,y) is the gradient weighted value of the last original image, and (x, y) is any coordinate of the fused weighted feature map.
6. The method according to claim 5, characterized in that Each layer of deconvolution operation is expressed as: Among them, F (l) represents the feature map of layer l, T (l) is the deconvolution kernel of the lth layer, * represents the convolution operation, Represents the feature map of the l+1th layer obtained by deconvolution.
7. The method according to claim 6, characterized in that Local gradient histogram H k(x,y) It is expressed as: H k (i,x,y)=∑ (x′,y′)∈N(x,y) δ(G k (x′,y′)) Among them, H k (i, x, y) is the i-th bin value of the local gradient histogram at position (x, y), which represents the gradient distribution of the pixel neighborhood. k (x′, y′) is the gradient amplitude of N(x, y); δ(*) is the impulse function.
8. The method according to claim 7, characterized in that Gradient weight W (x,y) The calculation formula is: N is the number of bins in the histogram.
Citation Information
Cited By
Local modification redrawing method and system for generating planning intention graph
CN120912724A