A method for fusing heterogeneous images in a traffic scenario

The method addresses computational inefficiencies and halo issues in traffic scene image fusion by using grid partitioning and separate network processing, achieving efficient and robust halo removal with enhanced detail in fused images.

CN115456922BActive Publication Date: 2025-07-15SOUTHEAST UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202211064526.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-01
Publication Date
2025-07-15
Estimated Expiration
2042-09-01

AI Technical Summary

Technical Problem

In the existing traffic scenes, infrared and visible image fusion algorithms have large amounts of calculation and are not universal in dealing with halo problems, resulting in a decline in image quality and cannot meet the needs of real-time traffic scenes.

Method used

The halo area and background area are processed separately by using the methods of image registration, meshing and network combination, and images are fusion using tensor addition and U-Net structure network groups to achieve independent training and rapid processing.

Benefits of technology

Improves the speed and quality of image fusion, suitable for a variety of traffic scenarios, providing high-quality monitoring data and a data source for vehicle anti-halo and safe driving.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115456922B_ABST
    Figure CN115456922B_ABST
Patent Text Reader

Abstract

The present invention provides a method for fusing heterogeneous images in a traffic scenario, including: acquiring heterogeneous visible light and infrared image data, performing heterogeneous image registration, dividing the source image into grids to form flare images and background images, designing a dual discriminator generative adversarial network group to classify and fuse the image blocks, restoring the positions of the image blocks, and outputting a fused image with anti-flare ability. The present invention effectively improves the overall quality of the fused image, improves the accuracy of matching points, and reduces the spatial deviation generated by registration in the image sequence. The present invention solves the problems of difficult data acquisition, slow calculation speed, and narrow application range of existing anti-flare algorithms, realizes convenient implementation and fast speed of the algorithm, is applicable to various traffic scenarios including flare, and provides a high-quality data source for traffic scenario applications such as monitoring data acquisition and anti-flare safe driving of vehicles.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of infrared and visible light image fusion, and particularly relates to a method for fusing heterogeneous images in a traffic scenario. Background Art

[0002] With the continuous improvement of computer technology and the continuous development of image processing technology, image-based traffic information perception applications are increasing day by day. Among them, visible light sensors have the following characteristics and are widely used: imaging is similar to human vision, and the data is intuitive; the data contains environmental information and has rich semantic features; high resolution and low deployment cost, etc. However, visible light sensors have poor imaging effects in low light, bad weather and other situations. Infrared image sensors based on thermal images detect the thermal radiation of objects to form images, and have the following characteristics: strong anti-interference ability and can work in harsh environments; all-weather working characteristics; good thermal target detection ability; the overall image formed is relatively blurred, with low resolution and not suitable for observation and understanding. Considering the imaging characteristics and complementary characteristics of infrared and visible light images, using the two to cooperate in imaging and fuse can obtain a fused image with prominent targets and rich background information at night, in bad weather and other situations, thereby improving the efficiency of traffic information perception.

[0003] Veiling glare in a traffic scenario refers to the blind area of vision caused by overexposure of night lights in a visible light sensor. Veiling glare from the driver's perspective will affect the driver's judgment of the road conditions and pose serious safety hazards; veiling glare from the monitoring perspective will cause the loss of traffic target details in the veiling glare area, affecting traffic data collection and road condition judgment. In order to reduce the impact of veiling glare and improve driving safety and the efficiency of monitoring data collection, in recent years, many studies have tried to solve the veiling glare problem from the perspectives of glare removal, heterogeneous image fusion, etc.

[0004] Heterogeneous image fusion based on visible light and infrared is a feasible and effective image processing technology. In existing research, Chinese Patent CN202111352352.X discloses an infrared and visible light fusion technology based on multi-scale features, and Chinese Patent CN202110901665.X discloses an infrared and visible light image fusion method based on frequency domain rules. However, the existing methods require manual calibration and design for specific scenarios, the entire process has a large amount of calculation and the algorithm has poor universality. Moreover, in a typical traffic scenario containing veiling glare, a large amount of single-source image information is lost, which has a negative impact on the quality of the fused image. Existing fusion algorithms rarely consider the situation of single-source information loss, resulting in a darker fused image or the loss of target details within the veiling glare in this scenario, and the fusion effect is not good.

[0005] The image anti-halos technology is a feasible and effective image processing technology. In existing research, Chinese Patent CN202110325989.3 proposes an anti-halos method for image fusion based on NSCT, and Chinese Patent CN201810617257.X proposes a method for removing halos based on the sparse representation of image brightness gradient, generating a fusion weight according to the image brightness gradient and exposure quality to remove halos in the image. Generally speaking, existing research has achieved certain results in the anti-halos technology. However, the halo removal technical route needs to use a low-exposure halo-free image as a reference image to extract the details of the halo area, which has a large image acquisition difficulty. And in the halo removal step, the local brightness value of the image is directly changed, and the contrast of the halo area in the output image is poor. The heterologous image fusion technical route mostly uses the idea of designing rules to fuse heterologous images to eliminate halos. The whole process has a large amount of calculation, cannot be applied to real-time traffic scenarios such as automotive anti-halos and monitoring traffic information collection, and has poor universality.

[0006] In view of this, a new method is needed to solve at least part of the above problems. Summary of the Invention

[0007] To overcome the deficiencies of the prior art, the present invention proposes a method for fusing heterologous images in a traffic scenario. First, image registration is performed on the heterologous images to make the spatio-temporal information of the source images consistent. Next, the source images are divided into grids, and the source images are cut into halo image blocks and background image blocks of different sizes. Then, the halo image blocks and background image blocks are respectively input into the halo network group and the background network group for image fusion. Finally, the fused image blocks are restored to the positions of the source images, and a fused image with anti-halos ability is output. This method has a simple and feasible process, fast processing speed, and strong robustness, and is applicable to various traffic scenarios containing halos.

[0008] Technical Solution: To solve the above technical problems, the present invention adopts the following technical solutions:

[0009] A method for fusing heterologous images in a traffic scenario, including:

[0010] S1: Obtain heterologous image data: Obtain source image data of the same time, same location, same shooting scene, and same shooting angle, including visible light images and infrared images, and the shooting scene is a common traffic scene;

[0011] S2: Register the heterologous image data: Taking the infrared source image as the reference image, use a registration algorithm based on the SURF operator and background robust association to complete the registration of the visible light and infrared source images, so that the spatio-temporal information of the visible light and infrared source images is consistent;

[0012] S3: Perform grid division on the source image: initially screen the veiling glare area in the image scene using pixel gradient weights, obtain the size information of the veiling glare area using the K-means clustering algorithm, and use a multi-resolution segmentation strategy to cut the source image into veiling glare image blocks and background image blocks of different sizes;

[0013] S4: Fuse the heterogeneous images: perform multi-resolution upsampling on the veiling glare image blocks and background image blocks to unify their sizes; use a veiling glare network group based on tensor addition and bilateral filtering to fuse visible light and infrared veiling glare image blocks; use a background network group based on U-Net and contour enhancement to fuse visible light and infrared background image blocks; perform multi-resolution downsampling on the fused visible light and infrared veiling glare image blocks, and visible light and infrared background image blocks to restore the size of the image blocks in the source image;

[0014] S5: Restore the image: restore the image blocks to their positions in the source image and output a fused image with anti-veiling glare ability.

[0015] Furthermore, for the heterogeneous image fusion method in the traffic scene of the present invention, the common traffic scenes in S1 include: road types at intersections, expressways, highways, and weaving areas under illumination conditions such as daytime, nighttime with streetlights, and nighttime without streetlights, and the shooting angles include aerial view, oblique view, and driver's view.

[0016] Furthermore, for the heterogeneous image fusion method in the traffic scene of the present invention, the specific steps for registering the heterogeneous image data in S2 include:

[0017] S2.1: Perform central symmetric scaling and cropping on the visible light image, and unify the pixel information of the visible light image and the infrared image without changing the field of view of the visible light image;

[0018] S2.2: Use the E-Net semantic segmentation algorithm to perform semantic segmentation on the source image, divide the road area and the background area, and use the SURF operator to find the feature corner points in the background area;

[0019] S2.3: Determine whether the feature corner point p i in the background area is in a dense corner point cluster. The judgment formula is as follows:

[0020]

[0021]

[0022] Where ||p i , p j ||2 represents the Euclidean distance between the feature point p i and p j , and Dis thD is the distance threshold for judging the corner density. th is the ratio threshold;

[0023] If the feature corner point p i is not in the dense corner point cluster, it is directly included in the candidate feature point set Cd;

[0024] If the feature corner point p i is in the dense corner point cluster, set the number of corner points in the dense corner point cluster to N0, and calculate the number of corner points N to be retained:

[0025]

[0026] At least 1 corner point and at most N - 1 corner points are retained in the dense corner point cluster;

[0027] After determining the number of corner points N, traverse all combinations of corner points with the number N in the corner point cluster, select the combination of corner points with the largest spacing as the corner points to be retained, and include them in the candidate feature point set Cd;

[0028] S2.4: For the candidate feature point set Cd, use the KNN algorithm to associate the feature point pairs in the visible light image and the infrared image, and represent the connection line between the associated matching point pairs as f(x n ) = kMx n + b, where M is an orthogonal 2*2 matrix with a determinant of 1, b is a 2*1 transformation vector, k is a proportionality coefficient, and minimize the following formula to find the optimal matching point pair connection line slope:

[0029]

[0030] where λ > 0 is the coordination proportionality coefficient, σ is the distribution standard deviation of the candidate feature point set, (x n , y n ) is the image coordinate of the feature point n;

[0031] Record the optimal connection line equation f(x n ) = kMx n + b as f(x n ) = Kx n + B, where K and B are the slope and constant of the straight line equation respectively. Select the point pairs whose connection line equation constants are within three standard deviations as the point pairs participating in registration, that is:

[0032] K - 3σ k ≤ K i ≤ K + 3σ k

[0033] B - 3σ b ≤ B i ≤ B + 3σ b

[0034] wherein, σ k is the distribution standard deviation of the proportionality coefficient k, and σ b is the distribution standard deviation of the transformation vector b, and K i , B i are constants of the point pairs participating in the registration;

[0035] S2.5: Obtain the set of point pairs participating in the registration, establish a perspective transformation matrix to generate a registration scheme, and generate the registered visible light image and infrared image.

[0036] Further, for the method for fusing heterogeneous images in a traffic scene of the present invention, the specific steps of grid-dividing the source image in S3 include:

[0037] S3.1: Calculate the pixel gradient weight value of the source image, extract the region with a large change in the gradient weight value as the candidate halo region, and frame it with a square bounding box;

[0038] S3.2: Count the bounding box sizes of the halo regions in the current image, perform clustering using the Kmeans algorithm, and obtain m types of bounding boxes in the current image, where the value of m is determined according to the traffic scene and the number of halo regions;

[0039] S3.3: Divide the source image into halo regions and background regions with fixed sizes by using a multi-resolution segmentation strategy.

[0040] Further, for the method for fusing heterogeneous images in a traffic scene of the present invention, the specific steps of the multi-resolution segmentation strategy in S3.3 include:

[0041] S3.3.1: Given the set of clustering sizes of the halo regions Lf{l, 2l, 2 2 l,..., 2 m l}, in the source image pair with a resolution of W*H, use CIoU to match the halo size corresponding to a certain halo region:

[0042]

[0043]

[0044]

[0045]

[0046] wherein, W is the number of pixels in the horizontal direction of the source image, H is the number of pixels in the vertical direction of the source image, A represents the bounding box of the halo region, B represents the bounding box of the clustering size with the same center point as A, b, b gtThey are the center points of A and B respectively, ρ is the Euclidean distance between the center points of A and B, and c is the diagonal distance of the minimum closed region of A and B;

[0047] 3.3.2: Update the size of the bounding box of the veiling glare area in the matched image. Denote the bounding box of a certain veiling glare area as l m as the size of the bounding box, and i as the number of the bounding box in the dataset with the same size;

[0048] 3.3.3: Traverse each bounding box of the veiling glare area and divide the background area of the row and column of the grid where it is located:

[0049]

[0050] Among them, dr represents the up, down, left, and right directions of the veiling glare bounding box, N dr represents the number of image patches after the background area in the corresponding direction is divided, and K ∈ {H, W} represents the boundary line of the wide side and the high side of the image;

[0051] 3.3.4: Calculate the boundary filling ratio r Padding :

[0052]

[0053] 3.3.5: Evaluate the boundary filling ratio. If a certain boundary filling ratio exceeds the filling threshold, that is, r Padding ≥r th , then modify the position, and return to 3.3.2 to recalculate the segmentation result until r Padding meets the requirements;

[0054] 3.3.6: Merge the background image patches to minimize the total number of segmented images and output the final segmentation result.

[0055] Furthermore, in the method for fusing heterogeneous images in the traffic scenario of the present invention, the veiling glare network group in S4 includes a generator network and a discriminator network. Among them, the generator network of the veiling glare network group is constructed by tensor addition, the discriminator network of the veiling glare network group is constructed by the PatchGAN structure, and the loss function of the veiling glare network group is designed through bilateral filtering weights;

[0056] Among them, the calculation formula of the bilateral filtering weight is as follows:

[0057]

[0058]

[0059]

[0060] Among them, p1 represents the weight of the visible light feature map, and p2 represents the weight of the infrared feature map. is the pixel weight coefficient, n is a certain pixel point in the feature map, the average of p1 and p2 is the pixel weight coefficient, k ∈ {i, v} represents that the parameter source is infrared or visible light, and C is a real constant. is the feature map after bilateral filtering.

[0061] Further, for the method for fusing heterogeneous images in a traffic scenario of the present invention, the loss function of the generator is as follows:

[0062]

[0063] Among them, represents the adversarial loss of the generator, which is given by the discriminator in the adversarial network, and L con is the content loss, and λ is the weight coefficient for adjusting the importance of the two losses.

[0064] Adversarial loss takes the cross-entropy of the discriminator's discrimination result for the pseudo-fused image, that is:

[0065]

[0066] Among them, D v (G(v,i)) represents the true discrimination probability of the discriminator D v for the pseudo-fused image G(v,i), and D i (G(v,i)) represents the true discrimination probability of the discriminator D i for the pseudo-fused image G(v,i);

[0067] Content loss L con is defined as follows:

[0068]

[0069]

[0070] Among them, is the SSIM loss of pixel p, which is expressed as:

[0071]

[0072] Among them, represents the expected image variance, represents the covariance between the expected image and the pseudo-fused image;

[0073] The loss function of the discriminator is as follows:

[0074]

[0075]

[0076] Among them, p1 and p2 are weight coefficients reflecting the information amount of the source image. p1 represents the visible light image, and p2 represents the infrared image, which are calculated by the bilateral filtering weight function.

[0077] Furthermore, for the method for fusing heterogeneous images in the traffic scenario of the present invention, the background network group in S4 includes a generator network and a discriminator network. Among them, a 7-layer U-Net is used to construct the generator network of the background network group, and a PatchGAN structure is used to construct the discriminator network of the background network group. The loss function of the background network group is designed through contour enhancement.

[0078] Furthermore, for the method for fusing heterogeneous images in the traffic scenario of the present invention, the loss function of the generator of the background network group is as follows:

[0079]

[0080] Among them, L SSIM is the structure loss, which is calculated by the structural similarity index SSIM. L L1 is the edge loss, which is calculated by the L1 norm. δ1 and δ2 are weight parameters for measuring the importance of the structure loss and the edge loss in the target network;

[0081] The calculation method of the structure loss L SSIM is as follows:

[0082]

[0083] Among them, is the SSIM loss of pixel p, expressed as:

[0084]

[0085] Among them represents the variance of the expected image, represents the covariance between the expected image and the pseudo-fused image;

[0086] The calculation method of the edge loss L L1 is as follows:

[0087]

[0088]

[0089] Among them, the weight parameter w(I k ) is the same as the weight function in the SSIM loss.

[0090] When the present invention adopts the above technical solutions compared with the prior art, the following technical effects are achieved:

[0091] 1. The method for fusing heterogeneous images in a traffic scenario of the present invention applies a dual discriminator generative adversarial network to image fusion to handle the problem of veiling glare in traffic scenarios, improving the problems of difficult data acquisition, slow calculation speed, and narrow application range in existing anti-veiling glare algorithms. The algorithm of the present invention is easy to implement and fast, and is applicable to various traffic scenarios with veiling glare, providing a high-quality data source for traffic scenario applications such as monitoring data acquisition and anti-veiling glare safe driving of vehicles.

[0092] 2. The method for fusing heterogeneous images in a traffic scenario of the present invention uses multi-resolution upsampling to unify the resolution of input image blocks before image fusion and multi-resolution downsampling to restore the image blocks to the size of the source image after image fusion, enabling the same network to process image inputs of different resolutions, and adjusting the regions of interest in the image during upsampling and downsampling through an adaptive network, broadening the input and output conditions of the image fusion process and providing conditions for the independent training of the veiling glare region and the background region.

[0093] 3. The method for fusing heterogeneous images in a traffic scenario of the present invention uses a veiling glare network group and a background network group to independently train veiling glare region image blocks and background region image blocks. In the veiling glare network group, a tensor addition network structure is used to eliminate the influence of veiling glare and restore the target details in the region. In the background network group, a U-Net structure is used to fuse the structural contour information in the heterogeneous image blocks, increasing the target contour details in the fused image and improving the visual perception of the fused image. At the same time, the method of independent training in different regions provides a new idea for the anti-veiling glare algorithm of image fusion and effectively improves the anti-veiling glare ability and global fusion effect of the fused image. BRIEF DESCRIPTION OF THE DRAWINGS

[0094] Figure 1 is a flowchart of the method for fusing heterogeneous images in a traffic scenario of the present invention;

[0095] Figure 2 is a schematic diagram of the generator network structure of the veiling glare network group of the method for fusing heterogeneous images in a traffic scenario of the present invention;

[0096] Figure 3 is a schematic diagram of the generator network structure of the background network group of the method for fusing heterogeneous images in a traffic scenario of the present invention;

[0097] Figure 4 is a schematic diagram of the dual discriminator network structure of the method for fusing heterogeneous images in a traffic scenario of the present invention;

[0098] Figure 5 is a diagram showing the effect of Example 1 of the method for fusing heterogeneous images in a traffic scenario of the present invention;

[0099] Figure 6 is a diagram showing the effect of Example 2 of the method for fusing heterogeneous images in a traffic scenario of the present invention. Detailed Implementation Modes

[0100] To further understand the present invention, the preferred implementation modes of the present invention will be described below in conjunction with embodiments. However, it should be understood that these descriptions are only for further explaining the features and advantages of the present invention, rather than limiting the claims of the present invention.

[0101] The description of this part only targets typical embodiments, and the present invention is not limited to the scope described in the embodiments. Combinations of different embodiments, mutual replacement of some technical features in different embodiments, and mutual replacement of the same or similar prior art means and some technical features in the embodiments are also within the scope of description and protection of the present invention.

[0102] A flowchart of a method for fusing heterogeneous images in a traffic scenario is as Figure 1 shown, and the specific steps are as follows:

[0103] Step 1: Obtain heterogeneous image data. In this example, two groups of visible light images and infrared images in the FLIR dataset are used for demonstration. Both groups of samples are from the driver's perspective, the shooting time is at night, and there is obvious veiling glare caused by the headlights of the vehicle in front in both groups of samples. The source images and the fusion effect are as Figure 5 and Figure 6 shown.

[0104] Step 2: Perform registration operations on the heterogeneous images. Taking the infrared image as the reference image, a registration algorithm based on the SURF (Speeded Up Robust Features) operator and background robust association is designed to complete the registration of the visible light and infrared source images, so that the spatio-temporal information of the source images is consistent. The specific steps are as follows:

[0105] 2.1, Perform central symmetric scaling and cropping on the visible light image to unify the pixel information of the visible light and infrared images without changing the field of view of the visible light image;

[0106] 2.2, Use E-Net (semantic segmentation algorithm) to perform semantic segmentation on the source images, divide the road area and the background area, and use the SURF operator to find the feature corner points in the background area;

[0107] 2.3, For the feature corner point p i in the background area, first use the following formula to determine whether p i is in a dense corner point cluster:

[0108]

[0109]

[0110] where ||p i ,p j||2 is the feature point p i and p j 's Euclidean distance, Dis th is the distance threshold D for judging the density of corner points th is the ratio threshold;

[0111] If p i is not in the dense corner point cluster, it is directly included in the candidate feature point set Cd. If p i is in the dense corner point cluster, let the number of corner points in the corner point cluster be N0, and calculate the number of corner points N to be retained:

[0112]

[0113] That is, at least 1 corner point and at most N - 1 corner points are retained in the dense corner point cluster;

[0114] After determining the number of corner points N, traverse all combinations of N corner points in the corner point cluster, select the combination with the largest spacing as the corner points to be retained, and include them in the candidate feature point set Cd;

[0115] 2.4. For the candidate feature point set Cd, use the KNN algorithm (K-Nearest Neighbor Classification Algorithm) to associate the feature point pairs in the visible light and infrared source images, and represent the connection line between the matching point pairs as f(x n ) = kMx n + b, where M is an orthogonal 2*2 matrix with a determinant of 1, t is a 2*1 transformation vector, k is a proportionality coefficient, and minimize the following formula to find the slope of the correct matching point pair connection line:

[0116]

[0117] where λ > 0 is the coordination proportionality coefficient, σ is the distribution standard deviation of the candidate feature point set, and (x n , y n ) is the image coordinate of the feature point n;

[0118] Record the optimal connection line equation f(x n ) = kMx n + b as f(x n ) = Kx n + B, where K and B are the slope and constant of the straight line equation, and select the point pairs with the connection line equation constant within the following range as the point pairs participating in registration:

[0119] K - 3σ k ≤ K i ≤ K + 3σ k

[0120] B - 3σ b ≤ B i ≤ B + 3σb

[0121] 2.5. Obtain the set of point pairs participating in registration, establish a perspective transformation matrix to generate a registration scheme, and generate the registered visible light and infrared source images.

[0122] Step 3: Divide the source image into grids, initially screen the veiling glare regions in the image scene using pixel gradient weights, cluster the size information of the veiling glare regions using the KMeans algorithm, and design a multi-resolution segmentation strategy to cut the source image into veiling glare image blocks and background image blocks of different sizes. The specific steps are as follows:

[0123] 3.1. Calculate the pixel gradient weights of the source image, extract the regions with large changes in gradient weights as candidate veiling glare regions, and frame the regions with square bounding boxes;

[0124] 3.2. Statistically analyze the bounding box sizes of the veiling glare regions in the current image, perform clustering using the Kmeans algorithm to obtain m types of bounding boxes in the current image, and the value of m depends on the traffic scene and the number of veiling glare regions;

[0125] 3.3. Design a multi-resolution segmentation strategy to divide the source image into fixed-size veiling glare regions and background regions. The specific steps are as follows:

[0126] 3.3.1. Given the set of veiling glare region clustering sizes Lf{l, 2l, 2 2 l,…, 2 m l}, in the source image pair with a resolution of W*H, use CIoU to match the veiling glare size corresponding to a certain veiling glare region:

[0127]

[0128]

[0129]

[0130]

[0131] A and B represent the circumscribed bounding box of the veiling glare region and the circumscribed bounding box of the clustering size with the same center point as A, b and b gt are the center points of A and B, and ρ is the Euclidean distance between the center points of A and B;

[0132] 3.3.2. Update the size of the circumscribed bounding box of the veiling glare region in the matched image. Denote the circumscribed bounding box of a certain veiling glare region as l m as the bounding box size, and i as the number of the bounding box in the dataset of the same size;

[0133] 3.3.3. Traverse each circumscribed bounding box of the veiling glare region Divide the background region of the grid rows and columns where it is located:

[0134]

[0135] where dr represents the up, down, left, and right directions of the flare bounding box, N dr represents the number of image patches after the background region in the corresponding direction is divided, and K ∈ {H, W} represents the boundary lines of the wide side and high side of the image;

[0136] 3.3.4, Calculate the boundary filling ratio:

[0137]

[0138] 3.3.5, Evaluate the boundary filling ratio. If the boundary filling ratio of a certain side exceeds the filling threshold, that is, r Padding ≥ r th , then stepwise modify the position, return to 3.3.2 to recalculate the segmentation result until r Padding meets the requirements;

[0139] 3.3.6, Merge the background image patches to minimize the total number of segmented images and output the final segmentation result.

[0140] Step 4: Implement heterologous image fusion. First, perform multi-resolution upsampling on the two types of image patches to unify the sizes of the flare image patches and the background image patches. Next, design a flare network group based on tensor addition and bilateral filtering to fuse the visible light and infrared flare image patches and improve the anti-flare ability of the fused image. At the same time, design a background network group based on U-Net (U-shaped network) and contour enhancement to fuse the visible light and infrared background image patches and improve the overall visual perception and texture quality of the fused image. Finally, perform multi-resolution downsampling on the image patches to restore the sizes of the image patches in the source images.

[0141] Among them, the flare network group has the following characteristics:

[0142] 1) The generator of the flare network group uses a network structure of tensor addition to generate the fused picture. The schematic diagram of the generator network structure is as Figure 2 shown, and the network structure parameter table is shown in Table 1. Specifically, for the input visible light and infrared training samples, the network first extracts the low-level features such as edges and corners and high-level features such as semantics of the samples through two convolutional layers (Conv1, Conv21 and Conv2, Conv22) respectively. Then, the proposed dual-source sample feature maps F21 and F22 are weighted and added to obtain the fused feature map F3, where the weights p1 and p2 are calculated through bilateral filtering weights. Finally, the fused image is reconstructed from F3 through three convolutional layers (Conv3, Conv4, Conv5).

[0143] Table 1 Parameters of the Glow Network Group Generator Network Structure

[0144]

[0145] 2) The weights of the dual-source convolutional layers before feature fusion in the glow network group generator are exactly the same, that is, "bound". The weight parameters of Conv1 and Conv2 are exactly the same. When adjusting the weight parameters at the end of each training, the changes in the weight parameters of Conv1 and Conv2 are also the same, and the same applies to Conv21 and Conv22. The advantages of this design are as follows: The network can learn exactly the same features from dual-source samples, enabling the weighted summation operation during fusion to have a theoretical basis; due to the additivity of features, the network tends to learn invariant features, which in this network are features insensitive to light, so that the network can handle overexposed and underexposed situations and generate fused images with rich details and resistance to adverse lighting conditions; binding weights reduces the number of parameters in the network, which can significantly improve the convergence speed of the network.

[0146] 3) The glow network group uses a bilateral filter to calculate the weights of the dual-source feature maps:

[0147]

[0148]

[0149]

[0150] where p1 represents the weight of the visible light feature map, and p2 represents the weight of the infrared feature map. is the pixel weight coefficient, n is a certain pixel point in the feature map, the feature map weights p1 and p2 are the average of the pixel weight coefficients, k ∈ {i, v} represents the parameter source as infrared or visible light, C is a real constant, and taking 70 has the best effect in this article through experiments. is the feature map after bilateral filtering and can be calculated by the following formula:

[0151]

[0152]

[0153]

[0154] where I BF (x, y) is the image after bilateral filtering, and (x′, y′) is the neighboring pixel of the pixel point (x, y). and are the Gaussian kernels in the geometric space and pixel intensity respectively.

[0155] 4) The loss function of the glow network group generator is expressed as follows:

[0156]

[0157] Denotes the generator adversarial loss, given by the discriminator in the adversarial network, \(L\) con is the content loss, and \(\lambda\) is the weight coefficient for adjusting the importance of the two losses.

[0158] Adversarial loss In this paper, the adversarial loss takes the cross-entropy of the discriminator's discrimination result for the pseudo-fusion image, that is:

[0159]

[0160] where \(D\) v (\(G(v, i)\)) represents the probability of the discriminator \(D\) v judging the pseudo-fusion image \(G(v, i)\) as real, and \(D\) i (\(G(v, i)\)) represents the probability of the discriminator \(D\) i judging the pseudo-fusion image \(G(v, i)\) as real.

[0161] The content loss is defined as follows:

[0162]

[0163]

[0164] where is the SSIM loss at pixel \(p\), which can be expressed as:

[0165]

[0166] where represents the variance of the expected image, represents the covariance between the expected image and the pseudo-fusion image;

[0167] 5) The loss function of the flare network group discriminator is as follows:

[0168]

[0169]

[0170] where \(p1\) and \(p2\) are weight coefficients reflecting the information content of the source images. \(p1\) represents the visible light image, and \(p2\) represents the infrared image, which are calculated by the bilateral filtering weight function.

[0171] Among them, the background network group has the following characteristics:

[0172] 1) In the generator of the background network group, a U-shaped network (U-Net) is used to generate pseudo-fusion images. The network structure is shown as Figure 3As shown in the figure, the network parameter table is shown in Table 2. Compared with the standard U-Net, the generator in this paper improves the network depth. For generating fused pictures, more precise expression ability is required, so the network depth is increased to 7 layers; the sampling method is changed, and a convolutional layer with a stride of 2 is directly used to reduce the resolution of the feature map, and the convolutional kernel is set to 4*4, and the padding is 1 to ensure that the size of the feature map and the input and output images are the same, avoiding information loss caused by additional operations such as cropping and mirroring; the activation function is changed to Leaky-Relu, which alleviates the problem that some neurons are in a silent state in the later stage of training to a certain extent.

[0173] Table 2 Network Structure Parameter Table of the Background Network Group Generator

[0174]

[0175]

[0176] 2) The discriminator network of the background network group is designed using the PatchGAN structure. The structure of the discriminator network is as Figure 4 shown. The network parameter table is shown in Table 3. For a single-channel input image, the network first uses three convolutional layers with a stride of 2 for downsampling to obtain a feature map after dimensionality reduction; then, zero-padding (padding) with a size of 2 is performed on the edges of the feature map to increase the attention to the edges of the feature map and adapt to the size of the network feature map; next, a convolutional layer with a stride of 1 is used to accurately extract features, and the feature values are normalized through BatchNorm (BN)+LeakyReLU to improve the network calculation speed and convergence; finally, after padding the feature map again, a convolutional layer with a convolutional kernel number of 1 is used to merge the feature map dimensions, and a score matrix of n*n is directly output through the tanh activation function. The average of the score matrices obtained by the discriminator network is taken to obtain the final true judgment probability of the discriminator.

[0177] Table 3 Network Structure Parameter Table of the Dual Discriminator

[0178]

[0179] 3) The loss function of the background network group generator is expressed as follows:

[0180]

[0181] Among them, L SSIM is the structural loss calculated by the structural similarity index SSIM, and L L1 is the edge loss calculated by the L1 norm. δ1 and δ2 are weight parameters that measure the importance of the structural loss and the edge loss in the target network;

[0182] L SSIM is calculated in the same way as in the halation network group. LL1 The calculation method is as follows:

[0183]

[0184]

[0185] Among them, the weight parameter w(I k ) is the same as the weight function in the SSIM loss.

[0186] Step 5: Restore the image, restore the image patches to the positions in the source image, and output the fused image with anti-halos ability.

[0187] The description and application of the present invention here are illustrative and not intended to limit the scope of the present invention to the above embodiments. The related descriptions of effects or advantages in the specification may not be reflected in actual experimental examples due to uncertainties in specific condition parameters or other factors, and the related descriptions of effects or advantages are not used to limit the scope of the invention. Modifications and changes to the disclosed embodiments here are possible, and substitutions and equivalents of various components are known to those of ordinary skill in the art. Those skilled in the art should clearly understand that the present invention can be implemented in other forms, structures, arrangements, proportions, and with other components, materials, and parts without departing from the spirit or essential characteristics of the present invention. Other modifications and changes can be made to the disclosed embodiments here without departing from the scope and spirit of the present invention.

Claims

1. A method for fusing heterogeneous images in a traffic scenario, characterized in that Including: S1: Obtain heterologous image data: Obtain source image data at the same time, in the same location, with the same shooting scene and the same shooting angle, including visible light images and infrared images, and the shooting scene is a common traffic scene; S2: Register the heterologous image data: Use the infrared source image as the reference image, and complete the registration of the visible light and infrared source images by using a registration algorithm based on the SURF operator and background robust association, so that the spatio-temporal information of the visible light and infrared source images is consistent; S3: Divide the source image into grids: Use the pixel gradient weight to initially screen the halation area in the image scene, use the K-means clustering algorithm to obtain the size information of the halation area, and use the multi-resolution segmentation strategy to cut the source image into halation image blocks and background image blocks of different sizes; S4: Fuse the heterologous images: Perform multi-resolution upsampling on the halation image blocks and background image blocks to unify the sizes of the halation image blocks and background image blocks; Use the halation network group to fuse the visible light and infrared halation image blocks. The halation network group includes a generator network and a discriminator network. Among them, use tensor addition and bilateral filtering to construct the generator network of the halation network group, and use the PatchGAN structure to construct the double discriminator network of the halation network group; Use the background network group to fuse the visible light and infrared background image blocks. The background network group includes a generator network and a discriminator network. Among them, use a 7-layer U-Net to construct the generator network of the background network group, and use the PatchGAN structure to construct the double discriminator network of the background network group; Perform multi-resolution downsampling on the fused visible light and infrared halation image blocks, and visible light and infrared background image blocks to restore the size of the image blocks in the source image; The loss function of the double discriminator is as follows: Among them, p1 and p2 are weight coefficients reflecting the information content of the source image. p1 represents the visible light image, and p2 represents the infrared image, which are calculated by the bilateral filtering weight function; S5: Restore the image: Restore the image blocks to the positions where the source images are located, and output a fused image with anti-halation ability.

2. The method for fusing heterogeneous images in a traffic scenario according to claim 1, wherein The common traffic scenes in S1 include: intersection roads, expressways, highways, and weaving area roads during daytime, at night with street lights, and at night without street light illumination conditions. The shooting angles include aerial perspective, oblique perspective, and driver's perspective.

3. The method for fusing heterogeneous images in a traffic scenario according to claim 1, wherein The specific steps for registering the heterologous image data in S2 include: S2.1: Perform central symmetric scaling and cropping on the visible light image, and unify the pixel information of the visible light image and the infrared image without changing the field of view of the visible light image; S2.2: Use the E-Net semantic segmentation algorithm to perform semantic segmentation on the source image, divide the road area and the background area, and use the SURF operator to find the feature corner points in the background area; S2.3: Determine whether the feature corner point p in the background area i is in a dense corner point cluster. The judgment formula is as follows: Among them, ||p i , p j ||2 represents the Euclidean distance between the feature point p i and p j . Dis th is the distance threshold for judging the density of corner points, and D th is the ratio threshold; If the characteristic corner point p i is not in the dense corner point cluster, it is directly included in the candidate feature point set Cd; If the feature corner point p i is in a dense corner point cluster, set the number of corner points in the dense corner point cluster as N0, and calculate the number of corner points N to be retained: At least 1 corner point and at most N - 1 corner points are reserved in the dense corner point cluster; After determining the number of corner points N, traverse all the corner point combinations with the number of N in the corner point cluster, select the corner point combination with the largest spacing as the corner points to be reserved, and record them in the candidate feature point set Cd; S2.4: For the candidate feature point set Cd, the KNN algorithm is used to associate the feature point pairs in the visible light image and the infrared image, and the connection line between the associated matching point pairs is represented as f(x n ) = kMx n + b, where M is an orthogonal 2*2 matrix with a determinant of 1, b is a 2*1 transformation vector, k is a proportionality coefficient, and the optimal matching point pair connection line slope is found by minimizing the following formula: where λ > 0 is the coordination ratio coefficient, σ is the distribution standard deviation of the candidate feature point set, and (x n , y n ) are the image coordinates of feature point n; Denote the optimal connection equation \(f(x n ) = kMx n +b\) as \(f(x n ) = Kx n +B\), where \(K\) and \(B\) are the slope and constant of the straight-line equation respectively. Select the point pairs within three times the standard deviation of the connection equation constant as the point pairs participating in registration, that is: K - 3σ k ≤K i ≤K + 3σ k B - 3σ b ≤ B i ≤ B + 3σ b Among them, σ k is the distribution standard deviation of the proportionality coefficient k, and σ b is the distribution standard deviation of the transformation vector b. K i , B i are constants for the point pairs participating in registration; S2.5: Obtain the point pair set participating in the registration, establish a perspective transformation matrix to generate a registration scheme, and generate the registered visible light image and infrared image.

4. A method for fusing heterogeneous images in a traffic scenario according to claim 1, characterized in that The specific steps for grid division of the source image in S3 are as follows: S3.1: Calculate the pixel gradient weights of the source image, extract the regions with large changes in gradient weights as candidate halation regions, and frame them with square bounding boxes; S3.2: Count the sizes of the bounding boxes of the halation regions in the current image, perform clustering using the Kmeans algorithm to obtain m sizes of bounding boxes in the current image, where the value of m is determined according to the traffic scene and the number of halation regions; S3.3: Use a multi-resolution segmentation strategy to divide the source image into halation regions and background regions with fixed sizes.

5. A method for fusing heterogeneous images in a traffic scenario according to claim 1 or 4, characterized in that The specific steps of the multi-resolution segmentation strategy in S3.3 are as follows: S3.3.1: The clustering size set Lf{l, 2l, 2 2 l, …, 2 m l} of the known veiling glare area. In the source image pair with a resolution of W*H, CIoU is used to match the veiling glare size corresponding to a certain veiling glare area: Among them, W is the number of pixels in the horizontal direction of the source image, H is the number of pixels in the vertical direction of the source image, A represents the bounding box of the veiling glare area, B represents the bounding box of the clustering size with the same center point as A, b and b gt are the center points of A and B respectively, ρ is the Euclidean distance between the center points of A and B, and c is the diagonal distance of the minimum closed region of A and B; 3.3.2: Update the size of the bounding box of the veiling glare area within the matched image. Denote the bounding box of a certain veiling glare area as l m as the size of the bounding box, and i as the number of the bounding box in the dataset of the same size; 3.3.3: Traverse the bounding box of each flare region Divide the background region of the grid rows and columns where it is located: Among them, dr represents the up, down, left, and right directions of the flare bounding box, and N dr represents the number of image patches after the background area in the corresponding direction is segmented, and K ∈ {H, W} represents the boundary line between the wide side and the high side of the image; 3.3.4: Calculate the boundary filling ratio r Padding : 3.3.5: Evaluate the boundary filling ratio. If the filling ratio of a certain boundary exceeds the filling threshold, that is, r Padding ≥r th , then modify the position, return to 3.3.2 to recalculate the segmentation result until r Padding meets the requirements; 3.3.6: Merge the background image blocks to minimize the total number of images after segmentation and output the final segmentation result.

6. A method for fusing heterogeneous images in a traffic scenario according to claim 1, characterized in that, The calculation formula for the bilateral filtering weights of the halation network group is as follows: Among them, p1 represents the weight of the visible light feature map, and p2 represents the weight of the infrared feature map. is the pixel weight coefficient, n is a certain pixel point in the feature map, the average of p1 and p2 is the pixel weight coefficient, k ∈ {i, v} represents that the parameter source is infrared or visible light, and C is a real constant. is the feature map after bilateral filtering.

7. A method for fusing heterogeneous images in a traffic scenario according to claim 1, characterized in that, The loss function of the halation network group generator is as follows: Among them, represents the generator adversarial loss, which is given by the discriminator in the adversarial network, and L con is the content loss, and λ is the weight coefficient for adjusting the importance of the two losses; Adversarial loss Take the cross-entropy of the discriminator's discrimination result for the pseudo-fused image, that is: Among them, D v (G(v,i)) represents the discriminator D v 's probability of judging the authenticity of the pseudo-fusion image G(v,i), and D i (G(v,i)) represents the discriminator D i 's probability of judging the authenticity of the pseudo-fusion image G(v,i); Content loss L con is defined as follows: Among them, is the SSIM loss of pixel p, expressed as: Among them, represents the expected image variance, represents the covariance between the expected image and the pseudo-fused image.

8. A method for fusing heterogeneous images in a traffic scenario according to claim 1, characterized in that, The loss function of the background network group generator is as follows: Among them, L SSIM is the structural loss, which is calculated by the structural similarity index SSIM. L L1 is the edge loss, which is calculated by the L1 norm. δ1 and δ2 are weight parameters that measure the importance of the structural loss and the edge loss in the target network; Structural loss L SSIM The calculation method is as follows: Among them, is the SSIM loss of pixel p, expressed as: wherein represents the expected image variance, represents the covariance between the expected image and the pseudo-fused image; Edge loss L L1 The calculation method is as follows: Among them, the weight parameter w(I k ) is the same as the weight function in the SSIM loss.

Citation Information

Patent Citations

  • A multi-exposure image fusion method for halo removal

    CN109035155B

  • Automobile anti-halation method based on improved NSCT

    CN113052779A

  • A method for fusion of infrared and visible light images

    CN113628151B

  • Infrared and visible light image fusion system and method

    CN114187214A