Traffic Detection Method and System Based on Visible Light and Infrared Image Fusion

By using adaptive filtering and multi-scale transformation registration technology in image fusion, combined with gradient weighted fusion and average fusion, the problems of poor image fusion effect and insufficient registration accuracy in the prior art are solved, and the accuracy and reliability of traffic detection are significantly improved.

CN119359568BActive Publication Date: 2025-06-10SHANDONG UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202411907005.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-24
Publication Date
2025-06-10
Estimated Expiration
2044-12-24

AI Technical Summary

Technical Problem

The prior art lacks adaptive adjustment in image fusion, resulting in the fused image being unable to accurately reflect the key information in the original image, affecting the target detection performance. At the same time, traditional feature matching algorithms are difficult to achieve high-precision registration when dealing with lighting changes, viewing angle differences and motion blur, and Gaussian filtering may lead to loss of image details during denoising.

Method used

Adaptive filtering technology is used to dynamically adjust the smoothing intensity, retaining the edge details of the image and removing noise. Use a combination rule of gradient-weighted fusion and average fusion, combining multi-scale transformation and registration, to improve registration accuracy and fusion effect.

Benefits of technology

It significantly improves the accuracy and reliability of traffic inspection, enhances the applicability of the system in complex environments, and provides a higher quality foundation for subsequent target inspections.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119359568B_ABST
    Figure CN119359568B_ABST
Patent Text Reader

Abstract

The present invention proposes a traffic detection method and system based on visible light and infrared image fusion, belonging to the field of image processing. It includes acquiring the original visible light image and infrared image within the same scene and performing preprocessing; based on adaptive filtering, constructing Laplacian pyramids for the preprocessed visible light image and infrared image respectively, and performing multi-scale fine registration on the two Laplacian pyramids; performing gradient weighted fusion on the high-level images of the two Laplacian pyramids to obtain a high-level fusion image, and performing dual-domain average fusion on the low-level images of the two Laplacian pyramids to obtain a low-level fusion image; performing layer-by-layer reconstruction based on the high-level fusion image and the low-level fusion image to obtain a reconstructed fusion image; performing traffic detection based on the reconstructed fusion image. The present invention uses adaptive filtering to retain edge details, and combines the advantages of visible light and infrared images through different weighted fusions, significantly improving the target recognition accuracy and robustness in traffic detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and particularly to a traffic detection method and system based on visible light and infrared image fusion. Background Art

[0002] With the acceleration of the urbanization process, traffic accidents occur frequently, especially at night and under adverse weather conditions. The fusion of visible light images and infrared dual-band images makes up for the deficiencies of a single sensor. The combination of the two combines the rich details of visible light and the stability of infrared images to achieve information complementarity, and can generate stable and information-rich fused images, thereby enhancing the system's perception of scene information and target recognition ability in complex environments such as at night, adverse weather, and poor lighting conditions.

[0003] However, the prior art rarely pays attention to the fusion effect of different modality images. For example, in the Chinese invention application with the application number 202311218824.1 and the name of "Traffic Event Detection Method and Device Based on Infrared Images and Visible Light Images", only common existing models are directly used for image fusion without making adaptive adjustments, resulting in the fused image being unable to accurately reflect the key information in the original image, thereby causing a decline in target detection performance. At the same time, the existing fusion technologies rely on traditional feature matching algorithms and rarely consider the combination of multi-scale transformation and registration. These algorithms often cannot achieve high-precision registration when dealing with lighting changes, perspective differences, and motion blur. Especially when dealing with image features of different resolutions and scales, features of different scales often have significant differences in importance and performance. Failing to make full use of this information will further affect the fusion quality. In addition, the current technology usually uses Gaussian filtering for image smoothing. This method may cause the loss of image details while eliminating noise. In the case of more noise or poor image quality, using Gaussian filtering will make the finally fused image appear blurred in detail presentation, thereby reducing the accuracy of subsequent processing (such as target detection). Summary of the Invention

[0004] To solve the above problems, the present invention proposes a traffic detection method and system based on visible light and infrared image fusion. By using adaptive filtering, it can dynamically adjust the smoothing intensity according to the local characteristics of the image, effectively retain the edge details of the image while removing noise, and provide a high-quality basis for subsequent image fusion. By adopting different weighted fusion strategies, that is, gradient weighted fusion for high-level images to enhance detail information, and average fusion for low-level images to maintain overall brightness and smoothness, this processing method can combine the advantages of visible light images and infrared images, improve the clarity and contrast of the reconstructed fusion image, and thus significantly enhance the accuracy and reliability of traffic detection. At the same time, processing and alignment are carried out at multiple scales to further improve the registration accuracy, optimize the fusion effect, enhance the applicability of the system in complex scenarios, and provide a higher-quality basis for subsequent object detection.

[0005] To achieve the above object, the present invention adopts the following technical solutions:

[0006] In the first aspect, the present invention provides a traffic detection method based on visible light and infrared image fusion, including:

[0007] Obtain the original visible light image and infrared image within the same scene and perform preprocessing;

[0008] Based on adaptive filtering, respectively construct Laplacian pyramids for the preprocessed visible light image and infrared image, and perform multi-scale fine registration on the visible light Laplacian pyramid and the infrared Laplacian pyramid;

[0009] Use gradient weighted fusion for the high-level images of the visible light Laplacian pyramid and the high-level images in the infrared Laplacian pyramid to obtain high-level fusion images, and use dual-domain average fusion for the low-level images of the visible light Laplacian pyramid and the low-level images in the infrared Laplacian pyramid to obtain low-level fusion images; perform layer-by-layer reconstruction based on the high-level fusion images and the low-level fusion images to obtain the reconstructed fusion image;

[0010] Perform traffic detection based on the reconstructed fusion image.

[0011] Preferably, the preprocessing includes:

[0012] After initially preprocessing the infrared image by two-point correction, blind pixel compensation, and median filtering, use platform histogram equalization to enhance the contrast, perform pixel quantization using maximum-minimum linear mapping, and then unify the image size through bilinear interpolation algorithm;

[0013] Perform scaling processing on the visible light image to ensure that the sizes of the visible light image and the infrared image match.

[0014] Preferably, based on adaptive filtering, Laplacian pyramids are respectively constructed for the preprocessed visible light image and infrared image, and the specific steps include:

[0015] Using the adaptive mean filtering method, smooth the preprocessed visible light image and infrared image;

[0016] Repeat downsampling for the filtered images to respectively construct a visible light Gaussian pyramid and an infrared Gaussian pyramid;

[0017] Perform layer-by-layer difference on the two Gaussian pyramids respectively to obtain a visible light Laplacian pyramid and an infrared Laplacian pyramid.

[0018] Preferably, the adaptive mean filtering method is specifically: dynamically adjust the smoothing intensity according to the local characteristics of the image, perform strong smoothing in the flat area of the image, and weaken the smoothing intensity in the edge area.

[0019] Preferably, the multi-scale fine registration of the visible light Laplacian pyramid and the infrared Laplacian pyramid specifically includes:

[0020] Obtain the images in the two Laplacian pyramids, and use the Sobel operator to calculate the gradients of the images in the two Laplacian pyramids respectively;

[0021] Divide the image into multiple small units, and calculate the gradient direction histogram in each unit;

[0022] Normalize the gradient direction histogram in each small unit, respectively reconstruct the normalized histograms into a visible light reconstructed image and an infrared reconstructed image, and calculate the HOG features of the visible light reconstructed image and the infrared reconstructed image respectively;

[0023] Calculate the distance between the feature points of the visible light reconstructed image and the infrared reconstructed image, and select the feature points with the smallest distance as matching pairs for feature matching, and transform the infrared image according to the estimated transformation matrix to align it with the visible light image.

[0024] Preferably, the low-layer images of the visible light Laplacian pyramid and the low-layer images of the infrared Laplacian pyramid are fused using dual-domain averaging to obtain a low-layer fused image, where the dual-domain averaging includes global average fusion and local average fusion, and specifically includes:

[0025] Fuse the low-layer images of the visible light Laplacian pyramid and the low-layer images of the infrared Laplacian pyramid using global average fusion to obtain a preliminary low-layer fused image;

[0026] Based on a preset window, divide the local areas of the two low-layer images, traverse the local areas for local average fusion, and obtain a local low-layer fused image;

[0027] Perform weighted summation on the preliminary low-level fusion image and the local low-level fusion image to obtain the final low-level fusion image.

[0028] Preferably, it further includes performing pseudo-color processing on the reconstructed fusion image by using the color transformation method.

[0029] In a second aspect, the present invention provides a traffic detection system based on visible light and infrared image fusion, including:

[0030] An image acquisition module, configured to acquire the original visible light image and the infrared image within the same scene and perform preprocessing;

[0031] A registration module, configured to respectively construct Laplacian pyramids for the preprocessed visible light image and the infrared image based on adaptive filtering, and perform multi-scale fine registration on the visible light Laplacian pyramid and the infrared Laplacian pyramid;

[0032] A fusion and reconstruction module, configured to use gradient weighted fusion to obtain a high-level fusion image for the high-level images of the visible light Laplacian pyramid and the high-level images in the infrared Laplacian pyramid, and use dual-domain average fusion to obtain a low-level fusion image for the low-level images of the visible light Laplacian pyramid and the low-level images in the infrared Laplacian pyramid; perform layer-by-layer reconstruction based on the high-level fusion image and the low-level fusion image to obtain a reconstructed fusion image;

[0033] A detection module, configured to perform traffic detection based on the reconstructed fusion image.

[0034] In a third aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the steps in a traffic detection method based on visible light and infrared image fusion described in the first aspect.

[0035] In a fourth aspect, the present invention provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, it implements the steps in a traffic detection method based on visible light and infrared image fusion described in the first aspect.

[0036] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0037] 1. The present invention fuses visible light and infrared images. The infrared technology can identify by sensing the temperature of objects under low light conditions, getting rid of the dependence of visible light sensors on light sources; while the visible light technology can capture clearer and higher-resolution images. After the two are fused, it can provide a more accurate and precise target recognition technology.

[0038] 2. The present invention optimizes by introducing an adaptive filter to replace the traditional Gaussian filter in the multi-level decomposition process of the Laplacian pyramid of the image. The smoothing intensity is dynamically adjusted according to the local characteristics of the image, enabling strong smoothing in the flat areas of the image while retaining more details in the edge areas, thereby better capturing the edge and texture information in the image and enhancing the retention of details of the structural information in the fused image.

[0039] 3. The present invention uses a combination rule of gradient-weighted fusion and average fusion in the fusion process of the Laplacian pyramid, optimizing the image fusion effect while taking into account both the image details and the global information.

[0040] 4. The present invention combines multi-scale transformation and multi-scale registration fusion, which can process and align image features at different pyramid levels, improving the registration accuracy and adaptability to complex scenes.

[0041] Advantages of additional aspects of the present invention will be partly given in the following description, partly will become obvious from the following description, or will be learned through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] The accompanying drawings forming a part of this specification are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute a limitation to the present invention.

[0043] Figure 1 It is a flow framework diagram of a traffic detection method based on visible light and infrared image fusion provided by an embodiment of the present invention;

[0044] Figure 2 It is a main flow chart of a traffic detection method based on visible light and infrared image fusion provided by an embodiment of the present invention;

[0045] Figure 3 It is a flow chart of image fusion based on multi-scale transformation provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0046] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0047] Embodiment 1

[0048] As Figure 1 shown, this embodiment discloses a traffic detection method based on visible light and infrared image fusion, including the following steps:

[0049] S1: Obtain the original visible light image and infrared image within the same scene and perform preprocessing;

[0050] S2: Based on adaptive filtering, construct Laplacian pyramids for the preprocessed visible light image and infrared image respectively, and perform multi-scale fine registration on the visible light Laplacian pyramid and the infrared Laplacian pyramid;

[0051] S3: Use gradient weighted fusion for the high-level images of the visible light Laplacian pyramid and the high-level images in the infrared Laplacian pyramid to obtain high-level fusion images, and use dual-domain average fusion for the low-level images of the visible light Laplacian pyramid and the low-level images in the infrared Laplacian pyramid to obtain low-level fusion images; perform layer-by-layer reconstruction based on the high-level fusion images and the low-level fusion images to obtain reconstructed fusion images;

[0052] S4: Perform traffic detection based on the reconstructed fusion images.

[0053] Next, combined with Figure 2 , a traffic detection method based on visible light and infrared image fusion disclosed in this embodiment will be described in detail.

[0054] In S1, within the traffic monitoring area, configure a visible light CCD sensor and an infrared focal plane array detector. For the same target scene, collect the original visible light image and the original infrared image of the traffic scene in real time, and extract the ROI regions from the visible light image and the infrared image. Specifically include:

[0055] Select appropriate visible light CCD sensors and infrared focal plane array sensors and install them on the same platform, and use the same optical elements (such as lenses) for imaging to ensure that the imaging fields of view of the two overlap to the greatest extent. Then, select appropriate lenses and adjust the optical system to ensure that the optical paths of visible light and infrared light coincide as much as possible. During installation, perform on-site adjustment through fine-tuning brackets or adjusting screws to ensure the precise alignment of the sensor positions. After the system is built, perform imaging tests to confirm the alignment of the two images, and perform position fine-tuning if necessary to verify the rough registration effect.

[0056] Use the visible light sensor and the infrared sensor to collect images of the same scene and target, and capture the original 8-bit visible light image and the original 14-bit infrared image. According to the selected ROI coordinates, crop the ROI region from the original image to form a new sub-image.

[0057] Perform preprocessing on the infrared image and the visible light image obtained under the same target scene respectively.

[0058] (1) Preprocess the infrared image. Perform three steps of preprocessing on the infrared image: correction, compensation, and denoising. Use the two-point correction method to improve the overall quality and contrast of the image, fill in the missing pixel values in the image through blind pixel compensation to improve the integrity of the image, and finally use median filtering to remove the noise in the image. Specifically:

[0059] S101: Perform three-step preprocessing on the infrared image, namely correction, compensation, and denoising.

[0060] First, use the two-point correction method to convert the original signal output by the sensor into more accurate grayscale values. Select two known grayscale values in the two images using a standard grayscale card, and record the corresponding sensor output values as and , set the target grayscale values as and , calculate the parameters and through linear regression, and establish a linear equation based on the known input and output relationship:

[0061]

[0062] Use the linear equation to correct each pixel value in the original image to obtain the corrected image :

[0063]

[0064] Check each pixel in the image, identify the pixels with abnormal values (completely 0 or extremely small values, etc.) and mark them as blind pixels, and compensate each blind pixel using the neighborhood information. For the blind pixel at position , calculate , , , , the pixel mean value at the position, and assign it to the blind pixel.

[0065] Then perform median filtering denoising on the corrected and compensated image, determine the filtering window size ( , , etc.), traverse each pixel point in the image, represents the pixel value of the image at the position , extract the pixel grayscale values within its surrounding window and sort them. The pixel values in the window are the set , and replace the pixel value at the center of the current window with the median value obtained after sorting. Repeat this process for each pixel in the image until the entire image is processed, thereby achieving the denoising effect. The pixel of the filtered image at position is as shown in Equation (3):

[0066]

[0067] S102: Use the method of platform histogram equalization to enhance the contrast of the infrared image, making the detailed information in the image clearer.

[0068] Calculate the image gray - level histogram and when the number of pixels at a certain gray - level exceeds the platform threshold perform clipping to avoid over - enhancement, as shown in Equation (4):

[0069]

[0070] where is the threshold - clipped histogram.

[0071] For each clipped gray - level, evenly distribute the number of removed gray - value pixels to other gray - levels to ensure that the total number of pixels remains unchanged. Then calculate the cumulative distribution function of the clipped histogram:

[0072]

[0073] Use the cumulative distribution function to map the gray - level value of each pixel in the pre - processed image to the equalized gray - level value, thereby realizing image enhancement. The formula is as follows:

[0074]

[0075] where is the minimum value of the cumulative distribution function, is the maximum value of the image gray - level (usually 256).

[0076] S103: Use the maximum - minimum linear mapping method to convert the infrared image pixel values from 14 - bit data to 8 - bit data, which is consistent with the visible - light image pixel values.

[0077] Traverse all pixels of the infrared image to find its minimum gray - level value and maximum gray - level value , and linearly map the pixel values of the original 14 - bit image from the range to the 8 - bit gray - level range , and the formula is as follows:

[0078]

[0079] S104: Unify the image size through the bilinear interpolation algorithm.

[0080] Due to the differences in the sizes of visible light images and infrared images, the bilinear interpolation algorithm is used to set their sizes to be the same. In the actual processing, the processing of visible light images is only carried out on the Y component (i.e., the luminance component), and the chrominance information (U and V components) is not processed to avoid affecting the visual quality of the images due to chrominance changes.

[0081] Take the preprocessed infrared image as the target image and use the bilinear interpolation algorithm to adjust the size of the visible light image. Since the processing of visible light images is only carried out on the Y component, it is necessary to first convert it from the RGB color space to YUV:

[0082]

[0083] Extract the Y component from the converted YUV image to form a separate luminance image, and calculate the width and height scaling factors according to the sizes of the target infrared image and the original Y component 、 。Create a new blank matrix for the Y component with the size of the target image, and traverse each new pixel position :

[0084]

[0085] Select four neighboring pixel points at the corresponding position and determine them as the interpolation points, namely 、 、 、 Perform bilinear interpolation calculation on each obtained pixel position, and assign the result to the corresponding position of the new Y component :

[0086]

[0087] Among them, is the pixel value at the target position after interpolation, 、 、 、 are the pixel values of the four neighboring integer pixel points.

[0088] Combine the processed Y component with the original U and V components and convert it back to the RGB space representation:

[0089]

[0090] Compared with visible light images, due to their imaging characteristics and environmental influences, infrared images often require more processing to improve image quality and usability. In this embodiment, the above method is used to preprocess infrared images, which can significantly enhance the overall quality and integrity of the images, reduce image defects through correction, compensation, and denoising, then enhance the contrast using platform histogram equalization to make image details clearer, and at the same time use the maximum-minimum linear mapping method to optimize pixel values to reduce data complexity, and unify the image size through the bilinear interpolation algorithm, providing more accurate, efficient, and convenient image data for subsequent image analysis, processing, and applications.

[0091] (2)Perform scaling processing on the visible light image to ensure that the sizes of the visible light image and the infrared image match.

[0092] The process from S2 to S3 is as Figure 3 shown. In S2, based on adaptive filtering, Laplacian pyramids are constructed for the preprocessed visible light image and infrared image respectively.

[0093] Before that, first, ensure that the two images are completely aligned spatially, so that the imaging fields of view of the visible light CCD sensor and the infrared focal plane array sensor are nearly the same to achieve rough registration.

[0094] S201: Construct a Gaussian pyramid, repeatedly perform adaptive smoothing filtering on the image using an adaptive filter, and downsample the smoothed image for each layer until an image pyramid with multiple resolution levels is generated.

[0095] First, construct a Gaussian pyramid, and input the preprocessed visible light image and infrared image respectively. Use an adaptive mean filter to smooth the initial layer of the image, and dynamically adjust the filtering intensity to adapt to different feature regions of the image to obtain the pixel values after adaptive mean filtering:

[0096]

[0097] Among them, is the pixel value of the original image at position , is a local neighborhood centered on (preferably or window), Z represents the normalization factor, is the local weight function used to dynamically adjust the filtering intensity.

[0098] For flat regions, that is, regions with low contrast, set a higher smoothing intensity:

[0099]

[0100] For the edge regions, i.e., the regions with higher contrast, set a lower smoothing intensity:

[0101]

[0102] where is the threshold for controlling edge sensitivity. represents the local contrast, as shown in the formula:

[0103]

[0104] where is the number of pixels in the neighborhood, is the average value in the neighborhood.

[0105] Then, perform 2:1 downsampling on the filtered image to generate the next layer of the image, and repeat this step until the Gaussian pyramid is completed.

[0106] In this embodiment, the adaptive mean filtering method is used instead of the traditional Gaussian filter, which can dynamically adjust the smoothing intensity according to the local characteristics of the image. In the flat regions of the image, the adaptive filtering can perform strong smoothing to effectively remove noise; while in the edge regions of the image, the adaptive filtering will weaken the smoothing intensity, thus retaining the detailed information of the image. By performing adaptive filtering on visible light images and infrared images, the image quality can be significantly improved, making it more suitable for subsequent image fusion and traffic detection tasks. The adaptive filtering helps to enhance the contrast of the image, making the target objects in the image clearer and easier to detect.

[0107] S202: Construct the Laplacian pyramid by performing layer-by-layer differencing on the Gaussian pyramid.

[0108] Upsample each layer of the image in the Gaussian pyramid, calculate the difference between the current layer and the upsampled image, so as to obtain each layer of the Laplacian pyramid. The formula is as follows:

[0109]

[0110] In the formula, represents the image of the th layer of the visible light Laplacian pyramid, represents the image of the th layer of the visible light Gaussian pyramid, is the visible light image obtained by upsampling the th layer of the image; represents the image of the th layer of the infrared light Laplacian pyramid, represents the image of the The image of the layer is the infrared image obtained by upsampling the image of the layer

[0111] Perform multi-scale fine registration on the visible light Laplacian pyramid and the infrared Laplacian pyramid, specifically including:

[0112] S211: Register the visible light Laplacian pyramid and the infrared Laplacian pyramid to obtain a fused Laplacian pyramid; specifically including:

[0113] Use the Sobel operator to calculate the gradient to obtain the gradient magnitude and direction of the visible light and infrared images:

[0114]

[0115] Calculate the gradient magnitude and direction :

[0116]

[0117] Divide the image into 8x8 pixel small cells, and calculate the gradient direction histogram in each cell. The gradient directions within each cell are divided into several bins (for example, 9 directions), and the gradient magnitude of each pixel is accumulated in the corresponding bin:

[0118]

[0119] where is the gradient magnitude of pixel , is the indicator function is the corresponding bin

[0120] Normalize the histogram in each small cell to improve the robustness to illumination changes. Combine the normalized histograms together, and reconstruct them into the visible light reconstructed image and the infrared reconstructed image respectively, and calculate the final HOG features of the visible light reconstructed image and the infrared reconstructed image. The normalization formula is as follows:

[0121]

[0122] where H represents the unnormalized HOG feature vector is a small constant to prevent division by zero

[0123] Calculate the distance between the feature points of the visible light reconstructed image and the infrared reconstructed image, and select the feature points with the smallest distance as the matching pairs for feature matching. Then, transform the infrared image according to the estimated transformation matrix to align it with the visible light image. Apply the transformation step by step at each layer of the Laplacian pyramid, and perform a fusion operation on the Laplacian pyramids of the visible light and infrared images to generate a fused Laplacian pyramid, achieving fine registration.

[0124] In S3, use the multi-scale decomposition fusion method based on Laplacian pyramid decomposition to perform image fusion on the registered images. The image fusion process based on multi-scale transformation includes the following steps:

[0125] For the higher-resolution levels, adopt the gradient weighted fusion method, which can retain the details of the visible light image and the heat source information of the infrared image; while for the lower-resolution levels, use the average fusion method to smooth the global features of the fused image and ensure the consistency of the image brightness and contour.

[0126] For the higher-resolution image layers at the lower part of the pyramid, use the gradient weighted fusion method. Calculate the fusion weights for each image according to the gradient magnitude. Regions with larger gradients have higher weights, and regions with smaller gradients have lower weights. The weight calculation formula is:

[0127]

[0128] Among them, is the weight of the visible light image, is the weight of the infrared image; represents the gradient of the visible light image, represents the gradient of the infrared image.

[0129] Use the gradient weighted fusion rule to perform weighted fusion on the high-resolution layers of the visible light and infrared images. The formula is:

[0130]

[0131] Among them, represents the high-level fused image, is the high-level image of the visible light Laplacian pyramid, is the high-level image in the infrared Laplacian pyramid.

[0132] For the low-resolution layers of the Laplacian pyramid, use dual-domain average fusion to obtain the final low-level fused image. Among them, dual-domain average fusion includes global average fusion and local average fusion.

[0133] First, perform global average fusion on the visible light and infrared images to obtain a preliminary fusion result:

[0134]

[0135] Among them, represents the preliminary low-level fused image, is the low-level image of the visible light Laplacian pyramid, is the low-level image in the infrared Laplacian pyramid. On the basis of the original global average fusion, the average fusion technology within the local area is added, so as to retain more detailed image features.

[0136] According to the global fusion result, select the local areas of the visible light and infrared images based on specific image features (such as edges or textures) for further processing, and use windows of size to define the local areas, and perform average calculation again within each local area, so as to retain more detailed image features:

[0137]

[0138] Among them, represents the local low-level fused image, are the pixel coordinates within the local area, is the local area in the low-level image of the visible light Laplacian pyramid, is the local area in the low-level image in the infrared Laplacian pyramid.

[0139] Combine the global average result and the local average result to obtain the final low-resolution layer fused image:

[0140]

[0141] Among them, is the adjustment parameter, which controls the ratio of global and local fusion.

[0142] In this embodiment, for high-level images, the gradient weighted fusion method can make full use of the edge and texture information in the high-level images, and obtain a clearer and more accurate high-level fused image through weighted fusion; for low-level images, the dual-domain average fusion method can smooth the noise and detail fluctuations in the images, so as to obtain a more stable low-level fused image. By combining the gradient weighted fusion and the average fusion methods, the complementary information in the visible light image and the infrared image can be fully utilized, and the quality and accuracy of the fused image can be improved. This helps the subsequent traffic detection tasks to be more accurate and reliable.

[0143] Among them, the dual-domain average fusion method first performs global average fusion, which can preliminarily fuse visible light and infrared images to obtain a relatively balanced preliminary fusion result, helping to integrate the overall information of the images. Secondly, local average fusion is based on global fusion and processes local regions of the infrared image according to specific image features, calculating the average using a window of a specific size, which can retain more detailed image features such as edges and textures. Finally, by adjusting parameters, the global and local fusion results are combined, achieving flexible control of the fusion ratio and enabling the fusion image to achieve good effects in both overall and detail aspects.

[0144] Perform layer-by-layer reconstruction on the fused high-level and low-level fusion images to obtain a fusion image containing the features of both modalities.

[0145] After all layers are fused, reconstruct the Laplacian pyramid layer by layer. Starting from the bottom layer of the pyramid, perform upsampling layer by layer, add the obtained image to the Laplacian low-resolution image of that layer to restore the reconstructed image of that layer. Repeat this process, reconstructing layer by layer up to the top until the size of the original image is restored. The formula is as follows:

[0146]

[0147] Among them, is the reconstructed fusion image of the th layer, is the fused Laplacian image of the th layer.

[0148] After obtaining the reconstructed fusion image, perform pseudo-color processing. The pseudo-color processing uses the gray-level-color transformation method. By establishing the mapping relationship between the gray levels of the gray image and various colors in the color space, the gray image is converted into a color image to further enhance its visual effect.

[0149] Use the gray-level-color transformation method to perform pseudo-color processing on the image after fusing the visible light image and the infrared image. First, perform normalization processing on the fusion image, normalizing each pixel value to the range. The formula is as follows:

[0150]

[0151] Among them, and are the minimum and maximum pixel values in the image respectively.

[0152] Use a color mapping table to map each pixel value of the gray-level image to a color value in the RGB space, convert it into a pseudo-color image and save it. The pseudo-color image The color of each pixel can be expressed as:

[0153]

[0154] where is a function that maps the grayscale value to an RGB color, represents the grayscale value of the pixel at the (x, y) position.

[0155] In S4, the fusion of visible light and infrared images is applied to traffic detection in various complex scenarios, which can perform real-time recognition and abnormal behavior analysis on pedestrians, vehicles, etc., and assist in the detection and early warning of traffic accidents, further improving traffic safety.

[0156] As an alternative implementation, a deep learning-based object detection model is used to detect the image to be detected, and the detected object information is output, including vehicle object information, non-motor vehicle object information, and pedestrian object information;

[0157] A deep learning-based object tracking model is used to track the detected object information, and the object trajectory data is output;

[0158] According to the output object information and object trajectory data, combined with the traffic event detection parameters pre-configured by the configuration unit, the parameter matching unit makes a determination. When the pre-configured traffic event detection parameters are met, the corresponding traffic event is determined, and the detected object information and traffic event are used as the detection result.

[0159] Furthermore, the deep learning-based object detection model adopts the object detection model YOLO, the single-shot multi-box detector SSD, or the faster region convolutional neural network model, namely the Faster RCNN model.

[0160] Furthermore, the deep learning-based object tracking model adopts the efficient convolutional neural network model for object tracking ECO, the Siamese region proposal network considering interference DaSiamRPN, or the multi-domain object tracking detection model, namely the MDNet model.

[0161] The visible light and infrared fusion image completed through the fusion process can be applied to various traffic detections, including the speed and position changes of vehicles, the illegal behaviors of pedestrians, etc. It is especially suitable for applications in low-light complex environments such as night scenes and bad weather, and can effectively reduce the interference caused by light changes. By using a deep learning model to perform object detection on the fusion image and conduct data training, it is possible to accurately identify and real-time detect traffic participants such as pedestrians and vehicles. On this basis, it is also possible to analyze the movement trajectories of the targets, monitor speed, direction changes, and relative positions, etc., and identify whether there are potential abnormal behaviors, such as sudden braking, lane deviation, pedestrians running through red lights, etc. At the same time, through real-time continuous traffic detection and setting judgment thresholds for abnormal behaviors, it is possible to timely analyze and predict possible accidents. Once an abnormal event is identified, the system will automatically issue an alarm and notify the relevant traffic management departments in a timely manner, reducing the possibility of accidents and improving the accuracy and efficiency of traffic accident detection.

[0162] Embodiment 2

[0163] This embodiment provides a traffic detection system based on the fusion of visible light and infrared images, including:

[0164] An image acquisition module, used to obtain the original visible light image and infrared image within the same scene and perform preprocessing;

[0165] A registration module, used to respectively construct Laplacian pyramids for the preprocessed visible light image and infrared image based on adaptive filtering, and perform multi-scale fine registration on the visible light Laplacian pyramid and the infrared Laplacian pyramid;

[0166] A fusion reconstruction module, used to use gradient weighted fusion to obtain a high-level fusion image for the high-level images of the visible light Laplacian pyramid and the high-level images in the infrared Laplacian pyramid, and use dual-domain average fusion to obtain a low-level fusion image for the low-level images of the visible light Laplacian pyramid and the low-level images in the infrared Laplacian pyramid; perform layer-by-layer reconstruction based on the high-level fusion image and the low-level fusion image to obtain a reconstructed fusion image;

[0167] A detection module, used to perform traffic detection based on the reconstructed fusion image.

[0168] Embodiment 3

[0169] This embodiment provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the steps in a traffic detection method based on the fusion of visible light and infrared images as described in Embodiment 1 above.

[0170] Embodiment 4

[0171] This embodiment provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps in a traffic detection method based on visible light and infrared image fusion as described in Embodiment 1 above are implemented.

[0172] The steps or modules involved in the above Embodiments 2 to 4 correspond to those in Embodiment 1. For specific implementation manners, reference may be made to the relevant description part of Embodiment 1. The term "computer-readable storage medium" should be understood to include a single medium or multiple media including one or more instruction sets; it should also be understood to include any medium that can store, encode, or carry an instruction set for execution by a processor and enable the processor to execute any method in the present invention.

[0173] The foregoing are only preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A traffic detection method based on visible light and infrared image fusion, characterized in that: include: The original visible light image and infrared image in the same scene are obtained and preprocessed, specifically: a visible light CCD sensor and an infrared focal plane array detector are configured in the traffic monitoring area; For the same target scene, the original visible light image and original infrared image of the traffic scene are collected in real time, and the ROI area is extracted from the visible light image and infrared image, including: Select appropriate visible light CCD sensors and infrared focal plane array sensors and install them on the same platform, use the same optical elements for imaging, and ensure that the imaging fields of the two overlap to the greatest extent; then select appropriate lenses and adjust the optical system to ensure that the optical paths of visible light and infrared light overlap as much as possible; during installation, make on-site adjustments through fine-tuning brackets or adjusting screws to ensure that the sensor positions are accurately aligned; after the system is built, perform imaging tests to confirm the alignment of the two images, and make fine adjustments to the positions to verify the rough registration effect; Use visible light sensor and infrared sensor to collect images of the same scene and target, capture the original 8-bit visible light image and the original 14-bit infrared image; according to the selected ROI coordinates, cut out the ROI area from the original image to form a new sub-image; Based on adaptive filtering, Laplacian pyramids are constructed for the preprocessed visible light image and infrared image respectively, and multi-scale fine registration is performed on the visible light Laplacian pyramid and the infrared Laplacian pyramid, including: Obtaining the images in the two Laplacian pyramids, and using the Sobel operator to respectively calculate the gradients of the images in the two Laplacian pyramids; Divide the image into multiple small units and calculate the gradient direction histogram in each unit; Normalize the gradient direction histogram in each small unit, reconstruct the normalized histogram into visible light reconstructed image and infrared reconstructed image respectively, and calculate the HOG features of the visible light reconstructed image and infrared reconstructed image respectively; Calculate the distance between the feature points of the visible light reconstructed image and the infrared reconstructed image, select the feature point with the smallest distance as the matching pair for feature matching, and transform the infrared image according to the estimated transformation matrix to align it with the visible light image; A high-level image of the visible light Laplace pyramid is fused with a high-level image in the infrared Laplace pyramid by using gradient weighted fusion to obtain a high-level fused image, and a low-level image of the visible light Laplace pyramid is fused with a low-level image in the infrared Laplace pyramid by using dual-domain average fusion to obtain a low-level fused image; based on the high-level fused image and the low-level fused image, layer-by-layer reconstruction is performed to obtain a reconstructed fused image; Traffic detection based on reconstructed fused images.

2. A traffic detection method based on visible light and infrared image fusion as claimed in claim 1, characterized in that: The pre-processing comprises: After the infrared image is preprocessed by two-point correction, blind pixel compensation and median filtering, the contrast is enhanced by platform histogram equalization, the pixel is quantized by maximum and minimum linear mapping, and the image size is unified by bilinear interpolation algorithm. The visible light image is scaled to ensure that the sizes of the visible light image and the infrared image match.

3. The traffic detection method based on visible light and infrared image fusion as claimed in claim 1, characterized in that: The method of constructing a Laplacian pyramid for the preprocessed visible light image and infrared image based on adaptive filtering comprises the following specific steps: Adaptive mean filtering method is used to smooth the preprocessed visible light image and infrared image; The filtered image is repeatedly downsampled to construct a visible light Gaussian pyramid and an infrared Gaussian pyramid respectively; The two Gaussian pyramids are respectively differentiated layer by layer to obtain a visible light Laplace pyramid and an infrared Laplace pyramid.

4. A traffic detection method based on visible light and infrared image fusion as claimed in claim 3, characterized in that: The adaptive mean filtering method specifically includes: dynamically adjusting the smoothing strength according to the local characteristics of the image, performing strong smoothing in the flat area of ​​the image, and weakening the smoothing strength in the edge area.

5. The traffic detection method based on visible light and infrared image fusion as claimed in claim 1, characterized in that: The low-level image of the visible light Laplace pyramid and the low-level image of the infrared Laplace pyramid are fused by dual-domain average fusion to obtain a low-level fused image, wherein the dual-domain average fusion includes global average fusion and local average fusion, specifically including: The lower-level image of the visible light Laplacian pyramid and the lower-level image of the infrared Laplacian pyramid are fused by global average to obtain a preliminary lower-level fused image; Divide the local areas of the two low-level images based on the preset window, traverse the local areas to perform local average fusion, and obtain a local low-level fusion image; The preliminary low-level fusion image and the local low-level fusion image are weightedly summed to obtain a final low-level fusion image.

6. The traffic detection method based on visible light and infrared image fusion as claimed in claim 1, characterized in that: It also includes using color transformation method to perform pseudo color processing on the reconstructed fused image.

7. A traffic detection system based on visible light and infrared image fusion, characterized in that: include: The image acquisition module is used to obtain the original visible light image and infrared image in the same scene and perform preprocessing. Specifically, in the traffic monitoring area, a visible light CCD sensor and an infrared focal plane array detector are configured; for the same target scene, the original visible light image and original infrared image of the traffic scene are collected in real time, and the ROI area of ​​the visible light image and the infrared image is extracted, which specifically includes: Select appropriate visible light CCD sensors and infrared focal plane array sensors and install them on the same platform, use the same optical elements for imaging, and ensure that the imaging fields of the two overlap to the greatest extent; then select appropriate lenses and adjust the optical system to ensure that the optical paths of visible light and infrared light overlap as much as possible; during installation, make on-site adjustments through fine-tuning brackets or adjusting screws to ensure that the sensor positions are accurately aligned; after the system is built, perform imaging tests to confirm the alignment of the two images, and make fine adjustments to the positions to verify the rough registration effect; Use visible light sensor and infrared sensor to collect images of the same scene and target, capture the original 8-bit visible light image and the original 14-bit infrared image; according to the selected ROI coordinates, cut out the ROI area from the original image to form a new sub-image; The registration module is used to construct Laplacian pyramids for the preprocessed visible light image and infrared image based on adaptive filtering, and perform multi-scale fine registration on the visible light Laplacian pyramid and the infrared Laplacian pyramid, specifically including: Obtaining the images in the two Laplacian pyramids, and using the Sobel operator to respectively calculate the gradients of the images in the two Laplacian pyramids; Divide the image into multiple small units and calculate the gradient direction histogram in each unit; Normalize the gradient direction histogram in each small unit, reconstruct the normalized histogram into visible light reconstructed image and infrared reconstructed image respectively, and calculate the HOG features of the visible light reconstructed image and infrared reconstructed image respectively; Calculate the distance between the feature points of the visible light reconstructed image and the infrared reconstructed image, select the feature point with the smallest distance as the matching pair for feature matching, and transform the infrared image according to the estimated transformation matrix to align it with the visible light image; A fusion and reconstruction module is used to fuse the high-level image of the visible light Laplace pyramid with the high-level image in the infrared Laplace pyramid by using gradient weighting to obtain a high-level fused image, and fuse the low-level image of the visible light Laplace pyramid with the low-level image in the infrared Laplace pyramid by using dual-domain average to obtain a low-level fused image; reconstruct the high-level fused image and the low-level fused image layer by layer to obtain a reconstructed fused image; The detection module is used for traffic detection based on the reconstructed fused image.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps in a traffic detection method based on visible light and infrared image fusion as described in any one of claims 1 to 6 are implemented.

9. A computer device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps in the traffic detection method based on visible light and infrared image fusion as described in any one of claims 1-6 are implemented.

Citation Information

Patent Citations

  • Traffic incident detection method and device based on infrared image and visible light image

    CN117173649A

  • Infrared and visible light fusion method, traffic monitoring device and storage medium

    CN112750095A

  • InSAR interferogram iteration adaptive filtering method based on Laplacian pyramid

    CN114066778A

  • Image registration method and apparatus, electronic device and storage medium

    WO2022100065A1