Dynamic multi-scale layered fusion infrared and visible light image fusion image enhancement method, system and medium
By using a dynamic multi-scale hierarchical fusion method, combined with dual-threshold Otsu and LBP detail density markers, and dynamically adjusting the number of Db4 wavelet decomposition layers, the problems of noise amplification in bright areas and missed detection of weak targets in dark areas in existing technologies are solved. This achieves enhancement of infrared thermal target details and dark area details, and improves the identification capability of low-altitude security monitoring.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-19
- Publication Date
- 2026-03-10
AI Technical Summary
Existing infrared-visible light image fusion technology cannot adapt to the differentiated characteristics of multiple regions in airport low-altitude safety monitoring, resulting in amplified noise in bright areas or missed detection of weak targets in dark areas, and cannot meet the requirements of high-precision, all-weather monitoring.
A dynamic multi-scale hierarchical fusion method is adopted. By jointly labeling the Otsu threshold with LBP detail density, the number of Db4 wavelet decomposition layers is dynamically adjusted. Combined with guided filtering and CLAHE algorithm, differential enhancement processing is performed to suppress noise and preserve details, thereby enhancing the details of infrared thermal targets and dark areas.
It improves the ability to identify and detect weak targets at low altitudes, enhances the quality of image fusion, and meets the high-precision, all-weather monitoring requirements for low-altitude safety monitoring at airports.
Smart Images

Figure CN121639484A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image processing technology, and in particular to the fusion technology of infrared and visible light images for low-altitude security monitoring. Specifically, it relates to a dynamic multi-scale layered fusion method for infrared and visible light image fusion and enhancement, which enhances the details of infrared thermal targets and dark areas, thereby improving the detection and recognition capabilities of weak targets at low altitudes. Background Technology
[0002] With the gradual opening of low-altitude airspace, the vitality of the low-altitude sector has been greatly stimulated, leading to an explosive growth in the application of drones, small aircraft, and low-altitude aircraft, bringing new challenges to airspace management. Various scenarios for low-altitude safety monitoring place stringent demands on monitoring technologies across multiple dimensions. Taking the area surrounding airports as an example, the requirements for low-altitude safety are extremely high. Any unauthorized low-altitude aircraft intruding into the airport's airspace could collide with commercial aircraft taking off or landing, causing flight safety accidents. Therefore, the low-altitude monitoring system around airports requires higher precision, comprehensive coverage, and real-time performance, capable of accurately detecting various objects from a distance and continuously and stably tracking their flight trajectories to enable timely and effective measures to ensure the normal operation of the airport and flight safety.
[0003] Traditional single-source infrared or visible light monitoring systems (such as the T2000 infrared thermal imager commonly used in civil aviation and visible light high-definition cameras) can capture weak thermal targets at low altitudes at night (such as small drones and birds with thermal radiation intensity of about 30-40℃) using single infrared imaging, but they have two major problems: First, they lose a lot of detail and cannot distinguish between "drone propellers and bird wings" (which appear as point-like thermal targets in infrared images), resulting in a false alarm rate as high as 23% (according to the statistics in the "White Paper on Low-Altitude Safety Monitoring Technology of Civil Aviation 2024"); Second, the inherent particle noise of infrared sensors (standard deviation of about 8-10) is more obvious in dark areas of airports (such as the shadows of grass next to the runway and the nighttime area of the apron), and noise is easily misjudged as small targets. While visible light imaging can provide rich details during the day (such as drone models), it is affected by uneven airport lighting at night (runway light brightness > 200 cd / m², grass shadow brightness < 10 cd / m²), resulting in overexposure in bright areas (pixel saturation of 255 in light areas) and loss of detail in dark areas (target grayscale value < 30 in shadows). This leads to a target detection rate of less than 60% at night, which is difficult to meet the civil aviation safety standard of "99% detection rate".
[0004] Currently, the mainstream equipment for low-altitude safety monitoring at airports typically uses pan-tilt-zoom (PTZ) multi-channel fusion camera terminals, such as visible light-infrared dual-channel camera terminals and visible light-infrared-laser ranging multi-channel camera terminals. These terminals use infrared sensors to detect the infrared radiation emitted by objects, enabling infrared imaging regardless of ambient light conditions. They can stably acquire thermal infrared images of targets even in complete darkness or adverse weather conditions. Visible light sensors capture the target's detailed texture, color, and other visual features, providing clear and intuitive images that facilitate direct observation and identification by operators. Applying infrared and visible light fusion technology to low-altitude safety monitoring has two advantages. First, it leverages the stable detection capabilities of infrared sensors in complex environments to achieve all-weather, all-time monitoring of low-altitude targets, effectively solving the problem of visible light monitoring being affected by light and weather. Second, by fusing the rich detail information of visible light images, it enables more accurate feature analysis and classification of targets detected by infrared, compensating for the shortcomings of infrared images in presenting target details. This significantly improves the overall performance of the low-altitude safety monitoring system, enabling more comprehensive, accurate, and reliable monitoring of low-altitude targets and meeting the needs of complex low-altitude safety monitoring scenarios.
[0005] Existing infrared-visible image fusion methods typically employ fixed-weight fusion (e.g., infrared 0.5 + visible light 0.5), single-scale wavelet fusion, and standard CLAHE enhancement. However, these methods suffer from a "one-size-fits-all" problem in airport scenarios, failing to adapt to the diverse characteristics of different areas within an airport, such as bright areas on runways, terminal transition zones, and dark areas in grass. Under fixed-weight fusion, bright areas on airport runways (where visible light noise is concentrated) experience noise amplification (standard deviation > 8) due to excessively high visible light weights, while dark areas in grass (dominated by infrared information) suffer from missed detections of weak targets (such as small drones weighing less than 2 kg) due to insufficient infrared weights. In standard CLAHE enhancement, traditional CLAHE uses a fixed clipLimit=2.0 (the default parameter in OpenCV). In dark areas of the airport (such as the shadows on the tarmac), it will forcibly stretch the grayscale range, which will amplify the sensor noise (the noise standard deviation will increase from 5 to 6.75). In bright areas (runway lights), due to the excessively high contrast limit, the grayscale difference between the light area and the background will be reduced from 20 to 8, which will blur the outline of the target (such as an intruder next to the runway). Summary of the Invention
[0006] In view of the defects and deficiencies of existing technologies, the purpose of this invention is to address the problems of weak target detection under complex lighting conditions, detail preservation in multiple scene backgrounds, and strong light noise suppression in airport low-altitude security monitoring. It proposes a dynamic multi-scale layered fusion infrared and visible light image enhancement method. Through dynamic multi-scale layered fusion and adaptive noise suppression, it achieves image fusion enhancement under complex lighting and scene backgrounds at airports. Layered differential enhancement improves the quality of the fused image, enhancing details of infrared thermal targets and dark areas, solving the problem of image fusion distortion in multiple brightness regions, and improving subsequent intrusion target identification, target tracking, event detection, and other recognition capabilities, especially the ability to identify and detect weak targets at low altitudes.
[0007] According to a first aspect of the present invention, a dynamic multi-scale layered fusion infrared and visible light image enhancement method is proposed, comprising the following steps:
[0008] Step 1: Obtain visible light image I VIS With infrared image I IR Input the image and perform inherent noise smoothing on each image, then output the smoothed visible light image I. VIS_smooth With infrared image I IR_smooth And for the smoothed visible light image I VIS_smooth Perform normalization and output the grayscale normalized image I. VIS_norm ;
[0009] Step 2: Traverse the smoothed visible light image I VIS_smooth Count the number of pixels for each grayscale value to generate a visible light image I. VIS_smooth The brightness histogram H(g) is given, where g is the gray value, ranging from 0 to 255, and H(g) represents the total number of pixels with a gray value of g.
[0010] Step 3: Based on the visible light image I VIS_smooth The brightness histogram H(g) is used to generate a brightness layering marker map M. bright (x,y) and detail density map D map (x,y); then the brightness layer marker map M is fused. bright (x,y) and detail density map D map (x,y) yields the joint labeling map M of the brightness-detail sublayer. bright−detail (x,y), where each pixel value is a combination of a luminance marker and a detail marker;
[0011] Step 4: Based on the smoothed infrared image I IR_smooth Based on the thermal target threshold T IR Perform thermal target region segmentation to obtain thermal target binary label map M. IR (x,y); and combined with the brightness-detail sublayer joint labeling map Mbright−detail (x,y) superimposed labels generate a multi-feature label map M. final (x,y), where each pixel value contains a brightness marker, a detail marker, and a feature marker for a thermal target;
[0012] Step 5: Based on the multi-feature marker map M final The three-dimensional features represented in (x,y) are used to determine the number of decomposition levels L in the dynamic wavelet decomposition, and a decomposition level mapping table is obtained; based on this, the multi-feature labeled map M is traversed. final For each pixel in (x,y), according to M final The pixel value of each pixel in (x,y) is matched with the corresponding decomposition level L(x,y) from the decomposition level mapping table to generate the level allocation map L. map (x,y);
[0013] Step 6: Normalize the grayscale image I VIS_norm And the smoothed infrared image I IR_smooth Wavelet decomposition is performed separately, and the decomposition process is assigned to the specified number of layers in Figure L. map (x,y) is dynamically adjusted, and the sets of low-frequency and high-frequency coefficients on the visible light side and the sets of low-frequency and high-frequency coefficients on the infrared side are obtained through wavelet decomposition.
[0014] Step 7: For the low-frequency and high-frequency coefficient sets of the visible light side and the infrared side, respectively, perform coefficient enhancement and align them according to the original decomposition level to ensure that the number of coefficient layers for each pixel is consistent.
[0015] Step 8: Assign infrared weights W based on different brightness levels. IR (x,y) and visible light weight W VIS (x,y), and the enhanced low-frequency and high-frequency coefficient sets of the infrared outer and visible light sides, are fused separately in the frequency domain, including low-frequency coefficient fusion and high-frequency coefficient fusion, and W is used in each brightness layer. IR (x,y)+W VIS (x,y)=1; and
[0016] Step 9: Based on the fused high-frequency and low-frequency coefficients, reconstruct the image spatial domain pixel by pixel using the inverse wavelet transform corresponding to the dynamic wavelet decomposition, and obtain the spatial domain infrared-visible fused image I. fusion (x,y).
[0017] According to a second aspect of the present invention, a computer system is provided.
[0018] One or more processors;
[0019] The memory stores operable instructions that, when executed by the one or more processors, cause the one or more processors to perform operations, including the flow of the dynamic multi-scale hierarchical fusion infrared and visible light image enhancement method of the foregoing embodiments.
[0020] In a second aspect of the invention, a computer-readable medium for storing software is provided, the software comprising instructions executable by one or more computers, the instructions causing the one or more computers to perform operations including the flow of the dynamic multi-scale layered fusion infrared and visible light image enhancement method of the foregoing embodiments.
[0021] The image enhancement method for infrared and visible light image fusion based on the dynamic multi-scale layered fusion of the above embodiments of the present invention addresses the complex characteristics of background and illumination in nighttime images in airport scenes. First, it uses dual-threshold Otsu and LBP detail density joint labeling to achieve dynamic layering of the three-peak brightness distribution of dark areas, transition areas, and bright areas in nighttime images of large airport scenes. Then, it combines LBP detail density to divide grayscale value brightness transitions into high and low detail sub-layers. The high and low detail sub-layers are used as a distinction for subsequent differentiated enhancement processing. For example, the high detail sub-layer focuses on detail enhancement, while the low detail sub-layer focuses on noise suppression.
[0022] In the method of this invention, based on dynamic brightness layering and high / low detail sub-layer capture, dynamic wavelet decomposition and hybrid enhancement are used to perform multi-scale image enhancement in different frequency domains. Unlike traditional wavelet transform with a fixed number of layers, this invention dynamically adjusts the number of Db4 wavelet decomposition layers (3-6 layers) according to the characteristics of dynamic brightness layering and sub-layers: 6 layers are decomposed for low detail sub-layers in dark areas to maximize the preservation of infrared thermal information, while only 3 layers are decomposed for low detail sub-layers in bright areas to suppress noise propagation. Simultaneously, for visible light images, guided filtering technology is used for high-frequency detail enhancement, calculating the detail gain map through optimization to achieve texture enhancement while preserving edges; while for low-frequency contrast enhancement, the CLAHE algorithm based on brightness layering is used, setting dynamic parameters for dark areas, transition areas, and bright areas to avoid the problems of noise amplification in dark areas or overexposure in bright areas caused by traditional fixed parameters.
[0023] Meanwhile, infrared images inherently suffer from grain noise (originating from the thermal noise of the infrared sensor). If directly involved in fusion, this can lead to speckled appearance in the final image. Therefore, this invention selectively attenuates the high-frequency coefficients of the infrared image: the high-frequency coefficients in the hot target region contain thermal details and are retained; the high-frequency coefficients in the non-hot target region are mostly noise, and attenuation reduces interference with the fused image; for the low-frequency coefficients of the infrared image, a slight Gaussian blur is used to smooth the grain noise while preserving the overall shape of the hot target, preventing the outline of the hot target from being distorted by noise interference at night.
[0024] It should be understood that all combinations of the foregoing concepts and the additional concepts described in more detail below may be considered part of the inventive subject matter of this disclosure, provided that such concepts do not contradict each other. Furthermore, all combinations of the claimed subject matter are considered part of the inventive subject matter of this disclosure.
[0025] The foregoing and other aspects, embodiments, and features of the teachings of the present invention will be more fully understood from the following description in conjunction with the accompanying drawings. Other additional aspects of the invention, such as features and / or beneficial effects of exemplary embodiments, will become apparent from the following description or may be learned through practice of specific embodiments according to the teachings of the present invention. Attached Figure Description
[0026] The accompanying drawings are not intended to be drawn to scale. In the drawings, each identical or nearly identical component shown in the various figures may be denoted by the same reference numeral. For clarity, not every component is labeled in each figure. Embodiments of various aspects of the invention will now be described by way of example and with reference to the accompanying drawings.
[0027] Figure 1 This is a flowchart illustrating the image enhancement method for dynamic multi-scale layered fusion of infrared and visible light images according to an embodiment of the present invention.
[0028] Figure 2 This is a schematic diagram of the process for generating a joint marker map of brightness-detail sublayers according to an embodiment of the present invention.
[0029] Figure 3 This is a schematic diagram illustrating the process of superimposing a binary marker map of the thermal target and a joint marker map of the brightness-detail sublayer according to an embodiment of the present invention.
[0030] Figure 4 This is a schematic diagram of the process for obtaining the generation layer allocation map of wavelet transform according to an embodiment of the present invention.
[0031] Figure 5 This is a schematic diagram illustrating wavelet decomposition of a grayscale normalized image and a smoothed infrared image according to an embodiment of the present invention.
[0032] Figure 6 This is a schematic diagram illustrating coefficient enhancement and alignment of the low-frequency and high-frequency coefficient sets on the visible light side according to an embodiment of the present invention.
[0033] Figure 7 This is a schematic diagram illustrating coefficient enhancement and alignment of the low-frequency and high-frequency coefficient sets outside the infrared region according to an embodiment of the present invention.
[0034] Figure 8This is a schematic diagram of obtaining global high-low frequency fusion coefficients by fusing high- and low-frequency coefficients from the visible light side and infrared measurement according to an embodiment of the present invention.
[0035] Figure 9 This is a schematic diagram of the process of obtaining a spatial domain fused image based on wavelet transform according to an embodiment of the present invention. Detailed Implementation
[0036] To better understand the technical content of the present invention, specific embodiments are described below in conjunction with the accompanying drawings.
[0037] Various aspects of the invention are described in this disclosure with reference to the accompanying drawings, which illustrate numerous illustrative embodiments. The embodiments of this disclosure are not necessarily intended to encompass all aspects of the invention. It should be understood that the various concepts and embodiments described above, as well as those described in more detail below, can be implemented in any of many ways, because the concepts and embodiments disclosed herein are not limited to any particular implementation. Furthermore, some aspects of the invention disclosed may be used alone or in any suitable combination with other aspects of the invention disclosed.
[0038] {Example 1}
[0039] Combined with appendix Figure 1 As shown, the image enhancement method for dynamic multi-scale layered fusion of infrared and visible light images according to an embodiment of the present invention includes the following steps:
[0040] Step 1: Obtain visible light image I VIS With infrared image I IR Input the image and perform inherent noise smoothing on each image, then output the smoothed visible light image I. VIS_smooth With infrared image I IR_smooth And for the smoothed visible light image I VIS_smooth Perform normalization and output the grayscale normalized image I. VIS_norm ;
[0041] Step 2: Traverse the smoothed visible light image I VIS_smooth Count the number of pixels for each grayscale value to generate a visible light image I. VIS_smooth The brightness histogram H(g) is given, where g is the gray value, ranging from 0 to 255, and H(g) represents the total number of pixels with a gray value of g.
[0042] Step 3: Based on the visible light image I VIS_smooth The brightness histogram H(g) is used to generate a brightness layering marker map M. bright (x,y) and detail density map D map (x,y); then the brightness layer marker map M is fused. bright(x,y) and detail density map D map (x,y) yields the joint labeling map M of the brightness-detail sublayer. bright−detail (x,y), where each pixel value is a combination of a luminance marker and a detail marker;
[0043] Step 4: Based on the smoothed infrared image I IR_smooth Based on the thermal target threshold T IR Perform thermal target region segmentation to obtain thermal target binary label map M. IR (x,y); and combined with the brightness-detail sublayer joint labeling map M bright−detail (x,y) superimposed labels generate a multi-feature label map M. final (x,y), where each pixel value contains a brightness marker, a detail marker, and a feature marker for a thermal target;
[0044] Step 5: Based on the multi-feature marker map M final The three-dimensional features represented in (x,y) are used to determine the number of decomposition levels L in the dynamic wavelet decomposition, and a decomposition level mapping table is obtained; based on this, the multi-feature labeled map M is traversed. final For each pixel in (x,y), according to M final The pixel value of each pixel in (x,y) is matched with the corresponding decomposition level L(x,y) from the decomposition level mapping table to generate the level allocation map L. map (x,y);
[0045] Step 6: Normalize the grayscale image I VIS_norm And the smoothed infrared image I IR_smooth Wavelet decomposition is performed separately, and the decomposition process is assigned to the specified number of layers in Figure L. map (x,y) is dynamically adjusted, and the sets of low-frequency and high-frequency coefficients on the visible light side and the sets of low-frequency and high-frequency coefficients on the infrared side are obtained through wavelet decomposition.
[0046] Step 7: For the low-frequency and high-frequency coefficient sets of the visible light side and the infrared side, respectively, perform coefficient enhancement and align them according to the original decomposition level to ensure that the number of coefficient layers for each pixel is consistent.
[0047] Step 8: Assign infrared weights W based on different brightness levels. IR (x,y) and visible light weight W VIS (x,y), and the enhanced low-frequency and high-frequency coefficient sets of the infrared outer and visible light sides, are fused separately in the frequency domain, including low-frequency coefficient fusion and high-frequency coefficient fusion, and W is used in each brightness layer. IR (x,y)+W VIS (x,y)=1; and
[0048] Step 9: Based on the fused high-frequency and low-frequency coefficients, reconstruct the image spatial domain pixel by pixel using the inverse wavelet transform corresponding to the dynamic wavelet decomposition, and obtain the spatial domain infrared-visible fused image I. fusion (x,y).
[0049] Therefore, the infrared and visible light image fusion image enhancement method of the present invention, based on dynamic multi-scale layered fusion and noise adaptive suppression, realizes image fusion enhancement under complex lighting and scene backgrounds of airports. It improves the quality of fused images through layered differential enhancement, enhances the details of infrared thermal targets and dark areas, solves the problem of image fusion distortion in multiple brightness regions, and improves the recognition capabilities of subsequent intrusion target identification, target tracking, event detection, etc., especially the ability to identify and detect weak targets at low altitudes.
[0050] Image Input and Preprocessing
[0051] As an optional implementation, in the airport video and security monitoring system platform, visible light images of nighttime scenes are acquired from video frame data transmitted back by deployed infrared-visible light photoelectric perimeter scanning cameras. VIS With infrared image I IR .
[0052] For visible light image I VIS (Grayscale format, grayscale range 0-255, resolution typically 512×512 or 1024×1024, suitable for nighttime monitoring scenarios such as airports and roads), inherent noise smoothing is performed, usually using a 3×3 Gaussian smoothing filter to suppress the inherent particle noise generated by the sensor due to low light, outputting a smoothed visible light image I. VIS_smooth .
[0053] Then, the smoothed visible light image I was further processed. VIS_smooth Perform normalization and output the grayscale normalized image I. VIS_norm To eliminate the interference of lighting fluctuations on detail judgment, the grayscale range remains 0-255 after normalization, ensuring that detail features are comparable under different lighting conditions.
[0054] In this example, the visible light image I is processed using the maximum-minimum gray value method. VIS_smooth Normalize.
[0055] Meanwhile, for infrared image I IR (Single-channel grayscale image, grayscale value is positively correlated with target thermal radiation intensity, resolution is related to I) VIS (Consistent) A 3×3 mean filter is also used for smoothing, effectively reducing particle noise without affecting the thermal target outline (compared to Gaussian filtering, it has less blurring of edges), outputting a smoothed infrared image I. IR_smooth.
[0056] As an optional implementation, in step 2, the smoothed visible light image I is traversed. VIS_smooth Count the number of pixels for each grayscale value (0-255) to generate a visible light image I. VIS_smooth The brightness histogram H(g) is given, where g is the gray value, ranging from 0 to 255, and H(g) represents the total number of pixels with a gray value of g.
[0057] Building upon this, the three-peak distribution characteristics of nighttime images are further determined through histogram peak detection. For example, a sliding window method is used, with a window size of 5 gray levels. The brightness histogram is slid along the window according to a single gray level step size to determine the "dark area concentrated peak" (gray value 20-60), the "transition area gentle peak" (gray value 60-140), and the "bright area light source peak" (gray value 140-220) of the nighttime image. In the embodiments of this invention, the above-mentioned different partitions are used to characterize the differentiated layering of nighttime images in lighting scenarios such as airports. The core feature that distinguishes nighttime images from daytime "single-peak / double-peak" distributions serves as the basis for subsequent brightness layering iterative optimization.
[0058] Luminance-detail sub-layer joint labeling based on luminance layering
[0059] As an optional implementation, in step 3, based on the visible light image I VIS_smooth The brightness histogram H(g) is used to generate a brightness layering marker map M. bright (x,y) and detail density map D map (x,y); then the brightness layer marker map M is fused. bright (x,y) and detail density map D map (x,y) yields the joint labeling map M of the brightness-detail sublayer. bright−detail (x,y), specifically including the following process:
[0060] Step 3-1: Based on the visible light image I VIS_smooth The brightness histogram H(g) is used to solve for the first threshold of optimal brightness stratification using Otsu iteration with dual threshold classification. With the second threshold and based on the first threshold With the second threshold Visible light image I VIS_smooth Perform pixel-by-pixel marking to obtain a brightness layering map M, which includes dark areas, transition areas, and bright areas. bright (x,y), where:
[0061] Dark Zone: I VIS_smooth (x,y)< ;
[0062] Transition zone: ≤I VIS_smooth (x,y)< ;
[0063] Bright area: I VIS_smooth (x,y)> ;
[0064] Step 3-2: Based on the visible light image I VIS_smooth The LBP encoded image is obtained by using neighborhood LBP binary encoding with a fixed window. map (x, y), where each pixel value ranges from 0 to 255, corresponding to different local grayscale modes; then, the image is encoded using LBP. map The jump in (x,y) yields a detail density map of the non-uniform mode. map (x, y), with the same resolution as the original image, and pixel values ranging from 0 to 8; where, for multiple pixels within each fixed window range, their LBP values are determined one by one. map Whether (x,y) is a non-uniform mode is determined by counting the number of pixels in the non-uniform mode, which is then used as the detail density D(x,y) of the center pixel within a fixed window. The non-uniform mode refers to LBP. map In the binary encoding of (x,y), the number of transitions between "0-1" and "1-0" is greater than or equal to 2.
[0065] Step 3-3: According to the brightness layering marking diagram M bright (x,y) and detail density map D of the non-uniform mode map (x,y), according to the set detail density threshold D th Perform high / low detail sub-layer division and labeling.
[0066] Dynamic brightness layering and marking
[0067] In an optional embodiment, the optimal brightness layer is solved by Otsu iteration. With the second threshold The process includes:
[0068] First, set the threshold search range: T1 candidate range is 20-80 (covering dark area - transition area boundary), T2 candidate range is 120-200 (covering transition area - bright area boundary).
[0069] Then, the inter-class variance is calculated iteratively:
[0070] 1) For each candidate T1 (step size 1) and candidate T2 (step size 1, and T2>T1), divide the image into three regions:
[0071] Dark Zone R0: I VIS_smooth (x,y)<T1;
[0072] Transition region R1: T1 (≤ I VIS_smooth (x,y)< ;
[0073] Bright area R2: I VIS_smooth (x,y)>T2;
[0074] 2) Calculate the pixel proportion ω of the three types of regions. k Average brightness μ k and global average brightness μ T :
[0075] ω k = ;
[0076] μ k = ;
[0077] k=0,1,2;
[0078] μ T =ω0μ0+ω1μ1+ω2μ2;
[0079] 3) Calculate the inter-class variance =ω0×(μ0-μ T ) 2 +ω1×(μ1-μ T ) 2 +ω2×(μ2-μ T ) 2 ;
[0080] Finally, determine if the current ,but And record T1 and T2 at this time as the first threshold. With the second threshold .
[0081] In this embodiment, after traversing all candidate thresholds, the optimal threshold for the night scene is finally obtained: the first threshold. =45, second threshold =155.
[0082] Then, based on the first threshold =45, Second Threshold =155, perform pixel-by-pixel marking:
[0083] Dark area (marked 0): I VIS_smooth (x,y)< This corresponds to shadow areas without light sources at night, such as grass not covered by streetlights, boundary areas far from the track lights, and shadows at the base of buildings.
[0084] Transition region (marked 1): 45≤I VIS_smooth (x,y)<155, corresponding to areas with normal lighting, such as roads under streetlights, and the lower walls of airport terminals;
[0085] Bright area (marked 2): I VIS_smooth (x,y)>155 corresponds to the area affected by the light source, such as the area illuminated by vehicle lights, the area haloed by streetlights, and the area haloed by running lights.
[0086] Based on this, a brightness layer marker map (with the same resolution as the original image and pixel values of 0 / 1 / 2) is generated to provide region localization for subsequent differential processing.
[0087] High / low detail sub-layer partitioning and labeling
[0088] Furthermore, in an optional embodiment, the visible light image I is viewed through a fixed 3×3 window. VIS_norm Starting from the top left corner (x=1, y=1), slide pixel by pixel (step size 1) until the entire I is covered. VIS_norm Image (edge pixels are mirrored to avoid loss of edge information, and the filling range is 1 pixel of the window radius).
[0089] Get the grayscale value g of the center pixel of the current window c =I VIS_norm (x, y) and its 8 neighboring pixel gray values (sorted clockwise from the top left corner) g0~g7:
[0090] g0=I VIS_norm (x-1, y-1), g1=I VIS_norm (x-1, y), ..., g7=I VIS_norm (x+1, y-1).
[0091] For each neighboring pixel, the grayscale relationship is determined using the sign function s(z):
[0092] s(g p -g c )=1,ifg p ≥g c p = 0~7;
[0093] s(g p -g c )=0,ifg p <g c p = 0~7;
[0094] Generate 8-bit binary LBP encoding: s(g0-g c )~s(g7-g c The LBP value of the current pixel, LBP(x,y), is obtained by sequentially constructing a binary number and converting it to decimal.
[0095] LBP map (x,y)= .
[0096] For example, if the gray values of all 8 neighboring pixels are greater than or equal to the gray value of the center pixel, then:
[0097] LBP(x,y=2) 0 +2 1 +2 2 +…+2 7 =255;
[0098] If the gray values of all 8 neighboring pixels are less than the gray value of the center pixel, then LBP(x,y)=0.
[0099] Based on this, we perform pattern classification on LBP(x,y), including:
[0100] Uniform mode: The number of transitions between "0-1" and "1-0" in binary encoding is ≤2 (e.g., "00011111" transitions twice, corresponding to a smooth grayscale change, mostly noise or low detail areas).
[0101] Non-uniform mode: Number of transitions > 2 (e.g., "00101110" transitions 4 times, corresponding to complex grayscale changes, which are valuable details such as edges and textures).
[0102] As an optional implementation, we count the number of pixels in non-uniform mode (i.e., the number of pixels in the window whose LBP(x,y) belongs to non-uniform mode) in each 3×3 fixed window, and denote it as detail density D(x,y). The value range is 0-8. The more non-uniform mode pixels in the window, the larger D(x,y) is, and the richer the detail.
[0103] Statistical method: For the 9 pixels (including the center pixel) in the current 3×3 fixed window, determine whether their LBP(x,y) is non-uniform. The cumulative number of pixels that meet the condition is the detail density D(x,y) of the center pixel of the window.
[0104] Finally, as an optional implementation, the brightness layering marker map and D are combined. th Each brightness layer is further divided into detail sub-layers, specifically:
[0105] Each luminance layer is further divided into high-detail sub-layers and low-detail sub-layers, where the detail density D(x,y)≥D thHigh detail is designated as a high detail sublayer (marked as A, such as tree branches in dark areas, building edges in transition areas, road markings in bright areas, etc.), otherwise it is designated as a low detail sublayer (marked as B, such as flat grass in dark areas, solid-color walls in transition areas, light halos in bright areas, etc.).
[0106] Based on this, the joint labeling map M of the brightness-detail sublayer was obtained. bright−detail (x,y), where each pixel value is a combination of a luminance marker and a detail marker.
[0107] In this embodiment, the aforementioned detail density threshold D th Configured to 3 for optimal differentiation between valuable details and noise / low details.
[0108] Brightness markers, detail markers, and multi-feature overlay and marking of thermal targets
[0109] As an optional implementation, in step 4, based on the smoothed infrared image I IR_smooth Based on the thermal target threshold T IR Perform thermal target region segmentation to obtain thermal target binary label map M. IR (x,y); and combined with the brightness-detail sublayer joint labeling map M bright−detail (x,y) superimposed labels generate a multi-feature label map M. final (x, y), where each pixel value contains a brightness marker, a detail marker, and a feature marker for a thermal target, specifically including the following steps:
[0110] Step 4-1: Based on the set weak heat target threshold T IR =175, for the smoothed infrared image I IR_smooth Perform binary segmentation:
[0111] R IR (x,y)=1, if I IR_smooth (x,y)>175, is the target thermal region;
[0112] R IR (x,y)=0, if I IR_smooth (x,y)≤175, is a non-thermal target region;
[0113] Combined with infrared image I IR_smooth The binary segmentation result, based on R IR All pixels (x,y)=0 constitute the binary marker map M of the thermal target. IR (x,y);
[0114] Step 4-2: Mark the thermal target binary image M IR (x,y) and joint labeling map of brightness-detail sublayer M bright−detail (x,y) superimposed labels generate the final multi-feature label map M.final (x, y) contains 12 label combinations, each of which includes feature labels for a brightness layer, a detail layer, and a thermal target; the labeling method is as follows:
[0115] If M IR (x,y) corresponds to R IR If (x,y)=1, it indicates a hot target area, so add an IR suffix after the original brightness-detail marker;
[0116] If M IR (x,y) corresponds to R IR (x,y)=0 indicates a non-thermal target area, in which case the original brightness-detail marker is preserved.
[0117] In this embodiment, we verified the marking accuracy by randomly selecting 100 thermal target areas (including weakly thermal targets) – ensuring T IR =175 The labeling accuracy of weak thermal targets is ≥95%, avoiding missed detections due to excessively high thresholds (the labeling accuracy of weak targets with traditional TIR=180 is only 82%).
[0118] Dynamic wavelet decomposition
[0119] As an optional implementation, in step 5, according to the multi-feature marker map M final The three-dimensional features represented in (x,y) are used to determine the number of decomposition levels L in the dynamic wavelet decomposition, and a decomposition level mapping table is obtained; based on this, the multi-feature labeled map M is traversed. final For each pixel in (x,y), according to M final The pixel value of each pixel in (x,y) is matched with the corresponding decomposition level L(x,y) from the decomposition level mapping table to generate the level allocation map L. map (x,y), specifically including the following steps:
[0120] Step 5-1: Based on the multi-feature marker map M final The feature labels of the brightness and detail layers at (x,y) are used to dynamically adjust the wavelet decomposition L, resulting in a decomposition layer mapping table, as follows:
[0121] Dark area 0 - low detail sublayer B: L=6;
[0122] Dark area 0 - High detail sublayer A: L=5;
[0123] Transition region 1 - Low detail sublayer B: L=4;
[0124] Transition region 1 - High detail sublayer A: L=5;
[0125] Bright area 2 - Low detail sublayer B: L=3;
[0126] Bright area 2 - High detail sublayer A: L=4;
[0127] Step 5-2: Traverse the multi-feature labeled map M final For each pixel in (x,y), according to M final The pixel value of each pixel in (x,y) is matched with the corresponding decomposition layer L(x,y) from the decomposition layer number mapping table to generate a layer allocation map L. map (x,y) is used for subsequent wavelet decomposition of the infrared and visible light images, decomposing each image into a low-frequency approximation coefficient A. L and the high-frequency detail coefficients D1 to D in group L L .
[0128] As an optional implementation, in step 6, the grayscale normalized image I is... VIS_norm And the smoothed infrared image I IR_smooth Wavelet decomposition is performed separately, and the decomposition process is assigned to the specified number of layers in Figure L. map (x,y) is dynamically adjusted, and the sets of low-frequency and high-frequency coefficients on the visible light side and the set of low-frequency and high-frequency coefficients on the infrared side are obtained through wavelet decomposition, including:
[0129] Step 6-1: Initialize the current decomposition layer number l, l=1, where 1 represents the highest frequency layer, corresponding to the finest details; L represents the lowest frequency layer, corresponding to the global contour.
[0130] Step 6-2: For each pixel (x, y), if the current layer number l ≤ L(x, y), then perform Db4 wavelet transform to decompose the coefficients of the current image into: low-frequency approximation coefficients A. l With high-frequency detail factor D l :
[0131] High-frequency detail factor D l =[ , , ], , , These correspond to detailed features in the horizontal, vertical, and diagonal directions, respectively.
[0132] Step 6-3: Let the current processing coefficient be A. l Accumulate l: l = l + 1, repeat step 6-2 until l > L(x,y), at which point A L(x,y) The final low-frequency coefficients of the pixel (x,y), D1~D L(x,y) The final high-frequency coefficients are used as the set of high-frequency coefficients.
[0133] Based on the processing in steps 6-1 to 6-3 above, the grayscale normalized image I is processed respectively. VIS_normAnd the smoothed infrared image I IR_smooth Perform dynamic wavelet decomposition separately to obtain the low-frequency coefficients on the visible light side. With high frequency coefficient set and the low-frequency coefficients of the outer red side. With high frequency coefficient set .
[0134] For example, for a pixel marked 0A, five wavelet decompositions are required to obtain A5 (low-frequency coefficients) and D1 to D5 (five sets of high-frequency coefficients); for a pixel marked 2B, only three wavelet decompositions are required to obtain A3 (low-frequency coefficients) and D1 to D3 (three sets of high-frequency coefficients).
[0135] Layered high / low frequency detail enhancement
[0136] As an optional implementation, in step 7, the low-frequency and high-frequency coefficient sets of the visible light side and the infrared side are respectively enhanced and aligned according to the original decomposition level to ensure that the number of coefficient layers for each pixel is consistent. Specifically, this includes:
[0137] Step 7-1: For the high-frequency coefficient set on the visible light side, normalize the image I with grayscale. VIS_norm As a guide map, guided filtering is applied to the high-frequency coefficients on the visible light side to enhance image edges and details, thus obtaining the enhanced visible light high-frequency coefficients. ;
[0138] For the low-frequency coefficients on the visible light side, dynamic contrast enhancement is performed using layered CLAHE to obtain the CLAHE-enhanced low-frequency coefficients of the visible light. ;
[0139] The enhanced visible light high-frequency coefficients and visible light low-frequency coefficients are aligned according to the original decomposition levels to ensure that the number of coefficient layers for each pixel is consistent.
[0140] Step 7-2: For the high-frequency coefficient set on the outer infrared side, according to the binary labeling map M of the thermal target... IR(x,y) Selective attenuation is performed to output the attenuated infrared high-frequency coefficient set. ;
[0141] For the low-frequency coefficients of the infrared outer region, Gaussian smoothing is applied for filtering, and the smoothed infrared low-frequency coefficients are output. ;
[0142] Align the attenuated infrared high-frequency coefficients and the smoothed infrared low-frequency coefficients according to the original decomposition levels to ensure that the number of coefficient layers for each pixel is consistent.
[0143] High-frequency coefficient guided filtering enhancement on the visible light side
[0144] For the high-frequency coefficient set on the visible light side It is divided into a high-frequency layer (l=1,2), a mid-to-high-frequency layer (l=3,4), and a low-to-high-frequency layer (l=5,l):
[0145] High-frequency layer: corresponds to the finest details (such as the edges of small targets), and needs to be enhanced in detail;
[0146] Mid-to-high frequency layer: for medium-level details (such as building edges and road markings), moderately enhanced;
[0147] Low-frequency and high-frequency layers: corresponding to coarse and fine details (such as the overall outline of the target), slightly enhanced to avoid noise amplification.
[0148] Accordingly, the size of the steering filter window, the regularization parameter λ, and the detail gain coefficient k are further determined by combining the high-detail sub-layer and the low-detail sub-layer. g :
[0149]
[0150] Therefore, the smaller the regularization parameter λ, the stronger the enhancement (high detail layers require strong enhancement, λ=0.01), and the larger the window, the better the smoothness (e.g., for low detail layer denoising, the window size is 7×7). The detail gain coefficient k... g It is positively correlated with detail density.
[0151] For each high-frequency coefficient on the visible light side Image I, l=1~L, is normalized to grayscale. VIS_norm As a guide, the filtered coefficients are obtained by solving an optimization problem. Achieve edge preservation and detail enhancement:
[0152]
[0153] Among them, Ω x,y This represents the filter window for the current pixel (x, y), with a size of 5×5 or 7×7; a x,y and b x,y The limiting coefficients within the respective windows are used to control gain intensity and brightness offset, respectively; λ represents the regularization parameter, with values of 0.01, 0.02, and 0.03, as shown in the table above, used to prevent a x,y Excessive gain can lead to uncontrolled gains.
[0154] Using existing linear solution methods, the mean, variance, and high-frequency coefficients of the guide graph within the window are calculated. Given the mean and covariance of the two, solve for a. x,y and b x,y .
[0155] Then the filtered high-frequency coefficients can be obtained. (x,y)=a x,y ×IVIS_norm (x,y)+b x,y ;l=1~L.
[0156] Among them, in the edge region a x,y ≈1, retaining details; in flat areas, a x,y ≈0, smooth noise.
[0157] Furthermore, the filtered coefficients are multiplied by the corresponding detail gain coefficient k. g To obtain the enhanced high-frequency coefficient :
[0158] = (x,y)×k g ;l=1~L.
[0159] CLAHE enhancement of low-frequency coefficients on the visible light side
[0160] In this embodiment, for the low-frequency coefficient on the visible light side Dynamic contrast enhancement was achieved through improved layered CLAHE, resulting in the CLAHE-enhanced low-frequency coefficient of visible light. .
[0161] First, from the multi-feature label map M final Extract brightness layer markers (0 = dark area, 1 = transition area, 2 = bright area) from (x,y) to generate a brightness layer mapping map M. bright_only (x,y);
[0162] Then, based on the luminance layer map, the key parameters for CLAHE enhancement are set, including the contrast cap clipLimit, sub-block size, and sub-block overlap rate, as follows:
[0163]
[0164] Among them, the smaller the clipLimit, the gentler the contrast enhancement (small value is needed for dark areas); the higher the overlap rate, the more natural the transition of sub-block boundaries (50% overlap rate in dark areas to avoid boundary effects); set the bright area sub-blocks to 16×16 to reduce overexposure in small areas.
[0165] As shown in the table design above, the low-frequency coefficients are set according to the defined sub-block size. normalization results The blocks are divided into sections, and edge sub-blocks that are not large enough are filled with mirror images to avoid losing edge information.
[0166] First, for each block, count the number of pixels with gray values from 0 to 255 and generate a sub-block histogram h(g); g represents the gray value.
[0167] Then, the sub-block histograms are clipped and redistributed to obtain the clipped sub-block histogram h. clip (g), where:
[0168] Clipping threshold T clip =clipLimit×(number of sub-block pixels / 256), meaning that the histogram portion exceeding this value is cropped to avoid overexposure;
[0169] The total number of pixels N cropped clip = ;
[0170] Then, the total number of pixels N that were cropped clip Distribute evenly across all gray levels, i.e., increase N for each gray level. clip / 256, obtain the clipped histogram h. clip (g)
[0171] Furthermore, the cumulative distribution function CDF(g) of the histogram of the clipped sub-blocks is calculated, CDF(g) = Based on this, grayscale mapping is performed to enhance contrast:
[0172] [ ( ] / [ - ]×255;
[0173] in, and Let represent the maximum and minimum values of the cumulative distribution function CDF(g), respectively.
[0174] After fusing the enhancement results of adjacent sub-blocks using bilinear interpolation, the global mean of the low-frequency coefficients before and after enhancement is calculated. If the global mean exceeds a set threshold, such as 10, it indicates that the brightness change is too large and affects visual naturalness. In this case, linear adjustment calibration is performed to obtain the calibrated visible light low-frequency coefficients. The details are as follows:
[0175] = ;
[0176] In the formula, These represent the global mean values of the low-frequency coefficients before and after enhancement, respectively.
[0177] Selective attenuation of infrared high-frequency coefficient
[0178] For the high-frequency coefficient set on the outer side of the red line , l= In non-thermal target area (M) IRThe infrared high-frequency coefficient of (x,y)=0 is multiplied by the attenuation coefficient α=0.5 (the infrared high frequency in non-thermal target areas is mostly particle noise, and attenuation reduces interference); for thermal target areas (M IR The infrared high-frequency coefficients of (x,y)=1 remain unchanged (preserving details of thermal targets, such as the thermal characteristics of UAV propellers):
[0179] ×0.5,ifM IR (x,y)=0;
[0180] );ifM IR (x,y)=1;
[0181] Based on this, the set of infrared high-frequency coefficients after selective attenuation is output. .
[0182] Infrared low-frequency coefficient Gaussian smoothing
[0183] In this embodiment, a 5×5 Gaussian filter is used to smooth the infrared low-frequency coefficients. σ=0.5, ensuring smoothing of particle noise with a small standard deviation without blurring the thermal target outline; outputting the smoothed infrared low-frequency coefficient. .
[0184] Dynamic weight optimization and frequency-based fusion
[0185] As an optional implementation, in step 8, the infrared weights W are determined according to different brightness levels. IR (x,y) and visible light weight W VIS (x,y), and the enhanced sets of low-frequency and high-frequency coefficients on the infrared outer and visible light sides, are fused in the frequency domain, including low-frequency coefficient fusion and high-frequency coefficient fusion, specifically including the following steps:
[0186] Step 8-1: Based on the basic infrared weights, detail adjustment terms, and thermal target gain terms in different brightness layers, determine the infrared fusion weight W for each pixel in different brightness layers. IR (x,y) and visible light weight W VIS (x,y), visible light weight W VIS (x,y)=1-W IR (x,y) generates an infrared weighted map W. IR_map With visible light weighting map W VIS_map Both have the same resolution as the original image.
[0187] Step 8-2: Based on the low-frequency coefficients of the enhanced infrared outer side and the visible light side, and the infrared weighting map W... IR_map With visible light weighting map W VIS_mapPerform low-frequency coefficient fusion and output the low-frequency fusion coefficients representing the global brightness and contour:
[0188] =W VIS (x,y)× +W IR (x,y)× ;
[0189] Step 8-3: Based on the enhanced high-frequency coefficient sets of the infrared outer side and the visible light side, the infrared weighting map W... IR_map With visible light weighting map W VIS_map Perform high-frequency coefficient fusion and output the high-frequency fusion coefficients representing global details and the furnace perimeter.
[0190] =W VIS (x,y)× +W IR (x,y)× ;
[0191] Where l = 1, 2, 3, ..., L.
[0192] In embodiments of the present invention, dynamic weight allocation is performed according to the following principles:
[0193] Dark areas (marked with 0 at the beginning): More visible light noise and less detail; the greater the noise, the higher the infrared weight; the richer the detail, the higher the visible light weight.
[0194] Transition zone (marked with 1): Visible light and infrared information are balanced, and weights are allocated according to local contrast. Areas with rich details are biased towards visible light.
[0195] Bright areas (marked with 2): These areas have more visible light details and more concentrated noise. The lower the noise, the higher the visible light weight. Areas with rich details further enhance the visible light weight.
[0196] In an optional embodiment, the dark area weight calculation is processed as follows (M final The markers contain "0": 0A, 0B, 0A-IR, 0B-IR).
[0197] ;
[0198] Where, N vis_norm (x,y) represents the high-frequency coefficients of the enhanced visible light. The normalized value of the local standard deviation represents the local noise index; K(x,y) represents the detail density factor, K(x,y) = D. map (x,y) / 8; S(x,y) represents the thermal target factor, S(x,y)=M IR (x,y), takes the value 1 or 0.
[0199] In an optional embodiment, the transition region weight calculation is processed as follows (M final (The markings include "1": 1A, 1A-IR, 1B-IR, etc.)
[0200] ;
[0201] Among them, C IR (x,y), C VIS (x,y) represents the local contrast between the infrared outer side and the visible light side, determined by the gradient magnitude:
[0202] C IR (x,y)= ;
[0203] C VIS (x,y)= ;
[0204] In an optional embodiment, the bright area weight calculation is processed as follows (M final (The markings contain "2": 2A, 2A-IR, etc.)
[0205] .
[0206] Inverse wavelet transform image reconstruction
[0207] As an optional implementation, in step 9, based on the fused high-frequency and low-frequency coefficients, the image spatial domain is reconstructed pixel by pixel according to the original decomposition level by performing an inverse wavelet transform corresponding to the dynamic wavelet decomposition, thereby obtaining the spatial domain infrared-visible fused image I. fusion (x,y), specifically including the following steps:
[0208] Step 9-1, from the lowest frequency coefficient Begin with the corresponding high-frequency coefficients Perform inverse wavelet transform to obtain the coefficients of layer l=L−1;
[0209] Step 9-2: Repeat step 9-1 above until l=1, to obtain the spatial domain infrared-visible fusion image I. fusion (x,y);
[0210] Step 3: Reconstruct the infrared-visible fused image I fusion (x,y) is normalized to a grayscale range of 0 to 255, and the output is I. fusion_norm (x,y):
[0211] I fusion_norm (x,y)=[I fusion (x,y)-min(Ifusion )] / [max(I fusion )-min(I fusion )]×255
[0212] Where max(I) fusion ), min(I fusion (I) represents the infrared-visible fusion image I. fusion The maximum and minimum gray values of (x,y).
[0213] Therefore, combining the infrared and visible light image processing processes of the above embodiments, in the image fusion process, considering the complex characteristics of nighttime image background and illumination, and based on the characteristics that dark areas in the same image rely more on infrared, bright areas rely more on visible light, detail areas retain more visible light, and noise areas rely more on infrared, a detail density-noise dual factor is introduced to construct a weighting mechanism that dynamically adjusts with the local characteristics of the image. For different brightness layers, the infrared fusion weight in different brightness layers is determined pixel by pixel based on the basic infrared weight, detail adjustment term, and thermal target gain term for each brightness. Then, fusion is performed in the low-frequency domain and the high-frequency domain respectively. The low-frequency domain fuses global brightness and contour information to ensure that the overall brightness of the image is balanced (dark areas are not too dark, and bright areas are not overexposed), which conforms to the visual habits of the human eye at night; while the high-frequency domain fuses details and edges to ensure that target details are clear (such as road markings, running guide lines, perimeter fence outlines, etc.) to meet the subsequent recognition requirements.
[0214] {Example 2}
[0215] In conjunction with the implementation of the dynamic multi-scale layered fusion infrared and visible light image enhancement method of the above embodiments, the present invention also proposes a computer system, comprising:
[0216] One or more processors;
[0217] Memory stores instructions that can be operated.
[0218] When the instructions are executed by the one or more processors, the one or more processors perform an operation, which includes the flow of the dynamic multi-scale hierarchical fusion infrared and visible light image enhancement method of the foregoing embodiments.
[0219] {Example 3}
[0220] In conjunction with the implementation of the image enhancement method for dynamic multi-scale hierarchical fusion of infrared and visible light images in the above embodiments, the present invention also proposes a computer-readable medium for storing software, the software including instructions executable by one or more computers, the instructions causing the one or more computers to perform an operation, the operation including the flow of the image enhancement method for dynamic multi-scale hierarchical fusion of infrared and visible light images in the foregoing embodiments.
[0221] While the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the invention. Those skilled in the art can make various modifications and refinements without departing from the spirit and scope of the invention. Therefore, the scope of protection of the present invention shall be determined by the claims.
Claims
1. A dynamic multi-scale hierarchical fusion infrared and visible image fusion image enhancement method, characterized in that, Comprising: Step 1, obtain a visible light image I VIS with the infrared image I IR input, and respectively perform inherent noise smoothing, output the smoothed visible light image I VIS_smooth with the infrared image I IR_smooth , to the smoothed visible light image I VIS_smooth perform normalization, output a gray scale normalized image I VIS_norm ; Step 2, traverse the smoothed visible light image I VIS_smooth , count the number of pixels of each gray value, and generate a luminance histogram H(g) of the visible light image I VIS_smooth , where g is a gray value, taking a value of 0-255, and H(g) represents the total number of pixels with a gray value of g. Step 3, generating a luminance layered mark map M VIS_smooth from the luminance histogram H(g) of the visible light image I bright (x,y) and a detail density map D map (x,y); then fusing the luminance layered mark map M bright (x,y) and the detail density map D map (x,y) to obtain a luminance-detail sublayer joint mark map M bright−detail (x,y) with each pixel value combining a luminance mark and a detail mark. Step 4, obtaining a thermal target binary mask M IR_smooth (x,y) based on the smoothed infrared image I IR (x,y) by thermal target region segmentation with a thermal target threshold T IR (x,y) and superimposing it with the brightness-detail sub-layer joint mask M bright−detail (x,y) to generate a multi-feature mask M final (x,y) where each pixel value contains the brightness mask, the detail mask and the thermal target feature mask. Step 5, according to the multi-feature mark graph M final (x,y) in the three-dimensional features embodied in, determine the decomposition level L of dynamic wavelet decomposition, obtain the decomposition level mapping table; and according to this, traverse the multi-feature mark graph M final (x,y) in each pixel, according to M final (x,y) in each pixel, the pixel value of each pixel is matched with the corresponding decomposition level L(x,y) from the decomposition level mapping table, and a level number allocation graph L is generated map (x,y) Step 6, normalizing the gray scale image I VIS_norm and the smoothed infrared image I IR_smooth performing wavelet decomposition respectively, the decomposition process allocating the image L according to the number of layers map (x, y) dynamically adjusting, obtaining the low frequency coefficient and high frequency coefficient set of the visible light side, and the low frequency coefficient and high frequency coefficient set of the infrared side through wavelet decomposition; Step 7, for the low-frequency coefficient and high-frequency coefficient set of the visible light side and the infrared side, respectively, the coefficient enhancement is carried out, and the original decomposition level is aligned to ensure that the coefficient layer number of each pixel is consistent; Step 8, Infrared weight W according to different luminance layers IR (x,y) and visible light weight W VIS (x,y), and the enhanced infrared side and visible light side low frequency coefficient and high frequency coefficient set, respectively, are fused in frequency domain, including low frequency coefficient fusion and high frequency coefficient fusion, W IR (x,y) + W VIS (x,y) = 1; And Step 9, based on the fused high frequency coefficients and low frequency coefficients, through the wavelet inverse transform operation corresponding to the dynamic wavelet decomposition, the image space domain is recovered by pixel by pixel reconstruction according to the original decomposition layer, and the spatial domain infrared-visible light fusion image I is obtained fusion (x,y).
2. The dynamic multi-scale hierarchical fusion infrared and visible image fusion image enhancement method according to claim 1, characterized in that, In step 3, based on the visible light image I VIS_smooth The brightness histogram H(g) is used to generate a brightness layering marker map M. bright (x,y) and detail density map D map (x,y); then the brightness layer marker map M is fused. bright (x,y) and detail density map D map (x,y) yields the joint labeling map M of the brightness-detail sublayer. bright−detail (x,y), specifically including the following process: Step 3-1, from the visible light image I VIS_smooth a luminance histogram H(g) and iteratively solving a first threshold of optimal luminance stratification with a bimodal threshold classification Otsu and a second threshold and based on the first threshold and the second threshold performing a pixel-wise labeling of the visible light image I VIS_smooth resulting in a luminance stratification labeled map M bright (x,y), wherein: Dark region: I VIS_smooth (x, y) < 0 ; Transition zone: ≤I VIS_smooth (x,y) ; bright region: I VIS_smooth (x, y) ; Step 3-2, obtaining a LBP coded image LBP VIS_smooth from the visible light image I map (x,y) using a fixed window neighborhood LBP binary coding map (x,y) based on the transitions in the LBP coded image LBP map (x,y), wherein for each of the plurality of pixels within the fixed window range, it is determined one by one whether its LBP value LBP map (x,y) is a non-uniform pattern, and the number of pixels with the non-uniform pattern is counted as the detail density D(x,y) of the center pixel within the fixed window range, the non-uniform pattern referring to the number of transitions of "0-1" and "1-0" in the binary coding of LBP map (x,y) being greater than or equal to 2 times. Step 3-3, marking the image M according to the luminance layering bright (x,y) and the detail density map D of the non-uniform pattern map (x,y) according to the set detail density threshold D th (x,y) according to the set detail density threshold D th (x,y) according to the set detail density threshold D bright−detail (x,y) according to the set detail density threshold D 3. The dynamic multi-scale hierarchical fusion infrared and visible image fusion image enhancement method according to claim 2, characterized in that, In step 4, according to the smoothed infrared image I IR_smooth , the thermal target region segmentation is performed based on the thermal target threshold T IR , to obtain a thermal target binary mask M IR (x, y); and the mask is superimposed with the brightness-detail sub-layer joint mask M bright−detail (x, y) to generate a multi-feature mask M final (x, y), wherein each pixel value contains the brightness mask, the detail mask and the feature mask of the thermal target, including the following steps: Step 4-1, according to the set weak heat target threshold T IR = 175, on the smoothed infrared image I IR_smooth perform binary segmentation: R IR (x,y) = 1, if I IR_smooth (x,y) > 175, hot target region; R IR (x,y) = 0, if I IR_smooth (x,y) < 175, for non-hot target regions; In combination with the binary segmentation result of the infrared image I IR_smooth , all pixels of R IR (x,y)=0 constitute a hot target binary label map M IR (x,y). Step 4-2, binary marking map M for hot target IR (x,y) and the luminance-detail sublayer joint marking map M bright−detail (x,y) and the luminance-detail sublayer joint marking map M final (x,y), which contains 12 kinds of marking combinations, each of which includes the feature marking of the luminance layer + the detail layer + the hot target; wherein the superimposed marking is as follows: If M IR (x,y) corresponds to R IR (x,y)=1, is marked as a hot target region, then an IR suffix is added after the original brightness-detail marker; if M IR (x,y) corresponds to R IR (x,y)=0, is marked as a non-hot target region, then the original brightness-detail marker is kept.
4. The dynamic multi-scale hierarchical fusion infrared and visible image fusion image enhancement method according to claim 2, characterized in that, In step 5, according to the multi-feature mark graph M final (x,y) in the three-dimensional features embodied in, determine the decomposition level L of dynamic wavelet decomposition, get the decomposition level mapping table; and according to the traversal of multi-feature mark graph M final (x,y) in each pixel, according to M final (x,y) in each pixel of the pixel value from the decomposition level mapping table matching corresponding decomposition level L(x,y), generate the level assignment graph L map (x,y), specifically comprising the following steps: Step 5-1, according to the multi-feature mark map M final (x, y) of the luminance layer and the detail layer of the feature mark, dynamically adjusting the decomposition L of the wavelet decomposition, obtaining a decomposition layer mapping table, and the details are as follows: Dark area-low detail sublayer: L=6; Dark area-high detail sublayer: L=5; Transition area-low detail sublayer: L=4; Transition area-high detail sublayer: L=5; Bright area-low detail sublayer: L=3; Bright area-high detail sublayer: L=4; Step 5-2, traversing the multi-feature mark graph M final (x,y), according to M final (x,y), from the decomposition level number mapping table, matching the corresponding decomposition level number L(x,y) of each pixel value in (x,y), to generate a level number assignment graph L map (x,y), for subsequent wavelet decomposition of the infrared image and the visible light image, decomposing the infrared image and the visible light image into 1 low-frequency approximation coefficient A L and L groups of high-frequency detail coefficients D1~D L .
5. The dynamic multi-scale hierarchical fusion infrared and visible image fusion image enhancement method according to claim 2, characterized in that, In step 6, the gray scale normalized image I VIS_norm and the smoothed infrared image I IR_smooth Wavelet decomposition is performed respectively, and the decomposition process allocates the image L map (x, y) is dynamically adjusted, and the low frequency coefficient and high frequency coefficient sets of the visible light side and the low frequency coefficient and high frequency coefficient sets of the infrared side are obtained through wavelet decomposition, including: Step 6-1, initialize the current decomposition layer number l, l=1, 1 represents the highest frequency layer, corresponding to the finest detail; L represents the lowest frequency layer, corresponding to the global contour; Step 6-2, for each pixel (x, y), if the current layer number l < L(x, y), perform Db4 wavelet transform, decompose the coefficients of the current image into: low frequency approximation coefficients A l and high frequency detail coefficients D l high frequency detail coefficients D l [ , , ], , , corresponding to horizontal, vertical, diagonal directional detail features, respectively Step 6-3, let the current processing coefficient be A l , increment 1: 1 = 1 + 1, repeat step 6-2 until 1 > L(x, y), at which time A L(x,y) is the final low frequency coefficient for the pixel (x, y), D1 ~ D L(x,y) are final high frequency coefficients, as the high frequency coefficient set; According to the processing of steps 6-1 to 6-3 above, the gray scale normalized image I VIS_norm and the smoothed infrared image I IR_smooth are respectively subjected to dynamic wavelet decomposition, to obtain the low frequency coefficient set and the high frequency coefficient set of the visible light side, and the low frequency coefficient set and the high frequency coefficient set of the infrared side, respectively.
6. The dynamic multi-scale hierarchical fusion infrared and visible image fusion image enhancement method according to claim 1, characterized in that, In the step 7, for the low-frequency coefficient and high-frequency coefficient set of the visible light side and the infrared side, respectively, the coefficient enhancement is carried out, and the original decomposition level is aligned to ensure that the coefficient layer number of each pixel is consistent, specifically comprising: Step 7-1, for the high frequency coefficient set of the visible light side, normalize the image I with gray scale VIS_norm As a guide map, the high frequency coefficient of the visible light side is guided filtering, the image edge and detail are enhanced, and the enhanced visible light high frequency coefficient is obtained ; For the low-frequency coefficients of the visible light side, dynamic contrast enhancement is performed by layered CLAHE to obtain the CLAHE-enhanced low-frequency coefficients of the visible light ; The enhanced visible light high-frequency coefficient and visible light low-frequency coefficient are aligned according to the original decomposition level to ensure that the coefficient layer number of each pixel is consistent; Step 7-2, for the high frequency coefficient set of the infrared side, according to the hot target binary label map M IR(x,y) selective attenuation is performed, and the attenuated infrared high frequency coefficient set is output ; For the low frequency coefficient of the infrared side, filtering processing is performed through Gaussian smoothing, and the smoothed infrared low frequency coefficient is output ; The attenuated infrared high-frequency coefficient and the smoothed infrared low-frequency coefficient are aligned according to the original decomposition level to ensure that the coefficient layer number of each pixel is consistent.
7. The dynamic multi-scale hierarchical fusion infrared and visible image fusion image enhancement method according to claim 1, characterized in that, In the step 8, the infrared weight W according to different brightness layers IR (x,y) and the visible light weight W VIS (x,y), and the enhanced infrared side and visible light side low-frequency coefficient and high-frequency coefficient set are respectively subjected to frequency domain fusion, including low-frequency coefficient fusion and high-frequency coefficient fusion, specifically including the following steps: Step 8-1, according to the base infrared weight, the detail adjustment term and the thermal target gain term in different luminance layers, determine the infrared fusion weight W in different luminance layers pixel by pixel IR (x,y) and the visible light weight W VIS (x,y), the visible light weight W VIS (x,y)=1-W IR (x,y), generate the infrared weight map W IR_map and the visible light weight map W VIS_map , both have the same resolution as the original image resolution; Step 8-2, perform low frequency coefficient fusion according to the enhanced low frequency coefficients of the infrared side and the visible light side, and the infrared weight map W IR_map and the visible light weight map W VIS_map to output low frequency fusion coefficients representing global brightness and contour: = W VIS (x, y) x + W IR (x, y) x ; Step 8-3, according to the enhanced high frequency coefficient set of the infrared side and the visible light side, the infrared weight map W IR_map and the visible light weight map W VIS_map perform high frequency coefficient fusion, and output high frequency fusion coefficients representing global details: = W VIS (x, y) x + W IR (x, y) x ; Wherein, l=1, 2, 3, …, L.
8. The dynamic multi-scale hierarchical fusion infrared and visible image fusion image enhancement method according to claim 7, characterized in that, In the step 9, according to the fused high frequency coefficients and low frequency coefficients, the image space domain is recovered by inverse wavelet transform corresponding to the dynamic wavelet decomposition, and the spatial domain infrared-visible light fusion image I is obtained. fusion (x,y), specifically comprising the following steps: Step 9-1, from the lowest frequency coefficient Start with the corresponding high frequency coefficient Perform inverse wavelet transform to obtain the coefficients of the l = L - 1 layer; Step 9-2, repeat the above step 9-1 until l = 1, to get the spatial domain infrared-visible light fusion image I fusion (x,y); Step 3, the reconstructed infrared-visible light fusion image I fusion (x,y) is normalized to the gray scale range of 0-255, and I is output fusion_norm (x,y): I fusion_norm (x,y)=[I fusion (x,y)-min(I fusion )] / [max(I fusion )-min(I fusion )]×255; where max(I fusion ), min(I fusion ) represent the maximum and minimum gray value of the infrared-visible light fusion image I fusion (x, y), respectively.
9. A computer system, characterized by Comprising: One or more processors; A memory storing instructions operable to, when executed by the one or more processors, cause the one or more processors to perform operations comprising the flow of the dynamic multi-scale hierarchical fusion infrared and visible light image fusion image enhancement method as claimed in any one of claims 1-8.
10. A computer readable medium storing software, characterized in that, The software comprises instructions executable by one or more computers, which, through such execution, cause the one or more computers to perform operations comprising the flow of the dynamic multi-scale hierarchical fusion infrared and visible light image fusion image enhancement method as claimed in any one of claims 1-8.