Cross-spectral fusion and recognition optimization method under night low infrared contrast conditions

By calculating the confidence weights of infrared and visible light images for matching and fusion, the problem of poor fusion of infrared and visible light images under low temperature difference or low illumination conditions is solved, and efficient target recognition effect is achieved.

CN121685285BActive Publication Date: 2026-05-12XIAN TIANMAO DIGITAL TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
XIAN TIANMAO DIGITAL TECH CO LTD
Filing Date
2026-02-10
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Existing technologies cannot effectively fuse infrared and visible light images, resulting in poor target recognition at night. Especially under conditions of low temperature difference or low illumination, ghosting or shadows appear at the edges of the images, and they are easily affected by noise in areas lacking clear texture features.

Method used

By acquiring coordinate point information from infrared and visible light images, confidence weights are calculated, image matching and fusion are performed based on the confidence weights, spatial reconstruction is carried out using the initial geometric correction displacement field, and target detection network is combined for recognition.

Benefits of technology

Accurate image fusion under low infrared contrast conditions is achieved, target information is preserved, the accuracy of target recognition is improved, and interference from noise and halo artifacts is avoided.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121685285B_ABST
    Figure CN121685285B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of image fusion, in particular to a cross-spectrum fusion and recognition optimization method under low infrared contrast conditions at night. The method is based on the edge strength of coordinate points on the infrared image, represents motion consistency based on the motion information reflected on the infrared image, and combines specific brightness attributes to obtain the confidence weight of each coordinate point. Based on the matching process, the initial geometric correction displacement field can be constructed by the reference points in the infrared image to the optimal matching points in the visible light image, filtering operations are performed in the vertical and horizontal directions to obtain a smooth geometric correction displacement field for spatial reconstruction, the feature information is filtered through the weighted fusion of the spatial correction thermal imaging image after spatial reconstruction, and a fusion image is obtained, and accurate and effective target recognition can be realized based on the fusion image. The present application improves the accuracy of night target recognition through effective image fusion.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image fusion technology, specifically to a cross-spectral fusion and recognition optimization method under low infrared contrast conditions at night. Background Technology

[0002] In applications such as nighttime security monitoring, assisted driving, and search and rescue, ambient lighting conditions are typically poor. To obtain complete target information, cross-spectral image fusion of infrared thermal imaging and low-light visible light imaging is often employed. Infrared thermal imaging, based on the thermal radiation difference of the target, can penetrate darkness and smoke. However, at long distances or under low temperature differences, the target outline is often blurred and diffused due to thermal diffusion effects, and it lacks texture details. Low-light visible light imaging, based on reflected light imaging, can capture scene textures. However, in low-light environments, the image is highly susceptible to high-frequency photon noise, and it produces bright halo artifacts at strong light sources such as vehicle headlights and streetlights.

[0003] Due to the fundamental differences in imaging mechanisms, the same physical target appears significantly differently in infrared and visible light images. For example, a thermal target in an infrared image typically appears as a diffuse heat patch, while the corresponding target in a visible light image may appear as a noisy spot or be covered by a halo. This phenomenological parallax prevents existing techniques from directly using rigid transformations to achieve precise alignment. This is because global rigid registration cannot correct non-rigid contour deviations caused by thermal diffusion effects, resulting in ghosting or blurring at the edges of the fused image. Furthermore, during the registration process, existing techniques are highly susceptible to convergence to incorrect extreme points in flat or noisy regions lacking clear texture features, leading to image distortion or tearing. This, in turn, results in poor quality fused images and interferes with subsequent target recognition. Summary of the Invention

[0004] To address the technical problem of existing technologies failing to effectively fuse infrared and visible light images, thus interfering with target recognition, the present invention aims to provide a cross-spectral fusion and recognition optimization method under low infrared contrast conditions at night. The specific technical solution adopted is as follows:

[0005] This invention proposes a cross-spectral fusion and recognition optimization method under low infrared contrast conditions at night, the method comprising:

[0006] Image information of the target area at night is obtained, including infrared images and visible light images; for any coordinate point in the target area, the confidence weight of each coordinate point is obtained based on the edge intensity of the coordinate point in the infrared image, the brightness in the visible light image, and the disorder of motion information between adjacent frames in the infrared image.

[0007] Using the coordinate points in the infrared image as reference points, the optimal matching point is searched in the visible light image. The matching process of the optimal matching point requires comparing the differences in motion features and edge information reflected by the coordinate points in the image information, and combining the confidence weight of the reference point to obtain the matching cost and perform the matching.

[0008] An initial geometric correction displacement field is constructed based on the coordinate deviations between all reference points and the corresponding optimal matching points; after performing horizontal and vertical filtering on the initial geometric correction displacement field, a smoothed geometric correction displacement field is obtained; the infrared image is spatially reconstructed based on the smoothed geometric correction displacement field to obtain a spatially corrected thermal image.

[0009] Based on the confidence weights of coordinate points, the spatially corrected thermal image and the visible light image are fused to obtain a fused image; the target detection network trained using the fused image and the confidence weights is used for nighttime target recognition.

[0010] Furthermore, the edge intensity is the gradient magnitude of the coordinate point on the infrared image.

[0011] Furthermore, the method for obtaining the motion information disorder includes:

[0012] Based on the image information between adjacent frames, an infrared velocity field is obtained; in the infrared velocity field, the average velocity vector within a preset neighborhood centered on the coordinate point is statistically analyzed, and the sum of squares is calculated based on the Euclidean distance between the average velocity vector and the velocity vector of the coordinate point to obtain the motion information disorder.

[0013] Further, obtaining the confidence weight for each coordinate point includes:

[0014] Based on the brightness information, the visible light image is divided into high-brightness and low-brightness regions. If the coordinate point is in the high-brightness region and the gray-scale gradient amplitude in the visible light image is less than the preset texture threshold, then the strong light interference flag value is set to 1; if the coordinate point is in the low-brightness region, then the strong light interference flag value is set to 0.

[0015] The confidence weight is obtained based on the strong light interference identifier value, the edge intensity, and the motion information disorder.

[0016] Further, obtaining the confidence weight based on the strong light interference identifier value, the edge intensity, and the motion information disorder includes:

[0017] For any coordinate point, the motion information disorder is compared with the overall motion information disorder of all coordinate points to obtain the disorder significance; the disorder significance is negatively correlated and normalized to obtain motion stability; the positive integer 1 minus the strong light interference identifier value is used as the interference masking term of the coordinate point; the product of the edge intensity, the motion stability and the interference masking term is used as the confidence weight.

[0018] Furthermore, the search for the optimal matching point in the visible light image includes:

[0019] In a visible light image, a search region is constructed with the coordinates corresponding to the reference point as the center and a preset size is used. The pixels in the search region of the visible light image are the points to be matched with the reference point. The matching cost between the reference point and each point to be matched is obtained, and the point to be matched with the minimum matching cost is selected as the optimal matching point.

[0020] Furthermore, the method for obtaining the matching cost includes:

[0021] For any point to be matched, obtain the lateral coordinate difference and the longitudinal coordinate difference between the point to be matched and the reference point; obtain the displacement constraint term based on the lateral coordinate difference and the longitudinal coordinate difference.

[0022] The edge information difference is obtained based on the absolute value of the cosine similarity between the gradient vector of the reference point in the infrared image and the gradient vector of the point to be matched in the visible light image; the infrared velocity field and the visible light velocity field are obtained based on the image information between adjacent frames, and the motion feature difference is obtained based on the difference between the velocity vector of the reference point in the infrared velocity field and the velocity vector of the point to be matched in the visible light image; the edge information difference and the motion feature difference are weighted and fused to obtain the feature information difference.

[0023] The confidence weight is used as the weight of the feature information difference, and the result of negative correlation mapping of the confidence weight is used as the weight of the displacement constraint term. The feature information difference and the displacement constraint term are weighted and summed to obtain the matching cost.

[0024] Furthermore, the method for acquiring the spatially corrected thermal image includes:

[0025] For any pixel in any infrared image, a translation search is performed based on the displacement features corresponding to the smooth geometric correction displacement field to obtain the source coordinates of the pixel. Then, the pixel values ​​within a preset neighborhood range are interpolated to obtain the replacement pixel value of the pixel.

[0026] For all pixels in the infrared image, the replacement pixel values ​​are used to replace the original pixel values ​​to obtain the spatially corrected thermal image.

[0027] Furthermore, the method for obtaining the fused image includes:

[0028] The spatially corrected thermal image and the visible light image are fused according to a preset fusion weight to obtain an initial fused image;

[0029] For each coordinate point, the confidence weight is used as the weight of the pixel value in the initial fused image, and the result of the negative correlation mapping of the confidence weight is used as the weight of the pixel value in the spatially corrected thermal image. The fused pixel value of each pixel in the fused image is obtained by weighted summation, thus obtaining the fused image.

[0030] Furthermore, the object detection network includes three visual channels and one physical attention channel. The visual channels are filled with the same fused image, and the physical attention channel is filled with a confidence weight distribution map formed by the confidence weights of all coordinate points.

[0031] The present invention has the following beneficial effects:

[0032] The goal of this invention in image fusion is to preserve as much important information as possible about the target to be identified, thereby facilitating subsequent target recognition. Therefore, to differentiate it from traditional algorithms that only calculate grayscale similarity between infrared and visible light images, this invention uses the edge intensity of coordinate points in the infrared image to represent structural saliency, and the motion information reflected in the infrared image to represent motion consistency. Combined with specific brightness attributes, a confidence weight is calculated for each coordinate point. This confidence weight represents the informational importance of the coordinate point in the image, ensuring that effective matching is performed only for targets with significant structure and stable motion in the infrared image during subsequent matching. Furthermore, the constraint of brightness attributes prevents infrared targets from being incorrectly matched to false optical artifacts. Based on the matching process, an initial geometric correction displacement field can be constructed by using the reference point in the infrared image to the optimal matching point in the visible light image. In order to ensure the spatial topological continuity of the reconstructed image, filtering operations are performed in the vertical and horizontal directions to obtain a smooth geometric correction displacement field for spatial reconstruction. The spatially corrected thermal image after spatial reconstruction can smoothly shrink and align infrared thermal targets with high image feature confidence to the visible light texture skeleton. Through weighted fusion operation, feature information is further filtered to obtain a fused image with clear moving target features. Based on this fused image, accurate and effective target recognition can be achieved. Attached Figure Description

[0033] To more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0034] Figure 1 This is a flowchart of a cross-spectral fusion and recognition optimization method under low infrared contrast conditions at night, provided as an embodiment of the present invention. Detailed Implementation

[0035] To further illustrate the technical means and effects adopted by the present invention to achieve its intended purpose, the following, in conjunction with the accompanying drawings and preferred embodiments, details the specific implementation, structure, features, and effects of a cross-spectral fusion and identification optimization method under low infrared contrast conditions at night, proposed according to the present invention. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. Furthermore, specific features, structures, or characteristics in one or more embodiments can be combined in any suitable form.

[0036] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.

[0037] The following description, in conjunction with the accompanying drawings, details a specific scheme for a cross-spectral fusion and recognition optimization method under low infrared contrast conditions at night, provided by the present invention.

[0038] Please see Figure 1 The diagram illustrates a flowchart of a cross-spectral fusion and recognition optimization method under low infrared contrast conditions at night, according to an embodiment of the present invention. The method includes:

[0039] Step S1: Obtain image information of the target area at night, including infrared images and visible light images; for any coordinate point in the target area, obtain the confidence weight of each coordinate point based on the edge intensity of the coordinate point in the infrared image, the brightness in the visible light image, and the disorder of motion information between adjacent frames in the infrared image.

[0040] In nighttime surveillance scenarios, infrared sensors and traditional surveillance cameras can be used to capture image information of the monitored area, i.e., the target area, at night. This image information includes infrared and visible light images of the same size and acquired at the same frequency.

[0041] In one specific implementation of this invention, considering that images under low-light conditions at night typically contain high-frequency photon noise, directly calculating subsequent motion features on the original resolution image easily yields meaningless random noise. Therefore, to obtain macroscopic and reliable regional motion trends, it is necessary to perform downscaling preprocessing on the obtained image information. In this specific implementation, Gaussian pyramid downsampling is chosen for downscaling, with two downsampling layers, reducing the resolution to one-quarter of the original. This effectively filters out single-pixel flicker noise while retaining the structural information of physical targets with a certain spatial scale. In other specific implementations of this invention, Gaussian smoothing filtering can also be used to eliminate noise preprocessing. Specifically, a 5×5 kernel with a standard deviation of 1.5 can be selected for processing. Image preprocessing is a well-known technique among those skilled in the art and will not be elaborated or limited here.

[0042] To distinguish between real moving targets and random noise backgrounds in an image, embodiments of this invention utilize the spatial consistency of motion information of a coordinate point in the image information for discrimination. This is because the neighborhood motion directions of real rigid targets tend to be consistent, while the motion directions of noise regions are usually chaotic. Therefore, this invention, for any coordinate point in the target region, determines the disorder of motion information of the coordinate point in the infrared image dimension through infrared image information between adjacent frames. That is, the greater the disorder of motion information, the more the coordinate point belongs to the noise region that is not the target region.

[0043] It should be noted that, because this embodiment of the invention requires obtaining the motion features of each coordinate point in the target area in the image information dimension, multiple consecutive frames of image information are needed for analysis. Therefore, in the initial stage of image acquisition, if the buffer is empty, i.e., there are no historical frames, motion information cannot be calculated. Thus, motion analysis can be omitted in this state, or the subsequent motion information disorder can be set to a preset identifier. This preset identifier is set to a large value to indicate that all motion information in the monitoring cold start phase is unreliable, thereby avoiding logical breaks in the process during the cold start. In a specific implementation of this embodiment, the preset identifier can be specifically set to 999, which can be considered as the motion information disorder of all current coordinate points being relatively large, i.e., the motion information is unreliable. The confidence weights for subsequent calculations can also be directly set to 0.

[0044] This invention aims to align infrared images onto visible light images using non-rigid deformation. The non-rigid registration algorithm relies on the obvious structural information displayed in the infrared image. Therefore, for a coordinate point, if it belongs to a clearly defined target, it should present a valid geometric edge in the infrared image. Compared to noise or halo areas, the target's edge should be more salient. Because visible light images often contain a large amount of high-frequency textures without thermal characteristics at night, such as road surface textures, leaf noise, and noisy edge information like artifact halos, if the edge intensity of the visible light image is used to analyze the coordinate point, and if the coordinate point does not belong to the monitored target (i.e., it has no specific edge information in the infrared image), the noisy edge information in the visible light image will cause incorrect alignment during subsequent fusion. Ultimately, the infrared background will be forcibly distorted to match the noise in the visible light image, leading to tearing or honeycomb distortion in the fused image. Therefore, this invention requires extracting the edge intensity of the coordinate point in the infrared image as an important confidence feature characterizing whether it belongs to a thermal target.

[0045] To eliminate interference from strong light sources and prevent incorrect registration of infrared targets to halo artifacts during subsequent matching, the brightness information of the coordinate points needs to be further considered. The brighter the coordinate point is in the visible light image, the more likely it is to be within a halo region. Therefore, this location should not be stretched or aligned during subsequent matching. This is because halos are often illusory structures formed by optical scattering and are not real object regions. If this information is not considered and the infrared and visible light images are directly fused, the real target in the infrared image will be forcibly stretched and distorted to match this illusory halo shape, leading to registration errors.

[0046] For each coordinate point in the target region, the edge intensity in the infrared image, the brightness in the visible light image, and the motion information disorder in the infrared image were analyzed. The lower the motion information disorder, the more likely the coordinate point belongs to a rigid target; the higher the edge intensity, the more obvious the structural features in the infrared image, and the more likely it is to belong to the thermal target outline in the infrared image; the lower the brightness in the visible light image, the less likely it is to be a significant halo region, and the stronger the infrared information reference value during the fusion process. Therefore, fusing these three features yields the confidence weight for each coordinate point. The confidence weight quantifies the physical reliability of the coordinate point in subsequent non-rigid geometric correction, physically characterizing whether the point simultaneously possesses a clear thermal radiation structure (not background), stable motion consistency (not random noise), and is not affected by strong visible light interference (not artifacts). Its core function is to act as an adaptive "gating" switch. In the subsequent process, in areas with high weight, it drives the algorithm to search for displacement vectors to correct thermal diffusion parallax and fuse textures; in areas with low weight (unreliable), it applies constraints to reduce information interference in the visible light image dimension, thereby preventing image tearing caused by noise and filtering out halo artifacts.

[0047] Preferably, in this embodiment of the invention, the edge intensity is the gradient magnitude of the coordinate point on the infrared image. The specific method for calculating the gradient magnitude can be obtained using the Sobel operator, which is a well-known technique and will not be elaborated upon here.

[0048] Furthermore, in a specific implementation of this invention, to unify the units and facilitate calculation, after obtaining the gradient magnitude, the obtained edge intensity can be normalized using a maximization normalization method. Specifically, a minimum gradient threshold is introduced, which is set to 3 in this embodiment, corresponding to the inherent thermal noise floor of the sensor. Then, the gradient magnitudes of all pixels in the infrared image are obtained. To prevent outliers from suppressing global features, all gradient magnitudes are arranged from largest to smallest, and the gradient magnitudes of the top 98% quantiles are selected as the global gradient magnitude. The global gradient magnitude is compared with the thermal noise floor, and the maximum value is selected as the denominator, while the edge intensity to be normalized is selected as the numerator. If the ratio is greater than 1, it is assigned a value of 1; if it is less than or equal to 1, the original value is retained. The normalized data is then obtained.

[0049] Preferably, in this embodiment of the invention, the method for obtaining motion information disorder includes:

[0050] The infrared velocity field is obtained based on the image information between adjacent frames. In this embodiment of the invention, a dense optical flow algorithm, such as the Farneback algorithm, can be selected, or an existing algorithm such as an optical flow estimation neural network can be selected to obtain the infrared velocity field of the current frame based on the information between the current frame infrared image and the previous frame infrared image. It should be noted that, similarly, the visible light velocity field of each frame can be obtained for visible light images, but the details are not elaborated here.

[0051] In the infrared velocity field, each coordinate position corresponds to a velocity vector. To quantify the local motion information disorder at a coordinate point, this embodiment of the invention first constructs a preset neighborhood centered on the coordinate point and analyzes the distribution of velocity vectors within the neighborhood. The average velocity vector within the preset neighborhood centered on the coordinate point is statistically analyzed, and the sum of squares of the Euclidean distances between the average velocity vector and the velocity vector at the coordinate point is calculated to obtain the motion information disorder. That is, for each velocity vector within the neighborhood, the Euclidean distance between each velocity vector and the average velocity vector is statistically analyzed. The obtained Euclidean distance is a dimensionless scalar; a larger value indicates a greater difference between the two vectors. By statistically analyzing the sum of squares of the Euclidean distances, the velocity vector distribution disorder in the current neighborhood can be determined. It should be noted that the sum of squares of the Euclidean distance can be compared to the variance obtained from scalar data statistics. The calculation method and significance of the variance, as well as the specific Euclidean distance calculation method, are well-known techniques to those skilled in the art and will not be elaborated here.

[0052] Preferably, in this embodiment of the invention, considering that the analysis of brightness in the visible light image aims to avoid the influence of halo regions during the calculation of confidence weights, halo regions can be directly identified in the visible light image and then marked to facilitate the calculation of confidence weights. Specifically, this includes:

[0053] Based on the brightness information, the visible light image is divided into high-brightness and low-brightness regions. If the coordinate point is in the high-brightness region and the gray-scale gradient amplitude in the visible light image is less than the preset texture threshold, then the strong light interference flag value is set to 1; if the coordinate point is in the low-brightness region, then the strong light interference flag value is set to 0.

[0054] In this embodiment of the invention, high-brightness regions and low-brightness regions are divided using a threshold segmentation method. The brightness threshold can be set to 98% of the maximum grayscale value of the visible light image, or it can be set to a fixed, large value, such as 245. Using the brightness threshold, the grayscale visible light image is segmented, with regions greater than the brightness threshold designated as high-brightness regions and regions less than the brightness threshold designated as low-brightness regions. The texture threshold can be set to 10% of the grayscale gradient magnitude of all pixels in the visible light image; that is, the texture threshold is a small value. Regions less than this texture threshold and considered high-brightness are identified as noise regions such as halos or artifacts without obvious target features.

[0055] The confidence weights can then be obtained by statistically analyzing the quantified strong light interference flag value, edge intensity, and motion information disorder. A higher strong light interference flag value (e.g., 1) indicates that the coordinate point is a halo region in the visible light image, and the confidence weight should be reset to a smaller value. Conversely, a higher edge intensity and lower motion information disorder indicate that the coordinate point is a more obvious thermal target in the infrared image, and therefore the corresponding confidence weight should be larger.

[0056] Further, the confidence weight is obtained based on the strong light interference identifier value, the edge intensity, and the motion information disorder, including:

[0057] For any given coordinate point, the disorder of motion information is compared with the overall disorder of motion information across all coordinate points to obtain the degree of disorder significance. This degree of disorder significance is then negatively correlated and normalized to obtain motion stability. In other words, the smaller the degree of disorder significance, the less disorder the coordinate point's motion information exhibits, which in turn represents greater motion stability, and the corresponding confidence weight for that coordinate point should be higher.

[0058] The positive integer 1 minus the strong light interference flag value is used as the interference masking term for the coordinate point. That is, the interference masking term has only two results: 0 and 1. If the coordinate point is a high-brightness area representing a halo region in the visible light image, the value of the interference masking term is 0. This term acts as a hard constraint, forcibly reducing the confidence weight of coordinate points in such cases. Subsequent fusion using multiplication further sets the confidence weight to 0, preventing the infrared target from being incorrectly attached to false optical artifacts during subsequent alignment and fusion.

[0059] The product of the edge intensity, motion stability, and interference shielding term is used as the confidence weight. That is, the greater the edge intensity and motion stability, and the greater the interference shielding term (1), the more likely the coordinate point is not in the halo region of the visible light image and is a significant thermal target in the infrared image. In this case, it can be aligned and fused with the visible light image, and the greater the confidence weight, the better.

[0060] As a specific example, in one implementation of this invention, the confidence weight is expressed by the formula:

[0061] ;in The confidence weights are for the coordinate point (x, y). Let be the edge strength at the coordinate point (x, y), and exp() be an exponential function with the natural constant as the base. The motion information at the coordinate point (x, y) is chaotic. This represents the lower bound of disorder corresponding to the static noise floor of the sensor. The average motion information disorder of all coordinate points The value represents the strong light interference indicator at the coordinate point (x, y), and max() is the maximum value filtering function.

[0062] In the above formula, choose As the denominator, the motion information disorder is standardized to a global minimum. This aims to make motion stability calculation scene-adaptable, allowing standardized calculations to be performed in different nighttime and motion scenarios. The reason for choosing the lower limit of disorder corresponding to static noise is to prevent minute sensor noise from being misjudged as violent motion in completely static scenes, leading to weight calculation overflow or oscillation. Therefore, static noise is chosen as one of the options for the denominator. The maximum value between static noise and global average motion information disorder is chosen as the denominator for significance comparison; that is, the ratio represents the degree of disorder significance. The higher the degree of disorder significance, the more significant the motion disorder value of the current coordinate point, and the more likely that the coordinate point is not the target coordinate point. In a specific implementation of this invention, the lower limit of disorder corresponding to static noise can be taken as 0.01. Then, an exponential function with the natural constant as the base is used to perform negative correlation mapping and normalization on the standardized motion information disorder. In the above formula, the three data terms are positively fused through multiplication, which allows each term to serve as a hard constraint to constrain the final value of the confidence weight. Motion stability, as a hard constraint, can reduce the confidence weight of coordinate points with significantly higher motion information disorder than the global level by mapping through an exponential function and multiplication, even to a small value or zero. Edge intensity, as a hard constraint, ensures that only coordinate points with significant thermal edges in the infrared image are allowed to deform and align in subsequent processes. If the coordinate point is a flat background location without significant edges or even edge information, the confidence weight will be reduced or even reduced to zero. Interference masking, as a hard constraint, can force the confidence of coordinate points in high-brightness areas, i.e., halo areas, in the visible light image to be set to 0, preventing the infrared target from being mistakenly attached to false optical artifacts during subsequent alignment.

[0063] Step S2: Using the coordinate points in the infrared image as reference points, search for the optimal matching point in the visible light image; the matching process of the optimal matching point requires comparing the differences in motion features and edge information reflected by the coordinate points in the image information, and combining the confidence weight of the reference point to obtain the matching cost and perform the matching.

[0064] To correct for the geometric displacement caused by thermal diffusion parallax, spatial correction of the infrared image is necessary before image information fusion. Therefore, this embodiment of the invention uses coordinate points in the infrared image as reference points. For each reference point, a search is performed in the visible light image to find the optimal matching point that best matches the reference point, thereby obtaining the offset of each reference point in the visible light image. Although the physical locations of the infrared heat source and the visible light source may not completely coincide, the motion trend of the rigid target is consistent within a local range. Therefore, during the matching process, it is necessary to compare the differences in motion characteristics of the coordinate points in the image information. Furthermore, to ignore the difference in grayscale amplitude between infrared and visible light, it is necessary to use the difference in contour direction information of the coordinate points between the two image information to evaluate the matching situation of the coordinate points. Therefore, edge information differences also need to be obtained. Finally, combining the differences in motion characteristics and edge information differences can reflect the differences in data characteristics between the reference point in the infrared image and the point to be matched in the visible light image. Furthermore, to prevent searching for incorrect optimal matching points in blurred regions such as halo areas in visible light images, a confidence weight is introduced. This ensures that reference points with low confidence weights do not search for offset optimal matching points in the visible light image, or that the coordinates of the optimal matching point (which can be considered a reference point) are the same as the reference point, thus avoiding erroneous offsets by reference points with low confidence weights. At this point, the matching cost between the reference point and the point to be matched in the visible light image can be obtained, and the optimal matching point can then be obtained through a matching operation.

[0065] Preferably, in this embodiment of the invention, searching for the optimal matching point in the visible light image includes:

[0066] For each reference point, its corresponding offset in the visible light image should be within a certain range around itself, without significant offset. Therefore, a global search is unnecessary in the visible light image. Instead, a search area is constructed in the visible light image with the coordinates corresponding to the reference point as the center, according to a preset size. The pixels within the search area in the visible light image are the matching points of the reference point. In a specific implementation of this invention, a 640×512 resolution image is used, and the preset size is set to 30, meaning the search area is a square centered on the coordinates of the reference point with a side length of 30.

[0067] The matching cost between the baseline point and each point to be matched is obtained, and the point to be matched with the minimum matching cost is selected as the optimal matching point. It should be noted that the process of traversing all points to be matched to determine the minimum matching cost is a technique well known to those skilled in the art, and will not be elaborated or limited here.

[0068] Furthermore, in one embodiment of the present invention, the specific process of obtaining the matching cost may include:

[0069] For any point to be matched, the lateral and longitudinal coordinate differences between the point to be matched and the reference point are obtained; a displacement constraint term is obtained based on the lateral and longitudinal coordinate differences. That is, the displacement constraint term is used to limit the matching cost to be matched only to points near the reference point coordinates, and the matching cost of the point to be matched is smaller the closer it is to the reference point coordinates, so as to avoid the reference point from having a large offset.

[0070] As a specific example, in this embodiment of the invention, the displacement constraint term can be expressed by the formula:

[0071] ; This is the displacement constraint term between the point to be matched and the reference point, where the difference in the horizontal coordinate is u and the difference in the vertical coordinate is v. The difference in the horizontal coordinate u is the result of subtracting the horizontal coordinate of the reference point from the horizontal coordinate of the point to be matched, and the difference in the vertical coordinate v is the result of subtracting the vertical coordinate of the reference point from the vertical coordinate of the point to be matched. The normalization adjustment factor is used to limit the value of the displacement constraint term to between 0 and 1; in this embodiment of the invention, it can be taken as 0.3. K is the side length of the search region. This formula eliminates the influence of negative values ​​by squaring and limits the value range by the normalization adjustment factor, so that the displacement constraint term of the matching point closer to the reference point is smaller, that is, the smaller the difference between the horizontal coordinate and the vertical coordinate, the smaller the displacement constraint term.

[0072] The edge information difference is obtained based on the absolute value of the cosine similarity between the gradient vector of the reference point in the infrared image and the gradient vector of the point to be matched in the visible light image. Since cosine similarity has positive and negative values, ranging from -1 to 1, with negative values ​​representing contrast reversal (e.g., the case of a white-hot exhaust pipe and ferrous metal), it should also belong to similar edge information. Therefore, this embodiment of the invention uses the absolute value as the final edge similarity. Furthermore, the edge information difference can be obtained by negatively mapping the absolute values ​​of cosine similarity; that is, the smaller the absolute value of cosine similarity, the greater the edge information difference.

[0073] As an example, in this embodiment of the invention, since the value range of the absolute value of cosine similarity is between 0 and 1, the positive integer 1 can be directly subtracted from the absolute value of cosine similarity as a negative correlation mapping process to obtain the edge information difference.

[0074] Based on image information between adjacent frames, an infrared velocity field and a visible light velocity field are obtained. The motion feature difference is obtained based on the difference between the velocity vector of the reference point in the infrared velocity field and the velocity vector of the point to be matched in the visible light image. It should be noted that the methods for obtaining the infrared and visible light velocity fields are well-known to those skilled in the art and can be obtained using optical flow methods, which will not be elaborated here. The infrared and visible light velocity fields are composed of velocity vectors at each coordinate position. Therefore, in the specific implementation of this invention, the Euclidean distance between two velocity vectors can be directly obtained. The motion feature difference is obtained by normalizing the Euclidean distance between the vectors. The normalization method can be a function mapping method; in this embodiment, it is set to a tangent hyperbolic function. The Euclidean distance between the vectors is multiplied by a preset sensitivity coefficient, and the result is substituted into the tangent hyperbolic function to obtain the motion feature difference. The sensitivity coefficient can be set to 0.1 to adjust the sensitivity to velocity differences.

[0075] The edge information differences and motion feature differences are then weighted and fused to obtain feature information differences. In this embodiment of the invention, the edge information differences and motion feature differences are averaged, i.e., the weights of the two features can be set to 0.5, to obtain feature information differences, thus achieving positive weighted fusion.

[0076] For benchmark points with low confidence weights, strong spatial correction should not be performed, or no spatial correction should be performed at all; for benchmark points with high confidence weights, effective spatial correction can be performed. Therefore, in this embodiment of the invention, the confidence weights are used as the weights of the feature information differences, and the result of negative correlation mapping of the confidence weights is used as the weights of the displacement constraint term. The feature information differences and the displacement constraint term are weighted and summed to obtain the matching cost.

[0077] In one specific implementation of this invention, since the confidence weight ranges from 0 to 1, the result of subtracting the confidence weight from the positive integer 1 can be used as the result of the negative correlation mapping, thereby obtaining the weight of the displacement constraint term. In high signal-to-noise ratio regions with high confidence weights, the matching cost is dominated by feature information, and the non-zero displacements of motion and structure can be aligned during the search process; in noisy or halo regions with low confidence weights, the matching cost is dominated by the displacement constraint term, and the minimum value is locked at the reference point coordinate position, thereby achieving halo anti-interference.

[0078] As an example, the above process of obtaining the matching cost can be expressed by the following formula:

[0079]

[0080] in, Let be the matching cost between the point to be matched and the reference point, where the difference in the horizontal coordinate is u and the difference in the vertical coordinate is v. The confidence weights are for the baseline point with coordinates (x, y). Information weights for differences in edge information. For edge information differences, Differences in motion characteristics, The displacement constraint terms between the point to be matched and the reference point, where the difference in horizontal coordinate is u and the difference in vertical coordinate is v. As described above, Set it to 0.5. The specific logic of the formula has been explained in the above description and will not be repeated here.

[0081] Step S3: Construct an initial geometric correction displacement field based on the coordinate deviations between all reference points and the corresponding optimal matching points; after performing horizontal and vertical filtering on the initial geometric correction displacement field, a smoothed geometric correction displacement field is obtained; based on the smoothed geometric correction displacement field, the infrared image is spatially reconstructed to obtain a spatially corrected thermal image.

[0082] Through the processing in step S2, each pixel in the infrared image can be used as a reference point, thereby finding the optimal matching point in the infrared image. An initial geometric correction displacement field can be constructed based on the coordinate deviation between the reference point and the optimal matching point. The initial geometric correction displacement field has the same size as the image information. In this embodiment of the invention, the data at each position in the initial geometric correction displacement field is the coordinate deviation between the optimal matching point and the reference point, and the coordinate deviation includes the difference in the horizontal coordinate and the difference in the vertical coordinate.

[0083] Because the initial geometric correction displacement field is calculated independently for each pixel in the infrared image, it is affected by local image noise or texture loss. The optimal displacement calculated for adjacent pixels may change abruptly, for example, adjacent points may point in opposite directions, leading to over-correction. This effect causes the initial geometric correction displacement field to be spatially discontinuous. Directly using this initial geometric correction displacement field for image reconstruction may result in image tearing or pixel stacking. Therefore, it needs to be smoothed.

[0084] Because the initial geometric correction displacement field contains both lateral and longitudinal offset information, horizontal filtering and vertical filtering are required during smoothing. Then, the two component matrices are merged to finally smooth the geometric correction displacement field.

[0085] In the specific implementation of this invention, the smoothing process uses a 5×5 or 7×7 filtering window. A window that is too small may not completely filter out large-area speckle noise, while a window that is too large may smooth out local deformation features. Median filtering can effectively filter out isolated noise displacements while maintaining the edge characteristics of the displacement field, preventing correction failure caused by excessive smoothing. The specific smoothing process is a well-known technique to those skilled in the art and will not be elaborated or limited here.

[0086] After obtaining the smoothed geometric correction displacement field, the infrared image can be spatially reconstructed based on the smoothed geometric correction displacement field to obtain a spatially corrected thermal image. This step can smoothly shrink and align regions with high confidence weights in the infrared image to the visible light texture skeleton; in low confidence regions, since the displacement field tends to 0, the original infrared geometry can be maintained in subsequent fusion processes.

[0087] Preferably, in this embodiment of the invention, the method for obtaining a spatially corrected thermal image includes:

[0088] For any pixel in any infrared image, since it has an offset displacement value in the smoothed geometric correction displacement field at the corresponding coordinates, the source coordinates of the pixel can be obtained by translation search based on the corresponding displacement features in the smoothed geometric correction displacement field. Because the obtained source coordinates may be non-integer, resulting in missing pixel information, the replacement pixel value of the pixel is obtained by interpolation based on the pixel value within a preset neighborhood range of the source coordinates.

[0089] As an example, in this embodiment of the invention, a bilinear interpolation algorithm can be used for interpolation. For a source coordinate, the nearest real pixel in the infrared image in the four directions (up, down, left, and right) is obtained, i.e., the pixel with integer coordinate values. The pixel values ​​of the four integer coordinate values ​​are weighted and summed based on their distances from the source coordinate to achieve bilinear interpolation and obtain the replacement pixel value. It should be noted that if the source coordinate exceeds the image boundary, the replacement pixel value should be directly set to the original pixel value to prevent out-of-bounds errors. The specific bilinear interpolation algorithm is a well-known technique to those skilled in the art and will not be described in detail here.

[0090] For all pixels in the infrared image, the replacement pixel values ​​are used to replace the original pixel values ​​to obtain the spatially corrected thermal image.

[0091] Step S4: Based on the confidence weights of the coordinate points, fuse the spatially corrected thermal image and the visible light image to obtain a fused image; use the fused image and the target detection network trained with the confidence weights to perform nighttime target recognition.

[0092] To preserve complementary information across the two bands in the fused image while automatically masking halo artifacts in the visible light channel, the spatially corrected thermal image and the visible light image are fused based on the confidence weight of each coordinate point within the target region. This results in a fused image, where areas with lower confidence retain infrared information, while areas with higher confidence achieve effective fusion of the two channels.

[0093] Preferably, in this embodiment of the invention, the method for obtaining the fused image includes:

[0094] The spatially corrected thermal image and the visible light image are fused according to a preset fusion weight to obtain an initial fused image. Since the spatially corrected thermal image has already undergone spatial correction, it can be directly superimposed and fused with the visible light image. The initial fused image can be obtained by fusing pixel values ​​at the same coordinate points. In a specific implementation of this embodiment, the fusion weight is set to 0.5, meaning the average pixel value at the same coordinate points is used as the new pixel value to obtain the initial fused image.

[0095] Because the initial fused image is a direct fusion result and does not consider whether the coordinate points belong to halo artifacts in the visible light image, it is necessary to further refer to the confidence weights to distinguish the information on the spatially corrected thermal image. For each coordinate point, the confidence weights are used as the weights of the pixel values ​​in the initial fused image, and the result of the negative correlation mapping of the confidence weights is used as the weights of the pixel values ​​in the spatially corrected thermal image. The fused pixel value of each pixel in the fused image is obtained by weighted summation, thus obtaining the fused image.

[0096] In this embodiment of the invention, since the confidence weights are data with values ​​between 0 and 1, negative correlation mapping can also be achieved by subtracting the confidence weight from a positive integer 1, and the final sum of the two weights is 1. By setting the confidence weights, the initial fused image result can be retained at positions with high confidence weights, that is, it simultaneously contains infrared thermal information and visible light texture information; at positions with low confidence weights, more space is retained for correcting the thermal imaging result on the thermal image, avoiding interference from noisy visible light information. The final fused image enables the system to automatically filter out noise information such as halo artifacts when encountering strong interference such as direct headlights, retaining only the thermal imaging information unaffected by illumination, thereby significantly improving the quality of the fused image.

[0097] Ultimately, target recognition can be performed based on the fused image. Common target recognition methods are implemented using target recognition networks. Considering that conventional target detection networks are usually single-channel grayscale images or simple RGB three-channel pseudo-color images, they focus more on information in visible light images. However, the embodiments of this invention have obtained confidence weights that can characterize the reference value of thermal target information. Therefore, the embodiments of this invention can use the fused image and the target detection network trained with the confidence weights to perform accurate nighttime target recognition on the fused image.

[0098] Preferably, in this embodiment of the invention, in order to improve the recognition ability of the target detection network, the target detection network includes three visual channels and one physical attention channel. The visual channels are filled with the same fused image, and the physical attention channel is filled with a confidence weight distribution map formed by the confidence weights of all coordinate points. The specific target detection network structure can adopt YOLOv5, SSD, or Faster R-CNN, etc., which are all well-known techniques to those skilled in the art and will not be elaborated upon here. Only a brief description of the network structure modification and fine-tuning training process in one embodiment of the present invention is provided below:

[0099] (1) First layer modification: Change the number of input channels of the first convolutional layer of the network from the default 3 to 4.

[0100] (2) Weight initialization: The convolution kernel weights of the first three visual channels are initialized to the default pre-trained weights, and the convolution kernel weights of the fourth physical attention channel are initialized to 0 or a tiny random value.

[0101] (3) Data set preparation: Construct a dedicated dataset containing raw infrared and visible light data and corresponding bounding boxes. During the training phase, a confidence weight distribution map needs to be constructed as the input of the fourth channel according to the method provided in the embodiment of the present invention.

[0102] (4) Fine-tuning the network: The network will automatically learn to use the weight information of the fourth channel as a spatial attention mask to suppress the region with a weight of 0 and focus on the region with a high weight.

[0103] Finally, by using the target detection network with the target's category label and bounding box coordinates as input, and through the suppression effect of physical attention, the false detection rate of false targets caused by road surface reflections and headlight halos can be significantly reduced.

[0104] In summary, this invention's embodiments characterize structural saliency based on the edge intensity of coordinate points in infrared images, characterize motion consistency based on motion information reflected in infrared images, and calculate confidence weights for each coordinate point in conjunction with specific brightness attributes. Based on the matching process, an initial geometric correction displacement field can be constructed using the reference point in the infrared image and the optimal matching point in the visible light image. Filtering operations are performed in the vertical and horizontal directions to obtain a smooth geometric correction displacement field for spatial reconstruction. The spatially reconstructed spatially corrected thermal imaging image is further filtered for feature information through a weighted fusion operation, resulting in a fused image with clearly defined moving target features. Based on this fused image, accurate and effective target recognition can be achieved. This invention improves the accuracy of target recognition at night through effective image fusion.

[0105] It should be noted that the order of the above embodiments of the present invention is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. The processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0106] The various embodiments in this specification are described in a progressive manner. The same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on describing the differences from other embodiments.

Claims

1. A method for cross-spectral fusion and recognition optimization under low infrared contrast conditions at night, characterized in that, The method includes: Image information of the target area at night is obtained, including infrared images and visible light images; for any coordinate point in the target area, the confidence weight of each coordinate point is obtained based on the edge intensity of the coordinate point in the infrared image, the brightness in the visible light image, and the disorder of motion information between adjacent frames in the infrared image. Using the coordinate points in the infrared image as reference points, the optimal matching point is searched in the visible light image. The matching process of the optimal matching point requires comparing the differences in motion features and edge information between the reference point and the coordinate points in the visible light image, and combining the confidence weight of the reference point to obtain the matching cost and perform the matching. An initial geometric correction displacement field is constructed based on the coordinate deviations between all reference points and the corresponding optimal matching points; after performing horizontal and vertical filtering on the initial geometric correction displacement field, a smoothed geometric correction displacement field is obtained; the infrared image is spatially reconstructed based on the smoothed geometric correction displacement field to obtain a spatially corrected thermal image. Based on the confidence weights of coordinate points, the spatially corrected thermal image and the visible light image are fused to obtain a fused image; the fused image and the target detection network trained by the confidence weights are used for nighttime target recognition. The edge intensity is the gradient magnitude of the coordinate point on the infrared image; The method for obtaining the motion information disorder includes: Based on the image information between adjacent frames, an infrared velocity field is obtained; in the infrared velocity field, the average velocity vector within a preset neighborhood centered on the coordinate point is statistically analyzed, and the sum of squares is calculated based on the Euclidean distance between the average velocity vector and the velocity vector of the coordinate point to obtain the motion information disorder.

2. The method for cross-spectral fusion and recognition optimization under low infrared contrast conditions at night, as described in claim 1, is characterized in that, The process of obtaining the confidence weight for each coordinate point includes: Based on the brightness information, the visible light image is divided into high-brightness and low-brightness regions. If the coordinate point is in the high-brightness region and the gray-scale gradient amplitude in the visible light image is less than the preset texture threshold, then the strong light interference flag value is set to 1; if the coordinate point is in the low-brightness region, then the strong light interference flag value is set to 0. The confidence weight is obtained based on the strong light interference identifier value, the edge intensity, and the motion information disorder.

3. The method for cross-spectral fusion and recognition optimization under low infrared contrast conditions at night, as described in claim 2, is characterized in that... The step of obtaining the confidence weight based on the strong light interference identifier value, the edge intensity, and the motion information disorder includes: For any coordinate point, the motion information disorder is compared with the overall motion information disorder of all coordinate points to obtain the disorder significance; the disorder significance is negatively correlated and normalized to obtain motion stability; the positive integer 1 minus the strong light interference identifier value is used as the interference masking term of the coordinate point; the product of the edge intensity, the motion stability and the interference masking term is used as the confidence weight.

4. The method for cross-spectral fusion and recognition optimization under low infrared contrast conditions at night, as described in claim 1, is characterized in that... The search for the optimal matching point in the visible light image includes: In a visible light image, a search region is constructed with the coordinates corresponding to the reference point as the center and a preset size is used. The pixels in the search region of the visible light image are the points to be matched with the reference point. The matching cost between the reference point and each point to be matched is obtained, and the point to be matched with the minimum matching cost is selected as the optimal matching point.

5. The method for cross-spectral fusion and recognition optimization under low infrared contrast conditions at night, as described in claim 4, is characterized in that... The method for obtaining the matching cost includes: For any point to be matched, obtain the lateral coordinate difference and the longitudinal coordinate difference between the point to be matched and the reference point; obtain the displacement constraint term based on the lateral coordinate difference and the longitudinal coordinate difference. The edge information difference is obtained based on the absolute value of the cosine similarity between the gradient vector of the reference point in the infrared image and the gradient vector of the point to be matched in the visible light image; the infrared velocity field and the visible light velocity field are obtained based on the image information between adjacent frames, and the motion feature difference is obtained based on the difference between the velocity vector of the reference point in the infrared velocity field and the velocity vector of the point to be matched in the visible light image; the edge information difference and the motion feature difference are weighted and fused to obtain the feature information difference. The confidence weight is used as the weight of the feature information difference, and the result of negative correlation mapping of the confidence weight is used as the weight of the displacement constraint term. The feature information difference and the displacement constraint term are weighted and summed to obtain the matching cost.

6. The method for cross-spectral fusion and recognition optimization under low infrared contrast conditions at night, as described in claim 1, is characterized in that... The method for acquiring the spatially corrected thermal image includes: For any pixel in any infrared image, a translation search is performed based on the displacement features corresponding to the smooth geometric correction displacement field to obtain the source coordinates of the pixel. Then, the pixel values ​​within a preset neighborhood range are interpolated to obtain the replacement pixel value of the pixel. For all pixels in the infrared image, the replacement pixel values ​​are used to replace the original pixel values ​​to obtain the spatially corrected thermal image.

7. The method for cross-spectral fusion and recognition optimization under low infrared contrast conditions at night, as described in claim 1, is characterized in that, The method for obtaining the fused image includes: The spatially corrected thermal image and the visible light image are fused according to a preset fusion weight to obtain an initial fused image; For each coordinate point, the confidence weight is used as the weight of the pixel value in the initial fused image, and the result of the negative correlation mapping of the confidence weight is used as the weight of the pixel value in the spatially corrected thermal image. The fused pixel value of each pixel in the fused image is obtained by weighted summation, thus obtaining the fused image.

8. The method for cross-spectral fusion and recognition optimization under low infrared contrast conditions at night, as described in claim 1, is characterized in that... The object detection network includes three visual channels and one physical attention channel. The visual channels are filled with the same fused image, and the physical attention channel is filled with a confidence weight distribution map formed by the confidence weights of all coordinate points.