Infrared remote sensing image enhancement method and device, computer equipment and storage medium
By globally aligning and fusing multiple frames of infrared remote sensing images, and combining this with a Kalman filter to determine the target location, an enhanced target image is generated. This solves the problem of unstable enhancement effect of infrared remote sensing images in existing technologies, and improves target recognition performance and signal-to-noise ratio.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ZHEJIANG LAB
- Filing Date
- 2026-01-06
- Publication Date
- 2026-05-08
AI Technical Summary
Existing infrared remote sensing image enhancement methods are limited by the global signal-to-noise ratio and imaging quality of a single image, resulting in unstable enhancement effects and difficulty in effectively improving the recognition performance of small targets.
By acquiring multiple consecutive frames of infrared remote sensing images, global alignment and fusion are performed to generate a globally enhanced image. A Kalman filter is then used to determine the location information of the target object, thereby generating a target enhanced image, which is then enhanced by combining local image information.
It achieves stability of image enhancement effect and improves target recognition performance, provides high signal-to-noise ratio data support, and alleviates the problem of unstable single-frame image enhancement effect.
Smart Images

Figure CN121458559B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, and more specifically, to a method, apparatus, computer device, and storage medium for enhancing infrared remote sensing images. Background Technology
[0002] Infrared remote sensing images are widely used in military reconnaissance, environmental monitoring and other fields. However, due to sensor limitations and complex background interference, they often face problems such as low signal-to-noise ratio, poor contrast and difficulty in identifying small targets. In order to improve image quality and support back-end tasks, researchers have proposed a variety of signal-to-noise ratio enhancement methods to improve target detection and recognition performance.
[0003] In related technologies, one method aims to fully exploit the statistical characteristics or structural components of the image itself, utilizing the inherent characteristics and structural information of infrared images to improve the signal-to-noise ratio (SNR) of the target. For example, by utilizing the difference between small targets and the background, and combining multi-scale gray-level difference operators with local image entropy, a weighted local image entropy is constructed, effectively suppressing background clutter and enhancing small targets. Another method employs image decomposition techniques, first decomposing the image into different components, then combining global and local mapping for contrast enhancement, and supplementing this with adaptive algorithms to strengthen detail components, thereby improving overall contrast while sharpening edges and suppressing noise. However, these enhancement algorithms use a single infrared remote sensing image as the information source, and the SNR enhancement effect is severely limited by the global SNR and imaging quality of this image, making the enhancement effect unstable. Summary of the Invention
[0004] In view of this, this application provides a method, apparatus, computer equipment, and storage medium for enhancing infrared remote sensing images.
[0005] Specifically, this application is implemented through the following technical solution:
[0006] In a first aspect, embodiments of this application provide a method for enhancing infrared remote sensing images, comprising:
[0007] Acquire consecutive multiple frames of infrared remote sensing images;
[0008] The multi-frame infrared remote sensing images are globally aligned to generate multi-frame registered infrared remote sensing images; the multi-frame registered infrared remote sensing images are then fused to generate a globally enhanced image.
[0009] After determining the initial position information of the target object in the global enhanced image, for each frame of the infrared remote sensing image, based on the initial position information of the target object, an initial local image of the target object is extracted from the infrared remote sensing image. Then, using a Kalman filter, the initial local image, and the position information of the target object corresponding to the previous moment in the infrared remote sensing image, the target position information of the target object in the infrared remote sensing image is determined. The target object is an object with a pixel size smaller than a preset size in the infrared remote sensing image.
[0010] Based on the target location information of the target object, a local target image corresponding to the target object is extracted from the infrared remote sensing image, and a target enhancement image is generated based on the global enhancement image and the local target image corresponding to the target object.
[0011] In one optional implementation, the method further includes:
[0012] The multiple frames of infrared remote sensing images are preprocessed to generate processed infrared remote sensing images; wherein the image preprocessing includes at least one of the following: non-uniformity correction processing, bad pixel repair processing, and preliminary background suppression processing.
[0013] Global alignment of the multiple frames of infrared remote sensing images to generate multiple registered infrared remote sensing images includes:
[0014] Global alignment is performed on multiple frames of the processed infrared remote sensing images to generate multiple frames of registered infrared remote sensing images.
[0015] In one optional implementation, the step of globally aligning the multiple frames of infrared remote sensing images to generate registered multiple frames of infrared remote sensing images includes:
[0016] Extract the initial gradient image for each frame of the infrared remote sensing image;
[0017] For each frame of infrared remote sensing image, the initial gradient image of the infrared remote sensing image is processed using the optical flow pyramid calculation model to generate multiple pyramid gradient images corresponding to the infrared remote sensing image, and the target number of control points are uniformly sampled from each pyramid gradient image.
[0018] Based on the multiple control points included in the multi-frame infrared remote sensing images, the affine transformation matrix parameters between the second infrared remote sensing image and the first infrared remote sensing image in the multi-frame infrared remote sensing images are determined, wherein the first infrared remote sensing image is a reference infrared remote sensing image in the multi-frame infrared remote sensing images, and the second infrared remote sensing image is another infrared remote sensing image other than the first infrared remote sensing image.
[0019] Using the affine transformation matrix parameters, the second infrared remote sensing image is geometrically corrected and resampled to generate a registered second infrared remote sensing image, wherein the first infrared remote sensing image and the registered second infrared remote sensing image constitute the multi-frame registered infrared remote sensing image.
[0020] In one optional implementation, fusing the multi-frame registered infrared remote sensing images to generate a globally enhanced image includes:
[0021] The weight information of the registered infrared remote sensing image is determined based on the image quality parameters of each frame of the registered infrared remote sensing image.
[0022] Based on the weight information corresponding to the registered infrared remote sensing images of the multi-frames, the pixel information of the same pixel position in the registered infrared remote sensing images of the multi-frames is summed to obtain the fused pixel information corresponding to the pixel position.
[0023] A globally enhanced image is generated based on the fused pixel information corresponding to each pixel position.
[0024] In one optional implementation, determining the weight information of the registered infrared remote sensing image based on the image quality parameters of each frame of the registered infrared remote sensing image includes:
[0025] For each frame of the registered infrared remote sensing image, the image signal-to-noise ratio of the registered infrared remote sensing image is determined based on the image mean and image variance of the registered infrared remote sensing image.
[0026] Based on the number of pixels in the gradient image corresponding to the registered infrared remote sensing image and the gradient information in multiple directions, the image gradient energy of the registered infrared remote sensing image is determined.
[0027] The second-order gradient value of the registered infrared remote sensing image is determined using the Laplacian operator, and the image blur of the registered infrared remote sensing image is determined based on the second-order gradient value.
[0028] The weight information of the registered infrared remote sensing image is determined based on the image signal-to-noise ratio, the image gradient energy, and the image blur.
[0029] In one optional implementation, determining the target location information of the target object in the infrared remote sensing image using a Kalman filter, the initial local image, and the target object's location information from the previous moment corresponding to the infrared remote sensing image includes:
[0030] The position information of the target object at the previous moment is corrected by the energy centroid coordinates in the neighborhood to generate the corrected position information;
[0031] Using the Kalman filter algorithm and the determined state vector at the current moment, the corrected position information is tracked and corrected to generate the predicted position information of the target object at the current moment;
[0032] The detection location information of the target object detected at the current moment is obtained, and the energy centroid coordinates in the neighborhood are corrected based on the pixel information of the initial local image to generate the corrected detection location information.
[0033] Based on the predicted location information and the corrected detection location information, the target location information of the target object in the infrared remote sensing image is determined.
[0034] In one optional implementation, determining the target location information of the target object in the infrared remote sensing image based on the predicted location information and the corrected detection location information includes:
[0035] Determine the position offset value between the predicted position information and the corrected detection position information;
[0036] When the position offset value is less than a preset offset value, the predicted position information is determined as the target position information of the target object;
[0037] When the position offset value is greater than or equal to the preset offset value, the mean position information between the predicted position information and the corrected detection position information is determined; the mean position information is determined as the target position information of the target object.
[0038] In one optional implementation, generating a target enhancement image based on the global enhancement image and the target local image corresponding to the target object includes:
[0039] The target local images corresponding to the target object are fused together to generate a fused local image;
[0040] Determine the adaptive fusion weights corresponding to the fused local image, wherein the adaptive fusion weights are positively correlated with the signal-to-noise ratio of the fused local image;
[0041] Based on the adaptive fusion weights, the fused local image is embedded into the global enhanced image to generate the target enhanced image.
[0042] In one optional implementation, determining the adaptive fusion weights corresponding to the fused local image includes:
[0043] Based on the pixel information of multiple target pixels belonging to the target object in the fused local image, the pixel mean is determined;
[0044] Based on the pixel information of other pixels in the fused local image besides the target pixel, the pixel standard deviation is determined;
[0045] Calculate the local signal-to-noise ratio based on the pixel mean and the pixel standard deviation;
[0046] Based on the local signal-to-noise ratio and the set weight extreme values, adaptive fusion weights are generated corresponding to the fused local image.
[0047] Secondly, embodiments of this application also provide an infrared remote sensing image enhancement device, comprising:
[0048] The acquisition module is used to acquire multiple consecutive frames of infrared remote sensing images;
[0049] The first generation module is used to globally align the multiple frames of infrared remote sensing images to generate multiple registered infrared remote sensing images; and to fuse the multiple registered infrared remote sensing images to generate a globally enhanced image.
[0050] The determination module is used to, after determining the initial position information of the target object in the global enhanced image, extract an initial local image of the target object from the infrared remote sensing image for each frame of the infrared remote sensing image based on the initial position information of the target object, and determine the target position information of the target object in the infrared remote sensing image using a Kalman filter, the initial local image, and the position information of the target object corresponding to the previous moment in the infrared remote sensing image, wherein the target object is an object with a pixel size smaller than a preset size in the infrared remote sensing image;
[0051] The second generation module is used to extract the target local image corresponding to the target object from the multi-frame infrared remote sensing images according to the target location information of the target object, and generate a target enhancement image based on the global enhancement image and the target local image corresponding to the target object.
[0052] This application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method for enhancing infrared remote sensing images.
[0053] This application provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the above-described method for enhancing infrared remote sensing images.
[0054] The method provided in this application generates a global enhanced image and a local image of the target object separately, utilizing both global and local target information from the entire image to achieve image enhancement, thus avoiding an increase in background signal-to-noise ratio and ensuring the image enhancement effect. Simultaneously, this application implements image enhancement using image information from multiple frames of infrared remote sensing images to obtain an enhanced target image, alleviating the instability of enhancement effects caused by relying on a single frame of remote sensing image, improving the enhancement effect of the enhanced target image, and providing high signal-to-noise ratio data support for downstream applications such as target selection and target localization. Attached Figure Description
[0055] The accompanying drawings, which are included to provide a further understanding of this specification and form part of this specification, illustrate exemplary embodiments and are used to explain this specification, but do not constitute an undue limitation thereof. In the drawings:
[0056] Figure 1 This is a flowchart illustrating an exemplary embodiment of an infrared remote sensing image enhancement method.
[0057] Figure 2 This is a schematic diagram of the global background registration process in an infrared remote sensing image enhancement method according to an exemplary embodiment of this application.
[0058] Figure 3 This is a schematic diagram illustrating the determination of target location information in an infrared remote sensing image enhancement method according to an exemplary embodiment of this application;
[0059] Figure 4 This is a schematic diagram illustrating the determination of adaptive fusion weights in an infrared remote sensing image enhancement method according to an exemplary embodiment of this application;
[0060] Figure 5 This is a schematic flowchart illustrating an infrared remote sensing image enhancement method according to an exemplary embodiment of this application;
[0061] Figure 6 This is a schematic diagram of an infrared remote sensing image enhancement device according to an exemplary embodiment of this application;
[0062] Figure 7 This is a schematic diagram of the structure of a computer device provided in this application. Detailed Implementation
[0063] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.
[0064] The terminology used in this application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. The singular forms “a,” “the,” and “the” used in this application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any or all possible combinations of one or more of the associated listed items.
[0065] It should be understood that although the terms first, second, third, etc., may be used in this application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, without departing from the scope of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to determination."
[0066] Infrared remote sensing images are widely used in military reconnaissance, environmental monitoring and other fields. However, due to sensor limitations and complex background interference, they often face problems such as low signal-to-noise ratio, poor contrast and difficulty in identifying small targets. In order to improve image quality and support back-end tasks, researchers have proposed a variety of signal-to-noise ratio enhancement methods. These methods can be proposed from the perspectives of multi-frame processing, local feature perception and image decomposition. These methods not only focus on the balance between noise suppression and signal enhancement, but also pay attention to maintaining edge structure and detail information in order to improve target detection and recognition performance.
[0067] In related technologies, one approach aims to fully exploit the statistical characteristics or structural components of the image itself, utilizing the inherent properties and structural information of infrared images to improve the signal-to-noise ratio (SNR) of the target. For example, by leveraging the difference between small targets and the background, and combining multi-scale gray-level difference operators with local image entropy, a weighted local image entropy is constructed, effectively suppressing background clutter and enhancing small targets. Another approach employs image decomposition techniques, first decomposing the image into different components, then combining global and local mapping for contrast enhancement, and supplementing this with adaptive algorithms to strengthen detail components, thereby improving overall contrast while sharpening edges and suppressing noise. However, these enhancement algorithms use a single infrared remote sensing image as the information source, and the SNR enhancement effect is severely limited by the global SNR and imaging quality of that image.
[0068] The study also found that some researchers improved the signal-to-noise ratio (SNR) by accumulating information from multiple frames and processing consecutive frame sequences. For example, by estimating subpixel displacements between images, distinguishing effective moving points from blind pixels, and performing temporal enhancement and neighborhood compensation respectively, the SNR of airborne infrared image sequences was effectively improved. Another example is establishing a multi-dimensional motion model of the target, achieving accurate accumulation of energy for weak targets through trajectory tracking, and filtering low SNR frames, resulting in significant target enhancement effects under complex motion models. Although these methods improve the SNR by utilizing multiple frames and motion information, they only utilize global information and global motion estimation of the entire image, without distinguishing the differences in SNR contributions between local targets and the global background. This can easily increase the SNR of the background while enhancing the target's SNR, leading to poor image enhancement results.
[0069] To alleviate the aforementioned problems, this application proposes an infrared remote sensing image enhancement method. After acquiring multiple consecutive frames of infrared remote sensing images, the registered frames are fused to generate a global enhanced image. After determining the initial position information of the target object in the global enhanced image, an initial local image of the target object is extracted from the infrared remote sensing image based on this initial position information. Using a Kalman filter, the initial local image, and the target object's position information from the previous moment in the infrared remote sensing image, the target position information of the target object in the infrared remote sensing image is determined. Then, based on the target position information, a corresponding target local image is extracted from the infrared remote sensing image. Finally, a target enhanced image is generated based on the global enhanced image and the target local image of the target object. It is evident that this application achieves image enhancement by separately generating a global enhanced image and a target local image of the target object, utilizing both global and local target information of the entire image, thus avoiding an increase in the background signal-to-noise ratio and ensuring the image enhancement effect. Meanwhile, this application realizes image enhancement using image information from multiple frames of infrared remote sensing images to obtain target enhanced images, which alleviates the instability of enhancement effect caused by relying on single-frame remote sensing images, improves the enhancement effect of target enhanced images, and provides high signal-to-noise ratio data support for downstream target screening, target positioning and other purposes.
[0070] To facilitate understanding of this embodiment, a detailed description of the infrared remote sensing image enhancement method disclosed in this disclosure is provided first. The execution entity of the infrared remote sensing image enhancement method provided in this disclosure is generally a computer device with a certain computing capability. This computer device may include, for example, a terminal device, a server, or other processing devices. The terminal device may be a user equipment (UE), mobile device, user terminal, personal digital assistant (PDA), handheld device, computing device, embedded device, etc. In some possible implementations, the infrared remote sensing image enhancement method can be implemented by a processor calling computer-readable instructions stored in memory.
[0071] See Figure 1 The diagram shows a flowchart of an infrared remote sensing image enhancement method provided in this embodiment of the present disclosure. The method includes steps S101 to S104, wherein:
[0072] S101. Acquire multiple consecutive frames of infrared remote sensing images;
[0073] S102. Globally align the multiple frames of infrared remote sensing images to generate multiple registered infrared remote sensing images; fuse the multiple registered infrared remote sensing images to generate a globally enhanced image.
[0074] S103. After determining the initial position information of the target object in the global enhanced image, for each frame of the infrared remote sensing image, based on the initial position information of the target object, an initial local image of the target object is extracted from the infrared remote sensing image, and the target position information of the target object in the infrared remote sensing image is determined using a Kalman filter, the initial local image, and the position information of the target object at the previous moment corresponding to the infrared remote sensing image, wherein the target object is an object with a pixel size smaller than a preset size in the infrared remote sensing image;
[0075] S104. Based on the target location information of the target object, extract the target local image corresponding to the target object from the multi-frame infrared remote sensing images, and generate a target enhancement image based on the global enhancement image and the target local image corresponding to the target object.
[0076] The processes S101-S104 are explained in detail below.
[0077] In S101, the number of infrared remote sensing images can be set as needed. For example, image enhancement can be performed based on two or three infrared remote sensing images. These infrared remote sensing images can be remote sensing images of any acquired scene.
[0078] In S102, after acquiring multiple consecutive frames of infrared remote sensing images, since the acquisition times of the multiple frames of infrared remote sensing images are different and the coordinate systems of the infrared remote sensing images are not unified, in order to generate target enhancement images more accurately, the multiple frames of infrared remote sensing images can be globally aligned and registered. That is, the multiple frames of infrared remote sensing images are transformed to the same coordinate system to generate registered infrared remote sensing images. For example, the image with the earliest acquisition time among the multiple frames of infrared remote sensing images can be used as a reference image, and the other infrared remote sensing images can be aligned to the coordinate system corresponding to the reference image to generate registered infrared remote sensing images, thus realizing the alignment of multiple frames of infrared remote sensing images.
[0079] Considering the potential presence of noise in infrared remote sensing images, which can result in poor image quality, preprocessing can be performed on multiple frames of infrared remote sensing images before global alignment to mitigate the poor image enhancement effect caused by these issues.
[0080] Optionally, the method further includes: performing image preprocessing on the multiple frames of infrared remote sensing images respectively to generate processed infrared remote sensing images; wherein the image preprocessing includes at least one of the following: non-uniformity correction processing, bad pixel repair processing, and preliminary background suppression processing.
[0081] In practice, at least one of the following techniques—non-uniformity correction, bad pixel repair, and preliminary background suppression—can be used to preprocess multiple frames of infrared remote sensing images. Non-uniformity correction can employ a two-point correction method to eliminate fixed pattern noise in the infrared focal plane array; bad pixel repair can replace dead and overheated pixels in the infrared remote sensing image based on neighborhood median filtering; and preliminary background suppression can use top-hat transformation to suppress large areas of slowly changing background components in the infrared remote sensing image, thereby highlighting potential target points.
[0082] In image preprocessing, which includes non-uniformity correction, bad pixel repair, and preliminary background suppression, preferably, for each frame of infrared remote sensing image, non-uniformity correction is first performed on the infrared remote sensing image to generate a first intermediate infrared remote sensing image; then, bad pixel repair is performed on the first intermediate infrared remote sensing image to generate a second intermediate infrared remote sensing image; finally, preliminary background suppression is performed on the second intermediate infrared remote sensing image to obtain the processed infrared remote sensing image.
[0083] Subsequently, the processed infrared remote sensing images can be globally aligned to generate multi-frame registered infrared remote sensing images.
[0084] Here, we first perform image preprocessing on multiple frames of infrared remote sensing images to improve the image quality of the processed infrared remote sensing images, so as to ensure the image enhancement effect of the subsequently generated target enhancement images.
[0085] In practice, traditional image alignment algorithms such as the Lucas-Kanade optical flow method can be used to globally align multiple frames of infrared remote sensing images.
[0086] Optionally, the step of globally aligning the multiple frames of infrared remote sensing images to generate multiple registered infrared remote sensing images includes:
[0087] Step a1: Extract the initial gradient image of each frame of infrared remote sensing image.
[0088] Step a2: For each frame of infrared remote sensing image, the initial gradient image of the infrared remote sensing image is processed using the optical flow pyramid calculation model to generate multiple pyramid gradient images corresponding to the infrared remote sensing image, and the target number of control points are uniformly sampled from each pyramid gradient image.
[0089] Step a3: Based on the multiple control points included in the multi-frame infrared remote sensing images, determine the affine transformation matrix parameters between the second infrared remote sensing image and the first infrared remote sensing image in the multi-frame infrared remote sensing images, wherein the first infrared remote sensing image is a reference infrared remote sensing image in the multi-frame infrared remote sensing images, and the second infrared remote sensing image is another infrared remote sensing image other than the first infrared remote sensing image.
[0090] Step a4: Using the affine transformation matrix parameters, perform geometric correction and resampling on the second infrared remote sensing image to generate a registered second infrared remote sensing image, wherein the first infrared remote sensing image and the registered second infrared remote sensing image constitute the multi-frame registered infrared remote sensing image.
[0091] In implementation, the initial gradient image of each frame of infrared remote sensing image can be extracted separately. The method for generating the initial gradient image can be set according to actual needs, and this application does not impose specific limitations on it. For example, the initial gradient image of the infrared remote sensing image can be generated by using operator modules such as the Sobel operator or the Scharr operator.
[0092] After obtaining the initial gradient images corresponding to multiple frames of infrared remote sensing images, for each frame, the initial gradient image can be processed using an optical flow pyramid calculation model to generate multiple pyramid gradient images of different sizes. Then, a target number of control points are uniformly sampled from each pyramid gradient image. In practice, after determining the number of targets, the sampling points can be determined based on the image size of the pyramid gradient images to sample the target number of control points.
[0093] Considering that infrared remote sensing images may include low-texture backgrounds such as the sky, they cannot detect enough stable feature points. For example, the background is relatively uniform under a sky background, and the cloud edges are blurred. Traditional feature point detectors such as Harris and SIFT can hardly find stable feature points in such environments, resulting in feature point extraction failure and the inability to perform subsequent alignment processes. To alleviate the above problems, this application uses the region-based pyramid LK optical flow method to estimate the global motion between two frames of infrared remote sensing images. After generating the pyramid gradient image, the traditional feature points are replaced by uniformly sampling control points so that the image alignment can be performed using the sampled multiple control points.
[0094] During implementation, a reference infrared remote sensing image can be set, and other infrared remote sensing images can be aligned with the reference infrared remote sensing image. For example, the earliest infrared remote sensing image acquired among multiple frames of infrared remote sensing images can be set as the reference infrared remote sensing image.
[0095] For example, the affine transformation matrix parameters between the second infrared remote sensing image and the first infrared remote sensing image in the multi-frame infrared remote sensing images can be determined by using multiple control points included in the multi-frame infrared remote sensing images. These parameters may include rotation parameters, translation parameters, and scaling parameters. The first infrared remote sensing image is a reference infrared remote sensing image in the multi-frame infrared remote sensing images, and the second infrared remote sensing image is another infrared remote sensing image other than the first infrared remote sensing image.
[0096] After obtaining the affine transformation matrix parameters between the second infrared remote sensing image and the first infrared remote sensing image, the second infrared remote sensing image is geometrically corrected and resampled using these parameters to generate a registered second infrared remote sensing image. The first infrared remote sensing image and the registered second infrared remote sensing image constitute a multi-frame registered infrared remote sensing image.
[0097] See Figure 2 As shown, Figure 2The reference image (i.e., the reference infrared remote sensing image) can be the earliest acquired infrared remote sensing image among multiple frames, while the floating image can be any infrared remote sensing image other than the reference image. In practice, the reference image (frame at time t-1) and the floating image (frame at time t) are processed separately to generate a gradient map of the reference image (frame at time t-1) and a gradient map of the floating image (frame at time t). These gradient maps are then input into an LK pyramid calculation model for further processing. For example, a large-window (e.g., 31×31) Lucas-Kanade pyramid calculation model can be used to process the reference image gradient map, constructing a 4-layer image pyramid (i.e., 4 reference pyramid gradient images). A target number (e.g., 64) control points (not feature points) are uniformly sampled in each layer of the reference pyramid gradient image. Similarly, multiple floating pyramid gradient images corresponding to the floating image, and multiple control points on each floating pyramid gradient image, can be obtained. Subsequently, starting from the top pyramid, the iterative reweighted least squares (IRLS) method is used to solve for the six parameters of the affine transformation matrix, estimating the global affine transformation matrix. This involves generating the affine transformation matrix between the reference image and the floating image using multiple control points corresponding to both the reference and floating images. Finally, based on the estimated global motion model and the affine transformation matrix, bilinear interpolation is used to resample the floating image and align it with the reference image, resulting in multi-frame registered images.
[0098] After obtaining multi-frame registered infrared remote sensing images, the multi-frame registered infrared remote sensing images can be globally multi-frame accumulated to obtain a globally enhanced image.
[0099] Optionally, fusing the multi-frame registered infrared remote sensing images to generate a globally enhanced image includes:
[0100] Step b1: Determine the weight information of the registered infrared remote sensing image based on the image quality parameters of each frame of the registered infrared remote sensing image.
[0101] Step b2: Based on the weight information corresponding to the registered infrared remote sensing images of the multiple frames, sum the pixel information of the same pixel position in the registered infrared remote sensing images of the multiple frames to obtain the fused pixel information corresponding to the pixel position.
[0102] Step b3: Generate a global enhanced image based on the fused pixel information corresponding to each pixel position.
[0103] During implementation, the weight information of the registered infrared remote sensing images can be determined based on the image quality parameters of each frame of the registered infrared remote sensing image. For example, infrared remote sensing images with good image quality have a larger weight, while infrared remote sensing images with poor image quality have a smaller weight.
[0104] After obtaining the weight information of each registered infrared remote sensing image frame, a weighted average of multiple registered infrared remote sensing images can be accumulated to obtain a globally enhanced image. Specifically, according to the weight information corresponding to the registered infrared remote sensing images, the pixel information at the same pixel position in multiple registered infrared remote sensing images can be summed to obtain the fused pixel information corresponding to that pixel position. Then, based on the fused pixel information corresponding to each pixel position, a globally enhanced image is generated.
[0105] For example, a weighted average can be accumulated from multiple registered infrared remote sensing images to generate a globally enhanced image. For instance, the registered infrared remote sensing images {I_registered1, I_registered2, ..., I_registered_N} can be accumulated using the following formula (1), where N is the number of frames in the infrared remote sensing images.
[0106] (1)
[0107] in, The weight information for the registered infrared remote sensing image is adaptively allocated based on the image quality (such as gradient energy) of the infrared remote sensing image to suppress the influence of blurred frames. Enhance the image globally.
[0108] This application's research found that during the global multi-frame accumulation process, weight information... The design of the image quality parameters is crucial, as it determines the contribution of each frame of infrared remote sensing image to the final fusion result. Therefore, to improve the image quality of the globally enhanced image, this application sets image quality parameters based on the infrared remote sensing images and determines the weight information of the registered infrared remote sensing images. This image quality-based adaptive weight allocation significantly improves the accumulation effect. Therefore, the weight information in this application... The design incorporates multiple dimensions of image quality parameters, such as the signal-to-noise ratio, image gradient, and image blur of the entire infrared remote sensing image. The image quality of the infrared remote sensing image is obtained by comprehensively calculating the quality assessment from multiple dimensions.
[0109] Optionally, determining the weight information of the registered infrared remote sensing image based on the image quality parameters of each frame of the registered infrared remote sensing image includes:
[0110] For each frame of the registered infrared remote sensing image, the image signal-to-noise ratio of the registered infrared remote sensing image is determined based on the image mean and image variance of the registered infrared remote sensing image.
[0111] Based on the number of pixels in the gradient image corresponding to the registered infrared remote sensing image and the gradient information in multiple directions, the image gradient energy of the registered infrared remote sensing image is determined.
[0112] The second-order gradient value of the registered infrared remote sensing image is determined using the Laplacian operator, and the image blur of the registered infrared remote sensing image is determined based on the second-order gradient value.
[0113] The weight information of the registered infrared remote sensing image is determined based on the image signal-to-noise ratio, the image gradient energy, and the image blur.
[0114] During implementation, the image signal-to-noise ratio, image gradient energy, and image blur of the registered infrared remote sensing image can be calculated separately. These parameters reflect the image quality of the infrared remote sensing image, and the weight information of the registered infrared remote sensing image can be obtained based on them.
[0115] For example, the weight information of the registered infrared remote sensing image can be calculated according to the following formula (2). :
[0116] (2)
[0117] in, This represents the overall signal-to-noise ratio (SNR) of the k-th registered infrared remote sensing image (i.e., the image SNR). Indicates the signal-to-noise ratio weight; This represents the overall gradient energy (i.e., image gradient energy) of the k-th registered infrared remote sensing image. This represents the sharpness weight, where gradient energy reflects the richness of image details; the more detailed the image frame, the higher the weight. This represents the overall blur (i.e., image blur) of the k-th registered infrared remote sensing image. This represents the deblurring factor, which is set here to suppress the impact of blurred frames caused by platform jitter or out-of-focus issues. In the formula above... , , It can be set based on prior experience, for example It can be set to 0.2. It can be set to 1.0. It can be set to 0.5.
[0118] The following explains the calculation of image signal-to-noise ratio, image gradient energy, and image blur.
[0119] During implementation, the image signal-to-noise ratio can be calculated according to formula (3). :
[0120] (3)
[0121] in, The average value of the registered infrared remote sensing image. The image variance of the registered infrared remote sensing image.
[0122] The image gradient energy can be calculated using formula (4). :
[0123] (4)
[0124] Where M represents the number of pixels in the gradient image corresponding to the k-th registered infrared remote sensing image, which can be obtained by multiplying the length of the gradient image by the image width; , This represents the gradient of the k-th registered infrared remote sensing image in the x and y directions. Dividing by M in the above formula is for normalization to calculate the average gradient energy, ensuring that the gradient energy is an energy value independent of image resolution.
[0125] Image blur can be calculated using formula (5). :
[0126] (5)
[0127] in This indicates that the second-order gradient is calculated using the Laplacian operator on the k-th registered infrared remote sensing image to identify regions or points of drastic change. Generally, in sharp images, the edges are sharp, and the Laplacian operator's response value is large (positive or negative) at the edges, with a wide overall range of response values. However, in blurred images, because the edges are smoothed and the intensity changes slowly, the Laplacian operator's response value is generally small, with a narrow overall range of response values.
[0128] Further use It can calculate the variance of the response value of the entire image after Laplacian filtering. This is a smaller value added to avoid the denominator having a value of 0. In the above formula (5), if the variance is large, it means that the Laplacian response values are widely distributed, with some being very large (sharp edges) and some being very small (flat areas), indicating that the image is clear. Conversely, if the variance is small, it means that the Laplacian response values are concentrated around 0, without any particularly prominent high-frequency details, indicating that the image is blurry.
[0129] This application can accurately determine the weight information of infrared remote sensing images through the above process, so as to better utilize the weight information to perform global image fusion and improve the image effect of the global enhanced image.
[0130] In S103, the initial position information of the target object in the global augmented image can be determined first; that is, target detection can be performed on the global augmented image to obtain the initial position information of the target object. For example, in the global augmented image... The initial position information of potential targets (i.e., target objects) is initially determined by applying low-rank sparse algorithms such as the Infrared Patch-Image Model (IPI) algorithm and the Sparse Representation with Weighted Sparsity (SRWS) algorithm. , ), i=1,2,...,n, where n represents the number of potential targets.
[0131] The target object in this application is a very small target, that is, an object whose pixel size in an infrared remote sensing image is smaller than a preset size. The preset size can be set according to requirements, such as 3 pixels × 3 pixels.
[0132] After determining the initial position information of the target object, a Kalman filter can be used to track the target object based on the initial position information to determine the target position information of the target object in each frame of infrared remote sensing image.
[0133] In implementation, for each frame of infrared remote sensing image, based on the initial position information of the target object, an initial partial image of the target object can be extracted from the infrared remote sensing image (or the processed infrared remote sensing image) according to a set target size. The target size can be set according to the size of the target object; for example, the target size can be a target multiple of the target object's size, or it can be determined based on a preset size. For instance, if the preset size is 3 pixels × 3 pixels and the target multiple is 3, then the target size can be 9 pixels × 9 pixels. This process ensures that the initial partial image can include the entire target object.
[0134] Furthermore, the target location information of the target object in the infrared remote sensing image can be determined by utilizing the Kalman filter, the initial local image, and the target object's position information from the previous moment in the infrared remote sensing image. For example, the initial local image can be used to perform centroid correction on the target object's position information from the previous moment to obtain the corrected position information. Then, the Kalman filter can be used to track and correct the corrected position information to generate the target object's position information at the current moment, thus obtaining the target location information of the target object in the infrared remote sensing image.
[0135] Optionally, determining the target location information of the target object in the infrared remote sensing image using a Kalman filter, the initial local image, and the target object's location information from the previous moment corresponding to the infrared remote sensing image includes:
[0136] Step c1: Correct the energy centroid coordinates within the neighborhood of the target object's position information at the previous moment to generate corrected position information;
[0137] Step c2: Using the Kalman filter algorithm and the determined state vector at the current moment, track and correct the corrected position information to generate the predicted position information of the target object at the current moment;
[0138] Step c3: Obtain the detection location information of the target object detected at the current time, and based on the pixel information of the initial local image, correct the energy centroid coordinates in the neighborhood of the detection location information to generate the corrected detection location information;
[0139] Step c4: Based on the predicted location information and the corrected detection location information, determine the target location information of the target object in the infrared remote sensing image.
[0140] For the current infrared remote sensing image (e.g., time t+1), in order to improve the prediction accuracy of the Kalman filter, the position information of the target object at the previous time (e.g., time t) can be corrected by adjusting the energy centroid coordinates in the neighborhood to generate more accurate corrected position information.
[0141] For example, the initial local image of the target object from the previous moment can be used to correct the energy centroid coordinates within the neighborhood. In practice, the position information of the target object is corrected based on the pixel information of each pixel in the initial local image to generate the corrected position information. For example, the corrected position information can be generated using the following formulas (6) and (7).
[0142] (6)
[0143] (7)
[0144] in( , () represents the corrected location information. , (This refers to the location information before correction.) This represents the pixel information of each pixel in the initial local image.
[0145] The above-mentioned correction of the energy centroid coordinates within the neighborhood achieves accurate local correction. Specifically, the local potential target is regarded as an energy point, and the energy centroid coordinates of the target object can be calculated using the above formulas (6) and (7), thus achieving robustness against noise in the local area. By calculating the energy centroid of the target within the initial local image, the sub-pixel level position information of the target object can be obtained more accurately, that is, the accuracy of the corrected position information is high.
[0146] Construct a motion model, for example, assuming the target object moves at a constant speed within a computation time (e.g., 2 seconds), and establish the state vector at the current moment. ,in The location information of the target object i (such as the corrected location information). Let be the velocity of target object i at the previous moment. Then, based on the state vector of target object i in the motion model, the Kalman filter algorithm is used to predict the predicted position information of the target object at the current moment.
[0147] Since infrared remote sensing images are continuously detected in real time, the current infrared remote sensing image can be acquired. Target detection is then performed on the current infrared remote sensing image to determine the target object's location at that moment. Based on the pixel information of the initial local image at the current moment, the acquired detection location information is corrected using the energy centroid coordinates within the neighborhood, generating corrected detection location information. The process of correcting the energy centroid coordinates within the neighborhood can be referred to the description above and will not be repeated here.
[0148] Then, using the predicted location information and the corrected detection location information, the target object's location information in the infrared remote sensing image at the current moment can be determined. For example, the predicted location information and the corrected detection location information can be averaged, and the average location can be used as the target location information. Alternatively, either the predicted location information or the corrected detection location information can be selected as the target object's location information.
[0149] By utilizing the energy centroid coordinate correction within the neighborhood and the Kalman filter, the target location information of the target object can be determined more accurately and efficiently, providing data support for subsequent determination of the local image of the target object.
[0150] Optionally, in step c4, determining the target location information of the target object in the infrared remote sensing image based on the predicted location information and the corrected detection location information includes: determining the position offset value between the predicted location information and the corrected detection location information; and determining the predicted location information as the target location information of the target object when the position offset value is less than a preset offset value.
[0151] When the position offset value is greater than or equal to the preset offset value, the mean position information between the predicted position information and the corrected detection position information is determined; and the mean position information is determined as the target position information of the target object.
[0152] During implementation, the positional offset between the predicted and corrected detection positions can be determined, such as the Euclidean distance between them. If the offset is less than a preset offset, it indicates a small deviation between the predicted and corrected detection positions, and the predicted position can be directly selected as the target position. The preset offset can be set according to actual needs; for example, it can be one pixel.
[0153] When the position offset value is greater than or equal to the preset offset value, it indicates that the deviation between the predicted position information and the corrected detection position information is large. In this case, the average position information between the predicted position information and the corrected detection position information can be determined, and then the average position information can be determined as the target position information of the target object to balance the predicted position information and the corrected detection position information and determine the target position information more accurately.
[0154] For example, see Figure 3 As shown, assume that target object 1 is detected in the image at time t, and its coordinates are ( , That is, the position information of the previous moment, after being corrected by the coordinates of the energy centroid within the domain, yields the corrected precise coordinates of target object 1 as ( , This refers to the corrected position information. Subsequently, the corrected position information is used as input to the Kalman filter algorithm to update the target state. And using the Kalman filter algorithm based on the updated target state The coordinates of target object 1 are predicted in the image at time t+1. , This refers to the predicted location information at the current moment. Also, it involves obtaining the coordinates of the target object 1 detected in a single frame of the infrared remote sensing image at time t+1. , That is, detecting location information, based on the coordinates of target object 1 ( , ), calculate the coordinates of the energy centroid at time t+1 ( , That is, the corrected detection location information.
[0155] Subsequently, based on the coordinates of target object 1 ( , ) and energy center of mass coordinates ( , ) Determine the target location information. For example, if the coordinates are ( , ) and energy center of mass coordinates ( , If the offset between () is sub-pixel (i.e., less than 1 pixel), then the coordinates () can be used directly. , This is used as the target location information. If it exceeds one pixel, the average of the two coordinates is calculated to obtain the target location information, i.e. .
[0156] In step S104, a partial image of the target object is extracted from multiple frames of infrared remote sensing images, centered on the target object's location information and according to a set size. Alternatively, a partial image of the target object can be extracted from multiple processed infrared remote sensing images. The set size is larger than the target object's size, such as 2 or 3 times the target object's size; or, the extracted size can be set according to a preset size, such as 5×5 pixels if the preset size is 2 pixels × 2 pixels.
[0157] For example, a local image of the target object can be embedded into a global enhancement image to generate a target enhancement image. This target enhancement image enhances the local area where the target object is located, while other background areas are not enhanced, thereby improving the signal-to-noise ratio of the small target.
[0158] In one optional implementation, generating a target enhancement image based on the global enhancement image and the target local image corresponding to the target object includes:
[0159] Step d1: Fuse the multiple frames of the target local image corresponding to the target object to generate a fused local image.
[0160] Step d2: Determine the adaptive fusion weights corresponding to the fused local image, wherein the adaptive fusion weights are positively correlated with the signal-to-noise ratio of the fused local image.
[0161] Step d3: Based on the adaptive fusion weights, embed the fused local image into the global enhanced image to generate the target enhanced image.
[0162] In step d1, multiple local images corresponding to the target object are aligned and then fused to generate a fused local image. For example, spatial alignment can be performed using the center position of the target local images as a reference. After aligning multiple target local images, the pixel values corresponding to the same pixel position in the multiple target local images are averaged to complete the fusion of local images, resulting in a fused local image. This fused local image is a locally enhanced image with a significantly improved signal-to-noise ratio, focused on the target.
[0163] In practice, the target local images of multiple frames corresponding to the target object can be fused according to the following formula (8) to generate a fused local image.
[0164] (8)
[0165] in, This represents the fused local image corresponding to target object i, where N is the number of frames in the infrared remote sensing image. This represents a local image k of the target extracted from an infrared remote sensing image k.
[0166] In step d2, the adaptive fusion weights corresponding to the fused local images are determined so that the fused local images can be embedded into the global enhanced image based on the adaptive fusion weights to achieve image enhancement.
[0167] The adaptive fusion weights are positively correlated with the signal-to-noise ratio (SNR) of the fused local images. That is, the higher the SNR, the more reliable the target is, and the greater the weight it is assigned during fusion.
[0168] Optionally, determining the adaptive fusion weights corresponding to the fused local image includes: determining the pixel mean based on the pixel information of multiple target pixels belonging to the target object in the fused local image; determining the pixel standard deviation based on the pixel information of other pixels in the fused local image besides the target pixels; calculating the local signal-to-noise ratio based on the pixel mean and the pixel standard deviation; and generating the adaptive fusion weights corresponding to the fused local image based on the local signal-to-noise ratio and the set weight extreme values.
[0169] During implementation, the pixel information of multiple target pixels belonging to the target object in the fused local image is determined, and the average value of the pixel information of the multiple target pixels is calculated to obtain the pixel mean. The standard deviation of the pixel information of other pixels in the fused local image, excluding the target pixels, is calculated to obtain the pixel standard deviation. Based on the pixel mean and pixel standard deviation, the local signal-to-noise ratio is calculated. For example, the pixel mean is subtracted from the pixel standard deviation, and the difference is divided by the pixel standard deviation to obtain the local signal-to-noise ratio.
[0170] See Figure 4 As shown, the size of the merged local image is 5 pixels × 5 pixels, and the size of the target object is 2 pixels × 2 pixels. For illustration, we can use the four target pixels included in the target object (e.g., Figure 4 The pixel mean is calculated using the pixel information of the yellow pixels in (b); and the pixel value is calculated using the pixel information of other pixels in the annular region surrounding the target object (e.g., yellow pixels). Figure 4 The pixel information of the pixels in the gray area (b) is used to calculate the pixel standard deviation. Then, the local signal-to-noise ratio is calculated based on the pixel mean and pixel standard deviation. .
[0171] The local signal-to-noise ratio is processed using an activation function to obtain activation values; the maximum value is subtracted from the minimum value in the weight extrema to obtain the extremum deviation; the extremum deviation is multiplied by the activation value, and the product is added to the minimum value to obtain the adaptive fusion weights corresponding to the target local image.
[0172] For example, the adaptive fusion weights are determined according to the following formulas (9) and (10):
[0173] (9)
[0174] (10)
[0175] in, The average pixel value. The standard deviation of pixels. This represents the local signal-to-noise ratio. It is the minimum value among the extreme values of the weights. It is the maximum value. and for The upper and lower limits, Slope factor The signal-to-noise ratio threshold. , , , The value can be set according to business needs, for example... It can be set to 1.25. It can be set to 0.2. It can be set to 0.25. It can be set to 0.1.
[0176] The above process generates adaptive fusion weights related to the signal-to-noise ratio, which can then be used to accurately enhance the image and ensure the image enhancement effect.
[0177] In step d3, after obtaining the adaptive fusion weights, the fused local image is embedded into a specific location in the global enhancement image according to the adaptive fusion weights to generate the target enhancement image, wherein the specific location of the embedding is determined based on the initial location information of the target object. For example, the fusion of the target local image and the global enhancement image can be achieved using the following formula (11):
[0178] (11)
[0179] in, This represents the pixel information of the pixels in the fusion region of the target enhanced image. For adaptive fusion weights, This indicates the fusion of pixel information from pixels in a local image. This represents pixel information that matches pixels in the fused local image within the fused region of the globally enhanced image.
[0180] For example, the fusion region of a globally enhanced image can be the region where the target object is located, determined based on the initial position information of the target object.
[0181] The following combination Figure 5 The method for enhancing infrared remote sensing images according to this application is illustrated by way of example, the method comprising:
[0182] Acquire multiple frames of infrared remote sensing images, and perform image preprocessing on the multiple frames of infrared remote sensing images to generate processed infrared remote sensing images. The image preprocessing includes at least one of the following: non-uniformity correction processing, bad pixel repair processing, and preliminary background suppression processing.
[0183] Next, global background matching is performed on the processed infrared remote sensing images. Specifically, after extracting the initial gradient image of the processed infrared remote sensing image, a multi-layer pyramid gradient image is constructed using the optical flow pyramid calculation model. Control points are uniformly sampled from each layer of the pyramid gradient image. Starting from the top layer pyramid gradient image, the six parameters of the affine transformation matrix are solved using the iterative reweighted least squares (IRLS) method. Then, using the estimated global motion model and the affine transformation matrix, other infrared remote sensing images are resampled using bilinear interpolation and aligned with the reference infrared remote sensing image to obtain multi-frame registered infrared remote sensing images.
[0184] The globally enhanced image is obtained by globally accumulating the registered infrared remote sensing images. The global multi-frame image accumulation process can be referred to the previous explanation of steps b1 to b3, and will not be detailed here.
[0185] Globally enhanced image Perform coarse detection of potential targets to determine the initial position information of minimal targets (i.e., target objects). For example, a low-rank sparse algorithm can be used for coarse detection of potential targets.
[0186] Next, candidate target original image patches are cropped based on the initial position information of the target object. That is, based on the initial position information of the target object, initial local images of the target object are extracted from the processed infrared remote sensing images. Then, a target motion model is established, and a Kalman filter is used to perform local precise tracking of the target object based on the target motion model and the initial local images to determine the target position information of the target object in the infrared remote sensing images.
[0187] After obtaining the target location information of the target object, based on this information and according to the set dimensions, a local image of the target object is extracted from the multi-frame processed infrared remote sensing image. Multiple local images of the same target object are then locally multi-frame accumulated to obtain the fused local image corresponding to the target object, i.e., the locally enhanced image. .
[0188] Local enhancement image With global enhancement image Adaptive smoothing fusion is performed to obtain the enhanced temporal image, i.e., the target enhanced image. The adaptive smoothing fusion process can be referred to the previous description of steps d1 to d3, and will not be described in detail here.
[0189] The target enhancement image obtained by this application can provide better data support for downstream processing. For example, downstream processes can perform target detection on the target enhancement image to obtain the location information of each target more accurately; or, the target enhancement image can be processed to achieve target tracking, localization, etc.
[0190] Corresponding to the embodiments of the aforementioned infrared remote sensing image enhancement methods, this application also provides embodiments of infrared remote sensing image enhancement devices. Figure 6 A schematic diagram of the infrared remote sensing image enhancement device provided in this application, specifically including:
[0191] The acquisition module 601 is used to acquire multiple consecutive frames of infrared remote sensing images;
[0192] The first generation module 602 is used to globally align the multi-frame infrared remote sensing images to generate multi-frame registered infrared remote sensing images; and to fuse the multi-frame registered infrared remote sensing images to generate a globally enhanced image.
[0193] The determining module 603 is used to, after determining the initial position information of the target object in the global enhanced image, extract an initial local image of the target object from the infrared remote sensing image for each frame of the infrared remote sensing image based on the initial position information of the target object, and determine the target position information of the target object in the infrared remote sensing image using a Kalman filter, the initial local image, and the position information of the target object corresponding to the previous moment in the infrared remote sensing image, wherein the target object is an object with a pixel size smaller than a preset size in the infrared remote sensing image;
[0194] The second generation module 604 is used to extract the target local image corresponding to the target object from the multi-frame infrared remote sensing images according to the target location information of the target object, and generate a target enhancement image based on the global enhancement image and the target local image corresponding to the target object.
[0195] In an optional embodiment, the apparatus further includes: a preprocessing module 605, configured to:
[0196] The multiple frames of infrared remote sensing images are preprocessed to generate processed infrared remote sensing images; wherein the image preprocessing includes at least one of the following: non-uniformity correction processing, bad pixel repair processing, and preliminary background suppression processing.
[0197] When the first generation module 602 performs global alignment on the multiple frames of infrared remote sensing images to generate multiple registered infrared remote sensing images, it is used to: perform global alignment on the multiple frames of processed infrared remote sensing images to generate multiple registered infrared remote sensing images.
[0198] In an optional implementation, when the first generation module 602 performs global alignment on the multiple frames of infrared remote sensing images to generate multiple registered infrared remote sensing images, it is used to:
[0199] Extract the initial gradient image for each frame of the infrared remote sensing image;
[0200] For each frame of infrared remote sensing image, the initial gradient image of the infrared remote sensing image is processed using the optical flow pyramid calculation model to generate multiple pyramid gradient images corresponding to the infrared remote sensing image, and the target number of control points are uniformly sampled from each pyramid gradient image.
[0201] Based on the multiple control points included in the multi-frame infrared remote sensing images, the affine transformation matrix parameters between the second infrared remote sensing image and the first infrared remote sensing image in the multi-frame infrared remote sensing images are determined, wherein the first infrared remote sensing image is a reference infrared remote sensing image in the multi-frame infrared remote sensing images, and the second infrared remote sensing image is another infrared remote sensing image other than the first infrared remote sensing image.
[0202] Using the affine transformation matrix parameters, the second infrared remote sensing image is geometrically corrected and resampled to generate a registered second infrared remote sensing image, wherein the first infrared remote sensing image and the registered second infrared remote sensing image constitute the multi-frame registered infrared remote sensing image.
[0203] In an optional implementation, when the first generation module 602 fuses the multi-frame registered infrared remote sensing images to generate a globally enhanced image, it is used to:
[0204] The weight information of the registered infrared remote sensing image is determined based on the image quality parameters of each frame of the registered infrared remote sensing image.
[0205] Based on the weight information corresponding to the registered infrared remote sensing images of the multi-frames, the pixel information of the same pixel position in the registered infrared remote sensing images of the multi-frames is summed to obtain the fused pixel information corresponding to the pixel position.
[0206] A globally enhanced image is generated based on the fused pixel information corresponding to each pixel position.
[0207] In an optional implementation, when the first generation module 602 determines the weight information of the registered infrared remote sensing image based on the image quality parameters of each frame of the registered infrared remote sensing image, it is used to:
[0208] For each frame of the registered infrared remote sensing image, the image signal-to-noise ratio of the registered infrared remote sensing image is determined based on the image mean and image variance of the registered infrared remote sensing image.
[0209] Based on the number of pixels in the gradient image corresponding to the registered infrared remote sensing image and the gradient information in multiple directions, the image gradient energy of the registered infrared remote sensing image is determined.
[0210] The second-order gradient value of the registered infrared remote sensing image is determined using the Laplacian operator, and the image blur of the registered infrared remote sensing image is determined based on the second-order gradient value.
[0211] The weight information of the registered infrared remote sensing image is determined based on the image signal-to-noise ratio, the image gradient energy, and the image blur.
[0212] In an optional implementation, when the determining module 603 determines the target position information of the target object in the infrared remote sensing image using the Kalman filter, the initial local image, and the target object's position information corresponding to the previous moment in the infrared remote sensing image, it is used to:
[0213] The position information of the target object at the previous moment is corrected by the energy centroid coordinates in the neighborhood to generate the corrected position information;
[0214] Using the Kalman filter algorithm and the determined state vector at the current moment, the corrected position information is tracked and corrected to generate the predicted position information of the target object at the current moment;
[0215] The detection location information of the target object detected at the current moment is obtained, and the energy centroid coordinates in the neighborhood are corrected based on the pixel information of the initial local image to generate the corrected detection location information.
[0216] Based on the predicted location information and the corrected detection location information, the target location information of the target object in the infrared remote sensing image is determined.
[0217] In an optional implementation, when determining the target location information of the target object in the infrared remote sensing image based on the predicted location information and the corrected detection location information, the determining module 603 is used to:
[0218] Determine the position offset value between the predicted position information and the corrected detection position information;
[0219] When the position offset value is less than a preset offset value, the predicted position information is determined as the target position information of the target object;
[0220] When the position offset value is greater than or equal to the preset offset value, the mean position information between the predicted position information and the corrected detection position information is determined; the mean position information is determined as the target position information of the target object.
[0221] In an optional implementation, when the second generation module 604 generates a target enhancement image based on the global enhancement image and the target local image corresponding to the target object, it is used to:
[0222] The target local images corresponding to the target object are fused together to generate a fused local image;
[0223] Determine the adaptive fusion weights corresponding to the fused local image, wherein the adaptive fusion weights are positively correlated with the signal-to-noise ratio of the fused local image;
[0224] Based on the adaptive fusion weights, the fused local image is embedded into the global enhanced image to generate the target enhanced image.
[0225] In one optional implementation, the second generation module 604, when determining the adaptive fusion weights corresponding to the fused local image, is used to:
[0226] Based on the pixel information of multiple target pixels belonging to the target object in the fused local image, the pixel mean is determined;
[0227] Based on the pixel information of other pixels in the fused local image besides the target pixel, the pixel standard deviation is determined;
[0228] Calculate the local signal-to-noise ratio based on the pixel mean and the pixel standard deviation;
[0229] Based on the local signal-to-noise ratio and the set weight extreme values, adaptive fusion weights are generated corresponding to the fused local image.
[0230] The specific implementation process of the functions and roles of each unit in the above device can be found in the implementation process of the corresponding steps in the above method, and will not be repeated here.
[0231] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this application according to actual needs. Those skilled in the art can understand and implement this without creative effort.
[0232] This application also provides a computer-readable storage medium storing a computer program that can be used to execute the infrared remote sensing image enhancement method described in the above embodiments.
[0233] This application also provides a computer device, see [link to relevant documentation] Figure 7The diagram shown is a structural schematic of the computer device provided in this application. At the hardware level, the computer device includes a processor, an internal bus, a network interface, memory, and non-volatile memory, and may also include other hardware required for business operations. The processor reads the corresponding computer program from the non-volatile memory into memory and then runs it to implement the infrared remote sensing image enhancement method described in the above embodiments. Of course, in addition to software implementation, this specification does not exclude other implementation methods, such as logic devices or a combination of hardware and software, etc. That is to say, the execution subject of the following processing flow is not limited to individual logic units, but can also be hardware or logic devices.
[0234] Those skilled in the art will understand that embodiments of this application can be provided as methods, apparatus, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0235] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0236] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0237] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1One or more processes and / or boxes Figure 1 The steps of the functions specified in one or more boxes. In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0238] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0239] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0240] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0241] This specification can be described in the general context of computer-executable instructions that are executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform a specific task or implement a specific abstract data type. This specification can also be practiced in distributed computing environments, where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.
[0242] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to interchangeably. Each embodiment focuses on describing the differences from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments.
[0243] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.
Claims
1. A method for enhancing infrared remote sensing images, characterized in that, The method includes: Acquire multiple consecutive frames of infrared remote sensing images; The multi-frame infrared remote sensing images are globally aligned to generate multi-frame registered infrared remote sensing images. For each registered infrared remote sensing image, the image signal-to-noise ratio (SNR) is determined based on the image mean and image variance. The image gradient energy is determined based on the number of pixels in the gradient image corresponding to the registered infrared remote sensing image and gradient information in multiple directions. The second-order gradient value of the registered infrared remote sensing image is determined using the Laplacian operator, and the image blurriness is determined based on the second-order gradient value. The weight information of the registered infrared remote sensing image is determined based on the image SNR, the image gradient energy, and the image blurriness. Based on the weight information corresponding to each of the multi-frame registered infrared remote sensing images, the pixel information at the same pixel position in the multi-frame registered infrared remote sensing images is summed to obtain the fused pixel information corresponding to the pixel position. A globally enhanced image is generated based on the fused pixel information corresponding to each pixel position. After determining the initial position information of the target object in the global enhanced image, for each frame of the infrared remote sensing image, based on the initial position information of the target object, an initial local image of the target object is extracted from the infrared remote sensing image. Then, using a Kalman filter, the initial local image, and the position information of the target object corresponding to the previous moment in the infrared remote sensing image, the target position information of the target object in the infrared remote sensing image is determined. The target object is an object with a pixel size smaller than a preset size in the infrared remote sensing image. Based on the target location information of the target object, a local target image corresponding to the target object is extracted from the infrared remote sensing image, and a target enhancement image is generated based on the global enhancement image and the local target image corresponding to the target object.
2. The method according to claim 1, characterized in that, The method further includes: The multiple frames of infrared remote sensing images are preprocessed to generate processed infrared remote sensing images; wherein the image preprocessing includes at least one of the following: non-uniformity correction processing, bad pixel repair processing, and preliminary background suppression processing. Global alignment of the multiple frames of infrared remote sensing images to generate multiple registered infrared remote sensing images includes: Global alignment is performed on multiple frames of the processed infrared remote sensing images to generate multiple frames of registered infrared remote sensing images.
3. The method according to claim 1, characterized in that, The step of globally aligning the multiple frames of infrared remote sensing images to generate multiple registered infrared remote sensing images includes: Extract the initial gradient image for each frame of the infrared remote sensing image; For each frame of infrared remote sensing image, the initial gradient image of the infrared remote sensing image is processed using the optical flow pyramid calculation model to generate multiple pyramid gradient images corresponding to the infrared remote sensing image, and the target number of control points are uniformly sampled from each pyramid gradient image. Based on the multiple control points included in the multi-frame infrared remote sensing images, the affine transformation matrix parameters between the second infrared remote sensing image and the first infrared remote sensing image in the multi-frame infrared remote sensing images are determined, wherein the first infrared remote sensing image is a reference infrared remote sensing image in the multi-frame infrared remote sensing images, and the second infrared remote sensing image is another infrared remote sensing image other than the first infrared remote sensing image. Using the affine transformation matrix parameters, the second infrared remote sensing image is geometrically corrected and resampled to generate a registered second infrared remote sensing image, wherein the first infrared remote sensing image and the registered second infrared remote sensing image constitute the multi-frame registered infrared remote sensing image.
4. The method according to claim 1, characterized in that, The step of determining the target location information of the target object in the infrared remote sensing image by using a Kalman filter, the initial local image, and the target object's location information corresponding to the previous moment in the infrared remote sensing image includes: The position information of the target object at the previous moment is corrected by the energy centroid coordinates in the neighborhood to generate the corrected position information; Using the Kalman filter and the determined state vector at the current moment, the corrected position information is tracked and corrected to generate the predicted position information of the target object at the current moment; The detection location information of the target object detected at the current moment is obtained, and the energy centroid coordinates in the neighborhood are corrected based on the pixel information of the initial local image to generate the corrected detection location information. Based on the predicted location information and the corrected detection location information, the target location information of the target object in the infrared remote sensing image is determined.
5. The method according to claim 4, characterized in that, Determining the target location information of the target object in the infrared remote sensing image based on the predicted location information and the corrected detection location information includes: Determine the position offset value between the predicted position information and the corrected detection position information; When the position offset value is less than a preset offset value, the predicted position information is determined as the target position information of the target object; When the position offset value is greater than or equal to the preset offset value, the mean position information between the predicted position information and the corrected detection position information is determined; the mean position information is determined as the target position information of the target object.
6. The method according to any one of claims 1-5, characterized in that, The step of generating a target enhancement image based on the global enhancement image and the target local image corresponding to the target object includes: The target local images corresponding to the target object are fused together to generate a fused local image; Determine the adaptive fusion weights corresponding to the fused local image, wherein the adaptive fusion weights are positively correlated with the signal-to-noise ratio of the fused local image; Based on the adaptive fusion weights, the fused local image is embedded into the global enhanced image to generate the target enhanced image.
7. The method according to claim 6, characterized in that, Determining the adaptive fusion weights corresponding to the fused local image includes: Based on the pixel information of multiple target pixels belonging to the target object in the fused local image, the pixel mean is determined; Based on the pixel information of other pixels in the fused local image besides the target pixel, the pixel standard deviation is determined; Calculate the local signal-to-noise ratio based on the pixel mean and the pixel standard deviation; Based on the local signal-to-noise ratio and the set weight extreme values, adaptive fusion weights are generated corresponding to the fused local image.
8. An infrared remote sensing image enhancement device, characterized in that, The device includes: The acquisition module is used to acquire multiple consecutive frames of infrared remote sensing images; The first generation module is used to globally align the multiple frames of infrared remote sensing images to generate multiple registered infrared remote sensing images; for each registered infrared remote sensing image, the module determines the image signal-to-noise ratio (SNR) of the registered infrared remote sensing image based on the image mean and image variance; determines the image gradient energy of the registered infrared remote sensing image based on the number of pixels in the gradient image corresponding to the registered infrared remote sensing image and gradient information in multiple directions; determines the second-order gradient value of the registered infrared remote sensing image using the Laplacian operator, and determines the image blur of the registered infrared remote sensing image based on the second-order gradient value; determines the weight information of the registered infrared remote sensing image based on the image SNR, the image gradient energy, and the image blur; sums the pixel information at the same pixel position in the multiple registered infrared remote sensing images based on the weight information corresponding to each pixel position to obtain the fused pixel information corresponding to the pixel position; and generates a globally enhanced image based on the fused pixel information corresponding to each pixel position. The determination module is used to, after determining the initial position information of the target object in the global enhanced image, extract an initial local image of the target object from the infrared remote sensing image for each frame of the infrared remote sensing image based on the initial position information of the target object, and determine the target position information of the target object in the infrared remote sensing image using a Kalman filter, the initial local image, and the position information of the target object corresponding to the previous moment in the infrared remote sensing image, wherein the target object is an object with a pixel size smaller than a preset size in the infrared remote sensing image; The second generation module is used to extract the target local image corresponding to the target object from the multi-frame infrared remote sensing images according to the target location information of the target object, and generate a target enhancement image based on the global enhancement image and the target local image corresponding to the target object.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, it implements the steps of the method according to any one of claims 1-7.
10. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, The processor performs the steps of the method according to any one of claims 1-7.
Citation Information
Patent Citations
Target tracking method and device, equipment and storage medium
CN115222774A
Infrared image enhancement method and system based on local phase correlation
CN120997061A
Image processing method and apparatus, and electronic device and medium
WO2023066147A1