A bipolar small target resolution adaptive local enhancement method

CN122820443APending Publication Date: 2026-09-25XI AN JIAOTONG UNIV +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611004827.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-07
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0007]本发明实施例的目的是提供一种双极性小目标分辨率自适应局部增强方法,通过建立尺度因子自适应确定形态学结构元素尺寸、分离构建亮暗双背景并独立提取双极性响应、结合几何约束筛选与软掩码局部融合,提高了多分辨率输入下亮暗小目标极性保持增强的稳定性和准确性,解决了现有形态学增强方法中参数固定难以跨分辨率迁移、暗目标极性失真以及复杂背景结构误增强的技术问题

Benefits of technology

1.通过建立尺度因子并根据输入图像分辨率自适应搜索最优形态学结构元素尺寸,同时将小目标几何筛选阈值与尺度因子进行联动缩放,本发明实现了算法参数在不同分辨率输入条件下的自动适配与稳定迁移,使得小目标的物理尺寸定义在像素域中保持一致,有效避免了传统形态学增强方法因采用固定结构元素尺寸和固定筛选阈值而导致的分辨率变化时增强效果不一致、参数难以统一配置的问题,显著提高了多分辨率场景下小目标增强的鲁棒性与工程适用性;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122820443A_ABST
    Figure CN122820443A_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of image processing, and discloses a bipolar small target resolution adaptive local enhancement method, which comprises the following steps: acquiring an input image and establishing a scale factor, and adaptively determining an optimal structure element size; constructing a bright-dark double background image and independently extracting a bright-dark response image; obtaining a candidate region through nonlinear enhancement and geometric constraint screening; generating a smooth soft mask and a gain image, determining an effective mask and a gain after conflict resolution; and fusing to obtain a polarity-maintained enhanced output image. Through scale adaptation, bright-dark double background separation, geometric constraint screening and polarity-maintained local fusion, the stability and accuracy of bipolar small target enhancement under multi-resolution input are effectively improved, and the technical problems that the parameters of the existing morphological enhancement method are difficult to migrate across resolutions, the polarity of dark targets is easy to distort, and the complex background structure is easy to be misenhanced are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of image processing and computer vision technology, and in particular to a bipolar small target resolution adaptive local enhancement method. Background Technology

[0002] In applications such as infrared early warning, long-range reconnaissance, and detection of small targets, the imaged target often appears as a bright or dark spot occupying very few pixels, with low contrast and easily submerged in complex backgrounds. Local enhancement processing is needed to improve the target's salience. Existing small target enhancement methods still have many shortcomings when dealing with complex backgrounds, multi-resolution inputs, and scenarios where bipolar targets coexist.

[0003] One existing approach is the full-image contrast enhancement method, such as histogram equalization, adaptive histogram equalization, and Retinex-like methods. These methods improve target visibility by adjusting the grayscale distribution or local contrast of the entire image. However, their enhancement effect covers the entire image, simultaneously changing background layers, edges, and textures. They tend to enhance non-target structures in the background as well, making it difficult to meet the engineering requirement of maintaining a stable background.

[0004] The second category consists of methods based on filter residuals or local differences, such as high-pass filtering, neighborhood difference, and local residual enhancement. These methods enhance suspected targets by extracting high-frequency or local abrupt change components, but they do not explicitly distinguish between bright and dark targets, and lack effective target geometric constraints. They are prone to misjudging large target edges and striped textures as targets, and their fixed parameters have poor adaptability to resolution changes.

[0005] Thirdly, there is the morphological Top-hat transform method, which constructs a background image using opening and closing operations and extracts the responses of bright and dark small targets by the difference between the original image and the background image. However, traditional implementations typically treat bright and dark responses as uniform positive enhancements without polarity preservation, leading to distortion of dark target attributes and making it difficult to achieve the enhancement effect of "brighter brighter and darker darker". In addition, this type of method lacks a connected component filtering mechanism based on geometric features such as area, width and height, aspect ratio, and fill rate, making large target edges and complex background structures prone to mis-enhancement; the size of structural elements and filtering thresholds are often fixed and cannot be adaptively adjusted with changes in input resolution, resulting in poor parameter transferability; and there is a lack of conflict resolution mechanism when bright and dark targets are spatially adjacent or overlap, affecting the stability of enhancement.

[0006] In summary, existing methods still have shortcomings in terms of bipolarity preservation, parameter adaptation, geometric constraint screening, and background preservation. There is a need to provide an enhancement method that can perform polarity preservation, local enhancement, geometric screening, and adaptive adjustment for small bright and dark targets respectively. Summary of the Invention

[0007] The purpose of this invention is to provide a resolution-adaptive local enhancement method for bipolar small targets. By establishing a scale factor adaptively to determine the size of morphological structural elements, separating and constructing bright and dark dual backgrounds and independently extracting bipolar responses, and combining geometric constraint screening and soft masking local fusion, this method improves the stability and accuracy of polarity preservation enhancement of bright and dark small targets under multi-resolution input. It solves the technical problems of fixed parameters making cross-resolution transfer difficult, dark target polarity distortion, and erroneous enhancement of complex background structures in existing morphological enhancement methods.

[0008] To address the aforementioned technical problems, a first aspect of this invention provides a bipolar small target resolution adaptive local enhancement method, comprising the following steps: Step S1: Acquire the input image and establish a scale factor based on the relationship between the resolution of the input image and the preset reference resolution. Use the scale factor to adaptively determine the optimal structural element size for morphological background modeling. Step S2: Based on the optimal structuring element size, perform morphological opening and morphological closing operations on the input image to construct a bright target background image and a dark target background image, and extract the bright target response image corresponding to the bright target background image and the dark target response image corresponding to the dark target background image. Step S3: Perform nonlinear enhancement processing on the bright target response map and the dark target response map respectively, and filter the candidate regions of bright targets and dark targets that meet the preset geometric conditions according to the geometric constraints of small targets; Step S4: Generate a smooth soft mask based on the bright target candidate region and the dark target candidate region, and construct a gain map corresponding to the smooth soft mask. When the bright target candidate region and the dark target candidate region have a spatial conflict, perform conflict resolution to determine the effective mask and the corresponding gain. Step S5: The input image is fused with the effective mask and the corresponding gain to apply positive local enhancement to bright small targets and negative local enhancement to dark small targets, so as to obtain an enhanced output image that preserves polarity.

[0009] Further, the step S1 of adaptively determining the optimal structuring element size for morphological background modeling using the scale factor includes the following sub-steps: Step S121: Select each candidate size in the candidate structural element size sequence after scaling by the scale factor, and perform morphological opening and morphological closing operations on the input image respectively to obtain the opening operation background estimation map and the closing operation background estimation map corresponding to each candidate size. Step S122: For adjacent candidate sizes, calculate the first average change between the background estimation maps of the opening operation and the second average change between the background estimation maps of the closing operation, and construct a comprehensive evaluation index of the stability of the background estimation based on the first average change and the second average change. Step S123: When the value of the comprehensive evaluation index is lower than the preset convergence threshold, it is determined that the background estimation has converged, the next candidate size is determined as the optimal structural element size and the search is terminated.

[0010] Furthermore, for adjacent candidate sizes k i and k i+1 The formula for calculating the comprehensive evaluation index is as follows: ; in, and respectively using candidate sizes k Structural elements for input image I Background estimation map obtained by performing morphological opening and closing operations. This indicates calculating the mean. ε These are preset non-zero small constants; When corresponding to candidate size k i With k i+1 Comprehensive evaluation indicators When the value is below the preset convergence threshold, the next candidate size k is selected. i+1 The optimal structural element size was determined.

[0011] Further, step S2, which involves performing morphological opening and closing operations on the input image based on the optimal structuring element size to construct a bright target background image and a dark target background image, and extracting the bright target response image corresponding to the bright target background image and the dark target response image corresponding to the dark target background image, includes the following sub-steps: Step S21: Using the optimal structural element size as a parameter, perform a morphological opening operation on the input image to obtain the bright target background image, which represents the smooth background after eliminating bright small structures; Step S22: Using the optimal structural element size as a parameter, perform a morphological closing operation on the input image to obtain the dark target background image, which represents the smooth background after filling in the dark small structures; Step S23: Subtract the gray values ​​of each pixel in the input image from the corresponding pixel gray values ​​in the bright target background image to obtain the bright target response map. The bright target response map is used to characterize the local prominence of the bright small target relative to its surrounding background. Step S24: Subtract the gray values ​​of each pixel in the dark target background image from the corresponding pixel gray values ​​in the input image to obtain the dark target response image. The dark target response image is used to characterize the degree of local concavity of the dark target relative to its surrounding background.

[0012] Furthermore, the bright target response map The calculation formula is: ; The dark target response map The calculation formula is: ; in, For the input image at pixel position grayscale value at that location The background image for the bright target. This is the background image of the dark target.

[0013] Furthermore, step S3, which involves filtering candidate regions for bright and dark targets based on the geometric constraints of small targets to obtain the predefined geometric conditions, includes the following sub-steps: Step S31: Threshold the response maps of the bright and dark targets after the nonlinear enhancement process to obtain the binary mask of the bright target and the binary mask of the dark target. Step S32: Perform connected component labeling on the bright target binary mask and the dark target binary mask respectively, and extract the geometric features of each connected component. The geometric features include at least two of the connected component's area, width, height, aspect ratio, and fill rate. Step S33: At least one of the area, width, and height values, scaled based on the scale factor, is combined with the aspect ratio threshold and fill rate threshold of the connected component to form a joint filtering condition. The width and height thresholds are scaled according to the scale factor s, and the area threshold is scaled according to the square of the scale factor s. 2 Scaling is performed, and the connected region is retained as either the bright target candidate region or the dark target candidate region only when the geometric features of a connected region simultaneously satisfy the joint screening conditions; Step S34: After step S33 is completed, if there are no retained connected components in the bright target binary mask, the entire bright target binary mask is used as the bright target candidate region; if there are no retained connected components in the bright target binary mask, the area threshold, aspect ratio threshold, or fill rate threshold is gradually relaxed according to a preset relaxation coefficient. If there are still no retained connected components after relaxation, the top N connected components in response intensity ranking and with an area less than a preset upper limit are retained as the bright target candidate regions, where N is a preset positive integer. If no connected components are preserved in the binary mask of the dark target, the entire binary mask of the dark target is used as the candidate region of the dark target. If no connected components are preserved in the binary mask of the dark target, the area threshold, aspect ratio threshold, or fill rate threshold is gradually relaxed according to a preset relaxation coefficient. If no connected components are preserved after relaxation, the top N connected components in response intensity ranking and with an area less than a preset upper limit are retained as the candidate regions of the dark target.

[0014] Further, step S4, which involves generating a smooth soft mask based on the bright target candidate region and the dark target candidate region, and constructing a gain map corresponding to the smooth soft mask, includes the following sub-steps: Step S41: Perform Gaussian convolution processing on the bright target candidate region and the dark target candidate region respectively to obtain a bright target soft mask and a dark target soft mask. The soft mask has the highest weight at the center of the target region and decays smoothly at the boundary of the target region. Step S42: Based on the bright target response map and dark target response map after the nonlinear enhancement processing, generate the initial gain map of the bright target and the initial gain map of the dark target respectively, and perform upper limit pruning on the initial gain map of the bright target and the initial gain map of the dark target respectively to obtain the gain map of the bright target and the gain map of the dark target, so as to prevent local over-enhancement; Step S43: Determine whether the bright target soft mask and the dark target soft mask overlap or conflict in spatial position. When a conflict occurs, compare the response intensity of the bright target and the response intensity of the dark target at the corresponding position, and retain the soft mask and gain corresponding to the one with the larger response intensity as the effective mask and corresponding gain. The criteria for determining overlap or proximity conflict are: at the same pixel position, the weight values ​​of the bright target soft mask and the dark target soft mask are both greater than a preset mask threshold, or the minimum spatial distance between the bright target candidate region and the dark target candidate region is less than a preset distance threshold.

[0015] Further, the step S5 of fusing the input image with the effective mask and corresponding gain includes: The input image, the bright target soft mask and bright target gain image in the effective mask, and the dark target soft mask and dark target gain image in the effective mask are weighted and fused to obtain a polarity-preserving enhanced output image. ; in, For the input image at pixel position grayscale value at that location and These are the bright target soft mask and the dark target soft mask in the effective mask, respectively. and These are the gain maps of the bright target and the dark target corresponding to the effective mask, respectively. and These are the preset enhancement weights for bright targets and dark targets, respectively.

[0016] Further, step S1, which involves acquiring the input image and establishing a scale factor based on the relationship between the resolution of the input image and a preset reference resolution, includes the following sub-steps: Step S111: Read the current pixel width and current pixel height of the input image, and calculate the scale factor based on the ratio of the current pixel width and current pixel height to the geometric mean of the reference pixel width and reference pixel height of the preset reference resolution; Step S112: Scale the preset set of candidate structural element sizes using the scale factor to obtain a sequence of candidate structural element sizes that is adapted to the resolution of the input image.

[0017] Furthermore, the scale factor The calculation formula is: ; in, and These are the current pixel height and the current pixel width, respectively. and These are the reference pixel height and the reference pixel width, respectively.

[0018] Accordingly, a second aspect of the present invention provides an electronic device, including: at least one processor; and a memory connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to cause the at least one processor to perform the above-described bipolar small target resolution adaptive local enhancement method.

[0019] Accordingly, a third aspect of the present invention provides a computer-readable storage medium having computer instructions stored thereon, which, when executed by a processor, implement the above-described bipolar small target resolution adaptive local enhancement method.

[0020] The above-described technical solutions of the embodiments of the present invention have the following beneficial technical effects: 1. By establishing a scale factor and adaptively searching for the optimal morphological structural element size according to the input image resolution, while linking the small target geometric screening threshold with the scale factor, this invention achieves automatic adaptation and stable migration of algorithm parameters under different resolution input conditions. This ensures that the physical size definition of small targets remains consistent in the pixel domain, effectively avoiding the problems of inconsistent enhancement effects and difficulty in uniform parameter configuration caused by the use of fixed structural element size and fixed screening threshold in traditional morphological enhancement methods when resolution changes. This significantly improves the robustness and engineering applicability of small target enhancement in multi-resolution scenarios. 2. By performing morphological opening and closing operations with optimal structuring element size to construct bright and dark target background images, and independently extracting bright and dark target responses, a polarity-preserving fusion strategy is adopted to apply positive local enhancement to bright targets and negative local enhancement to dark targets. This invention achieves the separate expression and independent processing of bright and dark bipolar targets throughout the entire process from background modeling to response extraction to enhancement output. It fundamentally changes the problem of dark target attribute distortion caused by the unified superposition of bright and dark responses in the traditional Top-hat method, and truly achieves the enhancement effect of "brighter bright targets and darker dark targets" with basically unchanged background radiation distribution. It has significant advantages in scenes with both bright and dark small targets in complex backgrounds. 3. By introducing joint geometric constraints on the area, width, height, aspect ratio, and fill rate of connected components after thresholding the response map, and setting a fallback strategy that retains the entire binary mask when no connected component meets the screening conditions, the risk of false responses generated by non-target structures in large target edges, striped textures, and complex background structures being mistakenly enhanced is effectively suppressed. At the same time, it avoids the situation where real small targets are wrongly rejected due to overly strict screening conditions. Based on this, combined with Gaussian smooth soft mask mapping and light-dark space conflict resolution mechanism, the local enhancement only acts on the region that truly meets the definition of small target and the enhancement boundary is smooth and natural, which significantly reduces the false alarm rate and improves the visual quality and stability of the enhanced output image. Attached Figure Description

[0021] Figure 1 This is a flowchart of the bipolar small target resolution adaptive local enhancement method provided in the embodiments of the present invention; Figure 2 This is a schematic diagram of the bipolar small target resolution adaptive local enhancement algorithm provided in an embodiment of the present invention. Detailed Implementation

[0022] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to specific embodiments and the accompanying drawings. It should be understood that these descriptions are merely exemplary and not intended to limit the scope of the invention. Furthermore, descriptions of well-known structures and techniques are omitted in the following description to avoid unnecessarily obscuring the concept of the invention.

[0023] Please refer to Figure 1 and Figure 2 The first aspect of this invention provides a bipolar small target resolution adaptive local enhancement method, comprising the following steps: Step S1: Acquire the input image and establish a scale factor based on the relationship between the resolution of the input image and the preset reference resolution. Use the scale factor to adaptively determine the optimal structural element size for morphological background modeling.

[0024] In applications such as long-range electro-optical reconnaissance, infrared early warning, and surveillance of small targets, the input images acquired by the imaging system often have different spatial resolutions. The apparent size of small targets in the pixel domain changes with the image resolution. To automatically match the structuring element size and geometric constraint threshold with the scale of the current input image, the input image is first read to obtain its width and height pixel values, and a scale factor associated with the input image resolution is introduced. This scale factor is established based on the geometric dimensional relationship between the current resolution of the input image and a pre-set reference resolution, which is typically selected as a standard resolution under typical imaging conditions. By introducing the scale factor, the originally fixed processing parameters are made capable of adaptively adjusting with the input resolution, thus overcoming the technical defect of inconsistent enhancement effects at different resolutions caused by fixed parameter settings in traditional morphological small target enhancement methods.

[0025] Based on the established scale factor, the optimal structuring element size for subsequent morphological background modeling is adaptively determined using this scale factor. In a typical implementation, a set of basic candidate structuring element sizes can be pre-defined, and these candidate sizes are scaled according to the scale factor to obtain a sequence of candidate sizes adapted to the current input image resolution. Subsequently, morphological opening and closing operations are performed on the input image using structuring elements of each candidate size to obtain the background estimation results at the corresponding sizes. By comparing the degree of difference between the background estimation maps obtained from adjacent candidate sizes, the convergence state of the background estimation results tending to stabilize as the structuring element size increases can be evaluated. When the change in background estimation between adjacent sizes is lower than a preset convergence threshold, it indicates that further increasing the structuring element size can no longer significantly improve the background estimation quality. At this point, the corresponding size can be determined as the optimal structuring element size under the current input image conditions. This adaptive search mechanism replaces the traditional method of relying on manual experience to preset fixed structuring element sizes, allowing the level of detail in background modeling to be automatically adjusted according to the background complexity of the image itself, significantly improving the robustness of the algorithm under multi-resolution and complex background conditions.

[0026] Step S2: Based on the optimal structuring element size, perform morphological opening and morphological closing operations on the input image to construct a bright target background image and a dark target background image, and extract the bright target response image corresponding to the bright target background image and the dark target response image corresponding to the dark target background image.

[0027] After obtaining the optimal structuring element size that matches the current input image scale, background estimation images for extracting small bright targets and small dark targets are constructed separately. Morphological opening has the property of eliminating bright details smaller than the structuring element while maintaining a smooth overall background distribution. Therefore, performing morphological opening on the input image using the optimal structuring element size can effectively filter out small bright targets as foreground details, and the resulting opening result is the bright target background image. This background image reflects the local background radiation level under the assumption of no small bright targets. Correspondingly, morphological closing has the property of filling dark recessed areas smaller than the structuring element while maintaining background continuity. Performing morphological closing on the input image using the same optimal structuring element size can fill in the dark target areas, thus obtaining a dark target background image, which characterizes the local background radiation distribution under the assumption of no small dark targets. By constructing background images for bright and dark targets respectively, the background estimation of two types of targets with different polarities is decoupled, laying the foundation for the independent extraction of bright and dark responses. Traditional morphological enhancement methods usually only construct a single background estimation model, which is difficult to take into account the differentiated background representation needs of bipolar targets.

[0028] After obtaining the dual background estimation maps, the bright target response map corresponding to the bright target background map and the dark target response map corresponding to the dark target background map are further extracted. Specifically, for the bright target response map, the response value at any pixel position is obtained by subtracting the estimated gray value of the corresponding position in the bright target background map from the original gray value of the input image at that position. Since the bright target background map has filtered out small bright structures and retained a smooth background, this difference operation can effectively highlight the regions in the original image that show a positive brightness protrusion relative to the local background, i.e., the candidate responses of small bright targets. For the dark target response map, its response value is obtained by subtracting the original gray value of the corresponding position in the input image from the estimated gray value of the dark target background map at that position. Since the dark target background map has filled in small dark structures and retained a smooth background, this difference operation can effectively highlight the regions in the original image that show a negative brightness concavity relative to the local background, i.e., the candidate responses of small dark targets. Through the above independent difference operation, the polarity separation expression of bright and dark targets is realized in the response extraction stage. This avoids the problem of dark target attribute distortion caused by treating bright and dark responses as a positive enhancement term in traditional methods, and provides an accurate bipolar response basis for subsequent polarity preservation enhancement.

[0029] Step S3: Perform nonlinear enhancement processing on the response maps of bright targets and dark targets respectively, and select candidate regions of bright targets and candidate regions of dark targets that meet the preset geometric conditions based on the geometric constraints of small targets.

[0030] After obtaining the bright target response map and the dark target response map, since small targets usually have low contrast in the original image, directly using the original response map for subsequent processing may not be enough to fully highlight the salience of weak targets. Therefore, nonlinear enhancement processing is first performed on the bright target response map and the dark target response map respectively. In an optional implementation, the nonlinear enhancement processing can adopt a piecewise mapping method, that is, dividing the response value into several intervals according to the intensity range, and applying different degrees of nonlinear stretching to the response values ​​of different intervals, so that the weak target signal in the low response interval receives a relatively stronger gain boost, while the region with a high response intensity maintains a moderate gain or is compressed, thereby selectively improving the local contrast of weak targets without significantly amplifying background noise.

[0031] The aforementioned piecewise mapping nonlinear enhancement process applies a relatively stronger nonlinear stretching gain to weak target signals with low response values, making them stand out from the background. For regions already possessing high response intensity, it maintains a moderate gain or performs limited compression to avoid excessive amplification of strong signal areas. In practical implementation, several response intensity breakpoints can be pre-defined based on the statistical distribution characteristics of the response map, dividing the dynamic range of the response values ​​into multiple continuous intervals, and configuring a corresponding Gamma correction parameter for each interval. For pixel positions with response values ​​in the lower interval, a smaller Gamma parameter value is used to produce a stronger stretching effect. For pixel positions with response values ​​in the middle or higher intervals, a gradually increasing or approaching 1 Gamma parameter value is used, gradually reducing the stretching until it approaches a linear mapping. Through this piecewise differentiated Gamma correction, weak target regions with originally low contrast in the response map can achieve a significant improvement in signal-to-noise ratio, while low-amplitude responses in the background caused by noise or non-target structures are not excessively amplified by uniform amplification. Furthermore, in applications processing continuous frame sequences, the response map of the previous frame after nonlinear enhancement can be selected as a reference base map. Based on the response distribution characteristics of the reference base map, the Gamma parameters of each segment of the current frame are adaptively fine-tuned. This effectively reduces fluctuations in the inter-frame enhancement effect caused by dynamic scene changes or noise fluctuations, maintaining the brightness stability and visual coherence of the output image sequence over time. Targeted enhancement of weak signal components in the original bipolar response avoids the adverse consequences of simultaneously amplifying background noise and strong edge structures, which can occur with a single linear amplification strategy. This provides higher-quality target saliency representation for subsequent thresholding and geometric constraint screening.

[0032] Piecewise Gamma nonlinear enhancement processing can be performed using the following formula framework, where R represents the original response value, R' represents the enhanced response value, t is the preset response intensity breakpoint, and γ is the Gamma parameter for the corresponding interval: ; in, Represents the set of breakpoints. This represents the gamma parameter. This represents the maximum value of the current response map, from which the enhanced bright target response map is obtained. Dark target response map .

[0033] This process continues until the last interval. This is used to normalize the response value to the zero-to-one range, perform a power transformation, and then map it back to the grayscale value range. The Gamma parameter corresponding to each range can be set according to the different requirements of weak signal enhancement and strong signal compression in the actual application scenario.

[0034] Furthermore, in applications involving processing sequential images, the enhancement response of the previous frame can be used as reference information to moderately adjust the nonlinear enhancement level of the current frame, thereby reducing fluctuations in the inter-frame enhancement effect and maintaining the stability of the output image sequence.

[0035] After nonlinear enhancement processing, candidate regions are further screened based on the geometric constraints of the small target to obtain candidate bright and dark target regions that meet preset geometric conditions. First, the enhanced bright and dark target response maps are binarized using preset response thresholds to generate corresponding binary masks, initially marking pixel locations with response intensities higher than the thresholds as potential target candidates. Then, connected component labeling analysis is performed on the binary masks to extract the geometric attributes of each connected region, including but not limited to the pixel area contained within the connected region, the width and height of the bounding rectangle, the width-to-height ratio, and the fill rate represented by the ratio of the connected region area to the bounding rectangle area. Based on the inherent properties of small targets in the physical world—small size and relatively compact shape—appropriate geometric thresholds can be set to constrain and screen connected regions, retaining only those connected regions that conform to the typical characteristics of small targets in terms of area, width, height, aspect ratio, and fill rate as the final output candidate bright and dark target regions. By using the geometric constraints described above, false target regions that have a strong response but clearly do not conform to the geometric definition of small targets, caused by large target edges, striped textures, or complex background structures, can be effectively eliminated, significantly reducing the risk of false alarms in subsequent enhancement processing. Simultaneously, the geometric thresholds involved can be synchronously scaled and adjusted according to the scale factor established in step S1, ensuring that the geometric definition of small targets remains physically consistent across input images at different resolutions.

[0036] Step S4: Generate a smooth soft mask based on the bright target candidate region and the dark target candidate region, and construct a gain map corresponding to the smooth soft mask. When there is a spatial conflict between the bright target candidate region and the dark target candidate region, perform conflict resolution to determine the effective mask and the corresponding gain.

[0037] After the geometric constraint screening in step S3, candidate regions for bright and dark targets that meet the definition of small targets have been obtained. If the original image is directly enhanced based on the binary hard boundary mask corresponding to the candidate region, obvious gray-level jumps and block artifacts are likely to occur at the boundary between the enhanced and unenhanced regions, thus destroying the overall visual harmony of the output image. To solve this problem, smooth soft masks are first generated based on the candidate regions for bright and dark targets respectively. The smooth soft mask is generated by performing a smoothing filter on the binary mask of the candidate region, so that the mask value has a high weight coefficient at the center of the candidate region, and gradually decays to zero from the center to the edge of the region. Unlike the hard boundary mask, which has abrupt value changes at the region boundary, the smooth soft mask can establish a gentle transition zone between the target region and the surrounding background, so that the subsequent local enhancement can be naturally integrated into the background, avoiding the artificial perception of enhancement traces.

[0038] While generating a smooth soft mask, corresponding gain maps are constructed based on the bright target response map and dark target response map after the nonlinear enhancement processing in step S3. The gain map reflects the enhancement intensity to be applied at different pixel locations, and its value is positively correlated with the response intensity; that is, the stronger the response of the target area, the stronger the enhancement gain will be. To further ensure the stability of the enhancement results, the upper limit of the gain amplitude can be clipped during the construction of the gain map to prevent undesirable visual effects such as local overexposure or underexposure due to abnormally high responses in individual pixels. In addition, in applications where the bright target candidate area and the dark target candidate area are spatially adjacent or partially overlapped, if a positive brightening operation and a negative darkening operation are performed simultaneously at the same pixel location, the two enhancement effects will cancel each other out or produce unexpected grayscale distortion. Therefore, when the above situation occurs, a bright-dark conflict resolution process is performed. By comparing the response intensity of the bright target and the response intensity of the dark target at the conflict location, only the smooth soft mask and its gain map corresponding to the side with the dominant response intensity are retained as the effective mask and corresponding gain. This ensures the spatial uniqueness and determinism of the local enhancement effect and improves the stability of the enhancement output under the condition of coexistence of bipolar targets.

[0039] Step S5: The input image is fused with the effective mask and the corresponding gain to apply positive local enhancement to bright small targets and negative local enhancement to dark small targets, so as to obtain an enhanced output image that preserves polarity.

[0040] After generating the smooth soft mask, constructing the gain map, and resolving brightness-darkness conflicts, the original input image is fused with the effective mask and corresponding gain determined in the above steps to generate the final polarity-preserving enhanced output image. During the fusion process, for pixels marked as bright target candidate regions and retained through conflict resolution, the bright target gain map is multiplied by the corresponding smooth soft mask weight, and then positively superimposed with the grayscale value of the original input image, thereby achieving a local brightening enhancement effect for small bright targets. For pixels marked as dark target candidate regions and retained through conflict resolution, the dark target gain map is multiplied by the corresponding smooth soft mask weight, and negatively subtracted from the grayscale value of the original input image, thereby achieving a local darkening enhancement effect for small dark targets. For regions that are neither bright nor dark target candidate regions, or that are eliminated during conflict resolution, their pixel grayscale values ​​remain unchanged from the state of the original input image.

[0041] Through the aforementioned fusion method, local enhancement that preserves the polarity of both bright and dark small targets is achieved. Polarity preservation involves performing a positive brightness boost on bright targets, making them appear brighter and more prominent in the output image; and a negative brightness reduction on dark targets, making them appear darker and more noticeable in the output image. Both types of enhancement only apply to a local spatial range that has been geometrically constrained and smoothed using soft masks, thus fully preserving the grayscale distribution and texture of the background region. Compared to existing technologies that uniformly treat bright and dark responses as positive enhancement terms, leading to distortion of the polarity of dark targets, or that apply enhancements indiscriminately to the entire image, resulting in background distortion, the polarity-preserving local fusion strategy employed in this invention can significantly improve the discernibility of bipolar small targets while maximally maintaining the authenticity and stability of the original image's background radiation distribution, providing high-quality input images for subsequent target detection, recognition, and tracking tasks.

[0042] As can be seen from the above, the method provided by this invention firstly utilizes the scale factor to adaptively determine the size of morphological structural elements, enabling the background modeling accuracy to automatically adjust with the input image resolution, thus solving the problem of insufficient cross-resolution adaptability caused by fixed parameters in traditional methods. Secondly, by constructing separate background images for bright and dark targets and extracting bipolar responses independently, the algorithm achieves separate representation of bright and dark targets at the front end, fundamentally changing the polarity distortion defect of dark targets caused by the uniform superposition of bright and dark responses in traditional morphological enhancement. Thirdly, by introducing a small target geometric constraint screening mechanism combined with scale scaling, and supplemented by smooth soft mask mapping and bright-dark conflict resolution processing, the erroneous enhancement of large target edges and complex background textures is effectively suppressed, while ensuring the natural transition between the enhancement boundary and the surrounding background and the output stability when bipolar targets coexist. Finally, through a local fusion strategy that preserves polarity, the enhancement effect of locally brightening bright targets and locally darkening dark targets is achieved while maintaining the overall radiation distribution of the background, significantly improving the salience, recognizability, and output image quality of bipolar small targets in complex multi-resolution backgrounds.

[0043] Specifically, step S1, which involves acquiring the input image and establishing a scaling factor based on the relationship between the input image's resolution and a preset reference resolution, includes the following sub-steps: Step S111: Read the current pixel width and current pixel height of the input image, and calculate the scale factor based on the ratio of the current pixel width and current pixel height to the geometric mean of the reference pixel width and reference pixel height of the preset reference resolution.

[0044] In typical applications such as long-range electro-optical reconnaissance and infrared early warning, the image resolution output by the imaging system often varies significantly due to factors such as sensor configuration, transmission bandwidth limitations, or region of interest cropping. The actual size of a small target in the pixel domain changes significantly with resolution. If the parameters related to spatial scale remain fixed in subsequent processing, a small target of the same physical size may occupy many pixels in a high-resolution image and be misjudged as the edge of a large target, while in a low-resolution image it may occupy too few pixels and be filtered out as noise, resulting in highly unstable enhancement effects. To address this cross-resolution adaptability problem, this invention first reads the current pixel width and current pixel height of the input image and establishes a correlation between these two parameters and a reference pixel width and reference pixel height corresponding to a pre-set reference resolution. This reference resolution is typically selected based on the standard field of view and detector specifications in typical imaging tasks, representing the optimal parameter calibration benchmark that the algorithm aims to achieve. After obtaining the current pixel width and height, the ratio of their product to the product of the reference pixel width and height is calculated, and the square root of this ratio is taken as the scaling factor. This scaling factor essentially reflects the geometric scaling ratio of the input image relative to the reference resolution under the current resolution conditions. Calculating the scaling factor in the form of a ratio of geometric means can integrate resolution change information in both width and height dimensions, making subsequent adaptive scaling operations well applicable to input images with different aspect ratios.

[0045] Step S112: Scale the preset set of candidate structural element sizes using a scaling factor to obtain a sequence of candidate structural element sizes that is adapted to the resolution of the input image.

[0046] After the scale factor is calculated, it is further used to linearly scale a pre-defined set of basic candidate structural element sizes. The basic candidate structural element size set typically encompasses a sequence of sizes from small to large, such as square or near-square structural elements calibrated as 3×3, 5×5, 7×7, or even larger at the reference resolution. The design principle of this sequence is to cover the typical pixel scale range that small targets might present under the reference resolution. Multiplying each basic candidate size by the scale factor yields a sequence of candidate structural element sizes that matches the actual resolution of the current input image. For example, if a basic candidate size is 5×5 pixels at the reference resolution, and the scale factor calculated based on the current input image size is 1.2, then the corresponding scaled candidate size will be adjusted to approximately 6×6 pixels, thus ensuring that the physical coverage of the structural element in the current image remains essentially consistent with the design intent at the reference resolution. After the above scaling process, the obtained candidate structuring element size sequence can accurately reflect the pixel scale distribution range that small targets may occupy under the current input resolution. This provides a search space that is strictly adapted to the image scale for the adaptive search of structuring elements and the determination of the optimal size in subsequent steps, laying the technical foundation for the stable operation of the entire method across resolutions from the parameter initialization level.

[0047] Furthermore, scale factor The calculation formula is: .

[0048] in, and These are the current pixel height and the current pixel width, respectively. and These are the reference pixel height and reference pixel width, respectively.

[0049] In establishing the aforementioned scaling factor, its specific value is determined by the square root of the ratio of the product of the current pixel height and width of the input image to the product of the reference pixel height and width at the preset reference resolution. This calculation method essentially measures the scaling ratio of the current input image relative to the reference baseline over the total area of ​​two-dimensional pixels, expressed as a one-dimensional length ratio. This allows the scaling factor to directly affect various parameters related to spatial dimensions. Because it employs a geometric mean ratio, this scaling factor can still provide a reasonable equivalent scaling scale even when the scaling ratios in the image width and height directions are inconsistent, providing a unified numerical basis for subsequent resolution adaptive operations such as structuring element size scaling and geometric threshold linkage adjustment.

[0050] Specifically, step S1, which adaptively determines the optimal structuring element size for morphological background modeling using a scale factor, includes the following sub-steps: Step S121: Select each candidate size in the candidate structuring element size sequence after scaling by the scale factor, and perform morphological opening and morphological closing operations on the input image respectively to obtain the opening operation background estimation map and the closing operation background estimation map corresponding to each candidate size.

[0051] After obtaining a sequence of candidate structuring element (SLE) sizes suitable for the current input image resolution, a trial calculation of morphological background estimation is performed sequentially for each candidate size in the sequence. For the currently selected candidate size, a corresponding morphological SLE is first constructed using that size as a parameter. The shape of the SLE is typically chosen to be square or disk-shaped, with its side length or diameter equal to the pixel span specified by the candidate size. Subsequently, morphological opening and closing operations are performed on the input image using this SLE. The morphological opening operation consists of two cascaded basic operations: erosion followed by dilation. It effectively filters out bright details smaller than the SLE in the image while retaining large-scale background undulation information. Therefore, the resulting background estimation map after opening represents the local background radiation distribution under the premise of eliminating bright small targets. The morphological closing operation consists of two cascaded basic operations: dilation followed by erosion. It effectively fills in dark recessed areas smaller than the SLE in the image while retaining large-scale background continuity features. Therefore, the resulting background estimation map after closing represents the local background radiation distribution under the premise of filling in dark small targets. For each candidate size in the candidate structuring element size sequence, the above opening and closing operations are performed to obtain a set of opening and closing background estimation maps that correspond one-to-one with each candidate size, providing a complete trial data basis for subsequent background estimation stability analysis.

[0052] Step S122: For adjacent candidate sizes, calculate the first average change between the background estimation maps of the opening operation and the second average change between the background estimation maps of the closing operation, and construct a comprehensive evaluation index for the stability of the background estimation based on the first average change and the second average change.

[0053] After completing background estimation trials for all candidate sizes, a quantitative analysis is performed on the degree of change in background estimation results between adjacent candidate sizes to determine whether the background estimation has stabilized with the increase of the structuring element size. For any two adjacent candidate sizes in the candidate size sequence, the first average change between their corresponding opening operation background estimation images and the second average change between their corresponding closing operation background estimation images are calculated. The first average change is obtained by calculating the absolute difference between the two opening operation background estimation images pixel by pixel and taking the average of the entire image. Its physical meaning is to measure the overall change in the background estimation result of bright targets after the structuring element size increases by one level. The second average change is obtained by calculating the absolute difference between the two closing operation background estimation images pixel by pixel and taking the average of the entire image. Its physical meaning is to measure the overall change in the background estimation result of dark targets. Based on obtaining the first and second average changes, they are further integrated to construct a comprehensive evaluation index that can fully reflect the stability of the dual-channel background estimation of opening and closing operations. The construction method of this comprehensive evaluation index takes into account both the first and second average changes, and uses the average gray level of the input image itself as the normalization benchmark, thereby eliminating the influence of the overall brightness difference of different images on the value of the change, so that the convergence judgment condition has consistent comparability between input images with different brightness levels.

[0054] Step S123: When the value of the comprehensive evaluation index is lower than the preset convergence threshold, it is determined that the background estimation has converged, the next candidate size is determined as the optimal structural element size and the search is terminated.

[0055] After constructing a comprehensive evaluation index for the stability of background estimation among adjacent candidate sizes, the value of this index is compared with a preset convergence threshold to determine whether the background estimation has reached a convergence state. The convergence threshold is typically set to a small positive real number based on empirical test results, representing that the relative change in the background estimation result as the structuring element size increases has decreased to a negligible level. When the value of the comprehensive evaluation index is higher than this convergence threshold, it indicates that increasing from the current candidate size to the next candidate size still causes a significant change in the background estimation result. At this point, the background estimation has not yet reached a stable state, and it is necessary to continue selecting the next candidate size in the sequence and repeat the above trial calculation and evaluation process. When the value of the comprehensive evaluation index falls below this convergence threshold for the first time, it indicates that continuing to increase the structuring element size can no longer substantially improve the background estimation result, and the background estimation has reached a convergence state. At this point, the current candidate size is determined as the optimal structuring element size suitable for the input image, and the subsequent trial search process for larger sizes is terminated. This adaptive search termination mechanism means that the determination of the size of the structuring element no longer depends on fixed values ​​preset by human experience, but is driven by the background complexity and resolution characteristics of the input image itself, automatically converging to a suitable scale that can sufficiently smooth background details without excessively blurring the target structure. This provides the structural parameter configuration that best matches the current image content for subsequent dual-background modeling and response extraction steps.

[0056] Furthermore, for adjacent candidate sizes k i and k i+1 The formula for calculating the comprehensive evaluation index is as follows: ; in, and respectively using candidate sizes k Structural elements for input image I Background estimation map obtained by performing morphological opening and closing operations. This indicates calculating the mean. ε It is a preset non-zero small constant.

[0057] In the construction of the aforementioned comprehensive evaluation index, for any two adjacent candidate sizes in the candidate size sequence, the average value of the pixel-by-pixel absolute difference between the open operation background estimation images obtained using the previous and subsequent candidate sizes, and the average value of the pixel-by-pixel absolute difference between the closed operation background estimation images obtained using the previous and subsequent candidate sizes, are calculated. These two average values ​​reflect the absolute variation in the overall grayscale distribution of the bright target background estimation result and the dark target background estimation result under the condition of increasing the structuring element size by one level. Subsequently, the first and second average variations are summed, and the sum is divided by the sum of twice the grayscale mean of the input image and a preset non-zero small constant. Twice the grayscale mean of the input image is introduced into the denominator as a normalization factor. The purpose is to eliminate the influence of differences in the overall brightness level of different input images on the absolute value of the background estimation variation, so that the comprehensive evaluation index can maintain a consistent convergence judgment scale under different brightness scenes. The non-zero constant added to the denominator is used to prevent division by zero in extreme cases where the mean gray level of the input image approaches zero, thus ensuring the numerical stability of the calculation process. The comprehensive evaluation index obtained through the ratio of the numerator to the denominator essentially characterizes the relative rate of change of the background estimation result as the size of the structuring element increases. The smaller this value, the weaker the difference in background estimation between adjacent candidate sizes, indicating that the background modeling result has stabilized. This provides a quantifiable objective basis for determining the convergence threshold in step S123.

[0058] When corresponding to candidate size k i With k i+1 Comprehensive evaluation indicators When the value is below the preset convergence threshold, the next candidate size k is selected. i+1 The optimal structural element size was determined. Based on the construction of the above comprehensive evaluation index, the specific selection rules for the optimal structural element size after convergence determination were further clarified. When corresponding to the adjacent candidate size k i With k i+1 Comprehensive evaluation indicators When the value first falls below the preset convergence threshold, it indicates that size k is being used. i With k i+1 The differences between the obtained background estimation results have decreased to a negligible level, and the background modeling has stabilized. Further increasing the size of the structuring element can no longer substantially improve the quality of the background estimation. At this point, the next candidate size k is... i+1 The optimal structuring element size is determined under the current input image conditions. The reason for choosing the latter candidate size k is... i+1 Instead of the previous candidate size k iThis is because a comprehensive evaluation index below the convergence threshold means that both adjacent candidate sizes have effectively smoothed background details and preserved large-scale background structure, while size k... i+1 Compared to size k i It provides a more robust background smoothing, ensuring that the response of small targets is not compromised, while avoiding interference from residual local background undulations caused by selecting a smaller size in subsequent target response extraction. The above selection rules clarify the deterministic correspondence between candidate sizes and optimal sizes, eliminating ambiguity that may arise from the expression "current candidate size," and enabling the adaptive structuring element size search process to have clear and reproducible termination and output criteria.

[0059] Specifically, step S2, based on the optimal structuring element size, performs morphological opening and closing operations on the input image to construct a bright target background image and a dark target background image, and extracts the bright target response image corresponding to the bright target background image and the dark target response image corresponding to the dark target background image, including the following sub-steps: Step S21: Using the optimal structuring element size as a parameter, perform a morphological opening operation on the input image to obtain a bright target background image. The bright target background image represents the smooth background after eliminating bright small structures.

[0060] After adaptively determining the optimal structuring element size through the aforementioned steps, a morphological opening operation is performed on the input image using this optimal size as the structuring element parameter to construct a bright target background image. The morphological opening operation consists of two cascaded basic operations: erosion and dilation. Erosion replaces bright areas smaller than the structuring element with their darker neighboring pixels, while dilation performs morphological restoration on the eroded image, allowing for the reconstruction of the large-scale background structure. Since small bright targets in an image typically appear as isolated bright spots relative to the local background, their spatial scale is generally smaller than the scale of the surrounding background undulations. Therefore, during the opening operation, the region containing the small bright target is effectively eliminated in the erosion stage. The subsequent dilation stage can only recover the large-scale background contour and cannot restore the details of the eliminated small bright targets. The bright target background image obtained after the above processing is actually a smoothed estimate of the local background radiation distribution under the assumption that no small bright targets exist in the input image. The background image retains the large-scale grayscale fluctuation trend and illumination distribution information in the image, while filtering out bright details smaller than the optimal structuring element, thus providing a reliable background reference benchmark for the accurate extraction of subsequent bright target responses.

[0061] Step S22: Using the optimal structuring element size as a parameter, perform a morphological closing operation on the input image to obtain a dark target background image. The dark target background image represents the smooth background after filling in the dark small structures.

[0062] A morphological closing operation is performed on the input image using the same optimal structuring element size as a parameter to construct a dark target background map. The morphological closing operation consists of two cascaded basic operations: dilation and erosion. Dilation fills dark regions smaller than the structuring element with brighter neighboring pixels, while erosion restores the morphology of the dilated image, allowing for the reconstruction of large-scale background structures. Small dark targets typically appear as isolated dark spots or depressions relative to the local background, with a spatial scale generally smaller than the surrounding background undulations. Therefore, during the closing operation, the region containing the small dark target is filled by brighter neighboring pixels during the dilation stage. The subsequent erosion stage can only recover the large-scale background outline and cannot recreate the filled-in dark target depression. The resulting dark target background map is, in effect, a smoothed estimate of the local background radiation distribution, assuming no small dark targets exist in the input image. The background image also retains the large-scale grayscale fluctuation trend and illumination distribution information in the image, while filling in the dark concave structures smaller than the optimal structuring element, thus providing a reliable background reference benchmark for the accurate extraction of subsequent dark target responses.

[0063] By constructing a bright target background image and a dark target background image in steps S21 and S22 respectively, this invention achieves differentiated processing of two types of targets with different polarities in the background modeling stage, avoiding the trade-off distortion problem of traditional single background models when simultaneously representing two types of target backgrounds, such as bright and dark targets.

[0064] Step S23: Subtract the gray values ​​of each pixel in the input image from the corresponding pixel gray values ​​in the bright target background image to obtain the bright target response map. The bright target response map is used to characterize the local prominence of a bright small target relative to its surrounding background.

[0065] After obtaining the bright target background image, the bright target response image is calculated by subtracting the original grayscale values ​​at each pixel location in the input image from the estimated grayscale values ​​of the corresponding pixel locations in the bright target background image, pixel by pixel. Since the bright target background image has already filtered out bright structures smaller than the structuring element through morphological opening operations, preserving a large-scale smooth background trend, the original grayscale values ​​of regions in the image where bright targets actually exist are significantly higher than the estimated values ​​at the corresponding locations in the background image. Subtracting the two will produce a large positive response value, the magnitude of which quantitatively reflects the prominence of the bright target relative to its local background. For flat background regions in the image where no bright targets exist, the original grayscale values ​​are close to the estimated background values, and the response value obtained by subtracting them is close to zero. For dark areas or edge transition regions in the image, since the original grayscale values ​​are lower than the estimated background values, subtracting them will produce a negative response value, but subsequent processing only focuses on the positive bright target response. Through the above difference operation, the bright target response map effectively suppresses the interference of large-scale background undulations on the target saliency measurement while retaining the spatial location information of the bright small target. This allows the faint bright target that was originally submerged in the complex background to be separated from the background, providing a target saliency-oriented processing object for subsequent nonlinear enhancement and geometric screening.

[0066] Step S24: Subtract the gray values ​​of each pixel in the dark target background image from the corresponding pixel gray values ​​in the input image to obtain the dark target response image. The dark target response image is used to characterize the degree of local concavity of the dark target relative to its surrounding background.

[0067] The dark target response map is calculated by subtracting the estimated grayscale value of the background at each pixel location in the dark target background image from the original grayscale value at the corresponding pixel location in the input image pixel by pixel. The direction of subtraction here is opposite to step S23, using the method of subtracting the original grayscale value from the estimated background value. The purpose is to ensure that the response of the dark target is positive, thus maintaining consistency in numerical sign with the response of the bright target, facilitating the subsequent use of a unified positive response threshold processing framework. Since the dark target background image has already filled in the dark depressions smaller than the structuring element through morphological closing operations, preserving the large-scale smooth background trend, for areas in the image where small dark targets actually exist, the estimated value at the corresponding location in the background image is significantly higher than the original grayscale value. Subtracting the two will produce a large positive response value, the magnitude of which quantitatively reflects the degree of depression of the small dark target relative to its local background. For flat background areas in the image where there are no small dark targets, the estimated background value is close to the original grayscale value, and the response value obtained by subtracting the two is close to zero. For bright areas in the image, since the background estimate is lower than the original grayscale value, subtracting the two will produce a negative response value. However, subsequent processing only focuses on the positive dark target response. Through the above difference operation, the dark target response map effectively suppresses the interference of large-scale background undulations while preserving the spatial location information of small dark targets. This makes the faint dark targets that were originally mixed with the dark background stand out, providing a saliency expression of dark targets independent of bright targets for subsequent nonlinear enhancement and geometric screening. Steps S23 and S24 complete the independent extraction of target responses from the two polarities of light and dark, respectively. This allows the subsequent processing flow to perform differentiated enhancement and screening based on the characteristics of bright and dark targets, laying a key foundation for response separation to achieve the final bipolar preservation enhancement.

[0068] Furthermore, highlight the target response map. The calculation formula is: . Dark target response map The calculation formula is: . in, For the input image at pixel position grayscale value at that location To highlight the background image, Background image for dark targets.

[0069] In the generation of the response maps described above, the response value of the bright target response map at any pixel location is obtained by subtracting the estimated gray value of the corresponding bright target background image from the original gray value of the input image at that location. Since the bright target background image has already filtered out small bright structures through morphological opening operations while retaining a smooth background trend, the original gray value will be significantly higher than the estimated background value in areas where small bright targets actually exist. The difference between the two constitutes a positive response, and the magnitude of this response characterizes the strength of the bright target's prominence relative to its local background. The response value of the dark target response map at any pixel location is obtained by subtracting the original gray value of the corresponding input image from the estimated gray value of the dark target background image at that location. The setting of the subtraction direction here ensures that the response of the dark target region also exhibits a positive value, thus maintaining consistency with the bright target response in terms of numerical sign, facilitating the subsequent use of consistent thresholding and enhancement strategies. In areas where small, dark targets actually exist, the background image of the dark target has already filled in the dark depressions through morphological closing operations, so its estimated background value will be significantly higher than the original grayscale value. The positive response formed by the difference between the two represents the depth of the depression of the small, dark target relative to its local background. Through the difference operations defined above, the bright target response image and the dark target response image independently complete the quantitative expression of the salience of targets of different polarities, providing a clear numerical basis for the parallel processing and separate enhancement of bipolar responses in subsequent steps.

[0070] Specifically, step S3, which involves filtering candidate regions for bright and dark targets based on the geometric constraints of small targets to obtain candidate regions that meet preset geometric conditions, includes the following sub-steps: Step S31: Threshold the response maps of the bright and dark targets after nonlinear enhancement processing to obtain the binary mask of the bright target and the binary mask of the dark target.

[0071] After performing nonlinear enhancement processing on the bright and dark target response maps, the numerical values ​​at each pixel location in the response map reflect the relative probability and significance of the presence of a small bright or small dark target at that location. To convert the continuously distributed response values ​​into a binary representation that facilitates subsequent connected component analysis, thresholding is performed on both the enhanced bright and dark target response maps. The basic operation of thresholding is to compare the response value at each pixel location in the response map with a preset response threshold. When the response value is greater than or equal to the threshold, the pixel location is marked as a candidate target point and assigned a first logical value; when the response value is less than the threshold, the pixel location is marked as a background point and assigned a second logical value. After the above pixel-by-pixel comparison and assignment operations, the enhanced continuous response map is converted into a bright target binary mask and a dark target binary mask containing only two value states. In the bright target binary mask, the pixel location taking the first logical value indicates that it exhibits a strong positive response in the enhanced bright target response map and is initially identified as a potential region for a small bright target. The pixel position with the first logical value in the binary mask of a dark target indicates that it exhibits a strong positive response in the enhanced dark target response map and is initially identified as the region to which a potential small dark target belongs. The response threshold can be determined based on the overall statistical distribution characteristics of the response map, such as taking a multiple of the mean of the response map or a certain percentage of the maximum value of the response map, so that the threshold can automatically adapt to the overall level difference in response intensity in different images. Through thresholding, the originally continuously changing response distribution is transformed into a spatially discrete set of binary regions, providing a structured processing object for subsequent connected component labeling and geometric feature extraction.

[0072] Step S32: Perform connected component labeling on the binary mask of the bright target and the binary mask of the dark target respectively, and extract the geometric features of each connected component. The geometric features include at least two of the following: area, width, height, aspect ratio and fill rate of the connected component.

[0073] After obtaining the binary masks for bright and dark targets, the pixels with the first logical value in the binary masks often cluster spatially into several interconnected regions that are either separate or adjacent to each other. Each connected region represents a potential candidate target. To further identify the real target regions that conform to the physical properties of small targets from these candidate targets, connected component labeling is performed on the binary masks for bright and dark targets respectively. Connected component labeling merges spatially connected pixels with the first logical value into the same connected region by scanning the logical values ​​of each pixel in the binary mask and their connection relationships with neighboring pixels, and assigns a unique identifier to each connected region. After completing the connected component labeling, the geometric feature parameters of each labeled connected region are extracted. The extracted geometric features include at least two or more of the following: the area of ​​the connected region, the width and height of the bounding rectangle, the aspect ratio defined by the ratio of the width to the height, and the fill rate defined by the ratio of the area of ​​the connected region to the area of ​​its bounding rectangle. The area of ​​the connected region is usually measured by the total number of pixels contained in the connected region, directly reflecting the scale of the candidate target in the pixel domain. Width and height represent the pixel span of the connected component in the horizontal and vertical directions, respectively. The aspect ratio reflects the shape extension characteristics of the candidate target in the two-dimensional plane. The fill rate reflects the density of the connected component within its bounding rectangle. Small, isolated targets with compact shapes typically exhibit a higher fill rate, while strip-shaped or ring-shaped non-target structures show a lower fill rate. By extracting these multiple geometric features, each candidate connected component is assigned a set of multi-dimensional attribute description vectors, providing a quantitative basis for subsequent target-non-target discrimination based on geometric prior knowledge.

[0074] Step S33: At least one of the area, width, and height thresholds, scaled according to a scale factor, is combined with the aspect ratio threshold and fill rate threshold of the connected component to form a joint filtering condition. The width and height thresholds are scaled according to the scale factor s, and the area threshold is scaled according to the square of the scale factor s. 2 Scaling is performed, and a connected region is retained as a candidate region for either a bright or dark target only if the geometric features of the connected region simultaneously satisfy the joint screening conditions.

[0075] After obtaining multiple geometric feature parameters for each connected component, based on the prior properties of small targets in the physical world, such as small size, compact shape, and isotropic or approximately isotropic properties, joint geometric constraints are applied to each connected component to filter out non-target structures that, although having high response intensity, clearly do not conform to the geometric definition of small targets. Specifically, at least one of the area threshold, width threshold, and height threshold is first scaled proportionally based on the scale factor established in the preceding steps, so that these spatial scale-related geometric thresholds can be adjusted synchronously with changes in the input image resolution, thereby ensuring that the physical definition of the size of small targets remains constant under different resolution conditions. Here, since the scale factor s of the image resolution represents the scaling ratio of the input image relative to the reference resolution in the one-dimensional length direction, the one-dimensional geometric thresholds such as the width threshold and height threshold are linearly scaled according to the scale factor s, while the area threshold, as a two-dimensional scale measure, is scaled according to the square of the scale factor s. 2 Scaling is applied. The physical basis of this differentiated scaling method is that when the image resolution changes by a scale *s*, the one-dimensional pixel span of the same physical target in the image is proportional to *s*, while the pixel area is proportional to *s*. 2 The scaling factors are proportional to the scaling order of the area threshold. Therefore, only when the scaling order of the area threshold maintains a corresponding dimensional relationship with the width and height thresholds can the scaled geometric threshold set maintain consistency in the judgment results for targets of the same physical size under different resolution conditions. For example, when the input image resolution increases relative to the reference resolution, the scale factor increases accordingly, and the area threshold, width threshold, and height threshold also increase according to their respective scaling orders, so that small targets occupying more pixels in high-resolution images can still be correctly identified and not misjudged as edges of large targets. After completing the scale-adaptive scaling of the geometric thresholds, the scaled area threshold, width threshold, and height threshold are combined with the preset aspect ratio threshold and fill rate threshold to form a set of joint screening conditions. For each connected component obtained by connected component labeling, only when its area, width, height, aspect ratio, and fill rate, etc., all simultaneously satisfy the corresponding threshold constraints in the joint screening conditions, is the connected component identified as a candidate region that conforms to the geometric definition of a small target and retained. For example, the area of ​​a connected region must be less than or equal to the scaled area threshold, the width and height must be within the range defined by the scaled width threshold and height threshold, the aspect ratio must be close to the value of 1 to reflect the compactness of the form rather than a thin strip, and the fill rate must be higher than the preset fill rate lower limit to exclude hollow or broken non-target structures.

[0076] It should be noted that in actual infrared early warning or long-range reconnaissance images, small dark targets are affected by factors such as atmospheric transmission attenuation, detector non-uniformity, and complex background shadows. Their connected component morphology in the binary mask sometimes does not appear as an ideal compact mass, but may exhibit butterfly-shaped, bilobal, or other loosely structured, incompletely filled forms. If the same strict aspect ratio and fill rate thresholds are applied to dark targets as to bright targets, some genuine small dark targets may be incorrectly eliminated because their connected component morphology deviates from the typical compact assumption. Therefore, in setting the joint screening conditions, a more lenient aspect ratio constraint range and fill rate lower limit threshold can be configured separately for the candidate region of dark targets compared to bright targets. This improves the retention ability of atypically shaped small dark targets while maintaining a certain level of suppression of non-target structures, further ensuring the recall performance of the method for dark targets in complex real-world scenarios.

[0077] By combining the constraints of the above multidimensional geometric features, regions such as large target edges, striped textures, and background undulation structures that may also produce strong responses in the response map but whose geometric shapes are significantly different from those of small targets will be effectively eliminated, thereby greatly reducing the false alarm probability in subsequent enhancement processing.

[0078] Step S34: After step S33 is completed, if no connected components are preserved in the binary mask of the bright target, the entire binary mask of the bright target is used as a candidate region for the bright target. If no connected components are preserved in the binary mask of the bright target, the area threshold, aspect ratio threshold, or fill rate threshold is gradually relaxed according to a preset relaxation coefficient. If no connected components are preserved after relaxation, the top N connected components in response intensity ranking with an area less than a preset upper limit are retained as candidate regions for the bright target, where N is a preset positive integer.

[0079] If there are no preserved connected components in the binary mask of the dark target, the entire binary mask of the dark target is used as a candidate region of the dark target. If there are no preserved connected components in the binary mask of the dark target, the area threshold, aspect ratio threshold, or fill rate threshold is gradually relaxed according to a preset relaxation coefficient. If there are still no preserved connected components after relaxation, the top N connected components in response intensity ranking and with an area less than a preset upper limit are retained as the candidate regions of the dark target.

[0080] After the joint geometric constraint screening in step S33, a critical situation may arise where several candidate connected components generated by thresholding processing originally existed in the binary mask of a bright or dark target, but these candidate connected components are all eliminated because they cannot simultaneously satisfy all the geometric thresholds in the joint screening conditions. This situation usually occurs when the actual small target signal is extremely weak, its connected component area is too small, or its shape deviates slightly from typical parameters, or when the small target is temporarily absent in the current frame image. To avoid the irreversible deletion of small targets in the early stages of the enhancement process due to overly strict geometric screening rules when actual small targets exist, a fallback protection mechanism is set up.

[0081] After the joint geometric constraint filtering in step S33 is completed, a global check is performed on the filtering results to determine whether at least one connected component in the bright target binary mask is successfully preserved. If the check result shows that no connected components are preserved in the bright target binary mask, that is, the filtering operation in step S33 has eliminated all candidate connected components in the bright target binary mask, it means that the filtering condition may be too strict in the current image. In this case, instead of directly abandoning the geometric filtering results and preserving the original binary mask as a whole, a hierarchical fallback strategy is adopted for processing. First, one or more of the area threshold, aspect ratio threshold, or fill rate threshold are relaxed step by step according to the preset relaxation coefficient. After each relaxation level, the joint geometric constraint filtering in step S33 is re-executed to determine whether a connected component that meets the relaxed condition appears. The purpose of progressive relaxation is to gradually loosen the geometric definition boundaries of small targets in a controllable manner. This allows real small targets that were mistakenly excluded due to slight deviations from the typical compactness assumption to be recaptured within a moderately expanded geometric constraint range, while still retaining the ability to exclude large-area non-target structures that clearly do not conform to the characteristics of small targets. If no connected components are retained after progressive relaxation, it indicates that the small target signal in the current image may be extremely weak, or that there are temporarily no suspicious targets that meet the geometric definition in the frame. In this case, a response intensity-oriented retention strategy is adopted as a last-line fallback measure. That is, connected components are sorted from high to low response intensity, and N connected components with the highest response intensity and a single connected component area less than a preset area upper limit are retained as bright target candidate regions, where N is a preset positive integer. By replacing overall retention with response intensity sorting, even in the worst case where the geometric features do not meet the screening conditions, a small number of candidate regions with the highest probability of being small targets and controlled area can be submitted to subsequent processing steps. This significantly reduces the risk of completely missing real small targets while maximally suppressing the false enhancement of non-target structures.

[0082] Similarly, for a binary mask of a dark target, if it is confirmed that the filtering operation in step S33 failed to retain any connected components, the above-mentioned hierarchical fallback strategy is also executed in sequence. That is, the area threshold, aspect ratio threshold, or fill rate threshold are first relaxed step by step according to the preset relaxation coefficient. If no connected components are retained after relaxation, the top N connected components with the response intensity ranking and the area is less than the preset upper limit are retained as candidate regions of the dark target and output.

[0083] The aforementioned hierarchical and progressive fallback protection mechanism enables the present invention to achieve a more refined and reasonable balance between screening accuracy and recall completeness in edge cases where small target signals are extremely weak or morphological features are not typical. Compared with the fallback method that directly retains the entire binary mask, the strategy of gradually relaxing the fallback and retaining Top-N responses can controllably increase the capture probability of real small targets while maintaining the effective suppression capability of non-target structures. This avoids the irreversible loss of real small targets caused by strict screening conditions and prevents a large number of false responses from re-entering the subsequent enhancement process due to overly broad fallback conditions, thereby significantly improving the robustness and reliability of the overall method in complex and ever-changing real-world application environments.

[0084] Specifically, step S4, which generates a smooth soft mask based on the bright target candidate region and the dark target candidate region, and constructs a gain map corresponding to the smooth soft mask, includes the following sub-steps: Step S41: Perform Gaussian convolution on the candidate regions of bright and dark targets respectively to obtain the soft mask of bright and dark targets. The soft mask has the highest weight at the center of the target region and decays smoothly at the boundary of the target region.

[0085] After obtaining the candidate regions for bright and dark targets after geometric constraint screening, if the original binary boundaries of these candidate regions are directly used as the application range of the enhancement effect, a step change in gray value will be formed at the boundary between the enhanced region and the unenhanced background. This will produce obvious traces of artificial processing and block artifacts in the output image, which will seriously affect the visual quality of the enhanced image and the reliability of subsequent analysis and processing.

[0086] To address the aforementioned boundary abruptness issue, smoothing processing is performed on both bright and dark target candidate regions, transforming the binary candidate regions with hard boundary characteristics into soft masks with gentle transitions. Specifically, Gaussian convolution is first applied to the binary mask images of both bright and dark target candidate regions. Gaussian convolution employs a two-dimensional smoothing filter with a Gaussian kernel, exhibiting a spatial distribution where the weights are maximized at the center and gradually decrease towards the edges. When the Gaussian convolution kernel is applied to the binary candidate region, the pixels within the candidate region are weighted and averaged by their neighboring pixels, smoothly replacing the abrupt change from the first logical value to the second logical value at the candidate region boundary. At the center of the candidate region, since the surrounding neighboring pixels are all within the candidate region, the response value after convolution remains at a high level. At the edges and surrounding areas of the candidate region, because the convolution kernel spans both the inner and outer sides of the candidate region, the response value gradually decreases from the center outwards.

[0087] After Gaussian convolution, the convolution result is further normalized, linearly mapping the dynamic range of the convolution response values ​​to a preset numerical range, such as a closed interval from 0 to 1. Specifically, the normalization operation involves dividing each pixel value in the smoothed response map obtained after Gaussian convolution by the maximum pixel value within the global range of the smoothed response map, thus linearly mapping the weight range of the soft mask to a closed interval from zero to one. The advantage of using the global maximum value as the normalization benchmark is that it ensures that the center of each candidate target region or its most prominent part is assigned the highest enhancement weight of one, while the weights of other parts are determined according to their relative proportion to the maximum value. This makes the weight distribution between different candidate target regions comparable at a uniform scale and also facilitates dimensional consistency when weighted and combined with the subsequent gain map. The above soft mask generation process can be expressed as dividing the result of a convolution operation between a Gaussian kernel and a binary-preserving mask by the maximum value of the convolution result.

[0088] The formula for generating the soft mask S is as follows: ; in, To express the standard deviation as The 2D Gaussian convolution kernel, where * denotes the 2D convolution operation, and M represents the binary mask retained after geometric constraint filtering. This indicates retrieving the global maximum value.

[0089] The normalized results are the bright target soft mask and the dark target soft mask. In the above soft mask, the geometric center of the target area or the core area with the strongest response is assigned the highest weight, while the weight gradually and smoothly decreases from the center to the edge, decaying to zero at a certain distance from the target's periphery. This weight distribution characteristic allows the subsequent local enhancement effect applied based on the soft mask to achieve the strongest enhancement effect within the target area, while exhibiting a natural gradient enhancement intensity in the transition zone between the target and the background. This effectively eliminates the processing traces caused by hard boundary masks, creating a visually imperceptible smooth transition between the enhanced target area and the surrounding background.

[0090] Step S42: Based on the bright target response map and dark target response map after nonlinear enhancement processing, generate the initial gain map of the bright target and the initial gain map of the dark target respectively, and perform upper limit pruning on the initial gain map of the bright target and the initial gain map of the dark target respectively to obtain the gain map of the bright target and the gain map of the dark target, so as to prevent local over-enhancement.

[0091] While generating a smooth soft mask to define the spatial distribution range of the enhancement effect, a gain map is also constructed to control the enhancement intensity amplitude at each pixel location. The gain map is constructed based on the bright target response map and the dark target response map after nonlinear enhancement processing in step S3. The pixel values ​​in these two response maps reflect the salience of small bright or small dark targets at the corresponding locations. Target areas with stronger salience should receive a correspondingly stronger enhancement amplitude. Therefore, firstly, based on the bright target response map after nonlinear enhancement processing, its response values ​​are directly used as the initial gain values ​​at each pixel location in the initial gain map of bright targets. Similarly, based on the dark target response map after nonlinear enhancement processing, its response values ​​are used as the initial gain values ​​at each pixel location in the initial gain map of dark targets. In this way, the initial gain map maintains spatial consistency with the response map, and the initial gain value corresponds to the higher the target salience area, thus providing an intensity basis for subsequent differentiated enhancement. However, in actual images, due to local noise interference, residual background structures, or abnormal amplification of individual response values ​​during nonlinear enhancement, a few pixel locations in the initial gain map may have excessively high gain values. If the aforementioned abnormally high gain values ​​are directly applied to enhance the original image without restriction, undesirable visual effects such as brightness saturation, overexposure, or underexposure will occur at the corresponding pixel locations, damaging the dynamic range and detail of the output image. To prevent such local over-enhancement, this invention further performs upper limit clipping operations on the initial gain maps of bright and dark targets respectively. The specific implementation of upper limit clipping is as follows: a reasonable upper limit threshold is preset, and for pixel locations in the initial gain map where the gain value exceeds the upper limit threshold, the gain value is forcibly truncated to the upper limit threshold.

[0092] The upper limit of the gain threshold can be determined using a parameterized form associated with the maximum grayscale value of the image. Specifically, a maximum gain ratio parameter can be preset for both the gain map of bright targets and the gain map of dark targets. The upper limit of the gain threshold is then the product of this maximum gain ratio parameter and the maximum grayscale value of the image (usually 255). For example, if the maximum gain ratio for bright targets is set to 0.6, the gain value at any pixel position in the gain map of bright targets will not exceed 153 after cropping. By introducing a relative parameter, the maximum gain ratio, rather than an absolute value, to constrain the upper limit of the gain, the gain cropping strategy has a consistent relative control effect across images with different dynamic ranges, avoiding the tedious operation of repeatedly adjusting the absolute threshold due to differences in overall image brightness.

[0093] Target gain map G b Dark target gain graph G d The upper limit clipping operations are performed as follows: ; ; in, and These are the response maps of a bright target and a dark target after nonlinear enhancement processing, respectively. and These are the preset maximum gain ratios for bright targets and dark targets, respectively. This indicates taking the smaller of the two values.

[0094] For pixel locations where the gain value does not exceed the upper limit threshold, the original gain value remains unchanged. The bright and dark target gain maps obtained after upper limit cropping retain the relative enhancement intensity differences between target regions while constraining the amplitude of individual abnormally high gain points within a controllable range, thus ensuring the overall grayscale stability and natural visual appearance of the final enhanced output image. Through the coordinated operation of steps S41 and S42, this invention completes refined modeling of the local enhancement operation from two dimensions: the spatial distribution range of the enhancement effect and the amplitude control of the enhancement intensity. This provides the necessary control parameter basis for achieving smooth, natural, and controllable local target enhancement.

[0095] By quantitatively converting the target saliency intensity information contained in the continuous response map into an actual enhancement value that can be applied to the image grayscale value, and by applying a reasonable upper limit constraint on the gain amplitude, it prevents individual abnormally high response points caused by noise, background residue, or nonlinear enhancement process from dominating the local enhancement result. Thus, while effectively improving the contrast of the target area, it maintains the overall grayscale level stability and natural visual appearance of the output image, and avoids adverse enhancement side effects such as overexposure or underexposure.

[0096] Step S43: Determine whether the bright target soft mask and the dark target soft mask overlap or conflict in spatial position. When a conflict occurs, compare the response intensity of the bright target and the response intensity of the dark target at the corresponding position, and retain the soft mask and gain corresponding to the one with the larger response intensity as the effective mask and corresponding gain.

[0097] After obtaining the soft masks for bright and dark targets, along with their corresponding gain maps, the candidate regions for bright and dark targets are extracted independently in spatial distribution. However, in actual images, they may appear adjacent or even partially overlap. For example, in an infrared image with a complex background, a locally bright structure and its adjacent shadow area may be identified as candidate regions for bright and dark targets, respectively, causing their soft masks to overlap in the spatial transition zone. If a positive brightening enhancement operation and a negative darkening enhancement operation are performed simultaneously at such overlapping locations, the two opposing enhancement effects will superimpose at the same pixel location. This can result in either weakening the enhancement effect or, in severe cases, producing unexpected grayscale distortion or enhancement cancellation, seriously affecting the stability and predictability of the output image.

[0098] To address the spatial conflict issue associated with bipolar enhancement, this invention performs a pixel-by-pixel determination of whether bright and dark target soft masks overlap or conflict in spatial location. The determination is based on whether, at a given pixel location, the weight values ​​of both the bright and dark target soft masks are simultaneously greater than a preset minimum threshold. If both are significantly greater than zero, the pixel location is considered to be within a spatial conflict region of bright and dark target enhancement. When a spatial conflict is determined, the original response intensity values ​​of the bright and dark target response maps after nonlinear enhancement processing are retrieved and compared. The original response intensity values ​​reflect the relative significance of the bright and dark target candidate attributes at that location; a higher response intensity indicates a greater confidence that the location belongs to the corresponding polarity target.

[0099] Based on the comparison results above, at the conflict location, only the soft mask and its gain corresponding to the side with the larger response strength can be retained as the effective mask and corresponding gain at that location, while the soft mask weight of the side with the smaller response strength is set to zero at that location. Through the above conflict resolution process, the originally spatially competing light and dark bipolar enhancement effects are transformed into mutually exclusive unique enhancement directions, ensuring that each pixel location only receives a single polarity enhancement operation in the final enhancement fusion stage. This eliminates the output instability caused by mutual interference of enhancement effects and provides spatial arbitration guarantee for reliable enhancement in scenarios where bipolar targets coexist.

[0100] Furthermore, the criteria for determining overlapping or proximity conflicts are: at the same pixel location, the weight values ​​of both the bright target soft mask and the dark target soft mask are greater than a preset mask threshold, or the minimum spatial distance between the bright target candidate region and the dark target candidate region is less than a preset distance threshold.

[0101] After explaining the basic logic of bright-dark conflict resolution, the specific judgment conditions for "overlapping or proximity conflict" are further clearly defined. In actual images, the spatial relationship between candidate regions of bright and dark objects is not always an ideal conflict state of complete overlap. More often, it manifests as a non-completely overlapping proximity relationship, such as boundary adjacency, local intersection, or close juxtaposition. If only non-zero mask weights are used as the sole conflict criterion, edge situations that interfere with each other due to spatial proximity under enhancement may be missed. Therefore, this invention sets two complementary conflict judgment conditions, and satisfying either one is judged as a spatial conflict. The first judgment condition is a pixel-level overlap judgment based on soft mask weights, that is, at a certain pixel position, the weight values ​​of the soft mask of the bright object and the soft mask of the dark object are simultaneously greater than a preset mask threshold. Because soft masks exhibit a smooth attenuation characteristic at the boundary of the target area, the region with a weight value greater than the preset mask threshold actually defines the core influence range of the enhancement effect. When the core influence ranges of the bright and dark soft masks overlap at the same pixel location, it indicates that this location will simultaneously experience both positive brightening and negative darkening enhancement effects. Since these two effects are opposite in direction, they will inevitably cancel each other out or cause grayscale distortion, and therefore should be identified as a conflict region. The introduction of the preset mask threshold ensures that conflict determination is not affected by the small non-zero values ​​at the attenuation tail of the soft mask boundary. Conflict resolution is only triggered when both bright and dark enhancement effects reach a non-negligible level, improving the accuracy and noise resistance of the determination. The second judgment condition is based on proximity conflict determination according to the spatial distance of the candidate regions. That is, when the minimum spatial distance between the bright target candidate region and the dark target candidate region is less than the preset distance threshold, even if the core influence ranges of their soft masks do not overlap at the same pixel location, they are still considered to have a spatial conflict. This condition takes into account the smooth transition characteristics of soft masks, causing the enhancement effect to extend outwards along the target region boundary to a certain extent. When two candidate regions of opposite polarities are too close, their respective soft mask attenuation bands will spatially overlap, forming a low-intensity interaction zone of brightness enhancement within the transition region. Although it does not reach the level of complete pixel-level overlap, it is still enough to cause instability in the enhancement effect and grayscale fluctuations in the boundary region. The preset distance threshold provides a quantifiable control parameter for the scale of this proximity range, and its value can be adaptively set according to factors such as the size of the structuring element, the standard deviation of the soft mask Gaussian kernel, or the resolution of the input image. The above two complementary conflict judgment conditions cover various situations in which bright and dark targets may cause enhancement interference from two spatial scales: pixel-level overlap and region-level proximity. This makes the conflict resolution mechanism have more comprehensive coverage and more robust execution in real complex scenarios. Once a conflict is determined based on any of the above conditions, the response intensity of the bright target and the response intensity of the dark target at the corresponding position are compared. The soft mask and gain corresponding to the one with the larger response intensity are retained as the effective mask and corresponding gain, thereby achieving arbitration of the unique enhancement direction at the conflict position.

[0102] Specifically, step S5 involves fusing the input image with the effective mask and corresponding gain, including... The input image, the bright target soft mask and bright target gain map in the effective mask, and the dark target soft mask and dark target gain map in the effective mask are weighted and fused to obtain a polarity-preserving enhanced output image: .

[0103] in, For the input image at pixel position grayscale value at that location and These are the soft masks for bright targets and dark targets in the effective masking, respectively. and These are the gain maps of bright targets and dark targets corresponding to the effective mask, respectively. and These are the preset enhancement weights for bright targets and dark targets, respectively.

[0104] After resolving brightness and darkness conflicts and determining the effective mask and corresponding gain at each pixel location, the original input image is weighted and fused with the aforementioned effective mask and corresponding gain to generate the final polarity-preserving enhanced output image. The fusion process is performed pixel-by-pixel, and for each pixel location in the output image, its final grayscale value is determined by three components.

[0105] The first part is the original grayscale value of the input image at this pixel location. This component preserves the background radiation distribution and texture details of the image, ensuring that pixels not identified as target candidate regions remain unchanged in the output image. The second part is the bright target enhancement component. This component consists of the product of the weight value of the bright target soft mask in the effective mask at this pixel location, the gain value of the bright target gain map at this pixel location, and a preset bright target enhancement weight. This is added as a positive increment to the original grayscale value. The weight value of the bright target soft mask determines the spatial intervention of the bright target enhancement at this pixel location, the bright target gain value determines the magnitude of the enhancement intensity, and the bright target enhancement weight is a global adjustment factor used to uniformly control the overall intensity of bright target enhancement. The third part is the dark target enhancement component. This component consists of the product of the weight value of the dark target soft mask in the effective mask at this pixel location, the gain value of the dark target gain map at this pixel location, and a preset dark target enhancement weight. This is subtracted from the original grayscale value as a negative deduction. The weight value of the dark target soft mask determines the degree of spatial intervention of the dark target enhancement at that pixel location, the dark target gain value determines the magnitude of the darkening intensity, and the dark target enhancement weight is a global adjustment factor used to uniformly control the overall strength of the dark target enhancement.

[0106] Through the weighted fusion method described above, regions identified as bright target candidates and retained through conflict resolution receive a positive boost in output grayscale value compared to the original background, resulting in a brighter and more prominent visual appearance. Regions identified as dark target candidates and retained through conflict resolution receive a negative decrease in output grayscale value compared to the original background, resulting in a darker and more prominent visual appearance. For regions that are neither bright nor dark target candidates, or polar components eliminated during conflict resolution, their corresponding soft mask weights are zero, and the contribution of the enhancement component is zero; the grayscale value of the pixels in these regions remains completely consistent with the input image. This fusion mechanism achieves bipolar-preserving local enhancement—bright small targets are locally brightened, and dark small targets are locally darkened—while the grayscale levels and texture details of the background region are fully preserved. This significantly improves the recognizability of small targets while maximizing the authenticity and stability of the original image's background radiation distribution.

[0107] Accordingly, a second aspect of the present invention provides an electronic device, including: at least one processor; and a memory connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to cause the at least one processor to perform the above-described bipolar small target resolution adaptive local enhancement method.

[0108] Accordingly, a third aspect of the present invention provides a computer-readable storage medium having computer instructions stored thereon, which, when executed by a processor, implement the above-described bipolar small target resolution adaptive local enhancement method.

[0109] The embodiments of the present invention aim to protect a bipolar small target resolution adaptive local enhancement method, which has the following effects: 1. By establishing a scale factor and adaptively searching for the optimal morphological structural element size according to the input image resolution, while linking the small target geometric screening threshold with the scale factor, this invention achieves automatic adaptation and stable migration of algorithm parameters under different resolution input conditions. This ensures that the physical size definition of small targets remains consistent in the pixel domain, effectively avoiding the problems of inconsistent enhancement effects and difficulty in uniform parameter configuration caused by the use of fixed structural element size and fixed screening threshold in traditional morphological enhancement methods when resolution changes. This significantly improves the robustness and engineering applicability of small target enhancement in multi-resolution scenarios. 2. By performing morphological opening and closing operations with optimal structuring element size to construct bright and dark target background images, and independently extracting bright and dark target responses, a polarity-preserving fusion strategy is adopted to apply positive local enhancement to bright targets and negative local enhancement to dark targets. This invention achieves the separate expression and independent processing of bright and dark bipolar targets throughout the entire process from background modeling to response extraction to enhancement output. It fundamentally changes the problem of dark target attribute distortion caused by the unified superposition of bright and dark responses in the traditional Top-hat method, and truly achieves the enhancement effect of "brighter bright targets and darker dark targets" with basically unchanged background radiation distribution. It has significant advantages in scenes with both bright and dark small targets in complex backgrounds. 3. By introducing joint geometric constraints on the area, width, height, aspect ratio, and fill rate of connected components after thresholding the response map, and setting a fallback strategy that retains the entire binary mask when no connected component meets the screening conditions, the risk of false responses generated by non-target structures in large target edges, striped textures, and complex background structures being mistakenly enhanced is effectively suppressed. At the same time, it avoids the situation where real small targets are wrongly rejected due to overly strict screening conditions. Based on this, combined with Gaussian smooth soft mask mapping and light-dark space conflict resolution mechanism, the local enhancement only acts on the region that truly meets the definition of small target and the enhancement boundary is smooth and natural, which significantly reduces the false alarm rate and improves the visual quality and stability of the enhanced output image.

[0110] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0111] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0112] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0113] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0114] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the specific implementation of the present invention. Any modifications or equivalent substitutions that do not depart from the spirit and scope of the present invention should be covered within the scope of protection of the claims of the present invention.

Claims

1. A bipolar small target resolution adaptive local enhancement method, characterized in that, The steps include the following: Step S1: Acquire the input image and establish a scale factor based on the relationship between the resolution of the input image and the preset reference resolution. Use the scale factor to adaptively determine the optimal structural element size for morphological background modeling. Step S2: Based on the optimal structuring element size, perform morphological opening and morphological closing operations on the input image to construct a bright target background image and a dark target background image, and extract the bright target response image corresponding to the bright target background image and the dark target response image corresponding to the dark target background image. Step S3: Perform nonlinear enhancement processing on the bright target response map and the dark target response map respectively, and filter the candidate regions of bright targets and dark targets that meet the preset geometric conditions according to the geometric constraints of small targets; Step S4: Generate a smooth soft mask based on the bright target candidate region and the dark target candidate region, and construct a gain map corresponding to the smooth soft mask. When the bright target candidate region and the dark target candidate region have a spatial conflict, perform conflict resolution to determine the effective mask and the corresponding gain. Step S5: The input image is fused with the effective mask and the corresponding gain to apply positive local enhancement to bright small targets and negative local enhancement to dark small targets, so as to obtain an enhanced output image that preserves polarity.

2. The bipolar small target resolution adaptive local enhancement method according to claim 1, characterized in that, Step S1, which involves adaptively determining the optimal structuring element size for morphological background modeling using the scale factor, includes the following sub-steps: Step S121: Select each candidate size in the candidate structural element size sequence after scaling by the scale factor, and perform morphological opening and morphological closing operations on the input image respectively to obtain the opening operation background estimation map and the closing operation background estimation map corresponding to each candidate size. Step S122: For adjacent candidate sizes, calculate the first average change between the background estimation maps of the opening operation and the second average change between the background estimation maps of the closing operation, and construct a comprehensive evaluation index of the stability of the background estimation based on the first average change and the second average change. Step S123: When the value of the comprehensive evaluation index is lower than the preset convergence threshold, it is determined that the background estimation has converged, the next candidate size is determined as the optimal structural element size and the search is terminated.

3. The adaptive local enhancement method for bipolar small target resolution according to claim 2, characterized in that, For adjacent candidate sizes k i and k i+1 The formula for calculating the comprehensive evaluation index is as follows: ; in, and respectively using candidate sizes k Structural elements for input image I Background estimation map obtained by performing morphological opening and closing operations. This indicates calculating the mean. ε These are preset non-zero small constants; When corresponding to candidate size k i With k i+1 Comprehensive evaluation indicators When the value is below the preset convergence threshold, the next candidate size k is selected. i+1 The optimal structural element size was determined.

4. The adaptive local enhancement method for bipolar small target resolution according to claim 1, characterized in that, Step S2, which involves performing morphological opening and closing operations on the input image based on the optimal structuring element size to construct a bright target background image and a dark target background image, and extracting the bright target response image corresponding to the bright target background image and the dark target response image corresponding to the dark target background image, includes the following sub-steps: Step S21: Using the optimal structural element size as a parameter, perform a morphological opening operation on the input image to obtain the bright target background image, which represents the smooth background after eliminating bright small structures; Step S22: Using the optimal structural element size as a parameter, perform a morphological closing operation on the input image to obtain the dark target background image, which represents the smooth background after filling in the dark small structures; Step S23: Subtract the gray values ​​of each pixel in the input image from the corresponding pixel gray values ​​in the bright target background image to obtain the bright target response map. The bright target response map is used to characterize the local prominence of the bright small target relative to its surrounding background. Step S24: Subtract the gray values ​​of each pixel in the dark target background image from the corresponding pixel gray values ​​in the input image to obtain the dark target response image. The dark target response image is used to characterize the degree of local concavity of the dark target relative to its surrounding background.

5. The adaptive local enhancement method for bipolar small target resolution according to claim 4, characterized in that, The bright target response map The calculation formula is: ; The dark target response map The calculation formula is: ; in, For the input image at pixel position grayscale value at that location The background image for the bright target. This is the background image of the dark target.

6. The bipolar small target resolution adaptive local enhancement method according to claim 1, characterized in that, Step S3, which involves filtering candidate regions for bright and dark targets based on the geometric constraints of small targets to obtain the pre-defined geometric conditions, includes the following sub-steps: Step S31: Threshold the response maps of the bright and dark targets after the nonlinear enhancement process to obtain the binary mask of the bright target and the binary mask of the dark target. Step S32: Perform connected component labeling on the bright target binary mask and the dark target binary mask respectively, and extract the geometric features of each connected component. The geometric features include at least two of the connected component's area, width, height, aspect ratio, and fill rate. Step S33: At least one of the area, width, and height values, scaled based on the scale factor, is used as a geometric threshold. This threshold, along with the aspect ratio threshold and fill rate threshold of the connected component, forms a joint filtering condition. The width and height thresholds are scaled according to the scale factor s, and the area threshold is scaled according to the square of the scale factor s. 2 Scaling is performed, and the connected region is retained as either the bright target candidate region or the dark target candidate region only when the geometric features of a connected region simultaneously satisfy the joint screening conditions; Step S34: After step S33 is completed, if there are no retained connected components in the bright target binary mask, the entire bright target binary mask is used as the bright target candidate region; if there are no retained connected components in the bright target binary mask, the area threshold, aspect ratio threshold, or fill rate threshold is gradually relaxed according to a preset relaxation coefficient. If there are still no retained connected components after relaxation, the top N connected components in response intensity ranking and with an area less than a preset upper limit are retained as the bright target candidate regions, where N is a preset positive integer. If no connected components are preserved in the binary mask of the dark target, the entire binary mask of the dark target is used as the candidate region of the dark target. If no connected components are preserved in the binary mask of the dark target, the area threshold, aspect ratio threshold, or fill rate threshold is gradually relaxed according to a preset relaxation coefficient. If no connected components are preserved after relaxation, the top N connected components in response intensity ranking and with an area less than a preset upper limit are retained as the candidate regions of the dark target.

7. The adaptive local enhancement method for bipolar small target resolution according to claim 1, characterized in that, Step S4, which involves generating a smooth soft mask based on the bright target candidate region and the dark target candidate region, and constructing a gain map corresponding to the smooth soft mask, includes the following sub-steps: Step S41: Perform Gaussian convolution processing on the bright target candidate region and the dark target candidate region respectively to obtain a bright target soft mask and a dark target soft mask. The soft mask has the highest weight at the center of the target region and decays smoothly at the boundary of the target region. Step S42: Based on the bright target response map and dark target response map after the nonlinear enhancement processing, generate the initial gain map of the bright target and the initial gain map of the dark target respectively, and perform upper limit pruning on the initial gain map of the bright target and the initial gain map of the dark target respectively to obtain the gain map of the bright target and the gain map of the dark target, so as to prevent local over-enhancement; Step S43: Determine whether the bright target soft mask and the dark target soft mask overlap or conflict in spatial position. When a conflict occurs, compare the response intensity of the bright target and the response intensity of the dark target at the corresponding position, and retain the soft mask and gain corresponding to the one with the larger response intensity as the effective mask and corresponding gain. The criteria for determining overlap or proximity conflict are: at the same pixel position, the weight values ​​of the bright target soft mask and the dark target soft mask are both greater than a preset mask threshold, or the minimum spatial distance between the bright target candidate region and the dark target candidate region is less than a preset distance threshold.

8. The adaptive local enhancement method for bipolar small target resolution according to claim 7, characterized in that, Step S5, which involves fusing the input image with the effective mask and corresponding gain, includes: The input image, the bright target soft mask and bright target gain image in the effective mask, and the dark target soft mask and dark target gain image in the effective mask are weighted and fused to obtain a polarity-preserving enhanced output image. ; in, For the input image at pixel position grayscale value at that location and These are the bright target soft mask and the dark target soft mask in the effective mask, respectively. and These are the gain maps of the bright target and the dark target corresponding to the effective mask, respectively. and These are the preset enhancement weights for bright targets and dark targets, respectively.

9. The adaptive local enhancement method for bipolar small target resolution according to any one of claims 1-8, characterized in that, Step S1, which involves acquiring the input image and establishing a scaling factor based on the relationship between the resolution of the input image and a preset reference resolution, includes the following sub-steps: Step S111: Read the current pixel width and current pixel height of the input image, and calculate the scale factor based on the ratio of the current pixel width and current pixel height to the geometric mean of the reference pixel width and reference pixel height of the preset reference resolution; Step S112: Scale the preset set of candidate structural element sizes using the scale factor to obtain a sequence of candidate structural element sizes that is adapted to the resolution of the input image.

10. The bipolar small target resolution adaptive local enhancement method according to claim 9, characterized in that, The scale factor The calculation formula is: ; in, and These are the current pixel height and the current pixel width, respectively. and These are the reference pixel height and the reference pixel width, respectively.