An image enhancement-based weak and small target tracking precision optimization method and system

By employing a dynamic feedback mechanism and local enhancement processing, the problem of decreased tracking accuracy of UAVs targeting small targets in complex environments was solved, achieving effective differentiation between target signals and background noise and improving tracking accuracy.

CN121837667BActive Publication Date: 2026-06-02CHENGDU AERONAUTIC POLYTECHNIC +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHENGDU AERONAUTIC POLYTECHNIC
Filing Date
2026-03-06
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

In complex and ever-changing environments, existing image enhancement and target tracking technologies cannot effectively distinguish target signals from background noise, resulting in decreased tracking accuracy or even loss of small targets by drones in conditions such as fog and haze.

Method used

By introducing a dynamic feedback mechanism and local enhancement processing, feedback signals are generated based on the real-time tracking status of the target tracking module, and local enhancement processing is performed on the areas requiring enhancement, including local contrast stretching, local detail sharpening, and background suppression. The enhancement strategy is then adjusted based on the tracking status after re-evaluation.

Benefits of technology

It effectively avoids the noise amplification and artifact problems of traditional global augmentation methods under complex weather conditions, significantly improves the tracking accuracy and robustness of weak targets, and ensures continuous and stable tracking of targets in complex backgrounds.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121837667B_ABST
    Figure CN121837667B_ABST
Patent Text Reader

Abstract

The application discloses a weak and small target tracking precision optimization method and system based on image enhancement, relates to the technical field of image processing, and is used for solving the problem that existing image enhancement and target tracking technologies cannot effectively distinguish target signals from background noise in a complex and changeable environment. The method comprises the following steps: acquiring a video frame; performing basic denoising and brightness correction on the video frame; inputting the preliminarily processed video frame into a target tracking module to obtain a tracking state, wherein the tracking state comprises tracking reliability and position prediction deviation; when the tracking reliability decreases or the position prediction deviation increases, a feedback signal is generated, the feedback signal is used for performing local enhancement processing on an enhancement demand region, the video frame after the local enhancement processing is inputted into the target tracking module again, the tracking state after reevaluation is obtained through the target tracking module, and the strategy of the local enhancement processing is adjusted according to the tracking state after the reevaluation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and in particular to a method and system for optimizing the tracking accuracy of small targets based on image enhancement. Background Technology

[0002] Drones move at high speeds, and their flight attitude and observation angle are constantly changing. This movement, combined with uneven fog areas, creates an even more challenging situation. For example, a drone might fly from a thin fog area into a thick fog area, or when turning, the change in the camera's angle of view might cause some areas of the image to be covered by dense fog under sunlight, while other areas remain relatively clear. This results in significant differences in the degree of blur and noise levels between different areas within the same image frame. In this case, enhancement methods using globally uniform parameters become completely ineffective. They cannot differentiate the degradation levels in different areas of the image. For areas with dense fog, the enhancement may not be sufficient to make the target visible; while for areas with thin fog, the same enhancement may over-amplify noise, creating new interference. This non-uniform, rapidly changing image degradation characteristic poses a serious challenge to existing enhancement and tracking technologies, making the continuous and stable tracking of small targets under complex weather conditions an unresolved problem. Summary of the Invention

[0003] This application provides a method and system for optimizing the tracking accuracy of weak targets based on image enhancement. It aims to solve the problem that existing image enhancement and target tracking technologies cannot effectively distinguish between target signals and background noise in complex and variable environments, especially under the non-uniform and rapidly changing image degradation characteristics such as fog, haze, high-speed movement of UAVs, and changes in observation angle. This leads to a decrease in tracking accuracy or even target loss.

[0004] In a first aspect, to address the aforementioned technical problems, this invention provides a method for optimizing the tracking accuracy of small targets based on image enhancement. The method includes: acquiring video frames; performing basic denoising and brightness correction on the video frames to obtain pre-processed video frames; inputting the pre-processed video frames into a target tracking module to perform target tracking on the pre-processed video frames, obtaining a tracking state, which includes tracking reliability and position prediction deviation; generating a feedback signal when tracking reliability decreases or position prediction deviation increases, the feedback signal containing the enhancement requirement region; performing local enhancement processing on the enhancement requirement region based on the feedback signal, the local enhancement processing including local contrast stretching, local detail sharpening, and background suppression; and inputting the locally enhanced video frames back into the target tracking module to obtain a re-evaluated tracking state, and adjusting the local enhancement processing strategy based on the re-evaluated tracking state.

[0005] Secondly, this application provides a system for optimizing the tracking accuracy of small targets based on image enhancement. The system includes: a video frame acquisition module for acquiring video frames; a preliminary processing module for performing basic denoising and brightness correction on the video frames to obtain pre-processed video frames; a target tracking module for inputting the pre-processed video frames into the target tracking module to perform target tracking on the pre-processed video frames and obtain a tracking state, the tracking state including tracking reliability and position prediction deviation; a feedback signal generation module for generating a feedback signal when tracking reliability decreases or position prediction deviation increases, the feedback signal containing the enhancement requirement area; a local enhancement processing module for performing local enhancement processing on the enhancement requirement area based on the feedback signal, the local enhancement processing including local contrast stretching, local detail sharpening, and background suppression; and a strategy adjustment module for re-inputting the locally enhanced video frames into the target tracking module to obtain a re-evaluated tracking state, and adjusting the local enhancement processing strategy based on the re-evaluated tracking state.

[0006] This application has at least the following beneficial effects: It discloses a method for optimizing the tracking accuracy of small targets based on image enhancement. By introducing a dynamic feedback mechanism, it can intelligently determine when image enhancement is needed and the area to be enhanced based on the real-time evaluation of the tracking status (including tracking reliability and position prediction deviation) by the target tracking module. When tracking reliability decreases or position prediction deviation increases, the system generates a feedback signal containing the area requiring enhancement and performs targeted local enhancement processing on that area. This processing includes local contrast stretching, local detail sharpening, and background suppression. This local, on-demand enhancement strategy effectively avoids the noise amplification and artifact problems that may occur in traditional global enhancement methods under complex weather conditions (such as fog or haze), and overcomes the shortcomings of existing technologies where enhancement operations cannot distinguish between the target signal and background noise caused by fog scattering. Furthermore, this application can dynamically adjust the local enhancement processing strategy based on the re-evaluated tracking status, forming a closed-loop optimization process, thereby effectively solving the problem of continuous and stable tracking of small targets by UAVs in high-speed motion, changing observation angles, and non-uniform degradation environments. This method significantly improves the distinguishability of targets in complex backgrounds, enabling the tracking module to more accurately lock onto and continuously follow the target's position. It significantly enhances the accuracy and robustness of tracking weak targets, achieving unexpected technical results. Attached Figure Description

[0007] Figure 1 This is a flowchart illustrating a method for optimizing the tracking accuracy of small targets based on image enhancement, as provided in this application. Detailed Implementation

[0008] The technical solutions of this application will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of this application, and not all embodiments. The components of this application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.

[0009] It should be noted that similar reference numerals and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. Furthermore, in the description of this application, terms such as "first," "second," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0010] In traditional wide-area surveillance missions using unmanned aerial vehicles (UAVs), image quality for small targets deteriorates significantly under complex weather conditions, such as fog or haze, causing the target signal to be overwhelmed by background noise. Existing image enhancement strategies often employ globally uniform parameters, which cannot effectively distinguish between the target signal and background noise caused by fog scattering. They may even over-amplify noise, generating numerous artifacts, making it difficult for the target tracking module to accurately identify and continuously track the target, ultimately leading to tracking interruption. This non-uniform and rapidly changing image degradation characteristic poses a severe challenge to the continuous and stable tracking of small targets.

[0011] In view of the above problems, this application provides a method for optimizing the tracking accuracy of weak targets based on image enhancement. By introducing a dynamic feedback mechanism and local enhancement processing, the image enhancement strategy can be adjusted according to the real-time tracking status, effectively addressing the image degradation problem under complex weather conditions and significantly improving the accuracy and robustness of tracking weak targets.

[0012] The following specific embodiments will provide a detailed introduction and explanation of the image enhancement-based method and system for optimizing the tracking accuracy of small targets provided in this application.

[0013] Reference Figure 1 This application provides a method for optimizing the tracking accuracy of small targets based on image enhancement. The method may include the following steps:

[0014] S1. Obtain video frames.

[0015] A video frame is a static image extracted from a continuous video stream and is the basic unit for image processing and target tracking.

[0016] These video frames can be captured in real time by an electro-optical pod mounted on a drone, or they can be read from a pre-recorded video stream. For example, a standard video capture card or image sensor interface can be used to acquire continuous digital image data at a rate of 25 or 30 frames per second.

[0017] S2. Perform basic noise reduction and brightness correction on the video frames to obtain preliminarily processed video frames.

[0018] Basic denoising refers to removing random noise from an image using various algorithms (such as mean filtering, Gaussian filtering, median filtering, etc.) to improve image clarity. Brightness correction involves adjusting the overall brightness level of the image to maintain visual consistency under different lighting conditions, such as through gamma correction or histogram equalization. The pre-processed video frame refers to the video frame after basic denoising and brightness correction; its image quality is improved compared to the original video frame, but localized degradation issues may still exist.

[0019] Basic denoising can be achieved using various methods. For example, a Gaussian filter can be applied to smooth the image and reduce the impact of random noise; alternatively, a median filter can be used to effectively remove salt-and-pepper noise. Brightness correction can be achieved through histogram equalization, which makes the overall brightness distribution of the image more uniform, thereby improving the visual effect. For instance, for a video frame that is too dark, histogram equalization can remap its pixel values, thus improving the overall brightness of the image and enhancing contrast.

[0020] S3. Input the pre-processed video frame into the target tracking module so that the target tracking module can track the pre-processed video frame and obtain the tracking status, which includes tracking reliability and position prediction deviation.

[0021] The target tracking module is a software or hardware unit that integrates target detection, feature extraction, and motion estimation functions. It is used to identify and track the position and trajectory of a specific target in consecutive video frames. Tracking reliability is an indicator that measures the confidence level of the target tracker in the tracking results. For example, it can be comprehensively evaluated by factors such as the stability of target features, the smoothness of the motion trajectory, and the degree of background interference. Position prediction bias refers to the difference between the target position predicted by the target tracker and the actual target position, reflecting the accuracy of the tracking.

[0022] The target tracking module can employ correlation filtering-based tracking algorithms, such as KCF (Kernelized Correlation Filter) or MOSSE (Minimum Output Sum of Squared Error), which can quickly and accurately locate the target. During tracking, the module continuously evaluates the tracking quality. For example, tracking reliability can be measured by calculating the peak response to sidelobe ratio of the target region; a higher peak response and lower sidelobe ratio indicate higher tracking reliability. Position prediction error can be quantified by comparing the Euclidean distance between the predicted target position and the actual detected position in the current frame.

[0023] S4. When tracking reliability decreases or location prediction deviation increases, a feedback signal is generated, which includes the enhanced demand area.

[0024] The feedback signal is a control signal generated by the system when the tracking status is poor. It is used to indicate the areas that need image enhancement and the enhancement requirements.

[0025] When tracking reliability decreases or position prediction error increases, the system generates a feedback signal that includes the region requiring enhancement. When the target tracking module detects a decline in tracking quality, such as a significant drop in response peaks or severe jitter in the target bounding box, the system triggers the feedback mechanism. The feedback signal can be a binary flag indicating whether enhancement is needed, and it also includes a rectangular box or mask indicating the area requiring enhancement. For example, when the target tracking module determines that the target may be about to be lost, it generates a feedback signal and marks the current target bounding box and a certain surrounding area as the region requiring enhancement.

[0026] S5. Based on the feedback signal, perform local enhancement processing on the area requiring enhancement. Local enhancement processing includes local contrast stretching, local detail sharpening, and background suppression.

[0027] In this context, the enhancement requirement area refers to the specific region in a video frame that the target tracking module determines needs image enhancement processing, typically the target and its surrounding area. Local enhancement processing refers to performing image enhancement operations only on this specific region (i.e., the enhancement requirement area), rather than performing global enhancement on the entire image. This includes local contrast stretching, local detail sharpening, and background suppression. Local contrast stretching involves adjusting the distribution of pixel values ​​within the enhancement requirement area to expand the dynamic range of the image, thereby enhancing the contrast between the target and the background. Local detail sharpening involves enhancing image edge and texture information to make the target's outline and internal details clearer. Background suppression involves reducing the brightness or contrast of the background portion within the enhancement requirement area to further highlight the target.

[0028] Local enhancement processing includes local contrast stretching, local detail sharpening, and background suppression. Local contrast stretching can be achieved through local histogram equalization or adaptive gamma correction, adjusting only the pixel values ​​within the enhancement area to improve the contrast between the target and the background. Local detail sharpening can employ unsharp masking techniques, enhancing high-frequency information to highlight target edges and textures. Background suppression can further emphasize the target by applying a decay factor to background pixels within the enhancement area, reducing their brightness or contrast. For example, for a target that appears blurry in fog, local contrast stretching can make its outline clearer, local detail sharpening can enhance its internal texture, and background suppression can reduce the interference of surrounding fog on the target.

[0029] S6. Input the video frame after local enhancement processing back into the target tracking module to obtain the tracking status for re-evaluation through the target tracking module, and adjust the local enhancement processing strategy based on the re-evaluated tracking status.

[0030] The strategy adjustment module is a unit that dynamically adjusts local enhancement processing parameters or methods based on the re-evaluated tracking status to achieve adaptive image enhancement.

[0031] After local enhancement, the salience of the target in the image is improved. The target tracking module then tracks the frame again and re-evaluates the tracking status. If the re-evaluation shows improved tracking reliability and reduced position prediction error, the current enhancement strategy is considered effective. If the tracking status remains unsatisfactory, the strategy adjustment module dynamically adjusts the parameters of the local enhancement process based on the feedback signal and the re-evaluation of the tracking status. This includes adjusting the intensity of contrast stretching, the parameters of the sharpening filter, or the attenuation factor of background suppression to achieve the best enhancement effect. For example, if the target is still blurry after the first enhancement, the strategy adjustment module might increase the intensity of contrast stretching or the coefficients of the sharpening filter to further highlight the target.

[0032] In summary, this application achieves intelligent and adaptive image enhancement by closely integrating target tracking status with image enhancement strategies. This method not only solves the problem of low tracking accuracy of weak targets under complex weather conditions using traditional methods, but also significantly improves the robustness and tracking performance of the system through local processing and strategy adjustment, providing more reliable technical support for UAV wide-area surveillance missions.

[0033] In some embodiments, this application further proposes a more refined local enhancement processing method, which aims to effectively filter false features and enhance the saliency of the target ontology through steps such as multi-scale feature extraction, degraded information encoding, clean feature space projection and reconstruction, and saliency map generation, thereby significantly improving the tracking accuracy of weak targets in complex environments.

[0034] Specifically, based on the feedback signal, the above-mentioned local enhancement processing of the enhancement requirement area includes: extracting multi-scale features from the enhancement requirement area to obtain multi-scale features, and analyzing the local brightness distribution and blur degree in the enhancement requirement area to obtain encoded degradation information; projecting the multi-scale features and encoded degradation information into a clean feature space; reconstructing the multi-scale features in the clean feature space according to the encoded degradation information and target saliency requirements to filter out false features caused by atmospheric interference and instantaneous strong local reflections, and to enhance the saliency features of the target body; mapping the reconstructed clean saliency features back to the image space to generate a target saliency map, and using the target saliency map as the result of the local enhancement processing.

[0035] Specifically, multi-scale feature extraction is performed on the enhancement-demand region to capture the structural and textural information of the target at different granularities. For example, wavelet transform, Gaussian pyramids, or convolutional neural networks in deep learning can be used to extract features such as edges, corners, and textures at different scales. Simultaneously, the local brightness distribution and blur level within the enhancement-demand region are analyzed to quantify the degree of image degradation. Local brightness distribution reflects issues such as image contrast and uneven illumination, while blur level indicates image sharpness. These analytical results are encoded as degradation information to guide subsequent feature reconstruction.

[0036] Projecting multi-scale features and encoded degradation information into a clean feature space can be understood as transforming the features in the original image, which contain both target and degradation information, into an abstract space that is less affected by degradation through a certain mapping relationship. In this space, the features of the target ontology can be more clearly represented, while the degradation information is used to guide how to "cleanse" these features. In practical applications, this can be achieved by designing specific neural network layers or mathematical transformations, such as using models like autoencoders or variational autoencoders to learn the mapping from degradation features to clean features.

[0037] Furthermore, within the clean feature space, multi-scale features are reconstructed based on encoded degradation information and target saliency requirements. The aim is to accurately separate and filter false features caused by atmospheric interference and transient strong local reflections, and to selectively enhance the saliency features of the target ontology. Encoded degradation information provides prior knowledge about the type and degree of image degradation, while target saliency requirements indicate which features are crucial for target identification and tracking. By combining these two types of information, adaptive reconstruction mechanisms can be designed. For example, attention mechanisms or gating units can be used to assign different weights to different feature dimensions in the clean feature space, thereby suppressing false features and enhancing target ontology features.

[0038] Therefore, the reconstructed clean saliency features are mapped back to the image space to generate a target saliency map, which is then used as the result of local enhancement processing. The target saliency map is a grayscale image where the target region has higher pixel values, while the background region has lower pixel values, clearly highlighting the target's position and contour in the image. This mapping process typically involves deconvolution or upsampling operations to restore the abstract features in the clean feature space to an image representation with spatial correspondence.

[0039] This application's scheme, by introducing multi-scale feature extraction and degradation information encoding, can comprehensively capture the detailed information and degradation status of an image. It is precisely by projecting this information together into a clean feature space that the features of the target ontology and the false features caused by degradation can be effectively separated at a higher dimension. Based on this, by reconstructing the multi-scale features within the clean feature space according to the encoded degradation information and target saliency requirements, this application can accurately identify and suppress false features caused by factors such as atmospheric interference and transient strong local reflections, while selectively enhancing the unique saliency features of weak target ontology. Finally, the reconstructed clean saliency features are mapped back into the image space to generate a target saliency map. This saliency map clearly highlights the target, providing high-quality, high-signal-to-noise-ratio input for the subsequent target tracking module, thus effectively solving the problem that traditional local enhancement methods struggle to effectively distinguish targets from interference in complex environments.

[0040] In some embodiments, this application further proposes that the steps of projecting multi-scale features and encoded degradation information into a clean feature space include: performing color channel bias analysis on the image within the enhancement requirement area to quantify the attenuation degree of different color channels and identify local highlight overflow areas; constructing a multi-dimensional degradation representation vector by combining local brightness distribution, blur degree, color channel bias, and local highlight overflow areas; and dynamically selecting or adjusting the feature projection path in the clean feature space according to the multi-dimensional degradation representation vector to project multi-scale features and degradation information into the clean feature space.

[0041] Specifically, color channel deviation analysis of images within the enhancement requirement area refers to quantifying the degree of attenuation of different color channels due to factors such as atmospheric scattering, absorption, or uneven sensor response by calculating the statistical differences or relative intensities between different color channels (e.g., red, green, and blue channels). Simultaneously, identifying local highlight overflow areas refers to detecting regions in the image where the brightness values ​​have reached or are close to saturation. These areas typically suffer from information loss due to strong light sources or reflections, interfering with subsequent feature extraction and projection. The aim is to comprehensively capture the degradation information of the image in the color dimension, providing a basis for subsequent degradation compensation.

[0042] Specifically, by combining local brightness distribution, blur level, color channel deviation, and local highlight overflow areas, a multi-dimensional degradation characterization vector is constructed. This can be understood as integrating the various degradation indicators obtained from the above analysis into a comprehensive vector representation. This vector can comprehensively and precisely describe the degradation state of the image in different dimensions. For example, local brightness distribution reflects the brightness of the overall image or a local area, blur level reflects the image sharpness, color channel deviation reflects color distortion, and local highlight overflow areas indicate extreme degradation areas of information saturation. The purpose is to provide an accurate degradation "fingerprint" for the projection of the clean feature space, guiding the de-degradation processing of features.

[0043] In practical applications, based on the multi-dimensional degradation representation vector, the feature projection path is dynamically selected or adjusted in the pure feature space. Specifically, this means adaptively selecting the most suitable mathematical transformation or neural network mapping function based on the degradation characteristics of the current image, mapping multi-scale features and degradation information from the original image space or preliminary feature space to the pure feature space. For example, when the degradation representation vector indicates severe blurring of the image, the system may choose a projection path that emphasizes edge and texture restoration; when there is significant color deviation, it may choose a path that focuses more on color correction. The purpose is to ensure that the projection process can specifically remove various degradation effects, so that the features of the target ontology can be clearly and accurately represented in the pure feature space, thus laying the foundation for subsequent feature reconstruction and target saliency enhancement.

[0044] The proposed solution analyzes color channel deviations in the image within the enhancement requirement area and identifies local highlight clipping regions, enabling a more comprehensive acquisition of image degradation information. By combining local brightness distribution, blur level, color channel deviations, and local highlight clipping regions, a multi-dimensional degradation representation vector is constructed, resulting in a more refined and accurate representation of degradation information. This refined multi-dimensional degradation representation allows for the dynamic selection or adjustment of feature projection paths in a clean feature space based on this vector, thereby achieving precise projection of multi-scale features and degradation information and effectively removing the effects of degradation.

[0045] In some embodiments, this application further proposes the following steps for reconstructing multi-scale features in a clean feature space based on encoded degradation information and target saliency requirements to filter false features caused by atmospheric interference and transient strong local reflections, and to enhance the saliency features of the target ontology: using encoded degradation information as a multi-dimensional degradation representation vector, and dynamically adjusting the reference benchmark of the target ontology features based on the multi-dimensional degradation representation vector to adapt to the dynamic fine-tuning of subtle features caused by changes in viewing angle of small targets; dynamically generating multiple weight coefficients based on the multi-dimensional degradation representation vector and target saliency requirements; and selectively enhancing feature dimensions related to the target ontology and suppressing feature dimensions related to atmospheric interference and transient strong local reflections using an adaptive gating mechanism, thereby reconstructing the features to filter false features caused by atmospheric interference and transient strong local reflections and to enhance the saliency features of the target ontology.

[0046] Specifically, using encoded degradation information as a multi-dimensional degradation representation vector means using the comprehensive degradation information previously constructed by analyzing local brightness distribution, blurriness, color channel deviation, and local highlight overflow areas as the basis for guiding feature reconstruction. Based on this multi-dimensional degradation representation vector, the reference benchmark for the target ontology features is dynamically adjusted. The purpose is to enable the feature reconstruction process to flexibly adapt to subtle feature adjustments caused by changes in viewing angle for small targets. For example, when a drone's attitude changes, causing a slight change in the target's viewing angle, some features of the target ontology (such as edges and textures) may undergo slight deformation or brightness changes. By dynamically adjusting the reference benchmark, it can be ensured that even under these subtle changes, the true features of the target ontology can be accurately identified and enhanced, avoiding misclassification as false features.

[0047] In this system, multiple weighting coefficients are dynamically generated based on the multi-dimensional degradation representation vector and the target saliency requirement. This can be understood as the system intelligently calculating a set of values ​​to guide feature enhancement and suppression based on the current image degradation status (reflected by the multi-dimensional degradation representation vector) and the assessment of the target's importance (target saliency requirement). These weighting coefficients can be adjusted in real time according to different degradation types (such as haze, strong light, blur) and target features (such as size, shape, motion state).

[0048] In practical applications, adaptive gating mechanisms selectively enhance feature dimensions relevant to the target ontology using weighted coefficients, while suppressing feature dimensions related to atmospheric interference and transient strong local reflections. The aim is to achieve refined feature filtering and enhancement. The adaptive gating mechanism can be a neural network-based gating unit, such as the gating structure in a Long Short-Term Memory (LSTM) network or a Gated Recurrent Unit (GRU), or an attention-based module. This mechanism performs dimension-wise or grouped weighted processing on multi-scale features in the clean feature space based on dynamically generated weighted coefficients. Feature dimensions highly relevant to the target ontology are given higher weights for enhancement, while feature dimensions related to atmospheric interference (such as scattered light and fog textures) and transient strong local reflections (such as water surface reflections and glass reflections) are given lower weights for suppression or even complete filtering. This effectively filters out false features caused by these interferences and significantly enhances the salient features of the target ontology, thereby improving the accuracy of subsequent target tracking.

[0049] Through the above technical solution, this application overcomes the limitations of existing technologies in the reconstruction of features of small targets. Specifically, by dynamically adjusting the reference benchmark of the target's ontological features, this application can effectively adapt to the subtle dynamic fine-tuning of features caused by changes in viewing angle, avoiding feature misjudgment or loss due to minor changes in the target's appearance. Furthermore, by dynamically generating weight coefficients based on multi-dimensional degraded representation vectors and target saliency requirements, and combining this with an adaptive gating mechanism for selective enhancement and suppression, this application can more accurately and thoroughly filter out false features caused by atmospheric interference and transient strong local reflections, while significantly enhancing the saliency features of the target's ontological features. This refined and adaptive feature reconstruction capability enables the ontological features of small targets to be extracted more clearly and stably in complex and changing observation environments, thereby providing higher-quality input for subsequent target tracking modules, significantly improving the accuracy and robustness of small target tracking, and reducing the risk of tracking failure.

[0050] In some embodiments, the process of mapping the reconstructed clean saliency features back to the image space to generate a target saliency map, and using the target saliency map as the result of local enhancement processing, includes the following steps: adjusting geometric correction parameters based on the real-time attitude information of the UAV and the instantaneous observation parameters of the electro-optical pod, and performing geometric correction based on the adjusted geometric correction parameters to compensate for spatial geometric distortion caused by high-speed maneuvering of the aircraft and instantaneous drastic changes in the observation angle; fine-tuning the geometric correction parameters based on the residual atmospheric influence information encoded in the multi-dimensional degradation characterization vector to correct for small positional shifts and scale changes caused by atmospheric refraction and scattering; applying adaptive smoothing processing to the edge regions of the target saliency map to ensure a natural transition of the target contour and avoid jagged effects or artifacts introduced by feature reconstruction; mapping the clean saliency features after geometric correction, fine-tuning, and edge smoothing processing back to the image space to generate a target saliency map, and using the target saliency map as the result of local enhancement processing.

[0051] The process of adjusting geometric correction parameters based on the UAV's real-time attitude information and the instantaneous observation parameters of the electro-optical pod, and then performing geometric correction based on these adjusted parameters, involves using real-time attitude data (such as roll angle, pitch angle, and yaw angle) provided by the UAV's onboard sensors (e.g., IMU, GPS) and the electro-optical pod's own instantaneous observation parameters (e.g., gimbal angle, focal length, and field of view) to calculate and adjust the geometric correction parameters used for image spatial transformation. These parameters can include affine transformation matrices, perspective transformation matrices, or more complex nonlinear transformation models. The purpose is to accurately compensate for spatial geometric distortions caused by image jitter, changes in viewing angle, and the instantaneous and drastic adjustments in the observation direction by the electro-optical pod during high-speed maneuvers, ensuring that the spatial position and shape of the target in the image can be accurately reproduced.

[0052] Furthermore, fine-tuning the geometric correction parameters based on the residual atmospheric influence information encoded in the multi-dimensional degradation characterization vector refers to refining the adjusted geometric correction parameters on the basis of the initial geometric correction by utilizing the information specifically encoded in the aforementioned multi-dimensional degradation characterization vector (which comprehensively considers information such as local brightness distribution, blurring degree, color channel deviation, and local specular overflow areas during its construction) regarding residual effects such as atmospheric refraction and scattering. This fine-tuning aims to correct for minute positional offsets and scale variations caused by the complexity of the atmospheric environment (such as changes in refractive index at different altitudes, humidity levels, and temperatures). These variations may be difficult to completely eliminate in macroscopic geometric correction, but they have a significant impact on the tracking accuracy of weak targets.

[0053] Furthermore, applying adaptive smoothing to the edge regions of the target saliency map refers to using an algorithm that dynamically adjusts the smoothing intensity based on local features to address potential jaggedness, discontinuities, or artifacts in the edge regions after the target saliency map is generated. For example, the size and weight of the smoothing kernel can be adaptively selected based on information such as edge gradient, texture complexity, and local contrast. This removes unnatural artifacts while preserving the true details of the target contour to the maximum extent, ensuring natural and smooth edge transitions in the target saliency map, thereby improving visual quality and the robustness of subsequent processing.

[0054] This application's solution employs geometric correction by incorporating real-time UAV attitude information and instantaneous observation parameters from the electro-optical pod. This effectively compensates for spatial geometric distortions caused by high-speed aircraft maneuvers and drastic changes in observation angles, ensuring the spatial accuracy of the target saliency map. Furthermore, the geometric correction parameters are fine-tuned using residual atmospheric influence information encoded in the multi-dimensional degradation characterization vector, further eliminating minute positional shifts and scale changes caused by atmospheric refraction and scattering, resulting in more precise target positioning in image space. Simultaneously, adaptive smoothing of the target saliency map's edge regions effectively avoids jagged edges or artifacts that might be introduced during feature reconstruction, ensuring a natural transition and visual quality of the target contour. Therefore, this application's solution ensures that the target saliency map mapped back from the pure feature space to the image space possesses high precision and quality, providing a more reliable basis for subsequent local enhancement processing.

[0055] Specifically, the steps described above, which involve inputting the pre-processed video frame into the target tracking module, performing target tracking on the pre-processed video frame, and obtaining the tracking status, can be further refined into the following process.

[0056] The pre-processed video frames are input into the target tracking module, which then performs target tracking on these frames. The target tracking module can employ various existing tracking algorithms, such as correlation filter-based trackers, deep learning-based trackers (e.g., the Siamese network series), or tracking methods combining Kalman filtering and visual features, to achieve continuous localization of small targets.

[0057] The target tracking module analyzes the consistency of the target's motion trajectory across consecutive frames. Specifically, this analysis aims to assess whether the target's motion pattern over time is smooth and predictable, and whether there are sudden, unexpected changes in displacement or velocity. For example, the displacement vector of the target's center point across consecutive frames can be calculated, and the trends in the magnitude and direction of these vectors can be analyzed.

[0058] The target tracking module analyzes the stability of target appearance features. These features can include visual attributes such as color, texture, shape, and size. The analysis aims to determine whether these features remain relatively stable across consecutive frames, or whether there are significant fluctuations that could be caused by occlusion, lighting changes, or target deformation. For example, histograms, Local Binary Patterns (LBP), or depth features of the target region can be extracted, and their temporal similarity can be calculated.

[0059] The target tracking module analyzes the dynamic changes in background interference. Specifically, this analysis focuses on the background information of the area surrounding the target, identifying whether there are interfering objects that may be confused with the target, fast-moving background elements, or environmental changes (such as smoke, water vapor, etc.), which may negatively affect target tracking. For example, motion detection or feature analysis can be performed on local areas around the target to identify potential sources of interference.

[0060] The target tracking module determines tracking reliability based on the consistency of the motion trajectory, the stability of the target's appearance features, and dynamic changes in background interference, distinguishing between genuine tracking failures and temporary visual interference. Tracking reliability is a comprehensive indicator used to assess the credibility of the current tracking results. Tracking reliability decreases when the target's motion trajectory exhibits significant anomalies, its appearance features change drastically, or background interference increases sharply. By differentiating between genuine tracking failures (i.e., the target has been lost) and temporary visual interference (i.e., the target is temporarily occluded or affected by the environment, but tracking may still be possible), unnecessary tracking re-initialization or misjudgments can be avoided.

[0061] The target tracking module quantifies the position prediction error. Specifically, the position prediction error refers to the difference between the target position predicted by the target tracking module based on historical motion information and the actual observed target position. Through quantization, this error can be expressed in numerical form, such as pixel distance or normalization error, thus intuitively reflecting the tracking accuracy.

[0062] The tracking reliability and the quantized position prediction deviation are used as the tracking status. Therefore, the tracking status is a composite index that includes tracking reliability and position prediction deviation, comprehensively reflecting the current target tracking performance and potential problems.

[0063] This application's solution refines the target tracking process into a comprehensive analysis of motion trajectory consistency, target appearance feature stability, and dynamic changes in background interference, enabling a more comprehensive assessment of the target tracking module's current operational status. Specifically, by analyzing the target's motion trajectory across consecutive frames, abnormal motion patterns can be identified, such as sudden acceleration, deceleration, or changes in direction, which may indicate potential tracking problems. Simultaneously, analyzing the stability of target appearance features helps determine whether the target is occluded, deformed, or confused with the background, thus avoiding misjudgments due to appearance changes. Furthermore, analyzing dynamic changes in background interference allows for the timely detection and identification of interfering objects or environmental changes that may affect target tracking. Therefore, by comprehensively considering these factors, tracking reliability can be more accurately determined, effectively distinguishing between genuine tracking failure and merely temporary visual interference, thereby avoiding unnecessary local enhancement processing or premature abandonment of tracking. Furthermore, quantifying position prediction deviation provides a clear reflection of tracking accuracy fluctuations, offering a precise basis for subsequent adjustments to local enhancement processing strategies.

[0064] In some embodiments, this application further proposes that the target tracking module judges the tracking reliability based on the consistency of the motion trajectory, the stability of the target appearance features, and the dynamic changes of background interference, so as to distinguish between real tracking failure and temporary visual interference. This includes the following steps: analyzing the motion trajectory of the target in a continuous frame before it disappears or is occluded for a short time, and extracting the motion trend and periodic pattern of the motion trajectory; analyzing the appearance features of the target in a continuous frame before it disappears or is occluded for a short time, and extracting the statistical distribution of the main color, texture, and shape features of the appearance features; analyzing the dynamic changes of the background area surrounding the target in a continuous frame before it disappears or is occluded for a short time, and identifying the motion patterns of potential interference objects or occluders in the background; when the target reappears... After the occlusion is removed, the target motion trajectory in the current frame is compared with the motion trend and periodic pattern, the target appearance features in the current frame are compared with the statistical distribution, and the background interference pattern in the current frame is compared with the motion pattern to obtain the comparison results. According to the comparison results, if the target motion trajectory in the current frame is consistent with the motion trend, the target appearance features in the current frame are within the statistical distribution range, and the background interference pattern in the current frame does not change significantly, then the tracking reliability is judged as temporary disconnection. According to the comparison results, if the target motion trajectory in the current frame deviates significantly from the motion trend, or the target appearance features in the current frame exceed the statistical distribution range, or the background interference pattern in the current frame changes drastically, then the tracking reliability is judged as true tracking failure.

[0065] Specifically, before the target disappears or is occluded for a short period, the system continuously observes and accumulates data on the target. Analyzing the target's motion trajectory over a continuous series of frames before its disappearance or occlusion aims to extract the inherent patterns of the target's motion through processing historical motion data, such as using Kalman filtering, particle filtering, or deep learning-based motion prediction models. This includes its average velocity, acceleration, directional trends, and any possible periodic motion patterns. These patterns provide a benchmark for subsequent judgments regarding whether the target has actually deviated from its intended path.

[0066] Simultaneously, the appearance features of the target are analyzed within a series of frames before it disappears or is occluded. The aim is to quantify the target's main color, texture, and shape features using image processing and feature extraction techniques, such as SIFT, SURF, HOG descriptors, or deep features extracted by convolutional neural networks. Statistical analysis of these features allows for the establishment of a statistical distribution model of the target's appearance. This model reflects the range of appearance changes under normal tracking conditions, thus enabling the identification of whether abnormal changes have occurred in the target's appearance.

[0067] Furthermore, analyzing the dynamic changes in the background region surrounding the target within a series of frames before the target disappears or is occluded for a short period aims to identify the motion patterns or occurrence regularities of potential distractions (such as other moving objects, smoke, or changes in lighting) or occlusions (such as trees or buildings) in the background through background modeling, optical flow analysis, or target detection techniques. This helps distinguish between the target's own motion changes and visual interference caused by the background environment.

[0068] Once the target reappears in the video frame or the occlusion is removed, the system acquires the target's motion trajectory, appearance features, and background interference patterns for the current frame. Then, the target's motion trajectory in the current frame is compared with pre-extracted motion trends and periodic patterns to assess the consistency of its motion. Simultaneously, the target's appearance features in the current frame are compared with previously established statistical distributions to determine if its appearance remains within the normal range of variation. Furthermore, the background interference patterns in the current frame are compared with historical background motion patterns to identify whether the background environment has undergone drastic changes.

[0069] Based on the comparison results above, if the target's trajectory in the current frame remains highly consistent with historical motion trends, and the target's appearance features are still within the statistical distribution range, while the background interference pattern in the current frame has not changed significantly, then the tracking reliability is judged as a temporary loss of connection. This means that although the target has temporarily disappeared or been occluded, its core motion and appearance characteristics have not fundamentally changed, and the system still has the ability to recover tracking through local enhancements and other means. Conversely, if the target's trajectory in the current frame deviates significantly from historical motion trends, or the target's appearance features exceed the statistical distribution range, or the background interference pattern in the current frame changes drastically, then the tracking reliability is judged as a true tracking failure. This indicates that the target may have been completely lost or undergone irreversible changes, requiring the activation of a higher-level reacquisition mechanism.

[0070] This application's solution establishes a dynamic historical model of the target and its environment by pre-analyzing and extracting the trend and periodicity of its motion trajectory, the statistical distribution of its appearance features, and the dynamic change patterns of its surrounding background before the target briefly disappears or is occluded. When the target reappears or the occlusion is lifted, the system can accurately compare the target state of the current frame with this historical model. This dynamic comparison mechanism based on historical context enables the tracking module to more accurately determine whether the currently observed change in target state is a recoverable temporary loss of connection or an irrecoverable true tracking failure. It is precisely because of this refined judgment process that the system can avoid misjudging tracking failure due to temporary visual interference, thereby improving the accuracy and robustness of tracking decisions.

[0071] In some embodiments, this application further proposes a specific method for the target tracking module to analyze the consistency of the target's motion trajectory between consecutive frames, including: the target tracking module performs real-time analysis of the target's displacement vector between consecutive frames; the target tracking module performs geometric correction on the displacement vector based on the UAV's real-time attitude information and the instantaneous observation parameters of the electro-optical pod to compensate for spatial geometric distortion caused by high-speed maneuvering of the aircraft and instantaneous drastic changes in the observation angle; the target tracking module dynamically adjusts the sensitivity threshold for changes in the geometrically corrected displacement vector based on the blur degree and brightness distribution of the image's local degraded areas; the target tracking module compares the geometrically corrected displacement vector with historical motion trends and, in conjunction with the adjusted sensitivity threshold, judges the consistency of the motion trajectory to distinguish between the target's actual motion changes and trajectory jitter or interruption caused by local image degrade.

[0072] Specifically, real-time analysis of the target's displacement vector across consecutive frames refers to the target tracking module continuously calculating the target's positional changes between adjacent video frames and representing them as displacement vectors. These displacement vectors contain information about the target's direction of motion and distance on the image plane.

[0073] This process involves geometrically correcting the displacement vector based on the UAV's real-time attitude information and the instantaneous observation parameters of the electro-optical pod. The aim is to eliminate or reduce the impact of the observation platform's own motion on the target's apparent motion. The UAV's real-time attitude information can include the aircraft's position, velocity, acceleration, pitch angle, roll angle, and yaw angle, while the electro-optical pod's instantaneous observation parameters can include its gimbal's pitch angle, azimuth angle, and zoom magnification. These parameters allow for the establishment of a mapping from the image coordinate system to the world coordinate system or the stable platform coordinate system. This transforms the observed displacement vector into a reference frame independent of the platform's motion, compensating for spatial geometric distortions caused by high-speed aircraft maneuvers and sudden, drastic changes in the observation angle.

[0074] In practical applications, the sensitivity threshold for changes in the geometrically corrected displacement vector is dynamically adjusted based on the degree of blur and brightness distribution of locally degraded areas in the image. The aim is to enable the tracking system to more intelligently adapt to changes in image quality. When the degree of blur or uneven brightness distribution in locally degraded areas is high, target features may become unclear. In such cases, minute displacement changes may be caused by noise or degradation rather than genuine target motion. Therefore, the sensitivity threshold can be appropriately increased to avoid overreacting to these small, uncertain changes. Conversely, when image quality is good, the threshold can be decreased to improve the response sensitivity to genuine target motion. The degree of blur can be quantified using methods such as image gradient, edge strength, or Fourier transform, while brightness distribution can be evaluated using methods such as histogram analysis or local variance.

[0075] Furthermore, the geometrically corrected displacement vector is compared with the historical motion trend, and combined with an adjusted sensitivity threshold, to determine the consistency of the motion trajectory. The historical motion trend can be obtained by modeling the target's motion trajectory over a period of time (e.g., using a Kalman filter, particle filter, or simple averaging / weighted averaging). By comparing the geometrically corrected displacement vector of the current frame with the predicted historical motion trend, it can be assessed whether the current motion conforms to the expected pattern. Combined with the adjusted sensitivity threshold, if the deviation between the current displacement vector and the historical trend is within the threshold range, the motion trajectory is considered consistent; if the deviation exceeds the threshold, it may indicate that the target has undergone a real motion change, or, in the case of severe image degradation, there may be trajectory jitter or interruption. The purpose is to distinguish between real changes in target motion and trajectory jitter or interruption caused by local image degradation, thereby improving the robustness of tracking.

[0076] This application's solution significantly improves the accuracy of the target tracking module's judgment on the consistency of target motion trajectory by introducing mechanisms such as geometric correction, dynamic sensitivity threshold adjustment, and comparison with historical motion trends. First, the geometric correction step effectively eliminates the influence of the observation platform's own motion on the target's apparent displacement, ensuring that the analyzed displacement vector more realistically reflects the target's own motion. Second, dynamically adjusting the sensitivity threshold allows the system to intelligently adjust its response sensitivity to displacement changes based on real-time image quality conditions (such as blur level and brightness distribution), avoiding misjudgment of target motion due to minor noise or artifacts when the image quality is degraded, while maintaining a rapid response to real motion when the image quality is good. Finally, comparing the corrected displacement vector with historical motion trends, combined with the dynamic threshold, allows the system to judge the rationality of the current motion based on the target's long-term motion pattern, thereby effectively distinguishing between the target's true motion changes and trajectory jitter or interruption caused by local image degradation.

[0077] In some embodiments, this application further proposes a target tracking module to analyze the stability of target appearance features, specifically including: the target tracking module performing multi-scale feature extraction on a preset target area to obtain target appearance features; the target tracking module performing geometric correction on the target appearance features based on the real-time attitude information of the UAV and the instantaneous observation parameters of the electro-optical pod to compensate for the geometric distortion of the target appearance caused by the high-speed maneuvering of the aircraft and the instantaneous drastic change in the observation angle; the target tracking module dynamically adjusting the sensitivity threshold for changes in the geometrically corrected target appearance features based on the blur degree and brightness distribution of the local degraded areas of the image; the target tracking module comparing the geometrically corrected target appearance features with historical appearance feature trends; and the target tracking module combining the adjusted sensitivity threshold to determine the stability of the target appearance features to distinguish between changes in the true appearance of the target and feature distortion or background confusion caused by image degradation.

[0078] Specifically, the target tracking module performs multi-scale feature extraction on the preset target region, aiming to capture rich features of the target from image information of different resolutions and scales, such as edges, textures, and color distributions. Multi-scale feature extraction can employ various techniques, such as feature extraction methods based on wavelet transform, Gaussian pyramids, or deep learning networks (such as convolutional neural networks), with the goal of enhancing the robustness of target features to scale changes and ensuring effective recognition even when the target size changes.

[0079] The target tracking module performs geometric correction on the target's appearance features based on the UAV's real-time attitude information and the instantaneous observation parameters of the electro-optical pod. This step aims to compensate for spatial geometric distortions caused by high-speed aircraft maneuvers and sudden, drastic changes in the observation angle. For example, when the UAV performs pitch, roll, or yaw maneuvers, or when the electro-optical pod's observation angle is rapidly adjusted, the target's projection in the image will deform, change in scale, or shift in position. By utilizing the UAV's attitude sensor data (such as an inertial measurement unit, IMU) and the electro-optical pod's encoder data, these geometric transformation parameters can be accurately calculated. The extracted target appearance features are then inversely transformed and corrected to a unified reference coordinate system, thereby eliminating or significantly reducing the impact of these external factors on the stability of the target's appearance features.

[0080] Furthermore, the target tracking module dynamically adjusts the sensitivity threshold for changes in the geometrically corrected target appearance features based on the degree of blurring and brightness distribution in locally degraded areas of the image. Local image degradation, such as blurring, noise, or uneven brightness, directly affects the quality and reliability of the target's appearance features. For example, in blurred areas, detailed information about the features is lost, while in excessively bright or dark areas, the contrast of the features is reduced. By analyzing this degradation information, the "credibility" of the target's appearance features in the current image frame can be assessed. When the degree of degradation is high, the sensitivity threshold for feature changes should be appropriately relaxed to avoid misjudging minor feature fluctuations caused by degradation as true changes in the target's appearance; conversely, when the image quality is good, the threshold can be tightened to more accurately capture subtle changes in the target's appearance.

[0081] Based on this, the target tracking module compares the geometrically corrected target appearance features with historical appearance feature trends. Historical appearance feature trends can be obtained through statistical analysis, modeling (e.g., smoothing and predicting feature vectors using Kalman filtering or particle filtering), or averaging of target appearance features from multiple consecutive past frames. By comparing with historical trends, it can be determined whether the current changes in target appearance features are within the expected range or whether they exhibit some kind of anomaly.

[0082] Therefore, the target tracking module, combined with the adjusted sensitivity threshold, judges the stability of the target's appearance features to distinguish between changes in the target's true appearance and feature distortion or background blurring caused by image degradation. This judgment process comprehensively considers the geometric correction of features, the impact of image quality on sensitivity, and the comparison results with historical trends, thus enabling a more accurate assessment of the intrinsic stability of the target's appearance.

[0083] This application's scheme acquires rich target information through multi-scale feature extraction and eliminates interference from external motion and changes in observation angle through geometric correction, making the target appearance features more comparable across different frames. Simultaneously, by dynamically adjusting the sensitivity threshold, the system can intelligently adapt to changes in image quality, avoiding misjudgments caused by image degradation. Finally, by comparing with historical trends, this scheme can effectively distinguish between the target's true appearance changes and feature distortions caused by the environment or sensors, thereby ensuring the accuracy of the target appearance stability assessment.

[0084] The specific implementation of this application also discloses a weak target tracking accuracy optimization system based on image enhancement. The system includes: a video frame acquisition module for acquiring video frames; and a preliminary processing module for performing basic noise reduction and brightness correction on the video frames to obtain pre-processed video frames.

[0085] The target tracking module receives the pre-processed video frames and tracks them to obtain a tracking status, including tracking reliability and position prediction deviation. The feedback signal generation module generates a feedback signal when tracking reliability decreases or position prediction deviation increases. This feedback signal includes the enhancement requirement area. The local enhancement processing module performs local enhancement processing on the enhancement requirement area based on the feedback signal. This local enhancement processing includes local contrast stretching, local detail sharpening, and background suppression. The strategy adjustment module re-inputs the locally enhanced video frames to the target tracking module to obtain a re-evaluated tracking status and adjusts the local enhancement processing strategy based on this re-evaluated tracking status.

[0086] The image enhancement-based system for optimizing the tracking accuracy of small targets proposed in this application aims to address the problem in traditional UAV wide-area surveillance missions where the image quality of small targets deteriorates significantly under complex weather conditions, causing the target signal to be submerged by background noise. Existing image enhancement strategies cannot effectively distinguish the target signal from background noise, and may even over-amplify the noise, resulting in a sharp decline in the performance of the target tracking module and a high risk of tracking interruption. This system, by introducing a dynamic feedback mechanism and modular local enhancement processing, can adaptively adjust the image enhancement strategy according to the real-time tracking status, effectively addressing the image degradation problem under complex weather conditions and significantly improving the accuracy and robustness of small target tracking. The various modules work collaboratively to form a closed-loop control system, ensuring stable tracking of small targets even in harsh environments.

[0087] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.

Claims

1. A method for optimizing the tracking accuracy of weak targets based on image enhancement, characterized in that, The method includes: acquiring video frames; The video frames are subjected to basic noise reduction and brightness correction to obtain pre-processed video frames; The pre-processed video frame is input into the target tracking module to perform target tracking on the pre-processed video frame and obtain the tracking status, which includes tracking reliability and position prediction deviation. When the tracking reliability decreases or the position prediction deviation increases, a feedback signal is generated, the feedback signal containing an enhanced demand region; Based on the feedback signal, local enhancement processing is performed on the enhanced area, including local contrast stretching, local detail sharpening, and background suppression. The video frame after the local enhancement process is input back into the target tracking module to obtain a re-evaluated tracking state, and the strategy of the local enhancement process is adjusted based on the re-evaluated tracking state. The step of performing local enhancement processing on the enhanced demand region based on the feedback signal includes: Multi-scale feature extraction is performed on the enhanced demand region to obtain multi-scale features, and the local brightness distribution and blur degree in the enhanced demand region are analyzed to obtain encoded degradation information; The multi-scale features and the encoded degraded information are projected into a pure feature space; Within the pure feature space, the multi-scale features are reconstructed based on the encoded degradation information and target saliency requirements to filter out false features caused by atmospheric interference and transient strong local reflections, and to enhance the saliency features of the target body. The reconstructed pure saliency features are mapped back to the image space to generate a target saliency map, which is then used as the result of local enhancement processing. The step of projecting the multi-scale features and the encoded degradation information into a clean feature space includes: Color channel deviation analysis is performed on the image within the enhancement requirement area to quantify the attenuation degree of different color channels and identify local highlight clipping areas; By combining the local brightness distribution, the degree of blurring, the color channel deviation, and the local highlight overflow area, a multi-dimensional degradation characterization vector is constructed. Based on the multi-dimensional degradation representation vector, the feature projection path is dynamically selected or adjusted in the pure feature space to project the multi-scale features and degradation information into the pure feature space.

2. The method for optimizing the tracking accuracy of weak targets based on image enhancement according to claim 1, characterized in that, Within the pure feature space, the multi-scale features are reconstructed based on the encoded degradation information and target saliency requirements to filter out false features caused by atmospheric interference and transient strong local reflections, and to enhance the saliency features of the target entity, including: The encoded degradation information is used as a multi-dimensional degradation representation vector, and the reference benchmark of the target ontology features is dynamically adjusted according to the multi-dimensional degradation representation vector to adapt to the dynamic fine-tuning of the subtle features of weak targets caused by changes in perspective. Based on the multi-dimensional degradation characterization vector and the target saliency requirement, multiple weight coefficients are dynamically generated; Through an adaptive gating mechanism, the weighting coefficients are used to selectively enhance feature dimensions related to the target ontology and suppress feature dimensions related to atmospheric interference and transient strong local reflections, thereby reconstructing the features to filter out false features caused by atmospheric interference and transient strong local reflections and enhance the salient features of the target ontology.

3. The method for optimizing the tracking accuracy of weak targets based on image enhancement according to claim 2, characterized in that, The step of mapping the reconstructed clean saliency features back to the image space to generate a target saliency map, and using the target saliency map as the result of local enhancement processing, includes: Based on the real-time attitude information of the UAV and the instantaneous observation parameters of the electro-optical pod, the geometric correction parameters are adjusted, and geometric correction is performed based on the adjusted geometric correction parameters to compensate for the spatial geometric distortion caused by the high-speed maneuvering of the aircraft and the instantaneous and drastic changes in the observation angle. Based on the residual atmospheric influence information encoded in the multidimensional degradation characterization vector, the geometric correction parameters are fine-tuned to correct the small positional shifts and scale changes caused by atmospheric refraction and scattering. Adaptive smoothing is applied to the edge regions of the target saliency map to ensure a natural transition of the target contour and avoid jagged effects or artifacts introduced by feature reconstruction. The clean saliency features, after geometric correction, fine-tuning, and edge smoothing, are mapped back to the image space to generate a target saliency map, which is then used as the result of local enhancement processing.

4. The method for optimizing the tracking accuracy of weak targets based on image enhancement according to claim 1, characterized in that, The pre-processed video frame is input into the target tracking module to perform target tracking on the pre-processed video frame and obtain the tracking status, including: The pre-processed video frame is input into the target tracking module, and the target tracking module performs target tracking on the pre-processed video frame; The target tracking module analyzes the consistency of the target's motion trajectory across consecutive frames; The target tracking module analyzes the stability of the target's appearance features; The target tracking module analyzes the dynamic changes in background interference; The target tracking module determines the tracking reliability based on the consistency of the motion trajectory, the stability of the target appearance features, and the dynamic changes of background interference, so as to distinguish between real tracking failure and temporary visual interference. The target tracking module quantifies the position prediction deviation; The tracking reliability and the deviation from the quantized position prediction are taken as the tracking status.

5. The method for optimizing the tracking accuracy of weak targets based on image enhancement according to claim 4, characterized in that, The target tracking module determines the tracking reliability based on the consistency of the motion trajectory, the stability of the target's appearance features, and the dynamic changes in background interference, in order to distinguish between genuine tracking failures and temporary visual interference, including: Analyze the motion trajectory of the target in a series of frames before it disappears or is occluded in a short time, and extract the motion trend and periodic pattern of the motion trajectory; Analyze the appearance features of the target in a series of frames before it disappears or is occluded in a short time, and extract the statistical distribution of the main color, texture and shape features of the appearance features; Analyze the dynamic changes of the background region surrounding the target in a series of frames before the target disappears or is occluded in a short time, and identify the motion patterns of potential interference or occlusion in the background. When the target reappears or the occlusion is removed, the target motion trajectory in the current frame is compared with the motion trend and periodic pattern, the target appearance features in the current frame are compared with the statistical distribution, and the background interference pattern in the current frame is compared with the motion pattern to obtain the comparison results. Based on the comparison results, if the target motion trajectory in the current frame is consistent with the motion trend, and the target appearance features in the current frame are within the statistical distribution range, while the background interference pattern in the current frame has not changed significantly, then the tracking reliability is determined to be a temporary loss of connection. Based on the comparison results, if the target motion trajectory in the current frame deviates significantly from the motion trend, or the target appearance features in the current frame exceed the statistical distribution range, or the background interference pattern in the current frame changes drastically, then the tracking reliability is determined to be a true tracking failure.

6. The method for optimizing the tracking accuracy of weak targets based on image enhancement according to claim 4, characterized in that, The target tracking module analyzes the consistency of the target's motion trajectory across consecutive frames, including: The target tracking module performs real-time analysis of the target's displacement vector between consecutive frames; The target tracking module performs geometric correction on the displacement vector based on the real-time attitude information of the UAV and the instantaneous observation parameters of the electro-optical pod, in order to compensate for the spatial geometric distortion caused by the high-speed maneuvering of the aircraft and the instantaneous and drastic changes in the observation angle. The target tracking module dynamically adjusts the sensitivity threshold to changes in the geometrically corrected displacement vector based on the degree of blurring and brightness distribution of the local degraded areas of the image. The target tracking module compares the geometrically corrected displacement vector with the historical motion trend and, in conjunction with the adjusted sensitivity threshold, determines the consistency of the motion trajectory to distinguish between the actual motion changes of the target and trajectory jitter or interruption caused by local image degradation.

7. The method for optimizing the tracking accuracy of weak targets based on image enhancement according to claim 4, characterized in that, The target tracking module analyzes the stability of the target's appearance features, including: The target tracking module performs multi-scale feature extraction on the preset target region to obtain the target appearance features; The target tracking module performs geometric correction on the target's appearance features based on the UAV's real-time attitude information and the instantaneous observation parameters of the electro-optical pod, in order to compensate for the geometric distortion of the target's appearance caused by the high-speed maneuvering of the aircraft and the instantaneous and drastic changes in the observation angle. The target tracking module dynamically adjusts the sensitivity threshold to changes in the geometrically corrected target appearance features based on the degree of blurring and brightness distribution of the local degraded areas of the image. The target tracking module compares the geometrically corrected target appearance features with historical appearance feature trends; The target tracking module, in conjunction with the adjusted sensitivity threshold, determines the stability of the target's appearance features to distinguish between changes in the target's true appearance and feature distortion or background obfuscation caused by image degradation.

8. A system for optimizing the tracking accuracy of small targets based on image enhancement, characterized in that, The system includes: a video frame acquisition module for acquiring video frames; The preliminary processing module is used to perform basic noise reduction and brightness correction on the video frames to obtain pre-processed video frames. The target tracking module is used to input the pre-processed video frame into the target tracking module so that the target tracking module can perform target tracking on the pre-processed video frame to obtain the tracking status, the tracking status including tracking reliability and position prediction deviation; A feedback signal generation module is used to generate a feedback signal when the tracking reliability decreases or the position prediction deviation increases, the feedback signal including an enhancement demand region; The local enhancement processing module is used to perform local enhancement processing on the enhancement requirement area according to the feedback signal. The local enhancement processing includes local contrast stretching, local detail sharpening, and background suppression. The strategy adjustment module is used to input the video frame after the local enhancement processing back into the target tracking module so as to obtain the tracking state after re-evaluation through the target tracking module, and adjust the strategy of the local enhancement processing according to the tracking state after re-evaluation. The step of performing local enhancement processing on the enhanced demand region based on the feedback signal includes: Multi-scale feature extraction is performed on the enhanced demand region to obtain multi-scale features, and the local brightness distribution and blur degree in the enhanced demand region are analyzed to obtain encoded degradation information; The multi-scale features and the encoded degraded information are projected into a pure feature space; Within the pure feature space, the multi-scale features are reconstructed based on the encoded degradation information and target saliency requirements to filter out false features caused by atmospheric interference and transient strong local reflections, and to enhance the saliency features of the target body. The reconstructed pure saliency features are mapped back to the image space to generate a target saliency map, which is then used as the result of local enhancement processing. The step of projecting the multi-scale features and the encoded degradation information into a clean feature space includes: Color channel deviation analysis is performed on the image within the enhancement requirement area to quantify the attenuation degree of different color channels and identify local highlight clipping areas; By combining the local brightness distribution, the degree of blurring, the color channel deviation, and the local highlight overflow area, a multi-dimensional degradation characterization vector is constructed. Based on the multi-dimensional degradation representation vector, the feature projection path is dynamically selected or adjusted in the pure feature space to project the multi-scale features and degradation information into the pure feature space.