Unmanned aerial vehicle identification method and system for low-altitude security

By using a joint optimization model of ambient light intensity change rate and background motion vector field in drone recognition technology, combining time-domain deconvolution kernel and adversarial sample defense feature extraction, the problem of degradation of recognition accuracy caused by dynamic lighting and complex background motion is solved, and efficient drone recognition and adversarial attack defense is achieved.

CN120198658AActive Publication Date: 2025-06-24TIANJIN YUNXIANG UAV TECH CO LTD

Patent Information

Application Number
CN202510688228.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-27
Publication Date
2025-06-24
Estimated Expiration
2045-05-27

AI Technical Summary

Technical Problem

The existing drone recognition technology reduces recognition accuracy when dealing with dynamic lighting changes and complex background movements, and the adversarial training strategy has limited defense capabilities for new adversarial samples.

Method used

A joint optimization model based on ambient light intensity change rate and background motion vector field based on low-altitude security scenes is used to generate optical compensation parameters in real time, eliminate motion blur through time domain deconvolution kernel, extract adversarial sample defense features, and perform target confidence estimation through a multi-scale residual network of attention mechanism.

Benefits of technology

It significantly improves the imaging stability and target capture quality of drone recognition, enhances the defense ability against attacks, and improves recognition accuracy and robustness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120198658A_ABST
    Figure CN120198658A_ABST
Patent Text Reader

Abstract

The invention provides an unmanned aerial vehicle identification method and system for low-altitude security and protection. An optical compensation parameter set for unmanned aerial vehicle imaging optimization is generated in real time through a joint optimization model of an ambient light intensity change rate and a background motion vector field, and an original optical sequence in a target capture window is processed by using the parameter set. And constructing a time domain deconvolution kernel in combination with the motion characteristics of the unmanned aerial vehicle to generate an enhanced optical image resistant to motion blur. Non-linear weighted fusion is carried out through a disturbance intensity evaluation function, and a confrontation disturbance feature mask is formed. The enhanced optical image and the confrontation disturbance feature mask are subjected to airspace superposition operation, a multi-scale residual network is adopted to carry out target confidence estimation on the superposed image, an unmanned aerial vehicle recognition result is generated, and the unmanned aerial vehicle recognition accuracy and the anti-interference capacity in the low-altitude complex environment are remarkably improved through the technical scheme provided by the invention.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of drone recognition, and particularly to a drone recognition method and system for low-altitude security. Background Art

[0002] In low-altitude security scenarios, drone recognition faces complex environmental interference and adversarial attacks, such as dynamic light changes, background motion interference, and malicious adversarial sample attacks.

[0003] Existing technologies usually rely on deep learning-based object detection networks and combine optical flow estimation methods to track moving targets. This solution compensates for background motion through the optical flow field, extracts target features using a deep learning model for recognition, and introduces an adversarial training strategy to enhance the model's defense against adversarial samples. When dealing with dynamic light changes, existing solutions fail to effectively combine the environmental light intensity change rate for adaptive compensation, resulting in a significant decline in recognition accuracy in scenarios with drastic light changes. In addition, the optical flow estimation method is prone to errors when dealing with complex background motion, and the adversarial training strategy has limited defense capabilities against new types of adversarial samples, making it difficult to cope with diverse adversarial attacks. Summary of the Invention

[0004] Embodiments of this application provide a drone recognition method and system for low-altitude security to solve the problem of poor accuracy in existing technologies.

[0005] In a first aspect, embodiments of this application provide a drone recognition method for low-altitude security, including: Based on a joint optimization model of the environmental light intensity change rate and the background motion vector field in a low-altitude security scenario, a set of optical compensation parameters optimized for drone imaging is generated in real time; An original optical sequence is obtained within a target capture window through the set of optical compensation parameters, and a time-domain deconvolution kernel is constructed in combination with the motion characteristics of the drone to generate an enhanced optical image with anti-motion blur characteristics; Adversarial sample defense features are extracted from the enhanced optical image, and a non-linear weighted fusion of the adversarial sample defense features is performed by constructing a perturbation intensity evaluation function to form an adversarial perturbation feature mask; The enhanced optical image and the adversarial perturbation feature mask are subjected to a spatial domain superposition operation, and a multi-scale residual network based on an attention mechanism is used to estimate the confidence of the drone target in the superimposed image; According to the target confidence estimation result, a drone recognition result with robustness against environmental interference and adversarial attacks is generated.

[0006] Optionally, based on a joint optimization model of the environmental light intensity change rate and the background motion vector field in a low-altitude security scenario, generating a set of optical compensation parameters optimized for drone imaging in real time includes: The ambient light intensity data in the low-altitude security scenario is collected in real time by a multi-channel light intensity sensor, the frequency-domain features of the light intensity change rate are extracted based on time series analysis, and a light intensity change rate function model is constructed; Perform multi-scale pyramid decomposition on the light intensity change rate function model, calculate the initial motion vectors of sparse feature points on each layer of the pyramid, and form a background motion vector field; Couple the light intensity change rate function model with the background motion vector field, establish a joint optimization equation with the light intensity smoothing constraint and the minimization of the motion vector field energy as the objective function, and use the alternating direction multiplier method to iteratively solve the joint optimization equation until it converges to a stable state; According to the output parameters of the joint optimization equation, dynamically adjust the exposure time, gain coefficient, and filtering threshold of the optical compensation module to generate an optical compensation parameter set optimized for UAV imaging.

[0007] Optionally, obtain the original optical sequence within the target capture window through the optical compensation parameter set, construct a time-domain deconvolution kernel in combination with the UAV motion characteristics, and generate an enhanced optical image with anti-motion blur characteristics, including: Based on the optical compensation parameter set, perform channel-by-channel enhancement processing on N consecutive optical signals collected within the target capture window. After each optical signal is separated into RGB three channels, it is respectively subjected to a dot product operation with the gain coefficient matrix of the corresponding channel to generate the original optical sequence; Perform time-domain feature modeling on the original optical sequence, extract the brightness gradient features and chromaticity correlation features of each pixel point in the time dimension through a multi-scale dilation convolution kernel, and output a feature tensor with spatio-temporal correlation; Construct a time-domain deconvolution kernel based on the feature tensor and the UAV motion characteristics, construct a motion consistency constraint equation by minimizing the optical flow residual between adjacent frames, and introduce the UAV motion trajectory smoothness constraint, and use a constrained optimization algorithm to solve for the pixel-level time-domain deconvolution kernel matrix; Perform spatio-temporal joint deconvolution operation on the time-domain deconvolution kernel and the original optical sequence, where inverse projection calculations are performed along the time axis for each spatial position, and the phase response characteristics of the deconvolution kernel are adjusted through an iterative update strategy to generate an enhanced optical image with anti-motion blur characteristics.

[0008] Optionally, couple the light intensity change rate function model with the background motion vector field, establish a joint optimization equation with the light intensity smoothing constraint and the minimization of the motion vector field energy as the objective function, and use the alternating direction multiplier method to iteratively solve the joint optimization equation until it converges to a stable state, including: Design the coupling term parameters, fuse the light intensity smoothing constraint of the light intensity change rate function model and the background motion vector field through a weighting factor to form a joint optimization equation; Decompose the joint optimization equation into a first sub-problem of introducing the light intensity smoothing regular term of the light intensity change rate function model to constrain the light intensity gradient between adjacent pixels and a second sub-problem of introducing the energy term of the background motion vector field to constrain the spatial continuity of the motion vector; Construct augmented Lagrangian functions for the first sub-problem and the second sub-problem respectively, and introduce a Lagrange multiplier term with linear constraints and a quadratic penalty term into the objective function; After each variable update, synchronously update the augmented Lagrange multiplier, and dynamically adjust the weight coefficient of the quadratic penalty term through an adaptive step size adjustment strategy; Based on the quadratic penalty term, use the alternating direction multiplier method to iteratively solve the joint optimization equation, calculate the relative error norm in two adjacent iterations, and determine that it converges to a stable state when the relative error norm is less than a preset threshold.

[0009] Optionally, construct a time-domain deconvolution kernel based on the feature tensor and the UAV motion characteristics, construct a motion consistency constraint equation by minimizing the optical flow residual between adjacent frames, and introduce the UAV motion trajectory smoothness constraint, and use a constrained optimization algorithm to solve to obtain a pixel-level time-domain deconvolution kernel matrix, including: Based on the multi-scale spatio-temporal correlation features of the feature tensor, extract spatio-temporal local contrast features, generate an initial parameter set of the dynamic convolution kernel, and combine the high-speed translation, hover jitter and rotor periodic motion modes in the UAV motion characteristics to perform prior correction of the motion trajectory on the initial parameter set, and map the initial parameter set to the weight space of the deconvolution kernel to form a time-domain deconvolution kernel; Perform spatio-temporal joint analysis on the time-domain deconvolution kernel and the function of minimizing the optical flow residual between adjacent frames, introduce the UAV motion trajectory smoothness constraint, and construct a motion consistency constraint equation with the pixel displacement vector as the variable, where the UAV motion trajectory smoothness constraint includes the UAV acceleration upper limit and the heading continuity condition; Iteratively perform spatial domain and time domain optimization through a constrained optimization algorithm, adjust the weight distribution of the optical flow residual function according to the UAV motion trajectory smoothness constraint in each iteration until the mean square error of the optical flow residual converges to a preset accuracy range, and output a pixel-level time-domain deconvolution kernel matrix that satisfies the motion consistency constraint.

[0010] Optionally, perform time-domain feature modeling on the original optical sequence, extract the brightness gradient features and chromaticity correlation features of each pixel point in the time dimension through a multi-scale dilated convolution kernel, and output a feature tensor with spatio-temporal correlation, including: Decompose the time-domain features of the original optical sequence into a luminance gradient component that obtains a per-pixel luminance change matrix through adjacent-frame difference operations and a chromaticity correlation component that describes the cross-frame chromaticity correlation using the covariance matrix of the chromaticity channels; Perform depthwise separable convolution operations on the luminance gradient component and the chromaticity correlation component respectively through each branch of the multi-scale dilated convolution kernel; Based on the depthwise separable convolution operations, perform cross-scale correlation modeling on the multi-branch output features, align the convolution results with different dilation rates in the time dimension, and output a feature tensor with spatio-temporal correlation.

[0011] Optionally, extract adversarial sample defense features from the enhanced optical image, and perform non-linear weighted fusion on the adversarial sample defense features through constructing a perturbation intensity evaluation function to form an adversarial perturbation feature mask, including: Perform multi-scale spatial filtering on the enhanced optical image, combine an adaptive noise suppression algorithm and a local contrast enhancement operation to generate a denoised and normalized optical image; Perform feature extraction based on the denoised and normalized optical image, screen out high-frequency texture features and low-frequency structure features of the multi-rotor structure and metal reflection characteristics of the drone through a cross-channel attention mechanism, and construct an adversarial sample defense feature matrix, where the attention mechanism focuses on the key parts of the drone rotors and fuselage; Dynamically calculate the weight coefficients of each channel of the adversarial sample defense features using the perturbation intensity evaluation function, and adjust the weight distribution using a layer-by-layer reverse gradient accumulation strategy to achieve non-linear weighted fusion in the feature space; Output an adversarial perturbation feature mask by iteratively optimizing the semantic consistency between the boundary of the mask and the original image for the non-linearly weighted fused adversarial sample defense feature matrix.

[0012] Optionally, perform a spatial domain superposition operation on the enhanced optical image and the adversarial perturbation feature mask, and use a multi-scale residual network based on the attention mechanism to estimate the confidence of the drone target for the superimposed image, including: Normalize the adversarial perturbation feature mask so that the numerical range of the adversarial perturbation feature mask matches the pixel distribution of the enhanced optical image, and generate a superimposed image through a spatial domain superposition operation; Construct a multi-scale processing module to perform multi-scale residual network operations based on the attention mechanism on the superimposed image to generate a multi-scale fusion feature tensor; Input the multi-scale fusion feature tensor into a confidence estimation unit, perform local feature statistic calculation, and generate a drone target confidence estimation value.

[0013] Optionally, based on the estimated result of the UAV target confidence, generate a UAV recognition result with robustness against environmental interference and adversarial attacks, including: Based on the estimated result of the UAV target confidence, use a sliding window mechanism to calculate the environmental interference intensity index in real time. When it is detected that the confidence fluctuation of consecutive frames exceeds a preset threshold, start the adaptive threshold adjustment algorithm; Through the adaptive threshold adjustment algorithm, identify the interference pattern of the confidence anomaly area and generate an environmental interference feature mask; Start the adversarial sample detection module for low-confidence targets, analyze the sensitivity of feature space perturbation using the gradient backpropagation method, and generate a robustness evaluation coefficient; Combine the estimated result of the UAV target confidence, the environmental interference feature mask, and the robustness evaluation coefficient to generate a UAV recognition result with robustness against environmental interference and adversarial attacks.

[0014] In a second aspect, an embodiment of the present application provides a UAV recognition system for low-altitude security, including: A generation module, configured to generate a set of optical compensation parameters optimized for UAV imaging in real time based on a joint optimization model of the environmental light intensity change rate and the background motion vector field in a low-altitude security scenario; the generation module is further configured to obtain an original optical sequence within a target capture window through the set of optical compensation parameters, and construct a time-domain deconvolution kernel in combination with the UAV motion characteristics to generate an enhanced optical image with anti-motion blur characteristics; A processing module, configured to extract adversarial sample defense features from the enhanced optical image, and perform non-linear weighted fusion on the adversarial sample defense features by constructing a perturbation intensity evaluation function to form an adversarial perturbation feature mask; An estimation module, configured to perform a spatial domain superposition operation on the enhanced optical image and the adversarial perturbation feature mask, and perform target confidence estimation on the superimposed image using a multi-scale residual network based on an attention mechanism; The generation module is further configured to generate a UAV recognition result with robustness against environmental interference and adversarial attacks according to the target confidence estimation result.

[0015] In the embodiments of the present application, based on the joint optimization model of the environmental light intensity change rate and the background motion vector field in the low-altitude security scenario, an optical compensation parameter set optimized for UAV imaging is generated in real time; the original optical sequence is obtained within the target capture window through the optical compensation parameter set, and a time-domain deconvolution kernel is constructed in combination with the motion characteristics of the UAV to generate an enhanced optical image with anti-motion blur characteristics; anti-sample defense features are extracted from the enhanced optical image, and a perturbation intensity evaluation function is constructed to perform non-linear weighted fusion on the anti-sample defense features to form an anti-perturbation feature mask; the enhanced optical image and the anti-perturbation feature mask are subjected to a spatial domain superposition operation, and a multi-scale residual network based on the attention mechanism is used to estimate the target confidence of the superimposed image; according to the target confidence estimation result, a UAV recognition result with anti-environmental interference and anti-adversarial attack robustness is generated.

[0016] The technical solution of the present application has the following beneficial effects: By fusing the light intensity change rate and the background motion vector field, the present application realizes the adaptive dynamic adjustment of the optical compensation parameters, significantly improving the imaging stability and target capture quality in complex environments. The image blur caused by the high-speed movement of the target is eliminated through the time-domain deconvolution kernel, the detail information is restored, and the clarity and recognizability of the target area are improved. The anti-attack area is accurately detected and marked, enhancing the robustness of image processing and target recognition, and reducing the false detection rate caused by anti-attack. Through multi-scale feature extraction and the attention mechanism, the target detection accuracy and stability are improved, ensuring a high-confidence recognition result in complex scenarios. Combining the confidence-driven screening and multi-frame fusion strategies, a highly reliable UAV recognition result is generated, significantly enhancing the environmental adaptability and anti-interference ability of the low-altitude security system.

[0017] Further, the present application uses a multi-channel light intensity sensor to collect the environmental light intensity data in the low-altitude security scenario in real time, extracts the frequency domain features of the light intensity change rate based on time series analysis, and constructs a light intensity change rate function model; performs multi-scale pyramid decomposition on the light intensity change rate function model, calculates the initial motion vectors of sparse feature points on each layer of the pyramid to form a background motion vector field; couples the light intensity change rate function model and the background motion vector field, establishes a joint optimization equation with the light intensity smoothing constraint and the minimum energy of the motion vector field as the objective function, and uses the alternating direction multiplier method to iteratively solve the joint optimization equation until it converges to a stable state; according to the output parameters of the joint optimization equation, dynamically adjusts the exposure time, gain coefficient, and filtering threshold of the optical compensation module to generate an optical compensation parameter set optimized for UAV imaging.

[0018] These aspects or other aspects of the present application will be more clearly understood in the following description of the embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0020] Figure 1 The flowchart of a drone recognition method for low-altitude security provided by the present application is shown; Figure 2 The structural schematic diagram of a drone recognition system for low-altitude security provided by the present application is shown. Detailed implementation manners

[0021] To enable those skilled in the art of the present technology to better understand the solutions of the present application, the following will clearly and completely describe the technical solutions in the embodiments of the present application in conjunction with the drawings in the embodiments of the present application.

[0022] In some processes described in the specification, claims and the above drawings of the present application, a plurality of operations appear in a specific order. However, it should be clearly understood that these operations can be executed not in the order in which they appear in this article or in parallel. The serial numbers of the operations, such as 101, 102, etc., are only used to distinguish different operations, and the serial numbers themselves do not represent any execution order. In addition, these processes can include more or fewer operations, and these operations can be executed in sequence or in parallel. It should be noted that the descriptions such as "first", "second", etc. in this article are used to distinguish different messages, devices, modules, etc., do not represent a sequence, and do not limit that "first" and "second" are of different types.

[0023] The following will clearly and completely describe the technical solutions in the embodiments of the present application in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present application.

[0024] Figure 1 The flowchart of a drone recognition method for low-altitude security provided for the embodiments of the present application is as Figure 1 shown, and the method includes: 101. Based on the joint optimization model of the ambient light intensity change rate and the background motion vector field in the low-altitude security scenario, a set of optical compensation parameters optimized for drone imaging is generated in real time; In this step, the ambient light intensity change rate refers to the rate at which the light intensity changes over time in the low-altitude security scenario, which is usually calculated from the time-series data collected by the light intensity sensor.

[0025] The background motion vector field is a vector field that describes the motion direction and speed of background objects in the low-altitude security scenario. It is generated by optical flow estimation or feature matching methods and is used to characterize the dynamic interference of the scene background (such as tree shaking and cloud drifting).

[0026] The set of optical compensation parameters includes parameters such as exposure time and gain coefficient, which are used to adjust the optical system to cope with the change of ambient light intensity and background motion interference, to suppress ambient light interference and compensate for the imaging blur caused by the high-speed movement of the drone.

[0027] The light intensity data in the scene is collected in real time by the light intensity sensor, and the ambient light intensity change rate is obtained by calculating its time difference sequence. Optical flow estimation is performed on the video frame sequence, the background motion vector field is extracted, and abnormal vectors are removed by motion consistency clustering. A joint optimization model is constructed to couple the light intensity change rate with the background motion vector field. The objective function includes a light intensity smoothing constraint and a motion vector field energy minimization term. The joint optimization model is iteratively solved using the alternating direction multiplier method, and the set of optical compensation parameters generated in real time is output.

[0028] In the low-altitude security scenario, the drone monitoring system collects ambient light intensity data through the light intensity sensor and calculates its change rate; at the same time, optical flow estimation is performed on the video frame sequence to generate the background motion vector field. The two are input into the joint optimization model to generate a set of optical compensation parameters optimized for drone imaging in real time, which are used to adjust the exposure time and gain coefficient of the camera to cope with light changes and background motion interference.

[0029] 102. Obtain the original optical sequence within the target capture window through the set of optical compensation parameters, construct a time-domain deconvolution kernel in combination with the motion characteristics of the drone, and generate an enhanced optical image with anti-motion blur characteristics; The original optical sequence refers to the continuous frame image sequence collected within the target capture window.

[0030] The time-domain deconvolution kernel is a time-domain filter used to eliminate motion blur in the image sequence, and a clear image is restored through deconvolution operation.

[0031] Adjust the camera parameters using an optical compensation parameter set and collect the original optical sequence within the target capture window. Construct a time-domain deconvolution kernel in combination with the motion characteristics of the UAV (including high-speed translation, hovering jitter, and the periodic motion pattern of the rotor blades), estimate the motion blur trajectory based on the optical flow field, and introduce constraints on the smoothness of the UAV motion trajectory (such as the acceleration upper limit and the heading continuity condition). Optimize the kernel function parameters through regularization constraints. Perform time-domain deconvolution on the original optical sequence to generate enhanced optical images with anti-motion blur characteristics.

[0032] The UAV flies obliquely at a speed of 15 m / s, and the target capture window tracks its motion and captures 5 blurred images. Optical flow analysis shows that the target forms a zigzag trajectory in the time domain (maximum displacement of 8 pixels / frame). Combining the motion characteristics of the UAV, the deconvolution kernel models the PSF (point spread function) according to the zigzag trajectory and introduces constraints on the smoothness of the UAV motion trajectory to optimize the kernel function parameters. After regularized deconvolution, the clarity of the target rotor blade texture is significantly improved (the edge gradient increases from 0.3 to 0.7), and the length of the background tree ghosting is reduced by 70%.

[0033] 103. Extract anti-adversarial sample defense features based on the multi-rotor structure and metal reflection characteristics of the UAV from the enhanced optical image, and form an adversarial perturbation feature mask by constructing a perturbation intensity evaluation function to perform non-linear weighted fusion on the anti-adversarial sample defense features; Anti-adversarial sample defense features refer to the features extracted from the enhanced optical image that can characterize adversarial attacks, such as abnormal gradient distributions or local texture mutations, and are used to detect potential adversarial attack regions.

[0034] The perturbation intensity evaluation function is used to quantify the perturbation intensity of the anti-adversarial sample defense features, and generates an adversarial perturbation feature mask through non-linear weighted fusion to identify high-risk regions.

[0035] Extract anti-adversarial sample defense features from the enhanced optical image, including periodic texture features based on the multi-rotor structure of the UAV, highlight region features caused by metal reflection characteristics, as well as gradient magnitude, local contrast, and texture complexity, which characterize the abnormal regions of the image. Construct a perturbation intensity evaluation function, calculate the weight coefficients of each feature based on the physical characteristics of the UAV (such as rotor symmetry, metal reflection intensity) and feature sensitivity, quantify the potential risk of adversarial attacks, and generate an adversarial perturbation feature mask through non-linear weighted fusion to identify potential adversarial attack regions in the image, providing a risk prompt for subsequent processing.

[0036] For example, after generating the enhanced optical image, the system extracts its periodic texture features based on the multi-rotor structure of the drone and the highlight area features caused by the metal reflection characteristics, and combines the gradient magnitude and local contrast features. Through the perturbation intensity evaluation function, the weight coefficient is calculated to generate an adversarial perturbation feature mask, identifying the possible adversarial attack areas in the image. For example, when it is detected that the periodicity of the rotor texture in a certain area of the image is abnormal, the metal reflection intensity changes suddenly, and the gradient magnitude is abnormally high, the system determines that this area may be under an adversarial attack and marks it as a high-risk area in the mask.

[0037] 104. Perform a spatial domain superposition operation on the enhanced optical image and the adversarial perturbation feature mask, and use a multi-scale residual network based on the attention mechanism to estimate the confidence of the drone target for the superimposed image. The spatial domain superposition operation refers to pixel-level fusion of the enhanced optical image and the adversarial perturbation feature mask to generate a composite image containing adversarial attack information, providing more comprehensive input data for subsequent recognition.

[0038] The multi-scale residual network is a deep learning model based on the attention mechanism, used to extract multi-scale features of the image and estimate the target confidence, improving the reliability of the recognition result.

[0039] Perform a spatial domain superposition operation on the enhanced optical image and the adversarial perturbation feature mask to generate a composite image, fuse the target details and adversarial attack information, use a multi-scale residual network to extract features from the composite image, use the attention mechanism to focus on the key areas, enhance the saliency of the target features, and output the drone target confidence estimation result through the fully connected layer, evaluate the reliability of the recognition result, and provide a basis for subsequent screening.

[0040] For example, the system performs a spatial domain superposition on the enhanced optical image and the adversarial perturbation feature mask, inputs it into the multi-scale residual network for feature extraction, outputs the target confidence estimation result, evaluates the reliability of the drone recognition. When a certain area in the composite image is marked as a high-risk adversarial attack area, the system reduces the confidence score of this area to avoid misrecognition.

[0041] 105. Generate a drone recognition result with robustness against environmental interference and adversarial attacks according to the target confidence estimation result.

[0042] The target confidence estimation result refers to the reliability score of the recognition result output by the multi-scale residual network, used to screen high-confidence recognition results. The recognition result includes information such as the drone model, location, and threat level.

[0043] The robustness against environmental interference and adversarial attacks refers to the ability of the recognition system to maintain high accuracy under complex environments and adversarial attacks, ensuring the reliability of the recognition result.

[0044] According to the target confidence estimation results, high-confidence recognition results are screened, low-confidence misidentifications are eliminated, and recognition accuracy is improved. By fusing multi-frame recognition results, the system's robustness to the environment and adversarial attacks is enhanced, the impact of single-frame misidentification is eliminated, and the final drone recognition results are output to ensure its accuracy and reliability in complex environments.

[0045] The system screens high-confidence recognition results based on the target confidence estimation results, enhances robustness through multi-frame fusion, and finally outputs accurate drone recognition results. In complex lighting and adversarial attack environments, the system eliminates single-frame misidentification through multi-frame fusion to ensure the reliability of the final recognition results.

[0046] Steps 101 to 105 of the present invention jointly optimize the ambient light intensity change rate and the background motion vector field to generate a set of optical compensation parameters optimized for drone imaging in real time, effectively dealing with dynamic lighting and background motion interference; eliminating motion blur through a time domain deconvolution kernel to generate a clear enhanced optical image; extracting adversarial sample defense features and generating adversarial perturbation feature masks to identify potential adversarial attack areas; using a multi-scale residual network to estimate target confidence and screen high-confidence recognition results; generating drone recognition results that are robust against environmental interference and adversarial attacks, significantly improving recognition accuracy and reliability in low-altitude security scenarios.

[0047] In order to solve the limitations of the single parameter compensation method in dynamic and complex scenes, the problem of image quality degradation caused by the drastic changes in dynamic illumination and complex background motion interference in low-altitude security scenes is solved. By fusing the ambient light intensity change rate data collected in real time by the light intensity sensor and the background motion vector field data extracted from the video frame sequence, a joint optimization model is constructed, the objective function is designed by coupling constraints of multi-source data, and the dynamic compensation parameters are solved in real time by an alternating optimization algorithm to achieve dynamic adaptive adjustment of optical system parameters. In some embodiments, based on the joint optimization model of the ambient light intensity change rate and the background motion vector field in the low-altitude security scene, a set of optical compensation parameters optimized for drone imaging is generated in real time, including: 201. Use a multi-channel light intensity sensor to collect ambient light intensity data in low-altitude security scenes in real time, extract the frequency domain characteristics of the light intensity change rate based on time series analysis, and construct a light intensity change rate function model; In step 201, the multi-channel light intensity sensor refers to a sensor with an integrated multi-spectral photosensing unit, which can simultaneously collect multi-band light intensity data such as visible light and infrared light. The light intensity change rate is the light intensity fluctuation rate calculated based on time series data, reflecting the dynamic change characteristics of ambient light intensity. The light intensity change rate function model refers to a mathematical model constructed through frequency domain characteristics (such as main frequency, amplitude), which is used to quantify the regularity of light changes.

[0048] In the embodiments of the present application, ambient light intensity data is collected by a multi-channel light intensity sensor at a frequency of 100 Hz, and the time-series data is segmented by using a sliding window (window length: 3 seconds, overlap rate: 50%). The fast Fourier transform (FFT) is performed on the light intensity data within each window to extract the main frequency component (0.1 Hz to 10 Hz) and its amplitude, and a light intensity change rate function model is constructed.

[0049] 202. Perform multi-scale pyramid decomposition on the light intensity change rate function model, calculate the initial motion vectors of sparse feature points on each layer of the pyramid, and form a background motion vector field. In step 202, the multi-scale pyramid decomposition refers to decomposing the light intensity change rate function model into components of different spatial scales (such as high-frequency details and low-frequency trends). The sparse feature points are key points extracted by a corner detection algorithm and are used to characterize the stable structural features in the scene. The background motion vector field is a set of background object motion vectors calculated by an optical flow method and reflects the motion trend of non-target objects.

[0050] In the embodiments of the present application, a three-layer Gaussian pyramid decomposition (scale factor: 2) is performed on the light intensity change rate function model, and the Harris corner detection algorithm (threshold set to 0.01) is used to extract sparse feature points on each layer of the pyramid. Subsequently, the motion vectors of the feature points are calculated based on the Lucas-Kanade optical flow method, and the global motion model (such as affine transformation) is fitted by the RANSAC algorithm to eliminate abnormal vectors (such as birds and fallen leaves) with residuals exceeding 2 pixels, and a smooth background motion vector field is generated.

[0051] 203. Couple the light intensity change rate function model with the background motion vector field, establish a joint optimization equation with the light intensity smoothing constraint and the minimization of the energy of the motion vector field as the objective function, and use the alternating direction multiplier method to iteratively solve the joint optimization equation until it converges to a stable state. In step 203, the light intensity smoothing constraint refers to restricting the sudden change of the light intensity change rate to avoid overexposure or underexposure of the image caused by severe fluctuations in illumination. The minimization of the energy of the motion vector field is achieved by constraining the spatial continuity of the motion vectors to suppress abnormal jumps in the background motion. The joint optimization equation refers to a multi-objective optimization model that couples the light intensity and motion constraints and is used to generate stable compensation parameters.

[0052] In the embodiments of the present application, a joint optimization equation is constructed by combining the light intensity smoothing constraint (constraining the gradient of the light intensity change rate through the first-order difference) with the minimization of the motion vector field energy (ensuring motion continuity through total variation regularization). The alternating direction method of multipliers (ADMM) is used to decompose the optimization problem into a light intensity compensation sub-problem and a motion compensation sub-problem, and the solutions are obtained by alternating iterations. In each iteration, the parameters of the light intensity sub-problem are updated by the gradient descent method, and the vector field of the motion sub-problem is optimized by the conjugate gradient method until the residual norm is less than a preset threshold (such as 0.01).

[0053] 204. According to the output parameters of the joint optimization equation, dynamically adjust the exposure time, gain coefficient, and filtering threshold of the optical compensation module to generate an optical compensation parameter set optimized for UAV imaging.

[0054] In step 204, the exposure time is a parameter for adjusting the light-sensitive duration of the camera, which is used to balance the image brightness and dynamic blur. The gain coefficient is a parameter for controlling the magnification of the image signal, which is used to suppress noise in low-light environments. The filtering threshold is a parameter for dynamically adjusting the image processing algorithm (such as the median filtering window size), which is used to suppress noise and artifacts. The optical compensation parameter set is used to suppress ambient light interference and compensate for the imaging blur caused by the high-speed movement of the UAV.

[0055] In the embodiments of the present application, according to the output parameters of the joint optimization equation, the optical compensation module is dynamically adjusted. The exposure time is adjusted such that if the light intensity change rate > 100 lux / s, the exposure time is proportionally extended (such as adjusted from 1 / 1000 s to 1 / 500 s). The gain coefficient is adjusted to dynamically set the gain according to the average ambient light intensity (such as when the light intensity < 200 lux, the gain is increased to 12 dB). The filtering threshold is adjusted to adjust the median filtering window size based on the average speed of the motion vector field (such as when the speed > 5 pixels / frame, a 5×5 window is used). The generated optical compensation parameter set can adapt to scene changes in real time.

[0056] The following is a specific example: In the low-altitude security scenario at dusk, the light intensity drops from 1000 lux to 200 lux within 5 minutes, and the background cloud layer moves at a speed of 2 m / s. In step 201, the multi-channel light intensity sensor detects that the light intensity drops from 1000 lux to 200 lux, with a change rate of 160 lux / min. The system performs time series analysis on the light intensity data, extracts the main frequency component (0.5 Hz) through Fourier transform, constructs a light intensity change rate function model, and quantifies the dynamic fluctuations of light. In step 202, the light intensity change rate function model is decomposed by three-layer pyramid decomposition, and sparse feature points are extracted respectively (Harris corner detection, with the threshold set to 0.01). The Lucas-Kanade optical flow method is used to calculate the initial motion vector, and outliers such as birds or fallen leaves are removed through the RANSAC algorithm to generate a smooth background motion vector field, characterizing the trend of the cloud layer moving at a speed of 2 m / s. In step 203, a joint optimization equation is constructed, combining the light intensity smoothing constraint (suppressing the noise of sudden light changes) and the minimization of the energy of the motion vector field (ensuring motion continuity), and the ADMM algorithm is used for iterative solution. Each round of iteration takes less than 10 ms until the residual norm is less than the preset threshold (0.01), and the compensation parameters are output: the exposure time is 1 / 300 s, and the gain is 10 dB. In step 204, the optical system is dynamically adjusted according to the output parameters, the exposure time is adjusted from 1 / 1000 s to 1 / 300 s, the gain coefficient is increased to 10 dB, and at the same time, 3×3 median filtering is applied to suppress noise. The standard deviation of the image brightness after compensation drops from 35 to 12, the background blur is reduced by 50%, and the clarity of the target area is significantly improved.

[0057] In summary, through the joint optimization of the multi-channel light intensity sensor and the background motion vector field, this solution can generate a set of optical compensation parameters optimized for UAV imaging in real time, effectively solving the problem of image quality degradation caused by dynamic light changes and complex background motion interference in low-altitude security scenarios, significantly improving the imaging clarity and recognition stability of UAV targets, and providing a high-quality input data basis for subsequent anti-motion blur processing and anti-attack defense.

[0058] In order to improve the image blur problem caused by the high-speed movement of the target in the low-altitude security scenario, in this step, the imaging parameters (such as exposure time, gain) within the target capture window are dynamically adjusted through the set of optical compensation parameters to obtain the original optical sequence; based on the modeling of the target motion trajectory and combined with the motion characteristics of the UAV, a time-domain deconvolution kernel is constructed, and deconvolution operation is used to eliminate motion blur and restore details, aiming to solve the defect that traditional fixed-parameter imaging cannot adaptively compensate for motion blur in dynamic scenarios, and the clear image is inversely solved through the time-domain degradation model, providing high-definition and anti-blur optical input data for subsequent target detection and recognition.

[0059] In some embodiments, an original optical sequence is obtained within a target capture window through the set of optical compensation parameters, and a time-domain deconvolution kernel is constructed in combination with the motion characteristics of the drone to generate an enhanced optical image with anti-motion blur characteristics, including: 301. Based on the set of optical compensation parameters, perform channel-by-channel enhancement processing on N frames of optical signals continuously collected within the target capture window. After each frame of optical signal is separated into RGB three channels, perform a dot product operation with the gain coefficient matrix of the corresponding channel respectively to generate an original optical sequence; In step 301, the gain coefficient matrix is a parameter matrix designed for the RGB three channels respectively, and is used to dynamically adjust the brightness enhancement ratio of each channel. Channel-by-channel enhancement processing is to decompose each frame of optical signal into three independent channels of red, green, and blue, and perform gain adjustment respectively to optimize color balance. The original optical sequence is a set of consecutive frame images generated after channel enhancement processing, retaining time-domain correlation and spatial details.

[0060] In the embodiments of the present application, based on the gain coefficients in the set of optical compensation parameters (such as R channel gain 1.2, G channel 1.0, B channel 0.8), perform channel-by-channel processing on 5 consecutive frames of optical signals within the target capture window. The specific process is to decompose each frame of image into grayscale images of three independent channels of RGB, perform a dot product operation on the pixel values of each channel with the corresponding gain coefficient matrix respectively (such as R channel pixel value × 1.2), and recombine the enhanced three-channel images to generate an original optical sequence. For example, in a low-light scene, compensate for the attenuation of red light by increasing the R channel gain to avoid the overall image being bluish.

[0061] 302. Perform time-domain feature modeling on the original optical sequence, extract the brightness gradient feature and chromaticity correlation feature of each pixel point in the time dimension through a multi-scale dilated convolution kernel, and output a feature tensor with spatio-temporal correlation; In step 302, time-domain feature modeling is to analyze the brightness and chromaticity change rules of pixels in consecutive frames, and capture motion trajectories and light fluctuation features. The multi-scale dilated convolution kernel is a convolution kernel with different time spans (such as 3 frames, 5 frames, 7 frames), and is used to extract short-term mutation and long-term trend features. The spatio-temporal correlation feature tensor refers to a tensor structure containing pixel brightness gradient, chromaticity correlation, and temporal motion information.

[0062] In the embodiments of the present application, perform time-domain feature modeling on the original optical sequence, use convolution kernels with 3 dilation rates (dilation rate 1 corresponds to a 3-frame span, dilation rate 2 corresponds to a 5-frame span), slide along the time axis to calculate the brightness gradient (difference between adjacent frames) and chromaticity correlation (RGB channel covariance) of each pixel, splice feature maps of different scales into a three-dimensional tensor, and output a feature tensor with spatio-temporal correlation.

[0063] 303. Construct a time-domain deconvolution kernel based on the feature tensor and the motion characteristics of the UAV. Construct a motion consistency constraint equation by minimizing the optical flow residual between adjacent frames, introduce the smoothness constraint of the UAV motion trajectory, and use a constrained optimization algorithm to solve for the pixel-level time-domain deconvolution kernel matrix. In step 303, the time-domain deconvolution kernel refers to a filter that dynamically adjusts parameters according to spatio-temporal features and is used to eliminate motion blur. The motion consistency constraint equation ensures the continuity of motion vectors between adjacent frames by minimizing the optical flow residual. The pixel-level time-domain deconvolution kernel matrix is a deconvolution kernel designed independently for each pixel to adapt to local motion differences.

[0064] Among them, the motion characteristics of the UAV include the prior knowledge of the high-speed translation, hovering jitter, and periodic motion of the rotor of the UAV, which are used to guide the calculation of the optical flow residual and the optimization of the deconvolution kernel parameters.

[0065] In the embodiment of the present application, perform dense optical flow estimation on adjacent frames, calculate the optical flow field residual (the deviation between the predicted motion and the actual motion), and correct the smoothness of the motion trajectory of the optical flow residual in combination with the motion characteristics of the UAV (such as high-speed translation, hovering jitter, and periodic motion of the rotor); use the minimization of the optical flow residual as a constraint condition, and design an objective function in combination with the spatio-temporal feature tensor and the motion characteristics of the UAV, where the motion characteristics of the UAV are used to constrain the weight distribution of the optical flow residual; use a constrained optimization algorithm (such as the projection gradient method) to iteratively solve the deconvolution kernel matrix, and adjust the deconvolution kernel parameters according to the smoothness constraint of the UAV motion trajectory in each iteration. Each pixel is optimized independently to ensure that the deconvolution kernel adapts to local motion differences.

[0066] 304. Perform a spatio-temporal joint deconvolution operation on the time-domain deconvolution kernel and the original optical sequence, where inverse projection calculations are performed along the time axis for each spatial position, and the phase response characteristics of the deconvolution kernel are adjusted through an iterative update strategy to generate an enhanced optical image with anti-motion blur characteristics.

[0067] In step 304, the spatio-temporal joint deconvolution operation performs deconvolution synchronously in the spatial and time dimensions to restore the target details. The inverse projection calculation is to trace back the pixel motion trajectory along the time axis and correct the blurred pixel values. The adjustment of the phase response characteristics refers to dynamically optimizing the phase parameters of the deconvolution kernel to match the target motion frequency.

[0068] In the embodiments of the present application, spatio-temporal joint deconvolution operation is performed on the original optical sequence using a time-domain deconvolution kernel to generate a motion-blur-resistant enhanced optical image. First, back-projection calculations are performed separately for each spatial position along the time axis. By analyzing the time-domain information in the optical sequence, the characteristic distribution of motion blur is extracted. Secondly, an iterative update strategy is adopted to adjust the phase response characteristics of the deconvolution kernel. The parameters of the deconvolution kernel are optimized by minimizing the reconstruction error function to ensure that it can effectively compensate for motion blur. In each iteration, the phase response of the deconvolution kernel is updated, and joint optimization is performed in combination with spatial-domain information to gradually improve the image reconstruction accuracy. Finally, an enhanced optical image is generated through spatio-temporal joint deconvolution operation, significantly reducing the influence of motion blur, improving the clarity and detail restoration ability of the image, and providing high-quality data support for subsequent analysis and processing.

[0069] To achieve the adaptive dynamic adjustment of optical compensation parameters, aiming at the problem of imaging parameter mismatch caused by the coupling of sudden light changes and background motion interference in dynamic scenes, by fusing the light intensity change rate function model (quantifying the law of light intensity fluctuations) and the background motion vector field (characterizing the motion trend of non-target objects), a joint optimization equation is constructed. The light intensity smoothing constraint is introduced to suppress sudden light change noise, and the motion continuity is ensured by minimizing the energy of the motion vector field. The alternating direction method of multipliers (ADMM) is used to decompose and optimize the variables and solve them by alternating iteration. The dynamic balance of parameters is achieved through residual convergence determination, solving the defects of parameter conflict or slow convergence of traditional single-dimensional compensation methods in complex scenes. In the embodiments of the present application, the steps of performing spatio-temporal joint deconvolution operation are as follows: taking the current frame as the center, tracing the target motion trajectory forward and backward along the time axis to generate an initial back-projection path. According to the phase response characteristics of the deconvolution kernel matrix (such as high-frequency kernels preferentially correcting rotor blur), the pixel values in different regions are updated in stages. When the change rate of the mean square error of adjacent iteration results < 1%, the calculation is terminated, and an enhanced optical image is output.

[0070] In some embodiments, the light intensity change rate function model and the background motion vector field are coupled to establish a joint optimization equation with the light intensity smoothing constraint and the minimum energy of the motion vector field as the objective function, and the alternating direction method of multipliers is used to iteratively solve the joint optimization equation until it converges to a stable state, including: 401. Design the coupling term parameters, fuse the light intensity smoothing constraint of the light intensity change rate function model and the background motion vector field through a weighting factor to form a joint optimization equation; In step 401, the coupling term parameter is used to associate the weight factor of the light intensity change rate with the background motion vector, balancing the influence of the two types of constraints on the optimization result. The light intensity smoothing constraint restricts the gradient difference of the light intensity change rate between adjacent pixels, avoiding noise amplification caused by sudden changes in illumination. The weighted factor fusion integrates the two types of constraints into a unified objective function by presetting weight coefficients (such as 0.6 for light intensity constraint and 0.4 for motion constraint).

[0071] In the embodiment of the present application, according to the physical correlation between the light intensity sensor data and the background motion vector field, the coupling term parameter is designed. The light intensity change rate data and the motion vector field data are normalized to the same dimension (such as the 0~1 interval), and the weight factor is dynamically adjusted based on the scene characteristics (such as when the dynamic illumination is intense, the light intensity constraint weight is increased to 0.7). The light intensity smoothing term and the motion energy term are fused through the weighted factor to form a joint optimization equation, ensuring the synergistic effect of the two types of constraints. 402. Decompose the joint optimization equation into a first sub-problem that introduces the light intensity smoothing regular term of the light intensity change rate function model to constrain the light intensity gradient between adjacent pixels, and a second sub-problem that introduces the energy term of the background motion vector field to constrain the spatial continuity of the motion vector; In step 402, the first sub-problem is an optimization problem centered on the light intensity smoothing regular term, which is used to constrain the consistency of the light intensity gradient between adjacent pixels. The second sub-problem is an optimization problem centered on the motion vector field energy term, which is used to ensure the spatial continuity of the motion vector.

[0072] In the embodiment of the present application, the joint optimization equation is decomposed into two independently solvable sub-problems. Fix the motion vector field parameters, and use the gradient descent method to optimize the light intensity parameters. Update the light intensity distribution by calculating the gray level gradient difference between adjacent pixels (such as a 3×3 neighborhood). Fix the light intensity parameters, and use the conjugate gradient method to optimize the motion vector field. Update the motion field by constraining the direction consistency of adjacent vectors (such as cosine similarity > 0.9). Alternately iterate the two sub-problems, and transfer the optimization results through intermediate variables to achieve global convergence.

[0073] 403. Construct augmented Lagrangian functions for the first sub-problem and the second sub-problem respectively, and introduce Lagrangian multiplier terms and quadratic penalty terms with linear constraints in the objective function; In step 403, the augmented Lagrangian function introduces Lagrangian multipliers and quadratic penalty terms in the objective function, transforming the constrained optimization problem into an unconstrained form. The linear constraint term is used to enforce the physical correlation between the light intensity parameters and the motion parameters, such as the light intensity change needs to match the background motion trend.

[0074] In the embodiments of the present application, an augmented Lagrangian function is constructed for each sub-problem. A Lagrange multiplier term is added to the light intensity smoothing term to enforce the correlation between the light intensity parameters and the motion field. A quadratic penalty term is introduced into the motion energy term to balance the relaxation degree of the constraint conditions. The initial value of the Lagrange multiplier is set to 0, and the weight coefficient of the quadratic penalty term is initialized to 1.0.

[0075] 404. After each variable update, synchronously update the augmented Lagrange multiplier, and dynamically adjust the weight coefficient of the quadratic penalty term through an adaptive step size adjustment strategy; In step 404, the adaptive step size adjustment strategy dynamically adjusts the multiplier update step size according to the iterative residual, accelerates convergence and avoids oscillation. The weight coefficient of the quadratic penalty term is a parameter that controls the strictness of the constraint conditions. The larger the weight, the stronger the constraint.

[0076] In the embodiments of the present application, the parameters are synchronously updated after each iteration. According to the residual between the current solution and the constraint conditions, the Lagrange multiplier is updated proportionally (for example, when the residual increases by 10%, the step size increases by 20%). If the change rate of the iterative residual between two adjacent iterations exceeds the threshold (such as 5%), the weight coefficient of the quadratic penalty term is adaptively increased (for example, adjusted from 1.0 to 1.2). A weight upper limit (such as 2.0) is set to prevent over-punishment from causing model rigidity.

[0077] 405. Based on the quadratic penalty term, use the alternating direction method of multipliers to iteratively solve the joint optimization equation, calculate the relative error norm between two adjacent iterations, and determine that it converges to a stable state when the relative error norm is less than a preset threshold.

[0078] In step 405, the alternating direction method of multipliers is a decomposition-coordination optimization framework that achieves global optimality by alternately solving sub-problems. The relative error norm is a measure of the difference between the results of two adjacent iterations and is used to determine convergence.

[0079] In the embodiments of the present application, determine the main optimization objective (such as minimizing the bandwidth allocation cost) and the constraint conditions (such as the energy consumption not exceeding the threshold). Add a quadratic penalty term to the objective function to balance the accuracy and stability of the solution. Set the initial values of the main variables, auxiliary variables, and Lagrange multipliers, usually taking the zero vector or reasonable values based on prior knowledge. Set the penalty coefficient (used to control the intensity of the penalty term) and the convergence threshold (used to determine whether the algorithm converges). Fix the auxiliary variables and the Lagrange multiplier, and solve the sub-problem of the main variables. Fix the main variables and the Lagrange multiplier, and solve the sub-problem of the auxiliary variables. Update the Lagrange multiplier according to the residual between the main variables and the auxiliary variables to adjust the intensity of the penalty term. Calculate the difference between the main variables of the solutions of two adjacent iterations and divide it by the norm of the main variables of the current iteration solution to obtain the relative error norm. If the relative error norm is less than the preset threshold (such as one in a million), it is determined that the algorithm converges to a stable state.

[0080] When the UAV flies in sandy dust weather, the light intensity fluctuates by 300 lux per minute due to the change in the dust concentration, and the background dust particles move in irregular trajectories. The change rate of the light intensity detects that the main frequency of the fluctuation is 0.6 Hz (the dust concentration changes periodically), the motion vector field shows that the average speed of the dust particles is 3 m / s, and the weight factor is dynamically allocated (light intensity constraint 0.65, motion constraint 0.35), and a joint optimization equation is constructed. After decomposing the sub-problems, the first sub-problem optimizes the light intensity parameters by the gradient descent method, reducing the gradient difference between adjacent pixels by 40%; the second sub-problem optimizes the motion field by the conjugate gradient method, and the vector direction consistency is improved to 0.85. An augmented Lagrangian function is constructed, the initial penalty weight is set to 1.0, and the multiplier is initialized to 0. The penalty weight is dynamically adjusted to 1.3 according to the residual change rate, and the multiplier step size is increased by 15%. After 12 rounds of ADMM iteration, the relative error norm is reduced to 0.008, and it is determined to converge. Finally, the light intensity parameters (exposure time 1 / 400 s) and motion compensation parameters (filtering threshold 5×5) are output, the image noise caused by dust interference is reduced by 60%, and the clarity of the target contour is improved by 55%.

[0081] In summary, through the design of coupling term parameters, sub-problem decomposition, construction of augmented Lagrangian function and adaptive ADMM optimization, this solution solves the convergence speed and stability problems of the coupled optimization of light intensity and motion parameters in dynamic scenes.

[0082] In order to eliminate the image blur caused by the high-speed movement of the target, aiming at the problem of insufficient adaptability of traditional deconvolution kernels due to the complex and changeable motion trajectories of the target in dynamic scenes, a time-domain deconvolution kernel is constructed based on the spatio-temporal correlation information extracted from the feature tensor (such as luminance gradient, chromaticity correlation); a motion consistency constraint equation is designed by minimizing the optical flow residual between adjacent frames to ensure that the deconvolution kernel matches the actual motion trajectory of the target; a constrained optimization algorithm (such as the projected gradient method) is used to iteratively solve the pixel-level deconvolution kernel matrix to solve the performance bottleneck problem of the global deconvolution kernel in local motion difference scenes.

[0083] In some embodiments, a time-domain deconvolution kernel is constructed based on the feature tensor and the motion characteristics of the UAV, a motion consistency constraint equation is constructed by minimizing the optical flow residual between adjacent frames, and the smoothness constraint of the UAV motion trajectory is introduced, and a constrained optimization algorithm is used to solve the pixel-level time-domain deconvolution kernel matrix, including: 501. Based on the multi-scale spatio-temporal correlation features of the feature tensor, extract the spatio-temporal local contrast features, generate the initial parameter set of the dynamic convolution kernel, and combine the high-speed translation, hovering jitter and periodic motion mode of the rotor in the motion characteristics of the UAV to perform prior correction of the motion trajectory on the initial parameter set, and map the initial parameter set to the weight space of the deconvolution kernel to form a time-domain deconvolution kernel; In step 501, the spatio-temporal local contrast feature: By analyzing the brightness differences in the local regions of the image in the temporal and spatial dimensions (such as the change rate of pixel values in adjacent frames), it characterizes the detailed changes caused by the target movement.

[0084] Initial parameter set of the dynamic convolution kernel: The initial filter parameters generated based on spatio-temporal features, used to describe the blurred trajectories under different motion patterns.

[0085] Weight space mapping: Maps the initial parameters to the weight matrix of the deconvolution kernel through a non-linear transformation to adapt to different motion speeds and directions.

[0086] UAV motion characteristics: Include the high-speed translation, hovering jitter, and periodic motion patterns of the rotor of the UAV, used to perform prior correction of the motion trajectory on the initial parameter set to ensure that the deconvolution kernel matches the actual motion pattern of the UAV.

[0087] In the embodiment of the present application, first, the spatio-temporal local contrast feature is extracted from the feature tensor, and the spatio-temporal contrast (the average brightness difference between the current frame and the previous and next two frames) is calculated for the 3×3 neighborhood of each pixel, and the high-contrast regions are screened (the threshold is set to 0.2).

[0088] Secondly, an initial parameter set of the dynamic convolution kernel is generated based on the contrast feature, and combined with the high-speed translation, hovering jitter, and periodic motion patterns of the rotor in the UAV motion characteristics, prior correction of the motion trajectory is performed on the initial parameter set. For example: High-speed translation region: Corresponding to a short-time high-frequency kernel, adapting to the blur caused by fast motion; Hovering jitter region: Corresponding to a medium-time medium-frequency kernel, adapting to the blur caused by slight jitter; Rotor periodic motion region: Corresponding to a periodic kernel, adapting to the periodic blur caused by the rotation of the rotor.

[0089] Finally, the corrected initial parameters are mapped to the weight space of the deconvolution kernel through a fully connected neural network to generate a time-domain deconvolution kernel.

[0090] 502. Perform spatio-temporal joint analysis on the time-domain deconvolution kernel and the function of minimizing the optical flow residual between adjacent frames, introduce the smoothness constraint of the UAV motion trajectory, and construct a motion consistency constraint equation with the pixel displacement vector as the variable, where the smoothness constraint of the UAV motion trajectory includes the upper limit of the UAV acceleration and the heading continuity condition.

[0091] In step 502, the optical flow residual function is a function that measures the difference between the predicted motion vector between adjacent frames and the actual pixel displacement, and is used to evaluate the accuracy of motion estimation. The motion consistency constraint equation is an optimization equation constructed with the goal of minimizing the optical flow residual, which forces the deconvolution kernel to match the actual motion trajectory of the target. The smoothness constraint of the UAV motion trajectory includes the upper limit of the UAV acceleration and the heading continuity condition, which are used to constrain the optical flow residual function to ensure that the deconvolution kernel matches the actual motion mode of the UAV.

[0092] In the embodiments of the present application, first, the dense optical flow algorithm (such as the Farneback algorithm) is used to calculate the pixel displacement vectors of adjacent frames. The residual (such as the Euclidean distance) between the predicted displacement (derived from the deconvolution kernel parameters) and the actual optical flow vector is calculated for each pixel.

[0093] Secondly, the sum of the squares of the optical flow residuals is used as the objective function, and a constraint equation is constructed by combining the spatio-temporal local contrast feature and the smoothness constraint of the UAV motion trajectory (such as the upper limit of acceleration and the heading continuity condition). For example: Acceleration upper limit constraint: Restrict the change rate of the optical flow residual to avoid abnormal residuals caused by the high-speed movement of the UAV; Heading continuity constraint: Ensure the smooth transition of the optical flow residual when the UAV heading changes, and avoid the residual jump caused by the sudden change of the heading.

[0094] Finally, the motion consistency constraint equation is iteratively solved by an optimization algorithm (such as the gradient descent method) until the optical flow residual converges to the preset accuracy range.

[0095] 503. Iteratively execute the spatial domain and temporal domain optimizations through the constraint optimization algorithm, and adjust the weight distribution of the optical flow residual function according to the smoothness constraint of the UAV motion trajectory in each iteration until the mean square error of the optical flow residual converges to the preset accuracy range, and output the pixel-level temporal deconvolution kernel matrix that satisfies the motion consistency constraint.

[0096] In step 503, spatial domain optimization: Adjust the weight distribution of the deconvolution kernel in the image plane to enhance the recovery ability of local motion features, and dynamically adjust the weight attenuation coefficient of the high-motion area in combination with the smoothness constraint of the UAV motion trajectory (such as the upper limit of acceleration). Temporal domain optimization: Optimize the response characteristics of the deconvolution kernel on the time axis to match the periodic motion pattern of the target, and adjust the time span of the deconvolution kernel based on periodic motion features such as the rotation speed of the UAV rotor.

[0097] In the embodiments of the present application, first, set the initial deconvolution kernel parameters and the weight distribution of the optical flow residual function, and initialize the weight attenuation coefficient according to the smoothness constraint of the UAV motion trajectory (for example, the weight attenuation coefficient of the high-speed translation area is set to 0.8).

[0098] Secondly, the projection gradient method is used to update the spatial weights of the deconvolution kernel, giving priority to optimizing the weight distribution in high-contrast regions (such as the edges of UAV rotors), and dynamically adjusting the regional weights of the residual function according to the upper limit constraint of the UAV acceleration. For example: Acceleration overrun area: Reduce the weight decay coefficient to suppress the residual interference caused by abnormal motion; Heading mutation area: Increase the weight penalty term to avoid residual jumps caused by discontinuous headings.

[0099] Next, based on the target motion frequency (such as the rotor speed of 20 Hz), the time span of the deconvolution kernel is adjusted, and the time window length is set according to the periodic motion characteristics of the rotor (such as 5 frames corresponding to half a period of rotor rotation) to suppress low-frequency background interference.

[0100] Finally, when the change rate of the mean square error of the optical flow residual is less than 1% for three consecutive iterations, the optimization is terminated, and the pixel-level time-domain deconvolution kernel matrix is output.

[0101] The UAV flies at a speed of 20 m / s in heavy rain. The falling raindrops cause high-frequency background noise, and the rotor motion causes periodic blurring. Spatiotemporal local contrast features are extracted from the feature tensor. The average contrast of the rotor area is detected to be 0.5 (threshold 0.2), and an initial short-time high-frequency kernel (time span 3 frames) is generated; the average contrast of the fuselage area is 0.1, and a long-time low-frequency kernel (time span 7 frames) is generated. The optical flow residual calculation shows that the residual in the rotor area is 3.2 pixels (threshold 2.5 pixels). A motion consistency constraint equation is constructed, and the residual weight in the rotor area is given 1.5 times the weight. Through spatial domain optimization, the kernel weight in the rotor area is increased by 30%, and the time domain optimization matches the 20 Hz motion frequency of the rotor. After 15 iterations, the mean square error of the optical flow residual decreases from the initial 5.6 to 0.8, meeting the convergence condition. The clarity of the rotor texture is improved from 0.4 (SSIM) to 0.75, and the raindrop noise interference is reduced by 70%. The target tracking and positioning error is reduced from 6 pixels to 1.5 pixels, meeting the real-time detection requirements.

[0102] In summary, through spatiotemporal local contrast feature extraction, construction of a motion consistency constraint equation, and spatiotemporal joint optimization, this solution solves the problem of insufficient adaptive ability of the deconvolution kernel in complex dynamic scenarios.

[0103] To accurately capture the rapid movement of the target, aiming at the problems that the target movement speed is variable in the dynamic scene and the traditional single-time-domain convolution kernel is difficult to effectively capture spatio-temporal correlation features, by designing multi-scale dilated convolution kernels (such as a short-time 3-frame span and a long-time 7-frame span), the brightness gradient features (representing the light and dark changes caused by movement) and chromaticity correlation features (reflecting the color consistency under light fluctuations) of pixel points are extracted hierarchically in the time dimension, and the local and global features with different time spans are fused to construct a three-dimensional feature tensor with spatio-temporal correlation, so as to solve the defect of insufficient feature representation ability of traditional methods in complex motion patterns. In some embodiments, perform time-domain feature modeling on the original optical sequence, extract the brightness gradient features and chromaticity correlation features of each pixel point in the time dimension through multi-scale dilated convolution kernels, and output a feature tensor with spatio-temporal correlation, including: 601. Decompose the time-domain features of the original optical sequence into a brightness gradient component that obtains a per-pixel brightness change matrix through adjacent-frame difference operations and a chromaticity correlation component that describes the cross-frame chromaticity correlation using the covariance matrix of the chromaticity channel; In step 601, the per-pixel brightness change matrix is a matrix generated by calculating the brightness difference of corresponding pixels in adjacent frames, representing the light and dark changes caused by target movement. The chromaticity correlation component is based on the covariance matrix of the RGB channels and describes the color correlation between different frames (such as whether the change in the red channel is synchronized with the blue channel). The time-domain feature decomposition splits the original optical sequence into two independent dimensions of brightness and chromaticity, and analyzes the effects of movement and light separately.

[0104] In the embodiments of the present application, perform time-domain feature decomposition on the original optical sequence. For 5 consecutive frames of images, calculate the brightness difference of adjacent frames for each pixel (such as the difference between the t-th frame and the t + 1-th frame), generate a 4-layer brightness change matrix, extract the RGB channel data of each frame, calculate the cross-frame covariance matrix, such as the covariance value between the red channel and the blue channel in consecutive frames, quantify the chromaticity consistency, and output the brightness gradient component (4-layer matrix) and the chromaticity correlation component (3×3 covariance matrix) for subsequent multi-scale convolution processing.

[0105] 602. Perform depthwise separable convolution operations on the brightness gradient component and the chromaticity correlation component respectively through each branch of the multi-scale dilated convolution kernel; In step 602, the multi-scale dilated convolution kernel has convolution kernels with different time spans (such as a dilation rate of 1 corresponding to a 3-frame span and a dilation rate of 2 corresponding to a 5-frame span) for capturing short-time mutations and long-time trends. The depthwise separable convolution decomposes the standard convolution into a depth convolution (processing each channel separately) and a pointwise convolution (channel fusion), reducing the computational amount and improving the feature discrimination.

[0106] In the embodiments of the present application, depthwise separable convolution is performed on the luminance and chrominance components. Three convolutional branches with different dilation rates (dilation rates 1, 2, and 3) are set, corresponding to short-term, medium-term, and long-term feature extraction respectively. Layer-by-layer convolution is performed on the luminance gradient component (single channel) to extract motion features with different time spans; channel-independent convolution is performed on the chrominance correlation component (3 channels) respectively to retain color correlation. The outputs of each branch are fused in channels through 1×1 convolution. For example, short-term luminance features are cross-channel correlated with long-term chrominance features.

[0107] 603. Perform cross-scale correlation modeling on the multi-branch output features based on the depthwise separable convolution operation, align the convolution results with different dilation rates in the time dimension, and output a feature tensor with spatio-temporal correlation. In step 603, cross-scale correlation modeling aligns and stitches the feature maps with different time spans along the time axis to construct a spatio-temporal correlation tensor. The time dimension alignment makes the feature maps with different scales have the same time length through interpolation or cropping, which is convenient for subsequent fusion.

[0108] In the embodiments of the present application, cross-scale correlation modeling is performed on the multi-branch output. Linear interpolation is performed on the long-term feature map with a dilation rate of 3 to make its time dimension consistent with the short-term feature Figure 1 (for example, interpolating from 5 frames to 7 frames). The feature maps with different dilation rates are stitched along the channel dimension to form a three-dimensional tensor (time × space × channel). 3D convolution is used to further fuse spatio-temporal features, and a feature tensor containing global motion trends and local details is output.

[0109] The drone flies at a speed of 15 m / s in a sandstorm environment. The background sand and dust cause the light intensity to fluctuate by 500 lux per minute. Both the target and the background have irregular motion blur. The sand and dust occlusion causes the maximum difference in luminance between adjacent frames to reach 80 (in the range of 0 - 255), generating a high-dynamic luminance change matrix. The yellow tone of the sand and dust causes the covariance value between the red and green channels to increase to 0.8 (about 0.3 in a normal scene). Short-term dilated convolution (3-frame span) captures the rapid vibration of the drone's rotor (frequency 25 Hz). Long-term dilated convolution (7-frame span) extracts the slow diffusion trend of the sand and dust cloud (speed 0.5 m / s). The long-term feature map is interpolated from 7 frames to 9 frames and stitched with the short-term feature map. 3D convolution enhances the correlation between the high-frequency vibration of the rotor and the low-frequency motion of the sand and dust, and outputs a feature tensor. The spatio-temporal feature discrimination in the rotor area is increased by 50%, and the false detection rate caused by sand and dust interference is reduced by 65%. The prediction error of the target motion trajectory is reduced from 5 pixels to 1.2 pixels, meeting the real-time tracking requirements in complex environments.

[0110] To enhance the defense ability of the low-altitude security system against adversarial sample attacks and address the problem of decreased image recognition accuracy caused by adversarial sample attacks in low-altitude security scenarios, by extracting adversarial sample defense features (such as local gradient direction consistency, frequency domain energy concentration) from enhanced optical images, constructing a perturbation intensity evaluation function to quantify the potential risk of adversarial attacks, and adopting a non-linear weighted fusion strategy (such as a random forest model or an attention mechanism) to integrate multi-dimensional features to generate an adversarial perturbation feature mask and identify high-risk areas, thereby solving the defect of insufficient detection ability of traditional methods for new adversarial attacks. In some embodiments, adversarial sample defense features are extracted from the enhanced optical image, and non-linear weighted fusion is performed on the adversarial sample defense features by constructing a perturbation intensity evaluation function to form an adversarial perturbation feature mask, including: 701. Perform multi-scale spatial filtering on the enhanced optical image, combine an adaptive noise suppression algorithm with a local contrast enhancement operation to generate a denoised standardized optical image; In step 701, the multi-scale spatial filtering process uses filters of different sizes (such as 3×3, 5×5, 7×7) to perform hierarchical processing on the image, suppressing noise while retaining details. The adaptive noise suppression algorithm dynamically adjusts the filtering intensity according to the local noise level to avoid loss of details caused by over-smoothing. The local contrast enhancement operation improves the visibility of details in local regions of the image through histogram equalization or contrast stretching.

[0111] In the embodiments of the present application, multi-scale spatial filtering is performed on the enhanced optical image, and Gaussian filters of 3×3, 5×5, and 7×7 are respectively used to smooth the image, retaining detail information at different scales. The filtering intensity is dynamically adjusted based on the local noise variance (such as the standard deviation of pixel values within a 3×3 neighborhood). Strong filtering is used when the noise level is high, and weak filtering is used when the noise level is low. Local histogram equalization is performed on the filtered image to enhance the contrast of the target area (such as the drone rotor) to generate a denoised standardized optical image.

[0112] 702. Based on the denoised standardized optical image, perform feature extraction, and screen out high-frequency texture features and low-frequency structure features of the multi-rotor structure and metal reflection characteristics of the drone through a cross-channel attention mechanism to construct an adversarial sample defense feature matrix; In step 702, the cross-channel attention mechanism dynamically screens important features in the RGB channels through attention weights to enhance the distinguishability between high-frequency textures and low-frequency structures, where the attention mechanism focuses on key parts of the drone rotor and fuselage. The high-frequency texture features represent detailed information such as edges and corners in the image and are used to detect abnormal textures introduced by adversarial attacks. The low-frequency structure features describe the overall contour and regional distribution of the image and are used to evaluate the impact of adversarial attacks on the global structure.

[0113] In the embodiments of the present application, feature extraction is performed based on the denoised standardized optical image. The image is decomposed into three independent RGB channels, and the gradient magnitude and direction of each channel are calculated respectively. The weight coefficients of each channel are calculated through a cross-channel attention mechanism (such as the weight of the red channel is 0.6, the green channel is 0.3, and the blue channel is 0.1). The high-frequency texture and low-frequency structure features are dynamically fused, and the selected features are concatenated according to the channel dimension to generate an adversarial sample defense feature matrix for subsequent perturbation intensity evaluation.

[0114] 703. Dynamically calculate the weight coefficients of each adversarial sample defense feature channel using a perturbation intensity evaluation function, and adopt a layer-by-layer reverse gradient accumulation strategy to adjust the weight distribution to achieve non-linear weighted fusion of the feature space; In step 703, the perturbation intensity evaluation function calculates the potential risk level of the adversarial attack based on the feature matrix and assigns weights to different feature channels. The layer-by-layer reverse gradient accumulation strategy optimizes the weight distribution through backpropagation to enhance the detection accuracy of high-perturbation regions.

[0115] In the embodiments of the present application, a perturbation intensity evaluation function is used for feature weighted fusion, initial weights are assigned to each feature channel (such as the weight of the high-frequency texture is 0.7, and the weight of the low-frequency structure is 0.3), the gradient values of the feature channels are calculated through backpropagation, and the weight distribution is dynamically adjusted (such as the weight of the high-frequency texture is increased to 0.8). The weighted feature channels are fused according to a non-linear function (such as Sigmoid) to generate a perturbation intensity distribution map.

[0116] 704. Output an adversarial perturbation feature mask by iteratively optimizing the semantic consistency between the mask boundary and the original image for the non-linearly weighted fused adversarial sample defense feature matrix.

[0117] In step 704, iteratively optimize the mask boundary: adjust the mask edge through multiple iterations to make it consistent with the semantic information of the original image (such as the target contour). Semantic consistency: Ensure that the mask region is precisely aligned with the target region in the image (such as the drone fuselage and rotor).

[0118] In the embodiments of the present application, multi-dimensional information such as gradient magnitude, pixel intensity distribution, and texture features is extracted from adversarial samples. The features are weighted through a non-linear function (such as Sigmoid) to generate a fused feature matrix. A high-dimensional matrix containing the fused features is generated as the input for subsequent mask optimization. Based on the fused feature matrix, an initial mask is generated to cover the main area of the adversarial perturbation. The semantic similarity between the boundary region of the mask and the original image is calculated to evaluate the semantic consistency of the mask. If the semantic similarity is lower than a threshold (such as 0.9), the mask boundary is adjusted to expand or shrink the mask coverage range to ensure that the semantic information in the boundary region is aligned with the original image. The semantic consistency evaluation and boundary adjustment are repeated until the semantic similarity between the mask boundary and the original image reaches the preset threshold. The defense performance (such as the recognition accuracy of adversarial samples) and semantic consistency (such as the similarity between the boundary region and the original image) of the mask are tested on the validation set. The optimized adversarial perturbation feature mask is output.

[0119] In summary, this solution solves the problem of the decrease in image recognition accuracy caused by adversarial sample attacks through multi-scale filtering, cross-channel attention feature extraction, perturbation intensity evaluation, and mask optimization.

[0120] In order to improve the accuracy and robustness of object detection in the scenario of adversarial perturbation, aiming at the problems of the destruction of image semantic information and the decrease in object detection accuracy caused by adversarial perturbation in the low-altitude security scenario, by performing a spatial domain superposition operation on the enhanced optical image and the adversarial perturbation feature mask, the details of the original image and the perturbation region identification information are fused; a multi-scale residual network based on the attention mechanism is used to extract multi-scale features (such as local texture, global structure) from the superimposed image, and dynamically focus on the key regions (such as the drone rotor) through attention weights, and the object confidence estimation result is output, solving the problems of insufficient feature extraction and inaccurate object localization in the traditional method in the scenario of adversarial perturbation. In some embodiments, performing the spatial domain superposition operation on the enhanced optical image and the adversarial perturbation feature mask, and using the multi-scale residual network based on the attention mechanism to estimate the object confidence of the superimposed image includes: 801. Normalize the adversarial perturbation feature mask so that the numerical range of the adversarial perturbation feature mask matches the pixel distribution of the enhanced optical image, and generate a superimposed image through a spatial domain superposition operation; In step 801, the normalization process scales the pixel values of the adversarial perturbation feature mask to the same numerical range as the enhanced optical image (such as 0~255) for subsequent superposition operations. The spatial domain superposition operation fuses the normalized mask and the enhanced image pixel by pixel to generate a superimposed image containing the perturbation region identification.

[0121] In the embodiments of the present application, the adversarial perturbation feature mask is normalized, and the mask pixel values are linearly mapped from the range of 0 to 1 to 0 to 255 to match the pixel distribution of the enhanced optical image. The normalized mask and the enhanced image are pixel-weighted fused (such as mask weight 0.3 and image weight 0.7) to generate a superimposed image. The perturbed area is marked in semi-transparent red, and the non-perturbed area retains the details of the original image.

[0122] 802. Construct a multi-scale processing module to perform multi-scale residual network operations based on the attention mechanism on the superimposed image to generate a multi-scale fusion feature tensor; In step 802, the multi-scale processing module includes convolutional layers with different receptive fields (such as 3×3, 5×5, 7×7) for extracting local details and global structure features. The attention mechanism focuses on key areas (such as the drone rotor) through dynamic weight allocation to enhance the pertinence of feature extraction. The multi-scale fusion feature tensor stitches feature maps of different scales into a three-dimensional tensor to retain local and global information.

[0123] In the embodiments of the present application, a multi-scale processing module is constructed and residual network operations are performed. Convolutional kernels of 3×3, 5×5, and 7×7 are respectively used to extract local texture and global contour features. The weight coefficients of features at each scale are calculated through the attention mechanism (such as rotor area weight 0.8 and background area weight 0.2), and the features of key areas are dynamically fused. Feature maps of different scales are stitched according to the channel dimension to generate a multi-scale fusion feature tensor for subsequent confidence estimation.

[0124] 803. Input the multi-scale fusion feature tensor into a confidence estimation unit to perform local feature statistic calculation and generate a target confidence estimation value.

[0125] In step 803, the multi-scale processing module includes convolutional layers with different receptive fields (such as 3×3, 5×5, 7×7) for extracting local details and global structure features. The attention mechanism focuses on key areas (such as the drone rotor) through dynamic weight allocation to enhance the pertinence of feature extraction. The multi-scale fusion feature tensor stitches feature maps of different scales into a three-dimensional tensor to retain local and global information.

[0126] In the embodiments of the present application, a multi-scale processing module is constructed and residual network operations are performed. Convolutional kernels of 3×3, 5×5, and 7×7 are respectively used to extract local texture and global contour features. The weight coefficients of features at each scale are calculated through the attention mechanism (such as rotor area weight 0.8 and background area weight 0.2), and the features of key areas are dynamically fused. Feature maps of different scales are stitched according to the channel dimension to generate a multi-scale fusion feature tensor for subsequent confidence estimation.

[0127] The drone flies in a strong wind environment. Against the attack, high-frequency noise is injected into the image (PSNR = 28dB), resulting in the failure of target detection. The pixel values of the adversarial perturbation feature mask are mapped from 0~1 to 0~255. The mask is superimposed on the enhanced image with a weight of 0.3:0.7, and the perturbed area is marked in semi-transparent red. Convolution kernels of 3×3, 5×5, and 7×7 are used to extract the rotor texture and fuselage contour features. The weight of the rotor area is increased to 0.8, and the weight of the background area is decreased to 0.2. The feature mean and variance are calculated for the 3×3 neighborhood of the rotor area, and the mean is significantly higher than that of the background area. The confidence level of the rotor area is increased to 0.9, and the confidence level of the background noise area is decreased to 0.2. The detection accuracy of the rotor area is increased to 95%, and the false detection rate caused by the adversarial attack is reduced by 80%. The target detection accuracy is restored from 60% to 90%, meeting the low-altitude security requirements.

[0128] In summary, through normalization superposition, multi-scale residual network, and confidence estimation, this solution solves the problem of the decline in target detection accuracy in the scenario of adversarial perturbation.

[0129] In order to improve the target recognition accuracy and stability of the low-altitude security system in complex dynamic scenarios, aiming at the problem of the decline in the recognition accuracy of drones caused by environmental interference (such as sudden changes in light and background movement) and adversarial attacks (such as high-frequency noise and image tampering) in the low-altitude security scenario, based on the target confidence estimation results (such as local area confidence scores), through dynamic threshold screening and multi-frame fusion strategies, low-confidence false detection areas are eliminated, the recognition weight of high-confidence target areas is enhanced, and the recognition results are optimized by combining spatio-temporal context information, solving the defects of high false detection rate and poor robustness of traditional methods in complex scenarios. In some embodiments, according to the target confidence estimation results, drone recognition results with robustness against environmental interference and adversarial attacks are generated, including: 901. Based on the target confidence estimation results, a sliding window mechanism is used to calculate the environmental interference intensity index in real time. When it is detected that the confidence fluctuation of consecutive frames exceeds a preset threshold, an adaptive threshold adjustment algorithm is started; In step 901, the change trend of the target confidence is analyzed in real time through a sliding window (such as a window length of 5 frames and an overlap rate of 50%). The environmental interference intensity index quantifies the intensity of environmental interference based on the confidence fluctuation amplitude (such as the confidence difference between adjacent frames). The adaptive threshold adjustment algorithm dynamically adjusts the confidence threshold according to the interference intensity to screen high-confidence target areas. In the embodiments of the present application, based on the target confidence estimation results, the environmental interference intensity index is calculated, a sliding window analysis is performed on the confidence data of 5 consecutive frames, the mean and variance of the confidence within the window are calculated. If the confidence difference between adjacent frames exceeds a preset threshold (such as 0.2), it is determined that there is environmental interference, and the interference intensity index is set to the absolute value of the difference. When the interference intensity index > 0.3, the adaptive threshold adjustment algorithm is started, and the confidence threshold is increased from 0.6 to 0.7 to screen high-confidence areas.

[0130] 902. Identify the interference pattern of the confidence anomaly region through an adaptive threshold adjustment algorithm to generate an environmental interference feature mask; In step 902, the image region where the confidence fluctuation amplitude of the confidence anomaly region exceeds the preset threshold represents a potential interference region. The interference pattern recognition identifies the interference type (such as sudden light change, background movement) through feature analysis (such as texture, movement trajectory). The environmental interference feature mask is a binary mask used to mark the interference region, assisting subsequent recognition optimization.

[0131] In the embodiment of the present application, the interference pattern of the confidence anomaly region is recognized, the texture features (such as histogram of oriented gradients) and motion features (such as optical flow vectors) of the anomaly region are extracted, and the interference type is identified through a random forest model (including 100 decision trees). For example, the gradient directions in the sudden light change region are concentrated, and the optical flow vectors in the background movement region are consistent. An environmental interference feature mask is generated based on the recognition result to mark the interference region (for example, the sudden light change region is marked in yellow, and the background movement region is marked in green).

[0132] 903. Activate the adversarial sample detection module for low-confidence targets, and use the gradient backpropagation method to analyze the perturbation sensitivity of the feature space to generate a robustness evaluation coefficient.

[0133] In step 903, the adversarial sample detection module detects potential adversarial attack regions by analyzing the perturbation sensitivity of the feature space. The gradient backpropagation method calculates the sensitivity of the feature to the perturbation through backpropagation to quantify the impact of the adversarial attack. The robustness evaluation coefficient is an index characterizing the anti-adversarial attack ability of the target region, and the higher the coefficient, the stronger the robustness.

[0134] In the embodiment of the present application, for low-confidence targets, the adversarial sample detection is activated to calculate the gradient value of the feature with respect to the perturbation through backpropagation to quantify the perturbation sensitivity (for example, gradient magnitude > 0.5 indicates high sensitivity). A robustness evaluation coefficient is generated based on the gradient value. For example, the coefficient of the region where the gradient magnitude < 0.2 is set to 1.0, and the coefficient of the region where the gradient magnitude > 0.5 is set to 0.3. The robustness evaluation coefficient of each target region is output for subsequent recognition optimization.

[0135] 904. Combine the target confidence estimation result, the environmental interference feature mask, and the robustness evaluation coefficient to generate a UAV recognition result with anti-environmental interference and anti-adversarial attack robustness.

[0136] In step 904, through joint optimization, the confidence estimation result, the interference feature mask, and the robustness coefficient are integrated to generate a robustness recognition result. The anti-environmental interference and anti-adversarial attack robustness improve the stability of the recognition result in complex scenarios through a dynamic screening and fusion strategy.

[0137] In the embodiments of the present application, a robust recognition result is generated by combining multi-dimensional information. High-confidence target regions are screened according to a confidence threshold and a robustness coefficient (e.g., confidence > 0.7 and robustness coefficient > 0.8). The screening result is fused with the interference feature mask. For example, low-confidence targets in the regions with sudden illumination changes are ignored, and high-confidence targets in the regions with background motion are retained. A UAV recognition result with robustness against environmental interference and adversarial attacks is generated for subsequent tracking and decision-making.

[0138] The UAV flies in a rainstorm environment. Raindrops cause high-frequency noise in the image (PSNR = 28 dB), and the illumination intensity fluctuates by 300 lux per minute due to cloud occlusion. The average fluctuation amplitude of the confidence of 5 consecutive frames is 0.25 (threshold 0.2), and the interference intensity index is set to 0.25. The confidence threshold is increased from 0.6 to 0.7 to screen high-confidence regions. The gradient directions in the regions with sudden illumination changes are concentrated, and the optical flow vectors in the background raindrop regions are dispersed. The sudden illumination changes and raindrop interference are recognized to generate an environmental interference feature mask. The gradient amplitude of the raindrop region is 0.6, and the robustness coefficient is set to 0.3. Coefficient generation: The gradient amplitude of the UAV body region is 0.1, and the robustness coefficient is set to 1.0. Target regions with confidence > 0.7 and robustness coefficient > 0.8 are retained. Low-confidence targets in the raindrop regions are ignored to generate a robust recognition result. The UAV recognition accuracy is increased from 65% to 92%, and the false detection rate caused by raindrop interference is reduced by 85%. The false detection rate in the regions with sudden illumination changes is reduced by 70%, meeting the real-time detection requirements in complex environments.

[0139] In summary, through sliding window interference detection, adaptive threshold adjustment, adversarial sample detection, and multi-dimensional joint optimization, this solution solves the problem of the decline in UAV recognition accuracy caused by environmental interference and adversarial attacks in low-altitude security scenarios.

[0140] Figure 2 The following is a schematic structural diagram of a UAV recognition device provided for the embodiments of the present application, as Figure 2 shown. The device includes: A generation module 21, configured to generate a set of optical compensation parameters optimized for UAV imaging in real time based on a joint optimization model of the environmental light intensity change rate and the background motion vector field in the low-altitude security scenario; the generation module 21 is further configured to obtain an original optical sequence within a target capture window through the set of optical compensation parameters, and construct a time-domain deconvolution kernel in combination with the UAV motion characteristics to generate an enhanced optical image with anti-motion blur characteristics; A processing module 22, configured to extract adversarial sample defense features from the enhanced optical image, and perform non-linear weighted fusion on the adversarial sample defense features through constructing a perturbation intensity evaluation function to form an adversarial perturbation feature mask; The estimation module 23 performs a spatial domain superposition operation on the enhanced optical image and the adversarial perturbation feature mask, and uses a multi-scale residual network based on an attention mechanism to estimate the target confidence of the superimposed image; The generation module 21 is further configured to generate a UAV recognition result with robustness against environmental interference and adversarial attacks according to the target confidence estimation result.

[0141] Figure 2 The described UAV recognition device for low-altitude security can execute Figure 1 The described UAV recognition method for low-altitude security in the illustrated embodiment, and its implementation principle and technical effects will not be elaborated. For the UAV recognition device for low-altitude security in the above embodiment, the specific manners in which each module and unit perform operations have been described in detail in the embodiment related to the method, and will not be elaborated here.

[0142] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and are not intended to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. An unmanned aerial vehicle recognition method for low-altitude security, characterized in that, Including: Based on the joint optimization model of the environmental light intensity change rate and the background motion vector field in the low-altitude security scenario, a set of optical compensation parameters optimized for UAV imaging is generated in real time; Using the set of optical compensation parameters to obtain the original optical sequence within the target capture window, and combining the motion characteristics of the UAV to construct a time-domain deconvolution kernel to generate an enhanced optical image with anti-motion blur characteristics; Extracting adversarial sample defense features from the enhanced optical image, and performing non-linear weighted fusion on the adversarial sample defense features by constructing a perturbation intensity evaluation function to form an adversarial perturbation feature mask; Performing a spatial domain superposition operation on the enhanced optical image and the adversarial perturbation feature mask, and using a multi-scale residual network based on the attention mechanism to estimate the UAV target confidence of the superimposed image; According to the target confidence estimation result, a UAV recognition result with robustness against environmental interference and adversarial attacks is generated.

2. The method according to claim 1, wherein Based on the joint optimization model of the environmental light intensity change rate and the background motion vector field in the low-altitude security scenario, a set of optical compensation parameters optimized for UAV imaging is generated in real time, including: Real-time collecting the environmental light intensity data in the low-altitude security scenario through a multi-channel light intensity sensor, extracting the frequency domain features of the light intensity change rate based on time series analysis, and constructing a light intensity change rate function model; Performing multi-scale pyramid decomposition on the light intensity change rate function model, calculating the initial motion vectors of sparse feature points on each layer of the pyramid to form a background motion vector field; Coupling the light intensity change rate function model and the background motion vector field, establishing a joint optimization equation with the light intensity smoothing constraint and the minimum energy of the motion vector field as the objective function, and using the alternating direction multiplier method to iteratively solve the joint optimization equation until it converges to a stable state; According to the output parameters of the joint optimization equation, dynamically adjusting the exposure time, gain coefficient and filtering threshold of the optical compensation module to generate a set of optical compensation parameters optimized for UAV imaging.

3. The method according to claim 1, wherein Using the set of optical compensation parameters to obtain the original optical sequence within the target capture window, and combining the motion characteristics of the UAV to construct a time-domain deconvolution kernel to generate an enhanced optical image with anti-motion blur characteristics, including: Based on the set of optical compensation parameters, performing channel-by-channel enhancement processing on N frames of optical signals continuously collected within the target capture window. After each frame of optical signal is separated into RGB three channels, it is respectively multiplied by the gain coefficient matrix of the corresponding channel to generate the original optical sequence; Performing time-domain feature modeling on the original optical sequence, extracting the brightness gradient features and chromaticity correlation features of each pixel point in the time dimension through a multi-scale dilation convolution kernel, and outputting a feature tensor with spatio-temporal correlation; Based on the feature tensor and the motion characteristics of the UAV, constructing a time-domain deconvolution kernel, constructing a motion consistency constraint equation by minimizing the optical flow residual between adjacent frames, and introducing the smoothness constraint of the UAV motion trajectory, and using a constraint optimization algorithm to solve for the pixel-level time-domain deconvolution kernel matrix; Perform spatio-temporal joint deconvolution operation on the time-domain deconvolution kernel and the original optical sequence, where back-projection calculation is performed along the time axis for each spatial position, and the phase response characteristics of the deconvolution kernel are adjusted through an iterative update strategy to generate an enhanced optical image with anti-motion blur characteristics.

4. The method according to claim 2, characterized in that Couple the light intensity change rate function model with the background motion vector field, establish a joint optimization equation with the light intensity smoothing constraint and the minimization of the motion vector field energy as the objective function, and use the alternating direction multiplier method to iteratively solve the joint optimization equation until it converges to a stable state, including: Design the coupling term parameters, fuse the light intensity smoothing constraint of the light intensity change rate function model and the background motion vector field through a weighting factor to form a joint optimization equation; Decompose the joint optimization equation into a first sub-problem that constrains the light intensity gradient between adjacent pixels by introducing the light intensity smoothing regular term of the light intensity change rate function model and a second sub-problem that constrains the spatial continuity of the motion vector by introducing the energy term of the background motion vector field; Construct augmented Lagrangian functions for the first sub-problem and the second sub-problem respectively, and introduce a Lagrangian multiplier term with linear constraints and a quadratic penalty term in the objective function; After each variable update, synchronously update the augmented Lagrangian multiplier, and dynamically adjust the weight coefficient of the quadratic penalty term through an adaptive step size adjustment strategy; Based on the quadratic penalty term, use the alternating direction multiplier method to iteratively solve the joint optimization equation, calculate the relative error norm between two adjacent iterations, and determine that it converges to a stable state when the relative error norm is less than a preset threshold; 5. The method according to claim 3, wherein Construct a time-domain deconvolution kernel based on the feature tensor and the UAV motion characteristics, construct a motion consistency constraint equation by minimizing the optical flow residual between adjacent frames, and introduce the UAV motion trajectory smoothness constraint, and use a constrained optimization algorithm to solve for the pixel-level time-domain deconvolution kernel matrix, including: Based on the multi-scale spatio-temporal correlation features of the feature tensor, extract spatio-temporal local contrast features, generate an initial parameter set of the dynamic convolution kernel, combine the high-speed translation, hover jitter, and rotor periodic motion modes in the UAV motion characteristics, perform prior correction of the motion trajectory on the initial parameter set, and map the initial parameter set to the weight space of the deconvolution kernel to form a time-domain deconvolution kernel; Perform spatio-temporal joint analysis on the time-domain deconvolution kernel and the function of minimizing the optical flow residual between adjacent frames, introduce the UAV motion trajectory smoothness constraint, and construct a motion consistency constraint equation with the pixel displacement vector as the variable, where the UAV motion trajectory smoothness constraint includes the UAV acceleration upper limit and the heading continuity condition; Iteratively perform spatial domain and time domain optimization through a constrained optimization algorithm, adjust the weight distribution of the optical flow residual function according to the UAV motion trajectory smoothness constraint in each iteration until the mean square error of the optical flow residual converges to a preset accuracy range, and output the pixel-level time-domain deconvolution kernel matrix that satisfies the motion consistency constraint.

6. The method according to claim 3, wherein Perform time-domain feature modeling on the original optical sequence, extract the brightness gradient features and chromaticity correlation features of each pixel point in the time dimension through a multi-scale dilated convolutional kernel, and output a feature tensor with spatio-temporal correlation, including: Decompose the time-domain features of the original optical sequence into a brightness gradient component that obtains a per-pixel brightness change matrix through adjacent frame difference operations and a chromaticity correlation component that describes the cross-frame chromaticity correlation using the covariance matrix of the chromaticity channel; Perform depthwise separable convolution operations on the brightness gradient component and the chromaticity correlation component respectively through each branch of the multi-scale dilated convolutional kernel; Based on the depthwise separable convolution operations, perform cross-scale correlation modeling on the multi-branch output features, align the convolution results with different dilation rates in the time dimension, and output a feature tensor with spatio-temporal correlation.

7. The method according to claim 1, characterized in that, Extract adversarial sample defense features from the enhanced optical image, and perform non-linear weighted fusion on the adversarial sample defense features by constructing a perturbation intensity evaluation function to form an adversarial perturbation feature mask, including: Perform multi-scale spatial filtering on the enhanced optical image, combine an adaptive noise suppression algorithm and local contrast enhancement operations to generate a denoised and normalized optical image; Based on the denoised and normalized optical image, perform feature extraction, screen the high-frequency texture features and low-frequency structure features of the multi-rotor structure and metal reflection characteristics of the drone through a cross-channel attention mechanism, and construct an adversarial sample defense feature matrix, where the attention mechanism focuses on the key parts of the drone rotor and fuselage; Dynamically calculate the weight coefficients of each adversarial sample defense feature channel using the perturbation intensity evaluation function, and adjust the weight distribution using a layer-by-layer reverse gradient accumulation strategy to achieve non-linear weighted fusion in the feature space; Output an adversarial perturbation feature mask by iteratively optimizing the semantic consistency between the mask boundary and the original image for the non-linearly weighted fused adversarial sample defense feature matrix.

8. The method according to claim 1, characterized in that, Perform a spatial domain superposition operation on the enhanced optical image and the adversarial perturbation feature mask, and use a multi-scale residual network based on the attention mechanism to estimate the confidence of the drone target for the superimposed image, including: Normalize the adversarial perturbation feature mask so that the numerical range of the adversarial perturbation feature mask matches the pixel distribution of the enhanced optical image, and generate a superimposed image through a spatial domain superposition operation; Construct a multi-scale processing module to perform multi-scale residual network operations based on the attention mechanism on the superimposed image to generate a multi-scale fusion feature tensor; Input the multi-scale fusion feature tensor into a confidence estimation unit to perform local feature statistic calculations and generate a drone target confidence estimation value.

9. The method according to claim 1, characterized in that, Generate a drone recognition result with robustness against environmental interference and adversarial attacks based on the drone target confidence estimation result, including: Based on the drone target confidence estimation result, use a sliding window mechanism to calculate the environmental interference intensity index in real time. When it is detected that the confidence fluctuation of consecutive frames exceeds a preset threshold, start an adaptive threshold adjustment algorithm; Identify the interference pattern for the confidence anomaly region through the adaptive threshold adjustment algorithm to generate an environmental interference feature mask; Activate the adversarial sample detection module for low-confidence targets, analyze the sensitivity of feature space perturbation using the gradient backpropagation method, and generate a robustness evaluation coefficient; Combine the drone target confidence estimation result, the environmental interference feature mask, and the robustness evaluation coefficient to generate a drone recognition result with robustness against environmental interference and adversarial attacks.

10. An unmanned aerial vehicle recognition system for low-altitude security and protection, characterized in that, It includes: A generation module for generating a set of optical compensation parameters optimized for drone imaging in real time based on a joint optimization model of the environmental light intensity change rate and the background motion vector field in a low-altitude security scenario; The generation module is further configured to obtain an original optical sequence within the target capture window through the set of optical compensation parameters, construct a temporal deconvolution kernel in combination with the motion characteristics of the drone, and generate an enhanced optical image with anti-motion blur characteristics; A processing module for extracting adversarial sample defense features from the enhanced optical image, and performing non-linear weighted fusion on the adversarial sample defense features by constructing a perturbation intensity evaluation function to form an adversarial perturbation feature mask; An estimation module for performing a spatial domain superposition operation on the enhanced optical image and the adversarial perturbation feature mask, and estimating the target confidence of the superimposed image using a multi-scale residual network based on an attention mechanism; The generation module is further configured to generate a drone recognition result with robustness against environmental interference and adversarial attacks according to the target confidence estimation result.

Citation Information

Patent Citations

  • Intelligent decision making system based on activity characteristics for ankle joint ligament damage

    CN111820902A

  • Multi-interference-band laser directed infrared countermeasure turret

    CN112461052A

  • Pedestrian target detection physical anti-attenuation confrontation method robust to imaging main body change

    CN116384107A

  • AI video low-altitude target identification and real-time tracking method based on deep learning

    CN119723421A

  • Method for generating adversarial sample of multi-modal remote sensing image

    CN119851126A

Cited By

  • Unmanned aerial vehicle rapid identification method and system based on lightweight convolutional neural network

    CN120431527A

  • A method and system for rapid identification of drones based on lightweight convolutional neural networks

    CN120431527B

  • Security box video monitoring system and method with intelligent discrimination function

    CN120568026A

  • Security box video monitoring system and method with intelligent discrimination function

    CN120568026B

  • Monitoring video enhancement method for farm

    CN120634901A