Target tracking protection method based on low-frequency disturbance attack retrieval
By adopting a protection method based on low-frequency perturbation attack retrieval in the visual target tracking system, using DCT analysis and regression analysis to identify and correct low-frequency perturbation signals, the accuracy and stability of the system under adversarial attacks are solved, and higher robustness and security are achieved.
Patent Information
- Application Number
- CN202510008781.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-03
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-01-03
AI Technical Summary
The accuracy and stability of existing visual target tracking systems are difficult to guarantee when facing adversarial attacks, especially in complex and low-resource environments, and the existing defense measures are limited in effect.
The target tracking protection method based on low-frequency perturbation attack retrieval is adopted. The low-frequency perturbation signal in the video sequence is analyzed through discrete cosine transform (DCT), low-frequency perturbation detection and regression analysis are performed, and the target tracking model is dynamically adjusted to correct the perturbation signal.
Effectively identify and defend against adversarial attacks, improve the stability and robustness of the visual target tracking system, significantly improve the reliability and security of the system in actual applications, and has low computing overhead, making it suitable for high-performance and resource-constrained environments.
Smart Images

Figure CN119942440A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of visual target tracking, and in particular to a target tracking protection method based on low-frequency disturbance attack retrieval. Background Art
[0002] Visual object tracking is a key area in computer vision. It locates the target in each frame by continuously inferring the key frames in the video sequence to provide the trajectory and activity area of the target. This technology has been significantly improved in practical applications in recent years and is widely used in autonomous driving, biological behavior monitoring, aerial visual target tracking and other fields. In recent years, visual object tracking algorithms based on deep learning have made great progress. However, it also exposes the vulnerability of the system in the face of adversarial attacks. In order to improve the robustness of the system, researchers have proposed a variety of defense mechanisms, including adversarial training, input preprocessing, and model architecture optimization. However, in the face of complex and varied attack methods, the effectiveness of these defense measures is still limited.
[0003] In existing technologies, adversarial attacks have gradually become a serious problem. Adversarial attacks generate carefully designed disturbance signals to cause the target tracker to misjudge the target's true position, thereby affecting the accuracy and stability of tracking. This attack may have serious consequences in practical application scenarios such as autonomous driving, security monitoring, and biological behavior monitoring, resulting in a significant decrease in safety and reliability.
[0004] The importance of the defense system lies not only in improving the security and reliability of a single application, but also in providing a comprehensive and reliable technical framework that enables visual target tracking technology to operate stably in various complex environments and potential attacks. By implementing a target tracking protection method based on low-frequency perturbation attack retrieval, it can effectively cope with the challenges of adversarial attacks and provide necessary security guarantees and technical support for a wide range of application fields.
[0005] In real-world applications, autonomous driving systems need to work reliably in complex environments, including extreme weather, complex road conditions, and emergencies. Attackers can generate carefully designed disturbance signals to cause the prediction results of the visual target tracking model to deviate seriously from the actual target, resulting in system misjudgment and potential safety hazards. In real-time tracking systems, the accuracy of calibrating the target position is critical. In the field of autonomous driving, vehicles rely on highly accurate visual object tracking technology to achieve reliable navigation and obstacle avoidance in complex road environments. In security monitoring systems, surveillance cameras and intelligent detection systems are widely used for security protection in public places and important facilities. Adversarial attacks can cause security vulnerabilities by generating interference signals that prevent the monitoring system from correctly identifying and tracking intruders. Many defense methods are not stable and reliable enough in the face of rapidly changing environments or complex interference noises, which affects the system's defense capabilities and cannot guarantee the stable operation of the system in various complex scenarios.
[0006] In addition, existing defense methods often require high-performance hardware support when facing large-scale data and complex tasks. For application scenarios that require real-time processing, especially in low-resource environments, such as edge devices and embedded systems, a large amount of forward and backward propagation calculations are required, which greatly increases the response time and resource requirements of the system. The high computing power requirements of these methods directly affect the deployment and operation efficiency of the system. For application scenarios with high real-time requirements, such as autonomous driving and security monitoring, this has caused great limitations, making it difficult to guarantee the efficiency and effectiveness of current defense methods in actual deployment. Summary of the invention
[0007] In order to solve the above technical problems, the present invention proposes a target tracking protection method based on low-frequency disturbance attack retrieval, which ensures that judgments can be made quickly and corresponding measures can be taken when encountering adversarial attacks, thereby effectively protecting the security and stability of the visual target tracking system.
[0008] On the one hand, to achieve the above-mentioned purpose, the present invention provides a target tracking and protection method based on low-frequency disturbance attack retrieval, comprising:
[0009] Acquire an input video sequence, determine a central image of a sub-frame detection area of the video sequence, perform low-frequency disturbance detection using the central image of the sub-frame detection area, and obtain a low-frequency sampling result;
[0010] Inputting the low-frequency sampling result into the target tracking model for processing to obtain a corrected target tracking result;
[0011] The target tracking model quantifies the low-frequency sampling results through a linear regression method, and dynamically adjusts parameters to correct the input frame.
[0012] Preferably, obtaining the low-frequency sampling result includes:
[0013] Converting the central image of the sub-frame detection area into a grayscale image, dividing the grayscale image into a number of identical small-block images, performing discrete cosine transform on each of the small-block images, and converting the pixel values in the spatial domain into frequency domain representation;
[0014] The coefficients of the same part in each of the small-block images are selected to extract the low-frequency components, and all the low-frequency components are integrated to obtain the low-frequency sampling result.
[0015] Preferably, the process of inputting the low-frequency sampling result into the target tracking model for processing is:
[0016] A regression analysis is performed on the low-frequency sampling results to determine the nature and scope of the attack disturbance and quantify the characteristics of the disturbance signal.
[0017] Preferably, before the step of performing linear regression analysis on the low-frequency sampling results, the method comprises:
[0018] Calculate the mean and standard deviation based on the low-frequency sampling results;
[0019] A reference model is constructed, a threshold is set for the reference model, the mean and the standard deviation are compared with the threshold, and it is determined whether there is an abnormal disturbance signal. If there is an abnormal disturbance signal, a regression linear analysis is performed; if there is no abnormal disturbance signal, the low-frequency sampling result is input into the target tracking model as the actual image to be predicted.
[0020] Preferably, the process of performing the regression linear analysis is:
[0021] Linear regression is introduced to analyze the multi-directional detection disturbance characteristics, a disturbance characteristic curve is established, and the abnormal disturbance signal is quantified; wherein the disturbance characteristic curve is related to the offset, the disturbance slope and the disturbance error.
[0022] Preferably, the process of inputting the low-frequency sampling result into the target tracking model for processing further includes:
[0023] Through the low-frequency sampling results, the low-frequency domain repair signal is iteratively searched, specifically:
[0024] Generate a two-dimensional signal based on the low-frequency sampling result, and project the two-dimensional signal to a sub-frame detection area to obtain a test frame;
[0025] The time-domain-based IoU score is calculated through the test frame to evaluate the impact of the current video frame on the target tracking model, and the sample with the highest time-domain-based IoU score is selected as the actual sample to be predicted.
[0026] Preferably, the fixed number of times is based on the computing power of the device to select the most efficient setting.
[0027] On the other hand, to achieve the above-mentioned purpose, the present invention also provides a target tracking and protection system based on low-frequency disturbance attack retrieval, comprising:
[0028] A low-frequency sampling result acquisition module is used to acquire an input video sequence, determine a central image of a sub-frame detection area of the video sequence, perform low-frequency disturbance detection through the central image of the sub-frame detection area, and obtain a low-frequency sampling result;
[0029] Target tracking model processing module: used for inputting the low-frequency sampling result into the target tracking model for processing to obtain a corrected target tracking result;
[0030] Wherein, the target tracking model processing module includes:
[0031] A disturbance detection unit, used for analyzing low-frequency coefficients and detecting low-frequency disturbance signals;
[0032] A regression analysis unit, used for performing regression analysis on the detected disturbance signal;
[0033] The model adjustment unit is used to adjust and correct the target tracking model according to the regression analysis results.
[0034] The present invention also provides a computer device, including a memory, a processor and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the target tracking and protection method based on low-frequency disturbance attack retrieval.
[0035] The present invention also provides a computer-readable storage medium having a computer program / instruction stored thereon, which, when executed by a processor, implements the steps of the target tracking and protection method based on low-frequency disturbance attack retrieval.
[0036] Compared with the prior art, the present invention has the following advantages and technical effects:
[0037] The present invention analyzes low-frequency disturbance signals in input video sequences through discrete cosine transform (DCT), which can effectively identify and defend against potential adversarial attacks, and has significant advantages in improving the stability and robustness of the tracking system; through regression analysis and dynamic adjustment, it can adapt to different types of disturbance attacks and effectively improve the robustness of the model, significantly improving the reliability and safety of visual target tracking in practical applications.
[0038] The present invention constructs a comprehensive and highly integrated defense framework through a series of steps such as preprocessing of input data, discrete cosine transform, low-frequency disturbance detection, disturbance regression analysis and model adjustment. The entire defense strategy runs through the complete chain from detection to response, ensuring that judgments can be made quickly and corresponding measures can be taken when encountering adversarial attacks, thereby effectively protecting the security and stability of the visual target tracking system. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] The drawings constituting a part of the present application are used to provide a further understanding of the present application. The illustrative embodiments and descriptions of the present application are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0040] Figure 1 A flow chart of a target tracking and protection method based on low-frequency disturbance attack retrieval according to an embodiment of the present invention;
[0041] Figure 2 The present invention is a flowchart of iteratively searching a low-frequency domain repair signal based on a low-frequency sampling space according to an embodiment of the present invention. DETAILED DESCRIPTION
[0042] It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0043] It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and that, although a logical order is shown in the flowcharts, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0044] The present invention proposes a target tracking protection method based on low-frequency disturbance attack retrieval, such as Figure 1 ,include:
[0045] Obtain an input video sequence, determine a central image of a sub-frame detection area of the video sequence, perform low-frequency disturbance detection using the central image of the sub-frame detection area, and obtain a low-frequency sampling result;
[0046] The low-frequency sampling result is input into the target tracking model for processing to obtain a corrected target tracking result;
[0047] Among them, the target tracking model quantifies the low-frequency sampling results through a linear regression method, and dynamically adjusts the parameters to correct the input frame.
[0048] Specifically, in the image processing process, the input video frame is first denoised and preprocessed to ensure the quality and consistency of the data. Subsequently, the input image is converted to the frequency domain space through discrete cosine transform (DCT), and the low-frequency components are extracted in the frequency domain space to generate orthogonal vector space and sample perturbations. This method can effectively reduce the misleading of high-frequency components to images and video frames, and ensure the directionality and effectiveness of generated patches. In order to improve the efficiency and resource optimization of the generation process, this embodiment adopts a sampling reconstruction strategy. The key sampling points are selected through a greedy algorithm, and combined with the reconstruction algorithm, the iterative process is optimized under the background of ensuring the demand for restoration effect.
[0049] In the attack scenario, the adjustment signal is generated through iterative query and optimization. Using the query data, the direction and amplitude of the signal are continuously optimized to improve the accuracy and effectiveness of the defense. Finally, the generated repaired video frame is further reconstructed and used in the visual object tracking process.
[0050] Furthermore, low-frequency sampling results are obtained, including:
[0051] Convert the central image of the sub-frame detection area into a grayscale image, and divide the grayscale image into a number of identical small blocks of images, perform discrete cosine transform on each small block of images, and convert the pixel value in the spatial domain into a representation in the frequency domain;
[0052] The coefficients of the same part in each small block of the image are selected, the low-frequency components are extracted, and all the low-frequency components are integrated to obtain the low-frequency sampling result.
[0053] Specifically, for the input video sequence, T, W, H, and C represent the sequence, width, height, and number of color channels of the video frames, respectively. In the actual tracking process, the first frame is used to provide the characteristics of the target. Detection and defense start from the second frame. The video sequence is input into the target tracking model in sequence. For a system that can relocate after the object is lost, the target tracking model can skip frames and reset itself as the target is lost during the tracking process.
[0054] Input video sequence is a four-element two-dimensional vector, which is used to represent the clockwise position of the box coordinate system where the predicted target is located, starting from the upper left corner. Based on the center of the area, the image patch area is obtained, with a length and width of Four times. Calculate the center coordinates of the image patch area (x c ,y c ) to obtain the center image of the sub-frame detection area.
[0055] In this embodiment, the generation calculation of the low-frequency sampling space is performed based on the image of the determined area, specifically:
[0056] (1-1) Read the color pixel information in the area and convert it into a grayscale image. For a traditional three-channel RGB color image, perform channel-by-channel conversion Gray = 0.299·R+0.587·G+0.114·B, where R represents the red component of a single pixel, taking a value of [0, 1], G represents the green component of a single pixel, taking a value of [0, 1], and B represents the blue component of a single pixel, taking a value of [0, 1];
[0057] (1-2) Divide the grayscale image into 8*8 small blocks of equal size;
[0058] (1-3) Discrete cosine transform is applied to each small image block to convert the pixel value in the spatial domain into a representation in the frequency domain:
[0059]
[0060] Where cos() is the corresponding trigonometric function; F(u, ν) is a two-dimensional matrix used to describe the frequency domain coefficients after two-dimensional discrete cosine transform; M and N are the number of rows and columns of the image, respectively, and α(u) and α(ν) are normalization coefficients, defined as:
[0061]
[0062] Where, u is the frequency component in the horizontal direction of the two-dimensional DCT transform, ranging from 0 to N-1, and v is the frequency component in the vertical direction of the two-dimensional DCT transform, ranging from 0 to N-1;
[0063] (1-4) After the transformation calculation, most of the information in the small block image is concentrated in the low-frequency part. The low-frequency components are extracted by selecting several coefficients in the upper left corner of the small block image. In this embodiment, only the 2×2 or 4×4 area in the upper left corner is selected from the 8×8 DCT coefficient matrix;
[0064] (1-5) All extracted low-frequency components are integrated to generate a new image or feature space, which represents the low-frequency sampling results of the image.
[0065] Furthermore, the process of inputting the low-frequency sampling results into the target tracking model for processing is:
[0066] Linear regression analysis is performed on the low-frequency sampling results to determine the nature and scope of the attack disturbance and quantify the characteristics of the disturbance signal.
[0067] Furthermore, the step of performing linear regression analysis on the low-frequency sampling results includes:
[0068] Calculate the mean and standard deviation based on the low-frequency sampling results;
[0069] Set a threshold, compare the mean and standard deviation with the threshold, and determine whether there is an abnormal disturbance signal. If there is an abnormal disturbance signal, perform regression linear analysis. If not, input the low-frequency sampling result as the actual image to be predicted into the target tracking model.
[0070] Specifically, based on the low-frequency sampling result, a low-frequency coefficient matrix L=F(u, v)|0≤u, ν<K, K represents the value range of the low-frequency coefficient, and the mean and standard deviation are calculated;
[0071] The method to calculate the standard deviation is:
[0072] In the formula, σ L is the standard deviation, m is the number of rows of data, n is the number of columns of data, i and j are the indexes of specific rows and columns respectively, L ij is the value of the two-dimensional data in the i-th row and j-th column, μ L is the mean of the two-dimensional data. The formula for calculating the mean of the two-dimensional data is:
[0073] A reference model R is constructed to implement detection based on the threshold τ: ΔL=max(|LR|), where ΔL is used to represent the model difference coefficient. If ΔL>τ, there is an abnormal disturbance signal and a regression analysis is performed. If there is no abnormal disturbance signal, the low-frequency sampling result is used as the actual image to be predicted and input into the target tracking model.
[0074] Furthermore, the process of performing regression linear analysis is:
[0075] Linear regression is introduced to analyze the disturbance characteristics of multi-directional detection and establish a disturbance characteristic curve, in which the disturbance characteristic curve is related to the offset, disturbance slope and disturbance error.
[0076] Specifically, linear regression is introduced to analyze the multi-directional detection disturbance characteristics and the disturbance characteristic curve L is established. perturb :
[0077] L perturb ≈β0+β1x+∈;
[0078] Where β0 is the offset, β1 is the disturbance slope, and ∈ is the disturbance error.
[0079] The linear regression results are applied to the target tracking model to analyze the impact of the disturbance signal on the response of the target tracking model. The specific process is as follows:
[0080] Calculate the residuals and calculate the average. The method for calculating the residuals is:
[0081] ∈ i =L perturb,i -(β0+β1x i );
[0082] In the formula, ∈ i Represents the difference between the actual observed value and the predicted value, L perturb,i is the actual observed value of the ith data point, x i are the input variables used for prediction.
[0083] The mean of the residuals should be close to zero. If the mean of the residuals deviates significantly from zero, there may be abnormal disturbance signals.
[0084] If there are abnormal disturbance signals, try to search for repair patches to improve the accuracy of the target tracking model, such as Figure 2 ,include:
[0085] Randomly select a direction from the low-frequency sampling result, generate a signal and superimpose it on the sampling area to obtain a test frame, wherein for the corresponding pixels on the signal and the sampling area, the color values of the three RGB channels are superimposed one by one;
[0086] Compute a Gaussian signal with a maximum range based on the distribution function and variance:
[0087] The distribution function is
[0088] The variance is
[0089] In the formula, π is the circumference of a circle, x is the horizontal coordinate in the normal distribution, and μ is the center position of the normal distribution, which is also the average value of the data. It represents the square of the standard deviation and is used to measure the degree of dispersion of the data.
[0090] The above distribution function generates corresponding full-channel noise for each pixel within the image patch area.
[0091] The generated Gaussian signal is appended to the original frame as a candidate test frame.
[0092] The image space generated by sampling the image patch area is used to generate an orthogonal patch area, and a block discrete cosine transform is performed. The low-frequency coefficients are extracted from the transformed frequency domain signal and analyzed, specifically:
[0093] (2-1) Select the direction vector q from the low-frequency sampling result, and consider the negative vector q'=-q;
[0094] (2-2) According to the input vector x = [x1, x2, ..., x n ], generating a sequence Q = [x1, x 1+s , x 1+2s , ..., x 1+ks ], this sequence contains all possible alternative repair signals;
[0095] (2-3) Divide the generated sequence Q into L groups evenly according to the stride, and randomly select a signal from each group, selecting a total of L' signals (L' = L);
[0096] (2-4) The above selected signals are linearly combined according to the corresponding weighting coefficients to generate a linear combination signal v = α1DCT1 + α2DCT2 + ... + α n DCT n , where DCT represents the orthogonal discrete cosine transform vector and the coefficient α n is the corresponding weight when generating the signal, and n is used to represent the number of DCT basis functions of the signal;
[0097] (2-5) Convert the time domain signal into the frequency domain signal:
[0098]
[0099] Where DCT(υ) is the frequency domain signal after transformation, v j is the jth sample of the input time domain signal sequence Q, j is the sequence number, and its value is [0, N-1]; k is the index of the transformed frequency domain component, and N is the sample length of the input time domain signal sequence Q;
[0100] (2-6) According to and Convolution is performed with the grayscale image to obtain the horizontal gradient image g y and the vertical gradient image g x , and calculate the tangent direction
[0101] Calculate the direction orthogonal to the tangent direction as the auxiliary direction;
[0102] (2-7) A two-dimensional signal is generated according to the tangent direction and the auxiliary direction. This process is based on the accumulation of low-frequency components. δ = v + β1u1 + β2u2 + ... + β k u k ;
[0103] Where δ is the two-dimensional signal accumulated in the tangent direction and the auxiliary direction, v is the low-frequency part of the signal, and β k is the weight coefficient in the kth signal direction, u k is the signal component on the kth signal;
[0104] Based on the low-frequency sampling space, the low-frequency domain repair signal is iteratively searched to generate a test frame, and the intersection-over-union score based on the time domain is calculated through the test frame. In each iteration, by evaluating the impact of the current video frame on the target tracking model, the signal patch that maximizes the reduction of the disturbance effect is selected, the current frame is updated, and the sample with the highest tracking score is selected as the actual sample to be predicted based on a fixed number of queries.
[0105] Furthermore, based on the low-frequency sampling space, the low-frequency domain repair signal is iteratively searched, specifically:
[0106] (3-1) Use the selected direction to obtain the two-dimensional signal δ and superimpose it on the sampling area to generate the test frame I +q , I -q :
[0107]
[0108] In the formula, is the video frame after the repair signal generated for the k-1th time, which is used to superimpose the current signal, δ q represents the two-dimensional signal generated by the direction vector q, δ q′ Represents the two-dimensional signal generated by the direction vector q';
[0109] For each test frame, the target detection area is generated by the target tracking model.
[0110] Calculate the intersection-over-union score based on the time domain
[0111] Where λ is the weighting coefficient, which ranges from [0, 1] and is used to balance the weights of the current frame and the previous frame; is the predicted bounding box of the original frame; is the predicted bounding box of the test frame; is the predicted bounding box of the previous original frame;
[0112] (3-2) Add the intersection-combination ratio score to the set ψ k =ψ k ∪S IOU ,In each iteration, by evaluating the impact of the current video frame on the target tracking model, the signal patch that can maximize the reduction of the disturbance effect is selected to perform the search for the optimal greedy strategy;
[0113] (3-3) Update the current frame: I′tk=I′tk-1+ψk·η;
[0114] Where η is the patch update step size, I′tk is the video frame repaired by the signal patch generated in the kth iteration;
[0115] (3-4) Update the iteration count to maintain computational efficiency:
[0116] Based on the fixed number of queries, the sample with the highest tracking score is selected as the actual sample to be predicted and input into the target tracking model. In this embodiment, the fixed number of queries is based on the computing power of the device to select the setting with the best efficiency. Each time the target tracking model is input and the intersection-over-union score S is calculated IOU , select the single test frame with the highest score for reconstruction. Specifically, Complete reconstruction; I' n is the pixel value of the reconstructed image.
[0117] Specifically, the reconstruction function The build process consists of the following steps:
[0118] (4-1) Find the corresponding source coordinates in the original image from each target coordinate in the single test frame with the highest score:
[0119]
[0120] In the formula, x and y represent the horizontal and vertical coordinates of the image, src and dst represent the target positions before and after reconstruction, w and h represent the lengths after reconstruction, and x src is the horizontal coordinate in the source image, z dst To transform the image horizontal coordinate, w dst is the converted image width, w src is the source image width, y src is the ordinate of the source image, y dst To transform the image ordinate, h dst To convert the image height, h src is the source image height;
[0121] (4-2) Based on the four directions of the two-dimensional coordinate axis, find the four neighboring pixels in the source image that are closest to the calculated position;
[0122] (4-3) Calculate the weight of each neighbor pixel. These weights are calculated based on the relative distance between the target pixel position and the neighbor pixel position. The distance is proportional to the weight.
[0123] (4-4) Use the calculated weights to perform weighted average calculation on the values of the four neighboring pixels;
[0124] (4-5) For pixels whose calculated positions exceed the range of the source image, round them down to avoid out-of-bounds access and obtain the reconstructed image I′ n ; Finally, I′ n The target tracking model is input as the actual image to be predicted; the actual defense system does not rely on any working mechanism or calculation law of the specific target tracking model.
[0125] This embodiment also provides a target tracking and protection system based on low-frequency disturbance attack retrieval, including:
[0126] A low-frequency sampling result acquisition module is used to acquire an input video sequence, determine a central image of a sub-frame detection area of the video sequence, perform low-frequency disturbance detection through the central image of the sub-frame detection area, and obtain a low-frequency sampling result;
[0127] Target tracking model processing module: used for inputting the low-frequency sampling result into the target tracking model for processing to obtain a corrected target tracking result;
[0128] Wherein, the target tracking model processing module includes:
[0129] A disturbance detection unit, used for analyzing low-frequency coefficients and detecting low-frequency disturbance signals;
[0130] A regression analysis unit, used for performing regression analysis on the detected disturbance signal;
[0131] The model adjustment unit is used to adjust and correct the target tracking model according to the regression analysis results.
[0132] This embodiment provides a computer device, including a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of a target tracking and protection method based on low-frequency disturbance attack retrieval.
[0133] This embodiment also provides a computer-readable storage medium having a computer program / instruction stored thereon, which, when executed by a processor, implements the steps of a target tracking and protection method based on low-frequency disturbance attack retrieval.
[0134] The low-frequency perturbation attack retrieval target tracking protection method proposed in this embodiment analyzes the low-frequency perturbation signal in the input video sequence through discrete cosine transform (DCT) to effectively identify and defend against potential adversarial attacks. Experiments show that this method has significant advantages in improving the stability and robustness of the tracking system.
[0135] First, this method performs well in application scenarios such as autonomous driving and security monitoring that require high security and real-time response. Experimental data show that the system can effectively detect more than 90% of low-frequency perturbation attacks and implement defensive measures with early warning. In the 12,000-frame video test, the system's false detection rate remained below 2%, while the missed detection rate was less than 5%. This high detection rate and low false detection rate ensure that the system maintains stable tracking performance in complex environments. Secondly, through regression analysis and dynamic adjustment, it can adapt to different types of perturbation attacks and effectively improve the robustness of the model. In the face of adversarial attacks with an intensity of 0.05 to 0.1, the tracking accuracy only dropped by less than 10%, compared with a drop of more than 30% in the absence of defensive measures. This shows that it can effectively defend against different attack intensities, significantly improving the reliability and security of the visual target tracking system in practical applications.
[0136] In addition, the computational overhead in the detection and defense process is low, and the average processing time for single-frame detection is less than 10 milliseconds, which is 30% more efficient than traditional methods. The low computing resource requirements make this defense method not only suitable for high-performance computing environments, but also can run effectively in resource-constrained embedded systems, which further expands the breadth of its practical applications.
[0137] This defense method builds a comprehensive and highly integrated defense framework through a series of steps such as preprocessing of input data, discrete cosine transform, low-frequency perturbation detection, perturbation regression analysis and model adjustment. The entire defense strategy runs through the complete chain from detection to response, ensuring that judgments can be made quickly and corresponding measures can be taken when encountering adversarial attacks, effectively protecting the security and stability of the visual target tracking system.
[0138] The above are only preferred specific implementations of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by a person skilled in the art within the technical scope disclosed in the present application should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.
Claims
1. A target tracking and protection method based on low-frequency disturbance attack retrieval, characterized in that: include: Acquire an input video sequence, determine a central image of a sub-frame detection area of the video sequence, perform low-frequency disturbance detection using the central image of the sub-frame detection area, and obtain a low-frequency sampling result; Inputting the low-frequency sampling result into the target tracking model for processing to obtain a corrected target tracking result; The target tracking model quantifies the low-frequency sampling results by a linear regression method, and dynamically adjusts parameters to correct the input frame.
2. The target tracking and protection method based on low-frequency disturbance attack retrieval according to claim 1 is characterized in that: Obtaining the low-frequency sampling result includes: Converting the central image of the sub-frame detection area into a grayscale image, dividing the grayscale image into a number of identical small-block images, performing discrete cosine transform on each of the small-block images, and converting the pixel values in the spatial domain into frequency domain representation; The coefficients of the same part in each of the small-block images are selected to extract the low-frequency components, and all the low-frequency components are integrated to obtain the low-frequency sampling result.
3. The target tracking and protection method based on low-frequency disturbance attack retrieval according to claim 1 is characterized in that: The process of inputting the low-frequency sampling result into the target tracking model for processing is: A regression analysis is performed on the low-frequency sampling results to determine the nature and scope of the attack disturbance and quantify the characteristics of the disturbance signal.
4. The target tracking and protection method based on low-frequency disturbance attack retrieval according to claim 3 is characterized in that: The step of performing linear regression analysis on the low-frequency sampling results includes: Calculate the mean and standard deviation based on the low-frequency sampling results; A reference model is constructed, a threshold is set for the reference model, the mean and the standard deviation are compared with the threshold, and it is determined whether there is an abnormal disturbance signal. If there is an abnormal disturbance signal, a regression linear analysis is performed; if there is no abnormal disturbance signal, the low-frequency sampling result is input into the target tracking model as the actual image to be predicted.
5. The target tracking and protection method based on low-frequency disturbance attack retrieval according to claim 4 is characterized in that: The process of performing the regression linear analysis is: Linear regression is introduced to analyze the multi-directional detection disturbance characteristics, a disturbance characteristic curve is established, and the abnormal disturbance signal is quantified; wherein the disturbance characteristic curve is related to the offset, the disturbance slope and the disturbance error.
6. The target tracking and protection method based on low-frequency disturbance attack retrieval according to claim 3 is characterized in that: The process of inputting the low-frequency sampling result into the target tracking model for processing also includes: Through the low-frequency sampling results, the low-frequency domain repair signal is iteratively searched, specifically: Generate a two-dimensional signal based on the low-frequency sampling result, and project the two-dimensional signal to a sub-frame detection area to obtain a test frame; The time-domain-based IoU score is calculated through the test frame to evaluate the impact of the current video frame on the target tracking model, and the sample with the highest time-domain-based IoU score is selected as the actual sample to be predicted.
7. The target tracking and protection method based on low-frequency disturbance attack retrieval according to claim 6 is characterized in that: The fixed number of times is selected based on the computing power of the device to achieve the most efficient setting.
8. A target tracking and protection system based on low-frequency disturbance attack retrieval, characterized in that: include: A low-frequency sampling result acquisition module is used to acquire an input video sequence, determine a central image of a sub-frame detection area of the video sequence, perform low-frequency disturbance detection through the central image of the sub-frame detection area, and obtain a low-frequency sampling result; Target tracking model processing module: used for inputting the low-frequency sampling result into the target tracking model for processing to obtain a corrected target tracking result; Wherein, the target tracking model processing module includes: A disturbance detection unit, used for analyzing low-frequency coefficients and detecting low-frequency disturbance signals; A regression analysis unit, used for performing regression analysis on the detected disturbance signal; The model adjustment unit is used to adjust and correct the target tracking model according to the regression analysis results.
9. A computer device comprising a memory, a processor and a computer program stored in the memory, characterized in that: The processor executes the computer program to implement the steps of the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program / instruction stored thereon, characterized in that: When the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Low-altitude unmanned aerial vehicle surveillance system
CN110855936A
Optical alignment detection device and detection method thereof
CN112902835A
Detecting and tracking objects in digital images
US20110268319A1