An AI Video Low-Altitude Target Recognition and Real-Time Tracking Method Based on Deep Learning

By combining deep learning, optical flow sensing and adaptive threshold algorithms, AI video low-altitude target recognition and real-time tracking methods, the limitations of the existing technology in low-altitude target recognition and tracking in complex environments are solved, and the target tracking effect with high accuracy, low latency and robustness is achieved.

CN119723421BActive Publication Date: 2025-06-03HANGZHOU YOUXIN TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411862382.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-17
Publication Date
2025-06-03
Estimated Expiration
2044-12-17

AI Technical Summary

Technical Problem

The existing low-altitude target recognition and real-time tracking technologies have limitations in handling complex motion patterns, environmental changes and abnormal behavior detection, making it difficult to achieve high-precision and low-latency tracking effects, especially in dealing with camera jitter, background interference and abnormal movements.

Method used

A method of low-altitude target recognition and real-time tracking of AI video based on deep learning is proposed. Combined with deep learning, optical flu sensing and adaptive threshold algorithms, real-time identification and tracking of low-altitude targets are achieved through multi-level feature fusion and abnormal detection. This method improves the robustness and stability of the system through dynamic optical flow compensation and spatial and temporal multi-layer feedback optimization.

Benefits of technology

It improves the recognition accuracy and tracking stability of the target in complex environments, enhances the system's robustness to camera jitter and environmental changes, ensures the continuous tracking and positioning accuracy of the target, and has the advantages of high efficiency, precision, anti-interference and strong dynamic adaptability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119723421B_ABST
    Figure CN119723421B_ABST
Patent Text Reader

Abstract

The present invention discloses an AI video low-altitude target recognition and real-time tracking method based on deep learning, comprising the following steps: S1, collecting and processing data to construct a standardized data set; S2, separating foreground low-altitude targets to generate foreground image data; S3, calculating the inter-frame motion field using the Farneback optical flow method to obtain optical flow features and extracting convolutional features; S4, fusing the optical flow features and convolutional features; S5, inputting the fused features into a multi-scale optical flow perception network to calculate optical flow features at different scales and inputting them into a hybrid attention module to generate enhanced tracking features; S6, inputting the enhanced tracking features into an anomaly detection module to generate an anomaly detection result; S7, inputting the anomaly detection result into a dynamic optical flow compensation module to generate a compensated tracking result; S8, inputting the compensated tracking result into a spatio-temporal multi-layer feedback optimization module to optimize the optical flow parameters and the tracking window. The present invention adopts deep learning and optical flow method, etc. to realize the real-time recognition and tracking of low-altitude targets.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of deep learning and computer vision technology, and in particular to an AI video low-altitude target recognition and real-time tracking method based on deep learning. Background Art

[0002] In the existing technology, the recognition and tracking of low-altitude targets have wide applications in the fields of drone monitoring, security systems and intelligent transportation. Especially with the increase in low-altitude flying equipment and the increase in the demand for monitoring in complex environments, how to achieve accurate recognition and stable tracking of low-altitude targets in changing scenarios has become a key issue. Traditional low-altitude target recognition technology mainly relies on traditional image processing methods, combined with feature extraction and classification algorithms, to detect and identify based on static features such as color, shape, and texture of the target. However, this method has poor recognition accuracy and stability when dealing with fast-moving low-altitude targets, especially in environments with changing light, occlusion, and complex backgrounds, it is very easy to have inaccurate detection or tracking loss.

[0003] In recent years, with the development of deep learning, target detection methods based on convolutional neural networks have been widely used in low-altitude target recognition. This type of method can effectively extract high-level semantic features of the target through the hierarchical feature extraction capability of deep neural networks, thereby improving the accuracy and stability of recognition. However, deep learning technology has challenges in real-time and computational efficiency, especially in high-frame rate and long-time video processing, which requires high computing resources. In addition, it is difficult to cope with the motion changes of the target and interference factors in complex environments by relying solely on visual features in static images for target recognition. In actual scenarios, low-altitude targets are usually accompanied by violent movements, rotations, occlusions, etc., and the camera may be affected by jitter and environmental changes when collecting data in real time. These factors increase the demand for robustness and real-time performance of tracking algorithms.

[0004] To improve the accuracy of motion tracking of low-altitude targets, traditional technology introduced the optical flow method, which captures the target's motion trajectory by analyzing the motion vectors between adjacent frames. However, the optical flow method itself is easily disturbed by the background in scenes with high noise, and it is difficult to separate stable target motion features. When dealing with fast-moving low-altitude targets, the traditional optical flow method has high computational complexity and is not adaptable to large-scale and small-scale motion patterns, which leads to tracking instability or offset problems when the target moves quickly or the scene changes dramatically.

[0005] In addition, in order to enhance the stability and accuracy of low-altitude target tracking, in the existing technology, the convolutional neural network and the optical flow analysis method are usually combined, and by fusing image features and motion features, an attempt is made to improve the adaptability of target tracking. However, due to the insufficient compatibility of traditional methods with multi-scale motion patterns, it is often difficult to comprehensively describe the motion characteristics of the target only through optical flow analysis at a fixed scale. At the same time, when the existing target tracking system faces various complex environmental changes, such as sudden changes in environmental light, background motion, or violent motion of the target itself, it often lacks an adaptive adjustment mechanism, resulting in a significant decline in target tracking performance.

[0006] The current technology also has deficiencies in dealing with the abnormal motion detection of targets. During the motion process, low-altitude targets may exhibit abnormal motion behaviors due to external factors. Traditional methods use preset fixed thresholds to identify abnormal behaviors, but this method has obvious limitations in complex scenarios. Fixed thresholds are difficult to dynamically adapt to environmental changes and are prone to false alarms or missed detections. Even some adaptive detection methods can improve the detection performance in simple environments, but they are still difficult to effectively handle in complex low-altitude target tracking scenarios. Poor detection of abnormal motion not only affects the accuracy of real-time tracking but may even lead to the interruption or failure of tracking.

[0007] In addition, while achieving real-time performance and high precision, the existing technology often neglects the feedback mechanism of the system. Low-altitude target tracking in complex environments requires continuous error correction and parameter optimization, but traditional methods lack an effective feedback and adjustment mechanism. Even if some methods introduce a feedback mechanism, it is usually a simple threshold judgment and cannot perform refined correction on diverse errors in different scenarios. Long-term target tracking requires the system to be able to quickly respond and adjust parameters when tracking offsets are detected, but the existing methods lack the support of a multi-level feedback mechanism, resulting in insufficient stability and robustness of tracking.

[0008] Generally speaking, the existing low-altitude target recognition and real-time tracking technologies have limitations in dealing with complex motion patterns, environmental changes, and abnormal behavior detection. In various complex environments, it is difficult for the recognition system to achieve high-precision and low-latency tracking effects, especially when dealing with camera jitter, background interference, and abnormal motion. At the same time, the traditional feedback mechanism is insufficient and cannot achieve continuous error correction and system optimization, resulting in limitations in the stability of the system during long-term tracking tasks.

[0009] Therefore, how to provide an AI video low-altitude target recognition and real-time tracking method based on deep learning is an urgent problem that needs to be solved by those skilled in the art. Summary of the Invention

[0010] An object of the present invention is to propose an AI video low-altitude target recognition and real-time tracking method based on deep learning. The present invention combines deep learning, optical flow perception and adaptive threshold algorithm to achieve real-time recognition and tracking of low-altitude targets. Through multi-level feature fusion and anomaly detection, the recognition accuracy and tracking stability of targets in complex environments are effectively improved. Through dynamic optical flow compensation and spatio-temporal multi-layer feedback optimization, the robustness of the system to camera jitter and environmental changes is further enhanced, ensuring continuous tracking and positioning accuracy of the target, and having the advantages of high efficiency, accuracy, anti-interference and strong dynamic adaptability, and is applicable to various application scenarios such as intelligent monitoring and navigation of low-altitude targets.

[0011] A method for AI video low-altitude target recognition and real-time tracking based on deep learning according to an embodiment of the present invention includes the following steps:

[0012] S1. Collect image and video data of low-altitude targets, perform denoising, frame rate adjustment and resolution standardization processing, and construct a standardized data set based on data augmentation technology;

[0013] S2. Input the standardized data set into an adaptive Gaussian mixture model for background modeling and segmentation, separate foreground low-altitude targets, and generate foreground image data;

[0014] S3. Based on the foreground image data, use the Farneback optical flow method to calculate the inter-frame motion field to obtain optical flow features, and use a convolutional neural network to extract convolutional features of the image to generate a foreground motion and visual feature matrix;

[0015] S4. Input the foreground motion and visual feature matrix into the feature fusion module in the convolutional neural network, fuse the optical flow features and convolutional features, and generate fused features;

[0016] S5. Input the fused features into a multi-scale optical flow perception network, calculate optical flow features at different scales, and input them into a hybrid attention module to focus on the significant displacement area and predict the future position of the target, generating enhanced tracking features;

[0017] S6. Input the enhanced tracking features into an anomaly detection module, and use an adaptive threshold to detect abnormal motion behaviors based on the optical flow vector field and historical trajectory information to generate anomaly detection results;

[0018] S7. Input the anomaly detection results into a dynamic optical flow compensation module to perform real-time compensation for camera jitter, and combine a recurrent neural network to predict the target time series information to generate a compensated tracking result;

[0019] S8. Input the compensated tracking results into a spatio-temporal multi-layer feedback optimization module to perform inter-frame error correction and periodic adjustment, and optimize the optical flow parameters and tracking window when the detected deviation exceeds the threshold.

[0020] Optionally, S2 specifically includes:

[0021] S21. Input each frame of image in the standardized dataset into an adaptive Gaussian mixture model for background modeling, and establish an initial background model for each pixel. The initial background model is represented in the form of a mixture of multiple Gaussian distributions. The number of initial Gaussian distributions is set to N, and the probability density of each pixel value is expressed as:

[0022] ;

[0023] Wherein, represents the probability density function that the pixel value x belongs to the background, represents the weight of the k-th Gaussian distribution, satisfying ; represents the mean of the k-th Gaussian distribution, represents the standard deviation of the k-th Gaussian distribution, and exp represents the exponential function;

[0024] S22. Perform background matching on each pixel value x to determine whether the pixel value x matches a certain Gaussian distribution in the background model. The matching condition is that the pixel value x and the mean of a certain Gaussian distribution are within a certain range, that is, satisfying:

[0025] ;

[0026] If the condition is satisfied, it is considered that the current pixel value x matches a certain Gaussian distribution;

[0027] S23. For the matching pixel value x, update the mean, variance, and weight of the Gaussian distribution that matches the pixel value x;

[0028] S24. If the current pixel value x fails to match any Gaussian distribution in the background model, it is regarded as a foreground pixel, and the background model is replaced. The Gaussian distribution with the smallest weight is replaced with a new distribution with the current pixel value x as the mean, and the initial standard deviation is set to , and the smallest weight is given;

[0029] S25. Sort the Gaussian distributions in the background model from largest to smallest according to the weights, and determine the number of background distributions B with a weight threshold T, satisfying the condition:

[0030] ;

[0031] The Gaussian distributions that meet the conditions are set as the background, and the remaining Gaussian distributions are used as the foreground;

[0032] S26. Generate binary foreground image data according to the division results of the background distribution and the foreground distribution, mark the pixels in the foreground area as foreground, and output the foreground image data.

[0033] Optionally, the specific steps of S3 are as follows:

[0034] S31. Based on the foreground image data, input two consecutive frames of images into the Farneback optical flow algorithm, construct a polynomial expansion model for each frame of the image to represent the local neighborhood changes of each pixel point, and identify the movement of pixels between image frames.

[0035] S32. Find the similarity regions in the front and back frames of images through the polynomial expansion model, match the corresponding relationships of pixel blocks in the current frame and the adjacent frame, and calculate the motion vectors of the pixel blocks.

[0036] S33. In the Farneback optical flow method, use a hierarchical pyramid structure to decompose the image into multiple resolution levels, estimate the optical flow starting from the lower resolution level and refine it layer by layer upward to capture the large-scale overall motion and the small-scale motion at the detail level.

[0037] S34. During the hierarchical calculation process, calculate the position offset of each pixel block to generate an inter-frame motion field. The motion field represents the moving direction and speed of pixels between two frames, and extract the motion information as optical flow features to reflect the motion pattern of the target.

[0038] S35. Convert the obtained optical flow feature matrix into a foreground motion feature matrix. The matrix includes the displacement information and direction information of each pixel, and is used to accurately describe the motion trend of the target in space.

[0039] S36. Input the foreground image data into a convolutional neural network, extract the convolutional features of the image through multiple convolutional and pooling operations, and combine them with the foreground motion feature matrix to generate a foreground motion and visual feature matrix.

[0040] Optionally, the specific steps of S5 are as follows:

[0041] S51. Input the fusion features into a multi-scale optical flow perception network. The multi-scale optical flow perception network constructs a multi-layer convolutional structure to perform optical flow calculations on the fusion features at different scales and generate optical flow feature maps with different resolutions.

[0042] S52. In the low-scale layer, the multi-scale optical flow perception network is used to capture large-scale motion features; in the high-scale layer, the multi-scale optical flow perception network is used to extract fine motion information and generate optical flow features at different scales to adapt to various motion patterns of the target.

[0043] S53. Standardize the multi-scale optical flow feature maps to unify the value ranges of each optical flow feature map into the same interval. After standardization, the optical flow feature maps include the motion direction and speed information of the target at each scale, facilitating the collaborative analysis of features at different scales.

[0044] S54. Input the standardized multi-scale optical flow feature maps into a hybrid attention module. The hybrid attention module includes an optical flow attention mechanism and a spatio-temporal attention mechanism. The optical flow attention mechanism focuses attention based on the displacement regions of significant motions in the optical flow feature maps, highlighting the significant motion regions of the target with high weights and ignoring the background or unimportant regions with low weights.

[0045] S55. The spatio-temporal attention mechanism receives the time series information in the multi-scale optical flow feature maps, analyzes the continuous frame position changes of the target, captures the motion trajectory trend of the target, and forms prediction and positioning information by analyzing the possible next frame position of the target through time series features.

[0046] S56. Perform weighted fusion on the results of the optical flow attention mechanism and the spatio-temporal attention mechanism to generate enhanced tracking features. The enhanced tracking features include attention weighted information of significant motion regions and spatio-temporal prediction positions, and are used to dynamically identify the core motion regions of the target in complex environments.

[0047] Optionally, S6 specifically includes:

[0048] S61. Input the enhanced tracking features into an anomaly detection module, extract the optical flow vector field of each frame of the image, obtain the motion vector information of each pixel, where the motion vector represents the horizontal displacement and vertical displacement of the pixel, forming an optical flow vector field matrix. The optical flow vector field matrix represents the motion direction and speed of the target between consecutive frames.

[0049] S62. Calculate the instantaneous motion speed V of the target based on the optical flow vector field matrix, using the Euclidean distance calculation formula:

[0050] ;

[0051] where M represents the total number of pixel points covered by the target within the frame, represents the horizontal displacement of the i-th pixel, represents the vertical displacement of the i-th pixel;

[0052] S63. Further analyze the instantaneous motion speed V and the motion direction of the target, and calculate the motion direction :

[0053] ;

[0054] S64. Establish a set of motion trajectory data of the target based on the motion information of historical frames , where represents the speed of the target in the t-th frame, represents the direction of the target in the t-th frame; perform time series analysis on the set of motion trajectory data to obtain the motion trend and trajectory changes;

[0055] S65. Conduct trend analysis on the motion trajectory of the target and calculate the speed change rate and the direction change rate :

[0056] ;

[0057] ;

[0058] where represents the speed of the target in the (t - 1)-th frame, represents the direction of the target in the (t - 1)-th frame;

[0059] S66. Define adaptive thresholds and , and dynamically adjust the thresholds according to the historical motion data of the target and environmental changes; when or is satisfied, mark the current frame as abnormal; the dynamic adjustment of the adaptive thresholds and is based on environmental fluctuations, target acceleration, and direction change rate to meet the adaptability of detection;

[0060] S67. If multiple consecutive frames meet the abnormal detection conditions, mark the target as abnormal motion, generate an abnormal detection result, and the abnormal detection result includes information such as the number of abnormal frames, speed change rate, and direction change rate, and finally output the abnormal detection result.

[0061] Optionally, the specific steps of S7 include:

[0062] S71. Input the abnormal detection result into the dynamic optical flow compensation module, analyze the motion characteristics of the target marked in the abnormal detection and the motion information in the optical flow vector field, identify the global motion offset caused by camera jitter, and initialize compensation parameters to correct the position deviation caused by camera jitter;

[0063] S72. Based on the initialized compensation parameters, perform position compensation on the tracking target of the current frame, apply the preliminarily calculated jitter compensation amount to the target position of the current frame, and adjust the coordinates of the target position in real time;

[0064] S73. Input the compensated target position into the recurrent neural network module. Use the recurrent neural network module to perform time series analysis on the historical motion data of the target, identify the motion patterns of the target in the past multiple frames, and predict the possible position of the target in the next frame.

[0065] S74. Fuse the predicted position output by the recurrent neural network module with the compensated position. By setting a fusion strategy, combine the jitter compensation result with the time series prediction result to generate a fused tracking position.

[0066] S75. According to the fused tracking position, adaptively adjust the target tracking window in the current frame. Adjust the size and position of the tracking window according to the actual motion situation of the target.

[0067] S76. Based on the tracking position after final compensation and prediction adjustment, generate a compensated tracking result.

[0068] Optionally, the specific content of S8 includes:

[0069] S81. Input the compensated tracking result into the spatio-temporal multi-layer feedback optimization module, compare the target positions in the current frame and the previous frame, and calculate the inter-frame error. Let the position of the current frame be and the position of the previous frame be : ;

[0070] where E represents the inter-frame error.

[0071] S82. Compare the inter-frame error with a preset threshold. When the inter-frame error exceeds the preset threshold, trigger a short-term error correction mechanism to adjust the position of the tracking frame in the current frame in real time and update the target position.

[0072] S83. Cumulatively analyze the error values of each frame, construct an error sequence, and perform periodic analysis on the error sequence to identify the long-term deviation trend. Evaluate the change of the overall tracking accuracy by calculating the mean and variance of the error sequence.

[0073] S84. If the periodic analysis result shows that the inter-frame error accumulates gradually and the deviation trend is significant, optimize the convolution kernel size and stride in the optical multi-scale optical flow perception network and dynamically adjust the optical flow parameters.

[0074] S85. After detecting the long-term inter-frame error accumulation, adaptively adjust the size and position of the tracking window:

[0075] ;

[0076] ;

[0077] where represents the adjusted window size, represents the current window size, and represents the window adjustment coefficient, represents the mean of the error sequence, represents the adjusted window position, represents the current window position, and represents the position adjustment coefficient, represents the variance of the error sequence, and tanh represents the hyperbolic tangent function;

[0078] S86. Apply the optimized optical flow parameters and tracking window configuration to the tracking result of the current frame, record the error correction result and parameter adjustment information, and form a feedback database.

[0079] The beneficial effects of the present invention are as follows:

[0080] First of all, the present invention combines multi-scale feature fusion of deep learning and optical flow analysis, enabling more accurate extraction and tracking of object features in different motion modes. Traditional methods usually analyze at a single scale, while the present invention can better handle the problems of fast movement of objects and tracking of objects of different sizes by introducing a multi-scale optical flow model. This method can accurately capture the dynamic changes of the object in the image, thereby improving the tracking stability of low-altitude objects in the cases of fast movement, rotation, and occlusion. This multi-scale fusion strategy can effectively avoid the problems of object loss and false detection easily encountered by traditional methods when dealing with the movement of low-altitude objects in complex backgrounds.

[0081] Secondly, the present invention adopts a deep learning model based on a hybrid attention mechanism, which can automatically focus on the key regions and dynamic changes in the image when processing multi-modal inputs. Compared with traditional object recognition methods, the attention mechanism can more intelligently screen out the important features related to the object and avoid the interference of redundant information. Especially in the cases where there is a strong contrast between the object and the background, fast movement of low-altitude objects, or high-noise environment, the attention mechanism can help the system better focus on the object area and improve the accuracy of object recognition and tracking.

[0082] In addition, the abnormal behavior detection module of the present invention adopts an adaptive threshold strategy, which can dynamically adjust the detection standard according to the changes in the actual environment. The fixed threshold method in traditional technologies often has difficulty dealing with the ever-changing actual scenarios, while the present invention can make the detection of abnormal movements more accurate by intelligently adjusting the threshold. Whether it is the severe jitter of low-altitude objects or the trajectory anomalies during high-speed movement, the present invention can achieve accurate recognition, reduce the occurrence of false alarms and missed detections, and thus ensure the continuity and stability of object tracking.

[0083] In terms of real-time performance, the spatio-temporal multi-layer feedback optimization mechanism proposed by the present invention effectively enhances the self-correction ability of the system. Different from the static tracking algorithms in traditional methods, the present invention can perform error correction in real time when the target motion trajectory deviates. Through a multi-level feedback mechanism, the system can adaptively adjust the algorithm parameters according to the changes between the front and back frames, quickly respond when the target is lost or the trajectory deviates, and re-guide the tracking of the target. This mechanism not only improves the real-time performance of the system, but also greatly enhances its stability in long-term tracking tasks.

[0084] In addition, the present invention significantly reduces the demand for computing resources through an optimized deep learning model and an optical flow analysis algorithm. Traditional deep learning algorithms usually require a large amount of computing resources and time for target detection and tracking, while the present invention can effectively reduce the computational complexity while maintaining high accuracy, and adapt to more application scenarios with high real-time requirements. BRIEF DESCRIPTION OF THE DRAWINGS

[0085] The drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention, and do not constitute a limitation to the present invention. In the drawings:

[0086] Figure 1 is the overall flowchart of an AI video low-altitude target recognition and real-time tracking method based on deep learning proposed by the present invention;

[0087] Figure 2 is the structural schematic diagram of an AI video low-altitude target recognition and real-time tracking method based on deep learning proposed by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0088] Now, the present invention will be further described in detail with reference to the drawings. These drawings are all simplified schematic diagrams, only showing the basic structure of the present invention in a schematic way, so they only show the components related to the present invention.

[0089] Referring to Figure 1 and Figure 2 , an AI video low-altitude target recognition and real-time tracking method based on deep learning includes the following steps:

[0090] S1. Collect image and video data of low-altitude targets, perform denoising, frame rate adjustment and resolution standardization processing, and construct a standardized data set based on data augmentation technology;

[0091] S2. Input the standardized data set into an adaptive Gaussian mixture model for background modeling and segmentation, separate the foreground low-altitude targets, and generate foreground image data;

[0092] S3. Based on the foreground image data, use the Farneback optical flow method to calculate the inter-frame motion field to obtain optical flow features, and use a convolutional neural network to extract the convolutional features of the image to generate a foreground motion and visual feature matrix;

[0093] S4. Input the foreground motion and visual feature matrix into the feature fusion module in the convolutional neural network to fuse the optical flow features and convolutional features to generate fused features;

[0094] S5. Input the fused features into the multi-scale optical flow perception network, calculate the optical flow features at different scales, and input them into the hybrid attention module to focus on the significant displacement regions and predict the future position of the target to generate enhanced tracking features;

[0095] S6. Input the enhanced tracking features into the anomaly detection module, and based on the optical flow vector field and historical trajectory information, use an adaptive threshold to detect abnormal motion behaviors to generate anomaly detection results;

[0096] S7. Input the anomaly detection results into the dynamic optical flow compensation module to perform real-time compensation for camera jitter, and combine a recurrent neural network to predict the target temporal information to generate compensated tracking results;

[0097] S8. Input the compensated tracking results into the spatio-temporal multi-layer feedback optimization module to perform inter-frame error correction and periodic adjustment, and optimize the optical flow parameters and tracking window when the detected deviation exceeds the threshold.

[0098] In this embodiment, the S2 specifically includes:

[0099] S21. Input each frame of image in the standardized dataset into the adaptive Gaussian mixture model for background modeling to establish an initial background model for each pixel. The initial background model is represented in the form of a mixture of multiple Gaussian distributions. The number of initial Gaussian distributions is set to N, and the probability density of each pixel value is expressed as:

[0100] ;

[0101] Among them, represents the probability density function that the pixel value x belongs to the background, represents the weight of the k-th Gaussian distribution, satisfying ; represents the mean of the k-th Gaussian distribution, represents the standard deviation of the k-th Gaussian distribution, and exp represents the exponential function;

[0102] S22. Perform background matching for each pixel value x to determine whether the pixel value x matches a certain Gaussian distribution in the background model. The matching condition is that the pixel value x and the mean of a certain Gaussian distribution are within a certain range, that is, satisfying:

[0103] ;

[0104] If the condition is satisfied, it is considered that the current pixel value x matches a certain Gaussian distribution;

[0105] S23. For the matched pixel value x, update the mean, variance, and weight of the Gaussian distribution that matches the pixel value x;

[0106] S24. If the current pixel value x fails to match any Gaussian distribution in the background model, it is regarded as a foreground pixel, and the background model is replaced. The Gaussian distribution with the smallest weight is replaced with a new distribution with the current pixel value x as the mean, and the initial standard deviation is set to , and the smallest weight is assigned ;

[0107] S25. Sort the Gaussian distributions in the background model from largest to smallest by weight, and determine the number of background distributions B with a weight threshold T, satisfying the condition:

[0108] ;

[0109] Set the Gaussian distributions that meet the conditions as the background, and the remaining Gaussian distributions as the foreground;

[0110] S26. Generate binary foreground image data according to the division result of the background distribution and the foreground distribution, mark the pixels in the foreground area as the foreground, and output the foreground image data.

[0111] In this embodiment, the specific steps of S3 include:

[0112] S31. Based on the foreground image data, input two consecutive frames of images into the Farneback optical flow algorithm, construct a polynomial expansion model for each frame of the image, represent the local neighborhood changes of each pixel point, and identify the movement of the pixels between the image frames;

[0113] S32. Find the similarity regions in the front and back two frames of images through the polynomial expansion model, match the corresponding relationship of the pixel blocks in the current frame and the adjacent frames, and calculate the motion vectors of the pixel blocks;

[0114] S33. In the Farneback optical flow method, use a hierarchical pyramid structure to decompose the image into multiple resolution levels, estimate the optical flow starting from the lower resolution level and refine it layer by layer upward, capturing the large-scale overall motion and the small-scale motion of the detail levels;

[0115] S34. During the hierarchical calculation process, calculate the position offset of each pixel block to generate an inter-frame motion field. The motion field represents the moving direction and speed of pixels between two frames, and extract the motion information as optical flow features to reflect the motion pattern of the target.

[0116] S35. Convert the obtained optical flow feature matrix into a foreground motion feature matrix. The matrix includes the displacement information and direction information of each pixel, and is used to accurately describe the motion trend of the target in space.

[0117] S36. Input the foreground image data into a convolutional neural network, extract the convolutional features of the image through multiple convolutional and pooling operations, and combine them with the foreground motion feature matrix to generate a foreground motion and visual feature matrix.

[0118] In this embodiment, the specific steps of S5 are as follows:

[0119] S51. Input the fusion features into a multi-scale optical flow perception network. The multi-scale optical flow perception network constructs a multi-layer convolutional structure to perform optical flow calculation on the fusion features at different scales, and generates optical flow feature maps with different resolutions.

[0120] S52. In the low-scale layer, the multi-scale optical flow perception network is used to capture large-range motion features; in the high-scale layer, the multi-scale optical flow perception network is used to extract fine motion information, and generate optical flow features at different scales to adapt to various motion patterns of the target.

[0121] S53. Perform normalization processing on the multi-scale optical flow feature maps to unify the value ranges of the optical flow feature maps into the same interval; after normalization processing, the optical flow feature maps include the motion direction and speed information of the target at each scale, which is convenient for collaborative analysis of features at different scales.

[0122] S54. Input the normalized multi-scale optical flow feature maps into a hybrid attention module. The hybrid attention module includes an optical flow attention mechanism and a spatio-temporal attention mechanism. The optical flow attention mechanism focuses attention based on the displacement regions of significant motion in the optical flow feature maps, highlighting the significant motion regions of the target with high weights and ignoring the background or unimportant regions with low weights.

[0123] S55. The spatio-temporal attention mechanism receives the time series information in the multi-scale optical flow feature maps, analyzes the continuous frame position changes of the target, captures the motion trajectory trend of the target, and forms prediction and positioning information by analyzing the possible next frame position of the target through time series features.

[0124] S56. Weightedly fuse the results of the optical flow attention mechanism and the spatio-temporal attention mechanism to generate enhanced tracking features, where the enhanced tracking features include attention-weighted information of significant motion regions and spatio-temporal predicted positions, and are used to dynamically identify the core motion regions of the target in complex environments.

[0125] In this embodiment, the S6 specifically includes:

[0126] S61. Input the enhanced tracking features into an anomaly detection module, extract the optical flow vector field of each frame of image to obtain the motion vector information of each pixel. The motion vector represents the horizontal displacement and vertical displacement of the pixel, forming an optical flow vector field matrix, and the optical flow vector field matrix represents the motion direction and speed of the target between consecutive frames;

[0127] S62. Calculate the instantaneous motion speed V of the target based on the optical flow vector field matrix, using the Euclidean distance calculation formula:

[0128] ;

[0129] where M represents the total number of pixel points covered by the target within the frame, represents the horizontal displacement of the i-th pixel, represents the vertical displacement of the i-th pixel;

[0130] S63. Further analyze the instantaneous motion speed V and the motion direction of the target, and calculate the motion direction :

[0131] ;

[0132] S64. Establish a set of motion trajectory data of the target based on the motion information of historical frames, where represents the speed of the target in the t-th frame, represents the direction of the target in the t-th frame; perform time series analysis on the set of motion trajectory data to obtain the motion trend and trajectory changes;

[0133] S65. Analyze the trend of the target's motion trajectory, and calculate the speed change rate and the direction change rate :

[0134] ;

[0135] ;

[0136] where, represents the speed of the target in the (t - 1)-th frame, Represents the direction of the target in the (t - 1)-th frame;

[0137] S66. Define an adaptive threshold and , and dynamically adjust the threshold according to the historical motion data of the target and environmental changes; when or is satisfied, mark the current frame as abnormal; the dynamic adjustment of the adaptive threshold and is based on environmental fluctuations, target acceleration, and direction change rate, meeting the self - adaptability of detection;

[0138] S67. If multiple consecutive frames meet the abnormal detection conditions, mark the target as having abnormal motion, generate an abnormal detection result, where the abnormal detection result includes information such as the number of abnormal frames, speed change rate, and direction change rate, and finally output the abnormal detection result.

[0139] In this embodiment, the specific steps of S7 are as follows:

[0140] S71. Input the abnormal detection result into the dynamic optical flow compensation module, analyze the motion characteristics of the target marked in the abnormal detection and the motion information in the optical flow vector field, identify the global motion offset caused by camera jitter, and initialize compensation parameters to correct the position deviation caused by camera jitter;

[0141] S72. Based on the initialized compensation parameters, perform position compensation on the tracking target of the current frame, apply the preliminarily calculated jitter compensation amount to the target position of the current frame, and adjust the coordinates of the target position in real - time;

[0142] S73. Input the compensated target position into the recurrent neural network module, use the recurrent neural network module to perform time - series analysis on the historical motion data of the target, identify the motion pattern of the target in the past multiple frames, and predict the possible position of the target in the next frame;

[0143] S74. Fuse the predicted position output by the recurrent neural network module with the compensated position, and by setting a fusion strategy, combine the jitter compensation result with the time - series prediction result to generate a fused tracking position;

[0144] S75. According to the fused tracking position, adaptively adjust the target tracking window of the current frame, and adjust the size and position of the tracking window according to the actual motion situation of the target;

[0145] S76. Generate a compensated tracking result based on the finally compensated and predicted - adjusted tracking position.

[0146] In this embodiment, the specific steps of S8 are as follows:

[0147] S81. Input the compensated tracking result into the spatio-temporal multi-layer feedback optimization module, compare the target positions of the current frame and the previous frame, and calculate the inter-frame error. Let the position of the current frame be and the position of the previous frame be : ;

[0148] where E represents the inter-frame error;

[0149] S82. Compare the inter-frame error with a preset threshold. When the inter-frame error exceeds the preset threshold, trigger the short-term error correction mechanism, adjust the position of the tracking box of the current frame in real time, and update the target position;

[0150] S83. Cumulatively analyze the error values of each frame, construct an error sequence, and perform a periodic analysis on the error sequence to identify the deviation trend of the long period. Evaluate the change in the overall tracking accuracy by calculating the mean and variance of the error sequence;

[0151] S84. If the result of the periodic analysis shows that the inter-frame error accumulates gradually and the deviation trend is significant, optimize the convolution kernel size and stride in the optical multi-scale optical flow perception network, and dynamically adjust the optical flow parameters;

[0152] S85. After detecting the long-period inter-frame error accumulation, adaptively adjust the size and position of the tracking window:

[0153] ;

[0154] ;

[0155] where, represents the adjusted window size, represents the current window size, and represents the window adjustment coefficient, represents the mean of the error sequence, represents the adjusted window position, represents the current window position, and represents the position adjustment coefficient, represents the variance of the error sequence, and tanh represents the hyperbolic tangent function;

[0156] S86. Apply the optimized optical flow parameters and tracking window configuration to the tracking result of the current frame, record the error correction result and parameter adjustment information, and form a feedback database.

[0157] Example 1:

[0158] To verify the feasibility of the present invention in implementation, the present invention is applied to the low-altitude airspace monitoring of an airport. Video surveillance cameras are set around the runway, apron and flight area to capture target data in the airspace all-weather. The output of the video surveillance system is transmitted to the analysis system in real time, including continuous video frame data. Based on the method of the present invention, the system processes the acquired video data, quickly identifies the targets and tracks them in real time, ensuring that the monitoring personnel can understand the changes in the airspace without interruption.

[0159] In this embodiment, the system first collects on-site image and video data through high-resolution video surveillance equipment. These data are processed by a preprocessing module for noise removal, frame rate adjustment and resolution standardization. Subsequently, the original data is enhanced by a data enhancement module, including rotation, mirroring, brightness adjustment, etc., to improve the diversity of the data and the robustness of the model. The preprocessed and enhanced image data is input into the background modeling module, which models the background through a Gaussian mixture model and performs foreground segmentation to extract the moving target area.

[0160] Next, the system extracts and fuses the features of the targets in the video through an optical flow calculation and convolutional neural network feature extraction module. The optical flow calculation uses the Farneback method to calculate the optical flow of the target area to obtain the motion information of the target. The convolutional neural network extracts the depth features of the target. After these features are combined with the optical flow information, multi-scale tracking and motion estimation are performed in the target tracking module.

[0161] In the low-altitude target tracking, the system calculates the motion trajectory of the target in real time, and continuously tracks the position and state of the target through a multi-scale optical flow perception module with a hybrid attention mechanism. If an abnormality occurs to the target, the system will immediately trigger the abnormality detection module to compare the target trajectory with the optical flow information to determine whether the target has deviated unexpectedly or there is a potential danger.

[0162] Finally, the system analyzes the feedback information on the target motion according to the spatio-temporal feedback optimization module to optimize the trajectory estimation and tracking accuracy of the target. When an error occurs in the target trajectory, the spatio-temporal feedback optimization compensates the tracking process to ensure the accuracy and stability of the target tracking.

[0163] During the implementation process, the low-altitude airspace of this airport is used as the test scenario. The test period is from June 1, 2024 to June 7, 2024, and the test location is the runway area and flight airspace of this airport. During the test period, the system processed approximately 150 hours of video data, covering target data under different weather conditions (sunny, cloudy, light rain, etc.) and different flight conditions (unmanned aerial vehicles, light aircraft, etc.).

[0164] Table 1 Comparison Table of Low-Altitude Target Recognition and Tracking Performance between the Method of the Present Invention and Traditional Monitoring Methods

[0165]

[0166] Table 1 shows the accuracy and real-time performance of the system in recognizing and tracking different low-altitude targets during the test period. By comparison with the traditional monitoring system, the method of the present invention shows significant advantages in aspects such as target recognition accuracy, tracking precision, response time, etc.

[0167] First of all, in terms of target recognition accuracy, the method of the present invention is 95.3%, significantly higher than 85.4% of the traditional method, indicating that the present invention can more accurately recognize low-altitude targets, especially performing excellently in complex environments. In terms of tracking precision, the present invention reaches 0.12 meters, with a 65.7% improvement in precision compared to 0.35 meters of the traditional method, showing that the present invention can more precisely track targets in a dynamic environment and reduce errors. In terms of real-time response time, the 80-millisecond response time of the present invention is improved by 63.6% compared to 220 milliseconds of the traditional method, enabling the system to quickly respond to target changes and improve the monitoring effect. In terms of the detection rate of abnormal targets, the present invention is 98.7%, significantly higher than 80.2% of the traditional method, indicating that the system can more sensitively identify abnormal targets and give early warnings of potential risks. Finally, in terms of the target tracking interruption rate, the present invention is only 2.4%, far lower than 15.6% of the traditional method, improving the stability and continuity of the system.

[0168] Table 2 Data Table of Low-Altitude Target Recognition and Tracking Performance under Different Weather Conditions

[0169]

[0170] According to the data in Table 2, the AI video low-altitude target recognition and real-time tracking method of the present invention shows excellent stability and reliability under different weather conditions. Under sunny, cloudy, and light rain conditions, the target recognition accuracy remains above 93%, with a maximum of 94.8%. Especially in the case of weak light or precipitation, the system can still accurately recognize low-altitude targets, showing strong adaptability. In terms of tracking precision, regardless of the weather conditions, the precision always remains between 0.13 meters and 0.15 meters, indicating that the system has stable and precise target tracking ability in complex environments, especially performing no less well in light rain weather. In terms of abnormal target detection, the detection rate of the system is above 96%, showing its efficient abnormal detection ability under different weather conditions and being able to respond to potential safety risks in a timely manner. The target tracking interruption rate is always lower than 3.2%, proving that the system can effectively reduce tracking loss and ensure the continuity of the monitoring process.

[0171] In this embodiment, by applying the AI video low-altitude target recognition and real-time tracking method of the present invention in the actual low-altitude target monitoring scenario, its excellent performance under complex weather conditions is demonstrated. The system uses deep learning algorithms to successfully identify and track low-altitude targets, and can maintain high-precision target detection and motion prediction under various environmental conditions. Experimental data show that this method has significant advantages in improving the accuracy of target recognition and the stability of tracking, can effectively cope with the challenges of complex environments, and provides an innovative solution for the field of low-altitude target monitoring.

[0172] The above are only the preferred specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, making equivalent substitutions or changes, shall be covered by the protection scope of the present invention.

Claims

1. A method for low-altitude target recognition and real-time tracking based on AI video based on deep learning, characterized in that: The steps include: S1. Collect image and video data of low-altitude targets, perform denoising, frame rate adjustment and resolution standardization, and build a standardized data set based on data enhancement technology; S2, inputting the standardized data set into an adaptive Gaussian mixture model to perform background modeling and segmentation, separate foreground low-altitude targets, and generate foreground image data; S3, based on the foreground image data, the Farneback optical flow method is used to calculate the inter-frame motion field to obtain the optical flow features, and the convolutional neural network is used to extract the convolution features of the image to generate the foreground motion and visual feature matrix; S4, inputting the foreground motion and visual feature matrix into a feature fusion module in a convolutional neural network, fusing the optical flow feature with the convolution feature, and generating a fusion feature; S5, inputting the fusion feature into a multi-scale optical flow perception network, calculating optical flow features of different scales, and inputting the fusion feature into a hybrid attention module, focusing on the significant displacement area and predicting the future position of the target, and generating enhanced tracking features; S6, inputting the enhanced tracking features into an anomaly detection module, detecting abnormal motion behaviors using an adaptive threshold based on the optical flow vector field and historical trajectory information, and generating an anomaly detection result; S7, inputting the abnormal detection result into a dynamic optical flow compensation module to perform real-time compensation for camera shake, combining a recurrent neural network to predict the target timing information, and generating a compensated tracking result; S8, inputting the compensated tracking result into a spatiotemporal multi-layer feedback optimization module to perform inter-frame error correction and periodic adjustment, and optimizing the optical flow parameters and tracking window when it is detected that the deviation exceeds a threshold.

2. According to the AI ​​video low-altitude target recognition and real-time tracking method based on deep learning according to claim 1, it is characterized in that: The S2 specifically includes: S21, input each frame image in the standardized data set into an adaptive Gaussian mixture model for background modeling, and establish an initial background model for each pixel. The initial background model is expressed as a mixture of multiple Gaussian distributions. The number of initial Gaussian distributions is set to N, and the probability density of each pixel value is expressed as: ; in, Represents the probability density function of the pixel value x belonging to the background, represents the weight of the kth Gaussian distribution, satisfying ; represents the mean of the kth Gaussian distribution, represents the standard deviation of the kth Gaussian distribution, and exp represents the exponential function; S22, perform background matching on each pixel value x to determine whether the pixel value x matches a Gaussian distribution in the background model. The matching condition is that the mean of the pixel value x and a Gaussian distribution is within a certain range, that is, it satisfies: ; If the conditions are met, the current pixel value x is considered to match a certain Gaussian distribution; S23. For the matched pixel value x, update the mean, variance and weight of the Gaussian distribution matching the pixel value x; S24. If the current pixel value x fails to match any Gaussian distribution in the background model, it is regarded as a foreground pixel and the background model is replaced by replacing the Gaussian distribution with the minimum weight with a new distribution with the current pixel value x as the mean and setting the initial standard deviation to , giving the minimum weight ; S25. Sort the Gaussian distributions in the background model by weight from large to small, and determine the number of background distributions B by weight threshold T, satisfying the condition: ; The Gaussian distribution that meets the conditions is set as the background, and the remaining Gaussian distributions are set as the foreground; S26, generating binary foreground image data according to the division result of the background distribution and the foreground distribution, marking the pixels in the foreground area as foreground and outputting the foreground image data.

3. According to the deep learning-based AI video low-altitude target recognition and real-time tracking method of claim 1, it is characterized in that: The S3 specifically includes: S31, based on the foreground image data, input two consecutive frames of images into the Farneback optical flow algorithm, construct a polynomial expansion model for each frame of image, represent the local neighborhood change of each pixel point, and identify the movement of pixels between image frames; S32, searching for similar regions in the two previous and next frames of images through a polynomial expansion model, matching the corresponding relationship between pixel blocks in the current frame and the adjacent frame, and calculating the motion vector of the pixel block; S33, in the Farneback optical flow method, a hierarchical pyramid structure is used to decompose the image into multiple resolution levels, and the optical flow is estimated starting from the lower resolution level and refined upward layer by layer to capture large-scale overall motion and small-scale motion at the detail level; S34, in the hierarchical calculation process, calculating the position offset of each pixel block, generating an inter-frame motion field, wherein the motion field represents the moving direction and speed of the pixel between two frames, and extracting the motion information as an optical flow feature to reflect the motion pattern of the target; S35, converting the acquired optical flow feature matrix into a foreground motion feature matrix, wherein the matrix includes displacement information and direction information of each pixel, and is used to accurately describe the motion trend of the target in space; S36, inputting the foreground image data into a convolutional neural network, extracting the convolution features of the image through multi-layer convolution and pooling operations, and combining them with the foreground motion feature matrix to generate a foreground motion and visual feature matrix.

4. According to the AI ​​video low-altitude target recognition and real-time tracking method based on deep learning in claim 1, it is characterized in that: The S5 specifically includes: S51, inputting the fused features into a multi-scale optical flow perception network, wherein the multi-scale optical flow perception network constructs a multi-layer convolution structure, performs optical flow calculations on the fused features at different scales, and generates optical flow feature maps with different resolutions; S52. In the low-scale layer, the multi-scale optical flow perception network is used to capture a wide range of motion features; in the high-scale layer, the multi-scale optical flow perception network is used to extract fine motion information and generate optical flow features of different scales to adapt to various motion modes of the target; S53, performing standardization processing on the multi-scale optical flow feature map, unifying the value range of each optical flow feature map into the same interval; after the standardization processing, the optical flow feature map includes the movement direction and speed information of the target at each scale, which is convenient for the collaborative analysis of features at different scales; S54, inputting the standardized multi-scale optical flow feature map into a hybrid attention module, wherein the hybrid attention module includes an optical flow attention mechanism and a spatiotemporal attention mechanism, wherein the optical flow attention mechanism focuses attention based on the displacement area of ​​significant motion in the optical flow feature map, highlights the target significant motion area with a high weight, and ignores the background or unimportant areas with a low weight; S55, the spatiotemporal attention mechanism receives the time series information in the multi-scale optical flow feature map, analyzes the continuous frame position change of the target, captures the movement trajectory trend of the target, analyzes the possible next frame position of the target through the time series characteristics, and forms predicted positioning information; S56. Perform weighted fusion on the results of the optical flow attention mechanism and the spatiotemporal attention mechanism to generate enhanced tracking features, wherein the enhanced tracking features include the attention weighted information and spatiotemporal predicted positions of the significant motion area, and are used to dynamically identify the core motion area of ​​the target in a complex environment.

5. According to the AI ​​video low-altitude target recognition and real-time tracking method based on deep learning in claim 1, it is characterized in that: The S6 specifically includes: S61, inputting the enhanced tracking feature into an anomaly detection module, extracting the optical flow vector field of each frame image, obtaining motion vector information of each pixel, wherein the motion vector represents the horizontal displacement and vertical displacement of the pixel, and forming an optical flow vector field matrix, wherein the optical flow vector field matrix represents the motion direction and speed of the target between consecutive frames; S62. Calculate the instantaneous motion speed V of the target based on the optical flow vector field matrix, using the Euclidean distance calculation formula: ; Among them, M represents the total number of pixels covered by the target in the frame, represents the horizontal displacement of the i-th pixel, Represents the vertical displacement of the i-th pixel; S63, instantaneous speed V and direction of movement of the target Further analysis is performed to calculate the direction of movement : ; S64: Establishing a target motion trajectory data set based on the motion information of the historical frame ,in represents the speed of the target in the tth frame, Indicates the direction of the target in the tth frame; performs time series analysis on the motion trajectory data set to obtain motion trends and trajectory changes; S65. Perform trend analysis on the target's motion trajectory and calculate the speed change rate and the rate of change of direction : ; ; in, represents the speed of the target in the t-1th frame, Indicates the direction of the target in the t-1th frame; S66. Define adaptive threshold and , dynamically adjust the threshold according to the target's historical motion data and environmental changes; when or When , the current frame is marked as abnormal; the adaptive threshold and The dynamic adjustment is based on environmental fluctuations, target acceleration and direction change rate to meet the adaptability of detection; S67, if multiple consecutive frames meet the abnormal detection condition, the target is marked as abnormal motion, and an abnormal detection result is generated. The abnormal detection result includes the abnormal frame number, speed change rate and direction change rate information, and finally the abnormal detection result is output.

6. According to the deep learning-based AI video low-altitude target recognition and real-time tracking method of claim 1, it is characterized in that: The S7 specifically includes: S71, inputting the anomaly detection result into a dynamic optical flow compensation module, analyzing the target motion features marked in the anomaly detection and the motion information in the optical flow vector field, identifying the global motion offset caused by camera shaking, and initializing compensation parameters for correcting the position deviation caused by camera shaking; S72, based on the initial compensation parameters, performing position compensation on the tracking target of the current frame, applying the initially calculated jitter compensation amount to the target position of the current frame, and adjusting the coordinates of the target position in real time; S73, inputting the compensated target position into a recurrent neural network module, using the recurrent neural network module to perform time series analysis on the target's historical motion data, identifying the target's motion pattern in the past multiple frames, and predicting the target's possible position in the next frame; S74, fusing the predicted position output by the recurrent neural network module with the compensated position, and combining the jitter compensation result with the timing prediction result by setting a fusion strategy to generate a fused tracking position; S75, adaptively adjusting the target tracking window of the current frame according to the fused tracking position, and adjusting the size and position of the tracking window according to the actual movement of the target; S76: Generate a compensated tracking result based on the final compensated and predicted adjusted tracking position.

7. According to the deep learning-based AI video low-altitude target recognition and real-time tracking method of claim 1, it is characterized in that: The S8 specifically includes: S81, input the compensated tracking result into the spatiotemporal multi-layer feedback optimization module, compare the target position of the current frame with that of the previous frame, and calculate the inter-frame error; set the current frame position to and the previous frame position is : ; Where E represents the inter-frame error; S82, comparing the inter-frame error with a preset threshold, and when the inter-frame error exceeds the preset threshold, triggering a short-term error correction mechanism, adjusting the tracking frame position of the current frame in real time, and updating the target position; S83, cumulatively analyzing the error values ​​of each frame, constructing an error sequence, and periodically analyzing the error sequence to identify long-term deviation trends; evaluating changes in overall tracking accuracy by calculating the mean and variance of the error sequence; S84. If the periodic analysis results show that the inter-frame error gradually accumulates and the deviation trend is significant, the convolution kernel size and stride in the optical multi-scale optical flow perception network are optimized, and the optical flow parameters are dynamically adjusted; S85, after detecting the long-period inter-frame error accumulation, adaptively adjust the size and position of the tracking window: ; ; in, Indicates the adjusted window size. Indicates the current window size. and represents the window adjustment coefficient, represents the mean of the error series, Indicates the adjusted window position, Indicates the current window position. and represents the position adjustment coefficient, represents the variance of the error sequence, tanh represents the hyperbolic tangent function; S86, applying the optimized optical flow parameters and tracking window configuration to the tracking result of the current frame, recording the error correction result and parameter adjustment information, and forming a feedback database.

Citation Information

Patent Citations

  • Target detection system and method based on self-adaption combined wave filtering and multilevel detection

    CN108154118A

  • Unmanned aerial vehicle aerial video moving small target real-time detection and tracking method

    CN109785363A