Unmanned aerial vehicle detection method based on bidirectional modern time convolutional network
By adopting a bidirectional modern time convolution network in drone detection, dynamic learning position coding and bidirectional depth separable convolution are introduced, which solves the problem of insufficient accuracy and robustness of drone target detection in complex environments, and achieves high-precision drone detection and recognition.
Patent Information
- Application Number
- CN202510653968.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-21
- Publication Date
- 2025-06-20
AI Technical Summary
The prior art has problems of insufficient accuracy and robustness in high-precision detection and identification of drone targets in complex environments.
Using a drone detection method based on bidirectional modern time convolution network, by constructing a radar data set for anti-UAV targets, dynamic learning position coding and bidirectional depth separable convolution are introduced, forward and reverse timing dependencies are captured, and mixed-scale convolution blocks are designed to cover multiple time scales.
It has improved the drone's mutation motion detection capabilities, achieved high-precision detection and identification of drone targets in complex environments, and provided technical support for safety protection in sensitive areas such as airports and urban airspace.
Smart Images

Figure CN120180243A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of unmanned aerial vehicles, and particularly to a method for detecting unmanned aerial vehicles based on a bidirectional modern time convolutional network. Background Art
[0002] With the continuous development of multi-rotor unmanned aerial vehicle technology and the further improvement of the marketization rate of unmanned aerial vehicles, unmanned aerial vehicles have been widely used in daily life, including civilian, commercial, and military applications. However, the popularization of unmanned aerial vehicles has also brought a series of safety problems, especially in sensitive areas such as airports and urban airspaces. Due to the lack of sufficient aviation safety awareness and loopholes in relevant supervision measures, the phenomena of "illegal flight" and "random flight" of unmanned aerial vehicles occur frequently, seriously threatening flight safety and public safety.
[0003] Traditional methods for identifying unmanned aerial vehicles mainly include identification based on optical sensors and methods based on radio frequency identification (RFID), etc. The identification method based on optical sensors will have a significant decline in performance in complex environments (such as bad weather, night, etc.), and its identification range is limited. For the method based on radio frequency identification, the unmanned aerial vehicle needs to be equipped with a corresponding radio frequency tag, but in fact, a large number of unmanned aerial vehicles do not have such conditions, and the radio frequency signal is easily interfered, making it difficult to accurately identify the unmanned aerial vehicle target.
[0004] As an important means of target detection, radar has the advantages of long detection range and being unaffected by light conditions, and is widely used for target detection and tracking in military and civilian fields. In the detection of unmanned aerial vehicles, radar can detect targets at a relatively long distance. However, radar point cloud data is often complex and contains a large amount of noise. Traditional target recognition methods based on radar data have a low recognition accuracy for small and low-observability targets such as unmanned aerial vehicles, and are affected by environmental factors (such as terrain, building reflections, etc.), resulting in blurred target features and making it difficult to accurately distinguish unmanned aerial vehicles from other aerial targets (such as birds, etc.).
[0005] Deep learning technology has achieved great success in fields such as time series and computer vision. It can automatically learn complex feature representations from a large amount of data and has shown excellent performance in target recognition tasks. The modern time convolutional network (Modern TCN) has achieved certain results in time series processing, but there are still the following limitations: 1) Limitations of fixed position encoding: It relies on predefined sine position encoding and cannot dynamically adapt to the feature distributions of different sequences.
[0006] 2) Insufficient one-way time series modeling ability: Traditional DWConv only processes data in chronological order and cannot capture reverse context dependencies, resulting in insufficient detection accuracy for sudden actions such as sudden stops and turns of unmanned aerial vehicles.
[0007] 3) Low feature fusion efficiency: Although a single large kernel expands the ERF, it has a high computational cost and may ignore local details such as high-frequency fluctuations in short-term predictions. Summary of the Invention
[0008] The technical problem to be solved by the present invention is how to provide a method that can improve the accuracy and robustness of time series modeling and achieve high-precision detection and recognition of UAV targets in complex environments.
[0009] To solve the above technical problem, the technical solution adopted by the present invention is: A UAV detection method based on a bidirectional modern time convolutional network, comprising the following steps: S1: Construct a radar data set in an anti-UAV scenario; S2: Construct a bidirectional modern time convolutional network model BiMTCN; S3: Use the data set processed in step S1 to train the bidirectional modern time convolutional network model BiMTCN constructed in step S2; S4: Use the optimal model trained in step S3 to detect UAVs from radar time series data; S5: Output and organize the results of UAV detection.
[0010] The beneficial effects produced by adopting the above technical solution are as follows: The method described in this application improves the UAV mutation action detection ability by introducing dynamic learnable position encoding and proposing a bidirectional depthwise separable convolution, capturing forward and reverse temporal dependencies simultaneously. By designing a hybrid scale convolution block, it covers multiple time scales while maintaining efficient computation, and can achieve high-precision detection and recognition of UAV targets in complex environments, providing technical support for the security protection of sensitive areas such as airports and urban airspaces. Brief Description of the Drawings
[0011] The present invention will be further described in detail below with reference to the drawings and specific embodiments.
[0012] Figure 1 is the main flowchart of the method described in the embodiment of the present invention; Figure 2 is the diagram of the bidirectional modern time convolutional network model in the method described in the embodiment of the present invention Figure 3 is the comparison diagram of partial radar point data prediction in the method described in the embodiment of the present invention. Detailed Description of the Embodiment
[0013] The following clearly and completely describes the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative work fall within the scope of protection of the present invention.
[0014] In the following description, many specific details are set forth in order to fully understand the present invention. However, the present invention can also be implemented in other ways different from those described herein. Those skilled in the art can make similar generalizations without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.
[0015] As Figure 1 shown, the embodiments of the present invention disclose a drone detection method based on a bidirectional modern temporal convolutional network. The method mainly includes the following steps: S1: Construct a radar dataset in an anti-drone scenario. This dataset includes radar point trace sequences labeled as drone categories and non-drone categories , and is divided into a training set , a validation set and a test set ; S2: Aiming at the limitations of fixed position encoding, insufficient unidirectional temporal modeling ability, high computational cost, etc. existing in the ModernTCN (Modern Temporal Convolutional Networks) model in the prior art, this application proposes an improved bidirectional modern temporal convolutional network BiMTCN (Bidirectional Modern Temporal Convolutional Networks).
[0016] The bidirectional modern temporal convolutional network BiMTCN model solves the problems of limitations of fixed position encoding, insufficient unidirectional temporal modeling ability, high computational cost, etc. existing in ModernTCN. It introduces dynamic learnable position encoding to embed position information in the input layer of each residual block to enhance temporal features. A bidirectional depthwise separable convolution is proposed. By changing the original single depthwise separable convolution (Depthwise Separable Convolution) in ModernTCN to a bidirectional depthwise separable convolution, it captures both forward and backward temporal dependencies simultaneously, improving the ability to detect sudden actions of drones. By designing a mixed-scale convolution block, it covers multiple time scales while maintaining efficient computation. The network model diagram is as Figure 2 shown.
[0017] S3: Use the processed dataset in step S1 to train the network model constructed in step S2.
[0018] S4: Use the optimal detection model trained in step S3 to detect drones in the radar time series data.
[0019] S5: Output and organize the results of drone detection.
[0020] The above steps will be described in detail below in combination with specific content. The specific implementation steps of S1 are as follows: S1-1, Construct an anti-drone target recognition dataset: The data is based on radar detection point track simulation data. The radar is a mechanical scan with a 360° azimuth and a phased scan in elevation.
[0021] S1-2: Sequence length normalization processing. To address the problems of inconsistent sequence feature lengths and loss of time information caused by padding in the dataset, a normalization method with dynamic timestamp extension is proposed to retain the original time feature information while unifying the sequence length.
[0022] For all samples in the dataset, determine the maximum time step : ; where is the original sequence length of the th sample, and is the total number of samples.
[0023] Dynamic padding and timestamp extension. For any sample , pad at the end of the sequence and synchronously extend the timestamp: For any sample, copy the last set of valid features at the end of the sequence until reaching . The padding formula is: ; where represents the last set of valid feature vectors of sample .
[0024] Based on the original last timestamp , generate new timestamps incrementally at a fixed step : ; where is the padded time feature value, and is to increase by 2 each time on the basis of the original last timestamp to maintain the time order.
[0025] S1-3, Sample equalization processing: Since the amount of data with drone labels in the dataset is small, the SMOTE oversampling technique is used to synthesize new sample data to increase the number of minority-class samples. The calculation formula is as follows: ; Among them, .
[0026] Furthermore, S2 specifically includes the following steps: S2-1, Introduce dynamic learnable positional encoding. By embedding positional information in the input layer of each residual block, the expression of key features is enhanced. Specifically, it includes the following steps.
[0027] Generate positional embeddings through the parameter matrix where is the length of the input sequence, is the feature dimension of the input sequence; Perform an addition operation on the dynamic learnable positional encoding and the input sequence , and its calculation method is: ; Obtain the enhanced input sequence , where .
[0028] S2-2, Design a bidirectional depthwise separable convolution core structure (Bi-DWConv) to capture both forward and backward temporal dependencies simultaneously and improve the ability to detect drone mutation actions.
[0029] Furthermore, S2-2 specifically includes the following steps: S2-2-1, Perform depth convolution on the sequentially enhanced input sequence in chronological order to extract forward temporal features: The calculation formula is: .
[0030] Among them, is the kernel size, is the dilation rate (to avoid conflicts with the feature dimension ).
[0031] is the channel index, represents the value of the channel of the enhanced input sequence at the time step .
[0032] S2-2-2, Perform depth convolution on the reverse-order enhanced input sequence to extract backward temporal features: The calculation formula is .
[0033] Among them, the forward and backward branches share the convolutional kernel parameters , ensuring efficient parameter quantity.
[0034] S2-2-3. Dynamically allocate the contribution weights of the forward and backward branches through a lightweight gating network. The gating weights are adjusted in real time according to the input characteristics to improve the robustness of the model to complex UAV scenarios. The specific steps are as follows: Concatenate bidirectional features: , and the concatenation operation integrates the forward and backward information together to provide a richer feature representation for subsequent processing.
[0035] To reduce the computational amount and extract key information, use a Convolution to compress features: , the input dimension is , and the output dimension is .
[0036] Among them, is a learnable weight matrix, and the parameter quantity is only .
[0037] Activation normalization: , and its output range is: , indicating the contribution weights of the forward and backward branches at each time step.
[0038] S2-2-4 proposes a dynamic gating weight fusion method. By dynamically adjusting the weights, it can better combine the forward and backward information and improve the performance of the model. Its calculation formula is:
[0039] Among them, is the weight of the forward branch at time step . : the weight of the backward branch at time step .
[0040] According to the input characteristics at different time steps, dynamically adjust the contribution ratio of the forward and backward features. At some time steps will be larger; while at some other time steps, will increase accordingly.
[0041] Through this dynamic adjustment, the model can better adapt to the complex flight scenarios of UAVs and improve the detection and recognition ability of targets.
[0042] S2-2-5, Pointwise convolutional layer for cross-channel information interaction:
[0043] Among them, is a 1×1 convolution kernel, is the number of output channels, which is decoupled from the input feature dimension decoupled.
[0044] S2-3: Set a mixed-scale convolution block that extracts high-frequency features and low-frequency features simultaneously. The small-scale convolution kernel extracts high-frequency features such as sudden stops and turning actions of the drone, and the large-scale convolution kernel captures global features such as the continuous flight trajectory of the drone. The specific steps are as follows: Furthermore, the S2-3 specifically includes the following steps: S2-3-1, the mixed-scale convolution block includes two small-scale convolution kernels with a convolution kernel size of 3×1 , and a large-scale convolution kernel with a convolution kernel size of 7×1 .
[0045] Among them , is the dimension of the concatenated features, and are the feature dimensions output by the small-scale and large-scale convolution kernels respectively.
[0046] S2-3-2, use the small-scale convolution kernel . Perform a convolution operation on to obtain the small-scale convolution output , and the calculation formula is: .
[0047] S2-3-3, use the large-scale convolution kernel to perform a convolution operation on to obtain the large-scale convolution output , and the calculation formula is: .
[0048] S2-3-4, calculate the weights for the small-scale convolution output . and the large-scale convolution output respectively, and generate weights and through the weight generation module. The weight generation module is a fully connected layer, and its input is , and the weights are obtained after being processed by the softmax activation function.
[0049] S2-3-5, perform weighted fusion on the small-scale convolution output . and the large-scale convolution output , and the fused output The calculation formula is: , where represents element-wise multiplication.
[0050] S2-4 introduces a self-attention module during the moving average process to capture the dependencies between different time steps in the input sequence.
[0051] Furthermore, the S2-4 specifically includes the following steps: S2-4-1, for the fused output , obtain the query matrix through a linear transformation , the key matrix and the value matrix , where , , are the dimensions of the query, key, and value respectively.
[0052] Calculate the attention score matrix :
[0053] Calculate the attention output :
[0054] S2-4-2, the multi-head self-attention mechanism includes parallel attention heads, and the outputs of each attention head are concatenated and then passed through a linear transformation layer to obtain the final self-attention output :
[0055] Among them, is the output of the -th attention head.
[0056] In the model of this application, a dynamic learnable position encoding layer is added to perceive the relative and absolute positions of elements in the sequence, so as to better capture the temporal characteristics and long-term dependencies of the sequence. A bidirectional depthwise separable convolution is proposed to reduce the computational amount and the number of parameters. In UAV detection, not only the current and past radar point track information is helpful for judging the UAV state, but the future point track information may also contain important clues. The bidirectional depthwise separable convolution can capture this bidirectional temporal dependency and improve the feature extraction ability of the model. Through the fusion of a lightweight gating network, in complex UAV flight scenarios, there may be problems such as noise, interference, or data loss. The lightweight gating network can suppress the influence of noise and interference features by adjusting the weights, enhance the adaptability of the model to complex scenarios, and thus improve the robustness and stability of the model. Design a hybrid scale convolution block to capture multi-scale feature information at the same time, enhance the adaptability of the model to different target sizes and shapes, and thus more comprehensively describe the features of the UAV.
[0057] Further, the specific implementation method of S3 includes the following steps: S3-1: Use the training set , validation set and test set obtained in step S1 to train the bidirectional modern time convolutional network BiMTCN model constructed in step S2. S3-2: Evaluate the model, calculate the model accuracy and the number of model parameters, and judge whether the model performance meets the requirements based on this. If the model accuracy does not meet the requirements, repeat the operation in step S3-1 and increase the number of training rounds.
[0058] Further, the specific implementation steps of step S4 are as follows: Perform the same data preprocessing operation as in step S1 on the radar time series data to be detected, and input it into the optimal detection model trained in step 3 for detection.
[0059] Figure 3 It is a comparison graph of the prediction of partial radar point data in the method described in the embodiments of the present invention. The meanings of the data in the graph are as follows: Target azimuth angle (°): The horizontal direction angle of the target obtained by the 360° mechanical scan of the radar, used to determine the target azimuth.
[0060] Target slant range (m): The straight-line distance between the radar and the target, reflecting the distance of the target.
[0061] Relative height (m): The vertical height of the target relative to the radar, assisting in determining the target's spatial position.
[0062] Radial velocity (m / s): The movement speed of the target along the direction of the line connecting the radar and the target, used to judge whether the target is approaching or moving away from the radar.
[0063] Measurement time (s): Record the time of each measurement, used to analyze the change of the target state over time.
[0064] RCS: Radar cross section, reflecting the scattering ability of the target to radar electromagnetic waves and affecting the detectability of the target.
[0065] Label: Mark the target category, 0 represents a non-UAV target, and 1 represents a UAV target.
[0066] Recognition result: The determination of whether the target is a UAV based on the method described in this application, compared with the label to evaluate the detection accuracy.
[0067] As can be seen from the figure, the method described in this application can accurately identify drone targets. This method first standardizes the radar point cloud dataset through the method of dynamic timestamp extension, and balances the number of samples of each category through the SMOTE oversampling technique. Then, a bidirectional modern time convolutional network BiMTCN is established, and the training set is input to train the network model. Finally, the optimal detection model is obtained and the drone detection result is output. BiMTCN introduces dynamic learnable position encoding to enhance temporal features. A bidirectional depthwise separable convolution is proposed to capture forward and backward temporal dependencies simultaneously, improving the ability to detect sudden actions of drones. By designing a hybrid-scale convolution block, multiple time scales are covered while maintaining efficient computation, providing an effective and accurate solution for drone detection in radar point cloud data.
Claims
1. A drone detection method based on a bidirectional modern temporal convolutional network, characterized in that The steps include: S1: Construct radar dataset for anti-UAV scenario; S2: Construct a bidirectional modern temporal convolutional network model BiMTCN; S3: Use the data set processed in step S1 to train the bidirectional modern temporal convolutional network model BiMTCN constructed in step S2; S4: Use the optimal model trained in step S3 to detect drones on radar time series data; S5: Output and organize the results of drone detection.
2. The drone detection method based on a bidirectional modern temporal convolutional network as claimed in claim 1 is characterized in that The S1 specifically includes the following steps: S1-1, construct an anti-UAV target recognition data set, the data in the data set is based on radar detection point trace simulation data; S1-2, sequence length normalization: for all samples in the data set, determine the maximum time step : ; in, For the The original sequence length of samples, is the total number of samples; Dynamic padding and timestamp extension for any sample , padding at the end of the sequence and synchronously extending the timestamp: For any Samples of, copy the last set of valid features at the end of the sequence Until reaching , the filling formula is: ; in, Representation sample The last set of valid eigenvectors of ; Original last timestamp As a benchmark, according to the fixed step size Generate a new timestamp incrementally: ; in, is the time characteristic value after filling, Increase the timestamp by 2 each time to maintain the time sequence; S1-3, sample equalization processing: Use the SMOTE oversampling method to synthesize new sample data to increase the number of class samples. The formula is: ; in, ; S1-4, sample division: Divide the processed samples into training sets , validation set and test set .
3. The drone detection method based on a bidirectional modern temporal convolutional network as claimed in claim 1, characterized in that The S2 specifically includes the following steps: S2-1: Introducing dynamic learnable position encoding to enhance the key feature expression by embedding position information in each residual block input layer; S2-2: Design a bidirectional depthwise separable convolution Bi-DWConv to capture both forward and reverse temporal dependencies; S2-3: A mixed-scale convolution block is set up to extract both high-frequency and low-frequency features, where the small-scale convolution kernel focuses on extracting the high-frequency features of the drone, and the large-scale convolution kernel is used to capture the global features of the drone; S2-4: A self-attention module is introduced in the moving average process to capture the dependencies between different time steps in the input sequence.
4. The drone detection method based on a bidirectional modern temporal convolutional network as claimed in claim 3, characterized in that: The S2-1 specifically includes the following steps: Through the parameter matrix Generate position embeddings where is the length of the input sequence, is the feature dimension of the input sequence; Combine the dynamic learnable positional encoding with the input sequence Perform the addition operation, and the calculation method is: ; Get the enhanced input sequence ,in .
5. The drone detection method based on a bidirectional modern temporal convolutional network as claimed in claim 4, characterized in that: The S2-2 specifically includes the following steps: S2-2-1: Sequentially enhanced input sequence Perform deep convolution in chronological order to extract forward time series features. The calculation formula is: ; in, is the convolution kernel size, is the expansion rate, is the channel index, Represents the enhanced input sequence No. Channel at time step The value of S2-2-2: Enhance the input sequence in reverse order , perform deep convolution to extract reverse time series features, and the calculation formula is: ; Among them, the forward and backward branches share the convolution kernel parameters ; S2-2-3: Dynamically allocate the contribution weights of the forward and backward branches through a lightweight gating network. The gating weights are adjusted in real time according to the input characteristics, including the following steps: Splicing bidirectional features: ,The concatenation operation integrates the forward and backward information to provide a richer feature representation for subsequent processing; Use one Convolution compression features: , the input dimension is , the output dimension is ; in, is the learnable weight matrix, and the number of parameters is only ; Activation Normalization: , its output range is: , represents the contribution weight of the forward and backward branches at each time step; S2-2-4 improves the performance of the model by dynamically adjusting weights and combining forward and backward information. The calculation formula is: ; in, is the time step The weight of the forward branch, is the time step The weight of the backward branch; S2-2-5, point-by-point convolution layer for cross-channel information interaction: ; in, is a 1×1 convolution kernel, is the number of output channels, which is the same as the input feature dimension Decoupling.
6. The drone detection method based on a bidirectional modern temporal convolutional network as claimed in claim 5, characterized in that: The S2-3 specifically includes the following steps: S2-3-1: The mixed-scale convolution block includes two small-scale convolution kernels with a kernel size of 3×1. And a large-scale convolution kernel with a convolution kernel size of 7×1 ,in , is the dimension of the concatenated features, and They are the feature dimensions of the small-scale and large-scale convolution kernel outputs respectively; S2-3-2: Use small-scale convolution kernels ,right Perform convolution operation to obtain small-scale convolution output , the calculation formula is: ; S2-3-3: Use large-scale convolution kernels right Perform convolution operation to obtain large-scale convolution output , the calculation formula is: ; S2-3-4: small-scale convolution output , and large-scale convolution output Calculate the weights separately and generate weights through the weight generation module and , the weight generation module is a fully connected layer, and its input is , the weight is obtained after being processed by the softmax activation function; S2-3-5: Output of small-scale convolution And large-scale convolution output Perform weighted fusion, the fused output The calculation formula is: ,in Represents element-wise multiplication.
7. The method for detecting drones based on a bidirectional modern temporal convolutional network as claimed in claim 6, characterized in that: The S2-4 specifically includes the following steps: For the fused output , the query matrix is obtained by linear transformation , key matrix Sum Matrix ,in , , The dimensions are query, key, and value respectively; Calculate the attention score matrix : ; Computing attention output : ; The multi-head self-attention mechanism includes The output of each attention head is concatenated and passed through a linear transformation layer. Get the final self-attention output : ; in, It is The output of an attention head.
8. The drone detection method based on a bidirectional modern temporal convolutional network as claimed in claim 1, characterized in that The S3 specifically includes the following steps: S3-1: Use the training set divided in step S1 , validation set and test set Train the bidirectional modern temporal convolutional network BiMTCN model constructed in step S2; S3-2: Evaluate the model, calculate the model accuracy and model parameter quantity, and use this as a basis to determine whether the model performance meets the requirements. If the model accuracy does not meet the requirements, repeat step S3-1 and increase the number of training rounds.
9. The drone detection method based on a bidirectional modern temporal convolutional network as claimed in claim 1, characterized in that The S4 specifically includes the following steps: The radar time series data to be detected The same data preprocessing operation as in step S1 is performed, and the data is input into the optimal detection model trained in step S3 for detection.
Citation Information
Patent Citations
Classification and identification method for low, slow small targets
CN112434643A
Method and system for predicting surrounding vehicle behavior of automatic driving emergency rescue vehicle based on multi-intention guiding comparative learning
CN118193953A
Composite material propeller blade damage identification method of large unmanned aerial vehicle propulsion system
CN118747469A
Unmanned ship navigation state multi-classification detection method based on dynamic coupling graph structure
CN119992474A