Detection method and detection device for real-time track initiation and track classification

The radar echo data is processed through deep learning models, and features are extracted using sparse neural networks and timing convolutional networks, and track classification is combined with transformer. The problems of slow detection speed and insufficient utilization of timing features in the existing technology are solved, and fast and accurate track detection in a strong clutter environment is achieved.

CN115656958BActive Publication Date: 2025-06-17四川启睿克科技有限公司 +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211391820.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-08
Publication Date
2025-06-17
Estimated Expiration
2042-11-08

AI Technical Summary

Technical Problem

The existing radar detection technology has slow detection speed in strong clutter environments, the traditional method has a large amount of calculation, the deep learning-based methods fail to fully utilize the timing characteristics of the data, and the track candidate set generation process takes a long time.

Method used

The deep learning model is used to process the echo dot data of continuous wave radar, and features are extracted using sparse neural networks and one-dimensional time-sequence convolutional networks. It does not rely on additional auxiliary information, and track classification and authenticity judgment are performed through transformers.

Benefits of technology

In a strong cluttered environment, it can quickly and accurately detect the tracks and categories of targets, which improves detection speed and efficiency and adapts to different detection needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115656958B_ABST
    Figure CN115656958B_ABST
Patent Text Reader

Abstract

The present invention discloses a detection method for real-time track initiation and track classification, including: inputting continuous wave radar echo dot data of N consecutive time steps, merging and then performing preprocessing; inputting the preprocessed data into a sparse neural network to preliminarily predict candidate tracks and corresponding target categories; first performing real track matching on the candidate tracks and then performing kinematic filtering to eliminate tracks that do not conform to the filtering rules; respectively inputting the filtered candidate tracks into a spatial feature extraction network to extract spatial features and into a temporal feature extraction network to extract temporal features; after merging the spatial features and temporal features obtained in the previous step, using a transformer for classification; after post-processing the output results, outputting the final classification results and corresponding tracks; the present invention also discloses a detection device for real-time track initiation and track classification; the present invention can efficiently and accurately detect the tracks and categories of targets.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of radar detection, and particularly to a detection method and a detection device for real-time track initiation and track classification. Background Art

[0002] Performing target track initiation on continuous wave radar detection data is the most basic and important task in multi-target tracking technology. Track initiation specifically refers to the process of determining the target track before the radar has stabilized target tracking; while track classification can refer to classifying the target type represented by the track or classifying the track movement pattern; generally speaking, track initiation and track classification are two separate processes. Classical track initiation can be divided into sequential processing methods and batch processing methods according to different data processing methods. Sequential processing methods include logical methods, heuristic rule methods, etc.; batch processing methods include Hough transform and its variant algorithms. The sequential processing method has a small computational amount and a fast operation speed, but the effect is average, and it is usually applicable to the detection environment with weak clutter; while the batch processing method has a large computational amount and a long operation time, and has a good effect and can be applicable to the detection environment with strong clutter. Classical track classification includes rule methods and traditional machine learning algorithms, etc. Generally, through the information provided in the radar echo signal, such as distance, azimuth, intensity, and amplitude, etc., features are artificially constructed and then rules or classifiers are used for detection. However, the modern radar detection environment is becoming increasingly severe, the strong clutter environment has become common, the requirements for detection speed and the increase in detection targets make it more and more difficult for traditional track initiation and track classification schemes to meet the requirements. In recent years, deep learning has achieved success in fields such as face recognition and natural language due to its advantages such as high accuracy, fast operation speed, and the ability to process complex data. Therefore, there has also been a preliminary exploration of using deep learning for radar data track initiation. Currently, the algorithms for using deep learning for track initiation can be roughly divided into two categories: methods based on convolutional neural networks and methods based on recurrent neural networks. The method based on convolutional neural networks converts the echo signal into a picture form, and then uses convolution to extract corresponding features for detection. The method based on recurrent neural networks extracts features according to the temporal properties of the signal for detection. For track classification, generally, SVM, MLP, or convolutional neural networks are used for classification.

[0003] Although there are many classical methods and preliminary deep learning algorithms in this field, these methods also face many problems in the actual application process. Summarized, there are the following points:

[0004] 1. Currently, radar detection is in a complex strong clutter environment, and the traditional track initiation methods (such as Hough transform, etc.) have too large a computational amount and the detection speed is not fast enough;

[0005] 2. The features used in the current track initiation methods based on convolutional neural networks are artificially designed and have limitations. At the same time, the temporal characteristics of the data are not utilized.

[0006] 3. The current methods based on recurrent neural networks have poor performance in a strong clutter environment and the detection speed is not fast enough.

[0007] 4. The current deep learning methods all need to use certain rules to generate a track candidate set, and this process takes a long time. Summary of the Invention

[0008] To solve the problems existing in the prior art, the object of the present invention is to provide a detection method and a detection device for real-time track initiation and track classification. The present invention directly processes the echo dot data of a continuous wave radar using a deep learning model, without using additional auxiliary information and without spending time generating a track candidate set, and can efficiently and accurately detect the track and category of the target.

[0009] To achieve the above object, the technical solution adopted by the present invention is: a detection method for real-time track initiation and track classification, comprising the following steps:

[0010] Step 1: Input the echo dot data of the continuous wave radar for N consecutive time steps, merge them and then perform preprocessing.

[0011] Step 2: Input the preprocessed data into a sparse neural network to preliminarily predict the candidate tracks and the corresponding target categories.

[0012] Step 3: First perform real track matching on the candidate tracks and then perform kinematic filtering to eliminate the tracks that do not conform to the filtering rules.

[0013] Step 4: Input the filtered candidate tracks into a spatial feature extraction network to extract spatial features and input them into a temporal feature extraction network to extract temporal features.

[0014] Step 5: After merging the spatial features and temporal features obtained in the previous step, use a transformer for classification.

[0015] Step 6: After post-processing the output results, output the final classification results and the corresponding tracks.

[0016] As a further improvement of the present invention, the preprocessing in Step 1 includes converting the distance, azimuth and elevation data provided in the radar echo into Euclidean coordinates in three-dimensional space.

[0017] As a further improvement of the present invention, in Step 2, the sparse neural network includes a backbone layer, a decoding layer and an output layer. The input data extracts features through the backbone layer, the features are decoded through the decoding layer and then enter the output layer to obtain the results.

[0018] As a further improvement of the present invention, step 2 is specifically as follows:

[0019] Input the three-dimensional coordinate data after converting the echo dot data into a sparse neural network. The sparse neural network outputs relevant information of the candidate track, specifically including the candidate track coordinates, the probability value of whether the track is a real track, and the object category information represented by the track. Then, perform a primary filtering based on the probability value of whether it is a real track.

[0020] As a further improvement of the present invention, in step 3, the real track matching includes selecting the dot closest to the predicted candidate track as the matched real track; the kinematic filtering includes speed screening, acceleration screening, and yaw angle screening.

[0021] As a further improvement of the present invention, in step 4, the extraction of spatial features specifically includes inputting the echo information corresponding to the candidate track into a spatial feature extraction network to obtain a spatial feature vector, where the spatial feature extraction network is a 3D sparse convolution network HDResNet; the extraction of temporal features specifically includes: arranging the echo information corresponding to the candidate track in chronological order and inputting it into a temporal feature extraction network to obtain a temporal feature vector, where the temporal feature extraction network is a one-dimensional dilated convolution network stacked with multiple layers, and residual connections are made between each layer.

[0022] As a further improvement of the present invention, the loss function for the sparse neural network to extract candidate tracks consists of a track coordinate loss, a probability loss of whether it is a real track, and an object category loss represented by the track. For the track coordinate loss, a smooth L1 loss function is used, and for the probability loss of whether it is a real track and the object category loss represented by the track, cross-entropy loss functions are respectively used; the loss for classifying candidate tracks using a transformer includes the probability loss of whether it is a real track and the object category loss represented by the track, both of which use cross-entropy loss functions; during training, optimize by maximizing this loss function; during the training process, when the loss value is not within a reasonable range, adjust the parameters and continue training until the loss value drops to within a reasonable range.

[0023] The present invention also discloses a detection device for real-time track initiation and track classification, including:

[0024] A dot data preprocessing module, used to merge the echo data of multiple consecutive steps of a continuous wave radar, and then perform 3D Euclidean coordinate conversion to obtain the standard input data for the next step;

[0025] A sparse network backbone module, composed of multiple layers of HDResNet networks, used to extract features from the input 3D coordinate data, and obtain a fixed-length backbone feature vector for each layer;

[0026] The sparse network decoding module, which consists of multiple layers of sparse convolutional networks, is used to input multiple feature vectors obtained from the backbone model into the decoding modules corresponding to their respective layers, and then merge the output results and input them into the output module of the sparse network;

[0027] The sparse network output module, which consists of a sparse convolutional network, is used to input the decoded vector into the output layer to obtain the output result;

[0028] The post-processing model of the sparse network is used to post-process the results of the output module. The post-processing filters according to whether the track obtained from the output layer is a valid track based on a threshold, and discards the results smaller than the threshold;

[0029] The kinematic filtering module first selects the closest true track to the candidate track as the output of the candidate track, and then performs rule filtering, which is divided into speed rules, acceleration rules, and yaw angle rules, to eliminate the tracks with abnormal speed, acceleration, and yaw angle, and obtain the final candidate track results;

[0030] The spatial feature extraction module and the temporal feature extraction module. The spatial feature extraction module consists of a sparse convolutional network HDResNet, and the temporal feature extraction module consists of an expanded convolutional network stacked with multiple layers, which respectively extract spatial features and temporal features from the echo data of the candidate track;

[0031] The Transformer classification module is used to merge the spatial features and temporal features and then input them into the transformer module for track classification and track authenticity judgment.

[0032] The beneficial effects of the present invention are:

[0033] 1. It can adapt to a strong clutter environment and can well detect the target tracks and their categories in four-period data;

[0034] 2. It uses a sparse neural network and one-dimensional temporal convolution to improve the detection speed, and has strong practicability;

[0035] 3. It does not use additional information, and at the same time, the model uses a modular design, which can be adjusted according to different scenarios to adapt to different detection requirements. Description of the Drawings

[0036] Figure 1 It is the flowchart of the detection method in the embodiment of the present invention;

[0037] Figure 2 It is the structural diagram of the detection device in the embodiment of the present invention. Detailed Embodiments

[0038] The embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0039] Embodiment

[0040] As Figure 1 shown, a detection method for real-time track initiation and track classification includes the following steps:

[0041] A. Merge and preprocess the input N-frame radar point data:

[0042] The number of frames N can be flexibly selected according to requirements. The preprocessing method includes converting the range, azimuth, and elevation data provided in the radar echo into Euclidean coordinates in 3D space.

[0043] B. Input into a sparse neural network for track extraction:

[0044] The sparse neural network consists of a backbone layer, a decoding layer, and an output layer. The input data extracts features through the backbone layer, and the features are decoded through the decoding layer and then enter the output layer to obtain the result.

[0045] The process of the sparse neural network extracting candidate tracks is as follows: input the three-dimensional coordinate data converted from the echo point data into the sparse neural network, and the sparse network outputs the relevant information of 128 or 256 candidate tracks, specifically including the four-point coordinates of the candidate track, the probability value of whether the track is a real track, and the object category information represented by this track. Then, the post-processing process performs a filtering based on the probability value of whether it is a real track.

[0046] C. Perform kinematic filtering on the identified candidate tracks:

[0047] The kinematic filtering includes the following steps: first, perform real track matching, and then perform speed screening, acceleration screening, and yaw angle screening; real track matching means selecting the point data closest to the predicted candidate track as the matched real track.

[0048] D. Extract spatial features and temporal features from the candidate track data:

[0049] The spatial feature extraction network is a 3D sparse convolutional network HDResNet; the process of extracting spatial features is to input the echo information corresponding to the candidate track, such as coordinates, signal strength, amplitude, etc., into the extraction network to obtain a 64-dimensional spatial feature vector;

[0050] The temporal feature extraction network consists of a one-dimensional dilated convolutional network stacked with multiple layers, and there are residual connections between each layer. The number of stacked layers can be flexibly selected according to requirements; the process of extracting temporal features is to input the echo information corresponding to the candidate track arranged in chronological order into the temporal feature extraction network to obtain a 64-dimensional temporal feature vector.

[0051] E. The combined features are classified using a transformer:

[0052] The combination of spatial and temporal features is obtained by direct concatenation. The combined features are input into a transformer to obtain the final result, which includes the probability of whether the track is real and the target classification represented by the track.

[0053] F. Output the final track and classification result:

[0054] After filtering the probability values output by the transformer according to the threshold, the final track and the classification result of the track are output.

[0055] Specifically, it also includes the setting of the model loss function, the setting of the method for iteratively updating the model parameters, the number of neural network layers of each module network and the number of neurons each time, the setting of the feature vector length, the setting of the probability filtering threshold, the initialization of the parameters of each layer in the model, the connection and alignment between each network layer, the selection of the model training parameters and the training, etc.

[0056] The loss function consists of multiple parts. The overall loss function is composed of the loss of extracting candidate tracks by the sparse neural network and the loss of classifying candidate tracks using a transformer. Among them, the loss function of the sparse neural network for extracting candidate tracks consists of the track coordinate loss, the probability loss of whether it is a real track, and the loss of the object category represented by this track. For the track coordinate loss, a smooth L1 loss function is used, and for the probability loss of whether it is a real track and the loss of the object category represented by this track, cross-entropy loss functions are respectively used. Among them, the loss of classifying candidate tracks using a transformer includes the probability loss of whether it is a real track and the loss of the object category represented by this track, both of which use cross-entropy loss functions. During training, the model is optimized by maximizing this loss function; during the training process, when the loss value is not within a reasonable range, the model parameters are adjusted and training continues until the loss value drops to within a reasonable range, and this model is used as the final track start and track classification model.

[0057] The real-time track start and track classification method integrating a sparse neural network and a temporal convolutional network does not require manual rule formulation and does not use additional auxiliary information. As long as there is enough training data, a suitable model can be trained to directly and quickly obtain the track and the object category information represented in the radar echo data, and has a wide range of application scenarios.

[0058] As Figure 2 shown, this embodiment also provides a detection device for real-time track start and track classification, including:

[0059] The dot data preprocessing module merges the echo data of multiple consecutive steps of the continuous wave radar, and then performs 3D Euclidean coordinate conversion to obtain the standard input data for the next step;

[0060] The sparse network backbone module consists of multiple layers of HDResNet networks, extracts features from the input 3D coordinate data, and obtains fixed-length backbone feature vectors at each layer.

[0061] The sparse network decoding module consists of multiple layers of sparse convolutional networks. The multiple feature vectors obtained from the backbone model are respectively input into the decoding modules corresponding to their layers, and then the output results are merged and input into the output module of the sparse network.

[0062] The sparse network output module consists of a sparse convolutional network, inputs the decoded vector into the output layer, and obtains the output result.

[0063] The post-processing model of the sparse network post-processes the results of the output module. The post-processing filters according to the threshold based on whether the output layer obtains a track, and discards the results smaller than the threshold.

[0064] The kinematic filtering module: Since the track coordinates obtained by the sparse network are not necessarily real values, first select the real track closest to the candidate track as the output of the candidate track, and then perform rule filtering, which is divided into speed rules, acceleration rules, and yaw angle rules, and eliminates the tracks with abnormal speed, acceleration, and yaw angle to obtain the final candidate track results.

[0065] The spatial feature extraction module and the temporal feature extraction module. The spatial feature extraction module consists of a sparse convolutional network HDResNet. The temporal feature extraction module consists of an enlarged convolutional network stacked with multiple layers. Spatial features and temporal features are respectively extracted from the echo data of the candidate tracks.

[0066] The Transformer classification module: Merges the spatial features and temporal features and inputs them into the transformer module for track classification and track authenticity judgment.

[0067] The above embodiments only represent the specific implementation manners of the present invention, and the description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the patent of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention.

Claims

1. A detection method for real-time track initiation and track classification, characterized in that, It includes the following steps: Step 1: Input the continuous wave radar echo point data of N consecutive time steps, merge them and then perform preprocessing; Step 2: Input the preprocessed data into a sparse neural network to preliminarily predict candidate tracks and corresponding target categories; Step 3: First perform real track matching on the candidate tracks and then perform kinematic filtering to eliminate tracks that do not meet the filtering rules; Step 4: Input the filtered candidate tracks into the spatial feature extraction network to extract spatial features and into the temporal feature extraction network to extract temporal features respectively; In Step 4, the extraction of spatial features specifically includes inputting the echo information corresponding to the candidate track into the spatial feature extraction network to obtain a spatial feature vector. Among them, the spatial feature extraction network is a 3D sparse convolutional network HDResNet; the extraction of temporal features specifically includes: arranging the echo information corresponding to the candidate track in chronological order and inputting it into the temporal feature extraction network to obtain a temporal feature vector. Among them, the temporal feature extraction network is a one-dimensional dilated convolutional network stacked with multiple layers, and residual connections are made between each layer; Step 5: After merging the spatial features and temporal features obtained in the previous step, use a transformer for classification; Step 6: After post-processing the output results, output the final classification results and corresponding tracks; The loss function for the sparse neural network to extract candidate tracks consists of a track coordinate loss, a probability loss of whether it is a real track, and a loss of the object category represented by this track. For the track coordinate loss, a smooth L1 loss function is used, and for the probability loss of whether it is a real track and the loss of the object category represented by this track, cross-entropy loss functions are used respectively; the loss for using a transformer to classify candidate tracks includes the probability loss of whether it is a real track and the loss of the object category represented by this track, and both use cross-entropy loss functions; during training, it is optimized by maximizing this loss function; during the training process, when the loss value is not within a reasonable range, adjust the parameters and continue training until the loss value drops to within a reasonable range.

2. The detection method for real-time track initiation and track classification according to claim 1, characterized in that, The preprocessing in Step 1 includes converting the distance, azimuth, and elevation data provided in the radar echo into Euclidean coordinates in three-dimensional space.

3. The detection method for real-time track initiation and track classification according to claim 2, characterized in that, In Step 2, the sparse neural network includes a backbone layer, a decoding layer, and an output layer. The input data extracts features through the backbone layer, and the features are decoded through the decoding layer and then enter the output layer to obtain the result.

4. The detection method for real-time track initiation and track classification according to claim 3, characterized in that, Step 2 is specifically as follows: Input the three-dimensional coordinate data converted from the echo point data into the sparse neural network. The sparse neural network outputs the relevant information of the candidate track, specifically including the candidate track coordinates, the probability value of whether this track is a real track, and the object category information represented by this track; then perform a first filtering according to the probability value of whether it is a real track.

5. The detection method for real-time track initiation and track classification according to claim 1 or 4, characterized in that, In Step 3, the real track matching includes selecting the point closest to the predicted candidate track as the matched real track; The kinematic filtering includes speed screening, acceleration screening, and yaw angle screening.

6. A detection device for real-time track initiation and track classification, characterized in that, It is implemented by using the real-time track initiation and track classification detection method described in any one of claims 1-5. The detection device includes: The dot data preprocessing module is used to merge the echo data of multiple consecutive steps of the continuous wave radar, and then perform 3D Euclidean coordinate conversion to obtain the standard input data for the next step; The sparse network backbone module, which consists of multiple layers of HDResNet networks, is used to extract features from the input 3D coordinate data, and each layer obtains a fixed-length backbone feature vector; The sparse network decoding module, which consists of multiple layers of sparse convolutional networks, is used to input the multiple feature vectors obtained by the backbone model into the decoding modules corresponding to the respective layers, and then merge the output results and input them into the output module of the sparse network; The sparse network output module, which consists of a sparse convolutional network, is used to input the decoded vector into the output layer to obtain the output result; The post-processing model of the sparse network is used to post-process the result of the output module. The post-processing filters according to the threshold based on whether the output layer obtains a track, and discards the results smaller than the threshold; The kinematic filtering module first selects the real track closest to the candidate track as the output of the candidate track, and then performs rule filtering, which is divided into speed rules, acceleration rules, and yaw angle rules, and eliminates the tracks with abnormal speed, acceleration, and yaw angle to obtain the final candidate track result; The spatial feature extraction module and the temporal feature extraction module. The spatial feature extraction module consists of a sparse convolutional network HDResNet, and the temporal feature extraction module consists of an enlarged convolutional network stacked with multiple layers, which respectively extract spatial features and temporal features from the echo data of the candidate track; The Transformer classification module is used to merge the spatial features and temporal features and then input them into the transformer module for track classification and track authenticity judgment.

Citation Information

Patent Citations

  • Adaptive high-speed network flow layered sampling and collecting method

    CN101420419A

  • Flight path prediction method based on graph neural network

    CN113505878A