Flight track identification and classification method based on multi-level residual cross attention
Through the multi-level residual cross attention structure and classification network, the shortcomings of traditional track recognition methods in feature extraction and long sequence data processing are solved, and efficient and stable track recognition and classification are achieved, meeting the high-precision and real-time requirements of aviation traffic management.
Patent Information
- Application Number
- CN202511106011.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-08
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2045-08-08
AI Technical Summary
When traditional track recognition and classification methods deal with complex track data, they have problems such as difficulty in feature extraction, limited ability to process long sequence data, insufficient generalization ability and high computational complexity, which is difficult to meet the needs of modern aviation traffic management.
The multi-level residual cross attention structure is used to extract and classify the global feature of the track data. Through the multi-level stacked residual cross attention module and classification network, the long-distance dependence and global features of the track data are automatically extracted, and the feature transfer stability is enhanced by combining residual connections.
It significantly improves the accuracy and robustness of track recognition and classification, can better meet the requirements of modern aviation traffic management for accuracy and real-time track recognition, and has good stability and efficient feature learning ability.
Smart Images

Figure CN120597100A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of signal processing, and specifically relates to a track recognition and classification method based on multi-level residual cross-attention, which is used to accurately identify target types and is particularly suitable for efficient and precise recognition and classification of target track data, thereby fully improving the efficiency and safety of air traffic management. Background Art
[0002] With the continuous increase in air traffic, traditional track identification and classification methods are no longer able to meet the needs of modern air traffic management. Existing track identification and classification methods mainly rely on traditional machine learning algorithms such as support vector machines (SVM) and random forests. These methods have the following problems when processing complex track data:
[0003] Difficulty in feature extraction: Traditional machine learning methods require manual design and extraction of features, which often makes it difficult to capture effective feature information when faced with high-dimensional and nonlinear track data.
[0004] Limited ability to process long sequence data: Track data is usually time series data. Traditional machine learning methods have difficulty capturing long-term dependencies in the sequence when processing long sequence data.
[0005] Insufficient generalization ability: Traditional machine learning methods have poor generalization ability when faced with unseen track patterns and are prone to overfitting.
[0006] High computational complexity: Traditional machine learning methods have high computational complexity when processing large-scale trajectory data, making it difficult to meet the needs of real-time processing. Summary of the Invention
[0007] To solve the above problems, the present invention proposes a track recognition and classification method based on multi-level residual cross attention, which overcomes the problems of insufficient long-range dependency modeling capability and insufficient global feature representation in traditional track recognition.
[0008] The technical solution adopted in the present invention is:
[0009] A track recognition and classification method based on multi-level residual cross attention includes the following steps:
[0010] S1: Obtain target track data, perform preprocessing, and divide it into training set and test set;
[0011] S2: Construct a multi-level residual cross attention structure to extract global features from the preprocessed target track data;
[0012] S3: Input the extracted global features into the classification network for target track classification and recognition, and map the input global features to the target category label;
[0013] S4: Use labeled training set samples to conduct supervised training on the constructed network model to obtain a trained classification model; the constructed network model includes a multi-level residual cross attention structure and a classification network;
[0014] S5: Input the target track data to be identified in the test set into the trained classification model for classification, and output the identification result of the target track.
[0015] Furthermore, the preprocessing steps in S1 are as follows:
[0016] S101: Eliminate outliers whose target position suddenly jumps to a position far away from the set range of the trajectory within a set time;
[0017] S102: For target track data with missing target position or velocity information, fill in the missing data based on adjacent time data by interpolation;
[0018] S103: Eliminate data points where the speed is continuously 0.
[0019] Furthermore, the multi-level residual cross attention structure constructed in S2 is composed of multiple residual cross attention modules of the same structure stacked in sequence. Each level of residual cross attention module includes a residual feature extraction module and a cross attention module. The overall structure is expressed as:
[0020] ;
[0021] Among them, ResAttBlock(l) represents the residual cross attention module of the lth layer, which is used to use F(l-1) output of the previous layer as input features for local semantic enhancement and global dependency modeling; L is the number of residual cross attention modules;
[0022] The specific processing process of the residual cross attention module on the input features is:
[0023] The input features are first enhanced by the residual feature extraction module. The residual feature extraction module includes two one-dimensional convolutional layers Conv1 and Conv2 and ReLU activation function, and introduces skip connection. The input and the convolution result are added to obtain the enhanced feature Fr(l). The calculation formula is as follows:
[0024] ;
[0025] After the initial feature enhancement, it enters the cross attention module, which maps the enhanced features Fr(l) into query Q, key K and value V vectors respectively, and then calculates the dependency between features based on the scaled dot product attention mechanism. , the calculation is as follows: ; ;
[0026] Where W Q 、W K and W V is the learnable weight matrix, d a is the attention feature dimension;
[0027] The output of the crisscross attention module Perform element-by-element addition with the output Fr(l) of the residual feature extraction module to form a residual connection, and obtain the output result of the residual cross attention module of the current layer. The calculation formula is as follows:
[0028] .
[0029] Furthermore, the goal of the classification network in S3 is to map the input global feature vector to the label of the target category. The classification network structure formula is as follows:
[0030] ;
[0031] Where, is a global feature, is the classification result; It is a classification network model, consisting of three fully connected layers and two activation layers, which learn data features layer by layer. The activation layers enable the network to have nonlinear mapping capabilities, and the Softmax function is used to convert the network output into the probability of each category for multi-classification.
[0032] Furthermore, the specific training process in S4 is as follows:
[0033] S401: Input the labeled training set samples into the established network model for supervised training, and output the label predictions for the training samples;
[0034] S402: Using the cross entropy loss function, calculate the loss function between the predicted label and the true label; assuming the number of target categories is , the probability distribution of the model output is , the true label is , the cross entropy loss function is expressed as:
[0035] ;
[0036] The network parameters are continuously adjusted and the network model is optimized through the back-propagation algorithm, so that the model can learn the key features of the target track data and the output results match the real target labels.
[0037] The advantages of the present invention compared to the prior art are:
[0038] This paper constructs a multi-level residual cross-attention structure to effectively exploit long-range dependencies and global feature expressions in track data, significantly improving the accuracy of target recognition and classification. This structure combines local feature extraction with global modeling, overcoming the limitations of traditional methods in terms of insufficient long-sequence modeling and global feature extraction. Furthermore, the residual connections between modules enhance the stability of feature transfer and training efficiency, making it suitable for track recognition tasks in complex dynamic scenarios, with excellent stability, robustness, and practical value.
[0039] Through multi-level stacking and residual connections, the proposed model effectively captures the multi-scale features and long-range dependencies in track data, improving feature expression and training stability. Compared with traditional methods, this approach eliminates the need for manual feature design and possesses stronger automatic feature learning and generalization capabilities. It also boasts high computational efficiency and robustness in processing large-scale track data, better meeting the dual requirements of modern air traffic management for accurate and real-time track recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 This is a flow chart of a track recognition and classification method based on multi-level residual cross attention provided by an embodiment of the present invention.
[0041] Figure 2 This is a schematic diagram of the principle of a track recognition and classification method based on multi-level residual cross attention provided by an example of the present invention. DETAILED DESCRIPTION
[0042] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0043] like Figure 1 As shown, the present invention provides a track recognition and classification method based on multi-level residual cross attention, which specifically includes the following steps:
[0044] S1: Obtain target track data, perform preprocessing to resolve outlier points, and divide the data into training and test sets.
[0045] In S1, the preprocessing of the input target track data to solve the problem of outlier points includes removing outliers, filling in points with missing data, and removing points with continuous speed of 0. The specific preprocessing steps are as follows:
[0046] S101: Eliminate outliers whose target position suddenly jumps to a position far away from the set range of the trajectory within a set time;
[0047] S102: For target track data with missing target position or velocity information, fill in the missing data based on adjacent time data by interpolation;
[0048] S103: Eliminate data points where the speed is continuously 0 to prevent target track data points with a stationary time of too long from affecting subsequent trajectory analysis.
[0049] S2: Construct a multi-level residual cross-attention structure to extract global features from the preprocessed target track data, comprehensively capture its long-distance dependencies and multi-scale features, and improve the global semantic modeling capabilities.
[0050] The multi-level residual cross attention structure constructed by the global feature extraction part of the target track data in S2 is composed of multiple residual cross attention modules with the same structure stacked in sequence. Each level of residual cross attention module contains a residual feature extraction module and a cross attention module. While taking into account both local feature modeling and global dependency modeling, it improves the feature expression ability of the network. The specific structure is as follows:
[0051] The multi-level residual crisscross attention structure is composed of multiple residual crisscross attention modules with the same structure stacked in sequence. Each level of residual crisscross attention module contains a residual feature extraction module and a crisscross attention module. The overall structure is expressed as:
[0052] ;
[0053] Among them, ResAttBlock(l) represents the residual cross attention module of the lth layer, which is used to use F(l-1) output by the previous layer as input features for local semantic enhancement and global dependency modeling. The first layer of residual cross attention module uses the preprocessed target track data as input features; L is the number of residual cross attention modules; through the continuous residual connection mechanism between modules, cross-layer information is effectively retained, the gradient vanishing problem in deep networks is alleviated, and the transmission efficiency and training stability of deep features are improved.
[0054] The specific processing process of the residual cross attention module on the input features is:
[0055] The input features are first enhanced by the residual feature extraction module. The residual feature extraction module includes two one-dimensional convolutional layers Conv1 and Conv2 and the ReLU activation function. It also introduces a skip connection to add the input to the convolution result to extract short-term local dependencies, thereby enhancing the underlying semantic expression and alleviating the gradient disappearance. The enhanced feature Fr(l) is obtained. The calculation formula is as follows:
[0056] ;
[0057] After the initial feature enhancement, it enters the cross attention module, which maps the enhanced features Fr(l) into query Q, key K and value V vectors respectively, and then calculates the dependency between features based on the scaled dot product attention mechanism. , the calculation is as follows:
[0058]
[0059] ;
[0060] Where W Q 、W K and W V is the learnable weight matrix, d a is the attention feature dimension;
[0061] The output of the crisscross attention module Perform element-by-element addition with the output Fr(l) of the residual feature extraction module to form a residual connection, and obtain the output result of the residual cross attention module of the current layer. The calculation formula is as follows:
[0062] .
[0063] Through the continuous residual connection mechanism between modules, cross-layer information is effectively retained, the gradient vanishing problem in deep networks is alleviated, and the transmission efficiency and training stability of deep features are improved;
[0064] The above structure achieves a layer-by-layer enhancement from low-level local feature modeling to high-level semantic representation, demonstrating excellent scalability and expressiveness. The feature representation output by the final Lth layer serves as the input to the subsequent track recognition or classification network, effectively improving the model's ability to identify complex track behaviors. In practical applications, this method fully combines the local perception capabilities of convolutional operations with the global dependency modeling capabilities of the attention mechanism. Compared to traditional structures, it can better capture the complex motion patterns and long-range behavioral dependencies present in track sequences, thereby improving the accuracy and robustness of track recognition tasks.
[0065] S3: The extracted global features are input into the classification network for target track classification and recognition, and the input global features are mapped to the labels of the target categories.
[0066] The classification network structure formula in S3 is as follows:
[0067]
[0068] ;
[0069] Where, is a global feature, is the classification result; This is a classification network model consisting of three fully connected layers and two activation layers, which learn data features layer by layer. The activation layers enable the network to have nonlinear mapping capabilities, and the Softmax function is used to convert the network output into the probability of each category for multi-classification.
[0070] The functions of each layer are as follows:
[0071] Input layer: receives latent feature vector;
[0072] The first fully connected layer (DenseLayer): This layer performs weighted processing on the input features and extracts features. Assume that this layer contains N1 neurons;
[0073] The second fully connected layer (DenseLayer): further processes the output of the previous layer and is set to N2 neurons;
[0074] The third fully connected layer (DenseLayer): This layer continues to extract higher-order features and is set to N3 neurons;
[0075] After each fully connected layer, a nonlinear activation function (such as ReLU) is used to introduce nonlinear mapping capabilities, enabling the network to learn more complex patterns.
[0076] Two layers of activation functions are added between the input layer and the output layer, and each activation layer is followed by a fully connected layer. The activation function formula is: ;
[0077] The last layer is a Softmax layer, which is used for multi-classification tasks and converts the output of the network into the probability of each category; the specific formula is:
[0078] ;
[0079] Where, and are the output values of the i-th category and the j-th category respectively. The Softmax function converts these outputs into category probabilities;
[0080] The goal of a classification network is to map an input feature vector to a target class label.
[0081] S4: Use labeled training set samples to perform supervised training on the constructed network model to obtain a trained classification model; wherein the constructed network model includes a multi-level residual cross attention structure and a classification network.
[0082] The specific training process in S4 is as follows:
[0083] S401: Input the labeled training set samples into the established network model for supervised training, and output the label predictions for the training samples;
[0084] S402: Using the cross-entropy loss function, calculate the loss function between the predicted label and the true label. This process is usually achieved by minimizing the loss function. For multi-category classification tasks, the commonly used loss function is the cross-entropy loss. Assume that the number of target categories is , the probability distribution of the model output is , the true label is , the cross entropy loss function is expressed as:
[0085] ;
[0086] The network parameters are continuously adjusted and the network model is optimized through the back-propagation algorithm, so that the model can learn the key features of the target track data and the output results match the real target labels.
[0087] S5: Input the target track data to be identified in the test set into the trained classification model for classification, and output the identification result of the target track.
[0088] In summary, the present invention effectively extracts long-range dependencies and global feature expressions in track data through a deeply stacked residual cross-attention module, thereby improving the model's ability to recognize complex track patterns. This method overcomes the problems of insufficient long-range dependency modeling capabilities and insufficient global feature representation in traditional track recognition, while enhancing the stability of feature transfer and training efficiency through the residual mechanism. Ultimately, efficient, stable, and accurate classification of track targets is achieved, meeting the high-precision and high-robustness requirements of modern air traffic management for intelligent recognition.
Claims
1. A track recognition and classification method based on multi-level residual cross attention, characterized in that: The specific steps include: S1: Obtain target track data, perform preprocessing, and divide it into training set and test set; S2: Construct a multi-level residual cross attention structure to extract global features from the preprocessed target track data; S3: Input the extracted global features into the classification network for target track classification and recognition, and map the input global features to the target category label; S4: Use labeled training set samples to conduct supervised training on the constructed network model to obtain a trained classification model; the constructed network model includes a multi-level residual cross attention structure and a classification network; S5: Input the target track data to be identified in the test set into the trained classification model for classification, and output the identification result of the target track.
2. A track recognition and classification method based on multi-level residual cross attention according to claim 1, characterized in that: The preprocessing steps in S1 are as follows: S101: Eliminate outliers whose target position suddenly jumps to a position far away from the set range of the trajectory within a set time; S102: For target track data with missing target position or velocity information, fill in the missing data based on adjacent time data by interpolation; S103: Eliminate data points where the speed is continuously 0.
3. The track recognition and classification method based on multi-level residual cross attention according to claim 1 is characterized in that: The multi-level residual cross attention structure constructed in S2 is composed of multiple residual cross attention modules of the same structure stacked in sequence. Each level of residual cross attention module contains a residual feature extraction module and a cross attention module. The overall structure is expressed as: ; Where ResAttBlock(l) represents the residual cross attention module of the lth layer, which is used to use F(l-1) output by the previous layer as input features for local semantic enhancement and global dependency modeling. The first layer of residual cross attention module uses the preprocessed target track data as input features; L is the number of residual cross attention modules; The specific processing process of the residual cross attention module on the input features is: The input features are first enhanced by the residual feature extraction module. The residual feature extraction module includes two one-dimensional convolutional layers Conv1 and Conv2 and ReLU activation function, and introduces skip connection. The input and the convolution result are added to obtain the enhanced feature Fr(l). The calculation formula is as follows: ; After the initial feature enhancement, it enters the cross attention module, which maps the enhanced features Fr(l) into query Q, key K and value V vectors respectively, and then calculates the dependency between features based on the scaled dot product attention mechanism. , the calculation is as follows: ; ; Where W Q 、W K and W V is the learnable weight matrix, d a is the attention feature dimension; The output of the crisscross attention module Perform element-by-element addition with the output Fr(l) of the residual feature extraction module to form a residual connection, and obtain the output result of the residual cross attention module of the current layer. The calculation formula is as follows: 。 4. The track recognition and classification method based on multi-level residual cross attention according to claim 1 is characterized in that: The goal of the classification network in S3 is to map the input global feature vector to the label of the target category. The classification network structure formula is as follows: ; Where, is a global feature, is the classification result; It is a classification network model, including three fully connected layers and two activation layers, which learns data features layer by layer; The activation layer enables the network to have nonlinear mapping capabilities, and the output of the network is converted into the probability of each category through the Softmax function for multi-classification.
5. The track recognition and classification method based on multi-level residual cross attention according to claim 1 is characterized in that: The specific training process in S4 is as follows: S401: Input the labeled training set samples into the established network model for supervised training, and output the label predictions for the training samples; S402: Using the cross entropy loss function, calculate the loss function between the predicted label and the true label; assuming the number of target categories is , the probability distribution of the model output is , the true label is , the cross entropy loss function is expressed as: ; The network parameters are continuously adjusted and the network model is optimized through the back-propagation algorithm, so that the model can learn the key features of the target track data and the output results match the real target labels.
Citation Information
Patent Citations
Chinese named entity recognition method based on multilevel residual convolution and attention mechanism
CN112926323A
Target trajectory recognition method based on residual network and attention mechanism
CN115048870A
Target track classification-based intention recognition method and device
CN119089293A
Flight track identification method based on deep residual neural network
CN119903417A
Model training and scene recognition method and apparatus, device, and medium
US20250239057A1
Cited By
Multi-weather robust target detection method and device for electric power inspection and storage medium
CN120807897A