Flight action recognition method and system based on encoder

Through the encoder-based flight action recognition method, the problems of low accuracy and low efficiency of flight action recognition in the prior art are solved, and higher recognition accuracy and stronger fault tolerance are achieved, which are suitable for complex flight parameter data processing.

CN120046033APending Publication Date: 2025-05-27NANJING LUKOU INT AIRPORT AIRPORT TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510060526.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The existing flight action recognition methods have low accuracy and low efficiency, are difficult to process complex flight parameter data, and are insufficient tolerant of data.

Method used

The flight action recognition method based on the encoder is adopted, and data preprocessing is performed by acquiring the original flight data, setting the time step and window width, dividing the data into data samples, and setting an encoder for each data sample. Then, the encoded data samples are input to the multi-head attention layer and the feedforward network layer, and iterates multiple times, and finally the probability distribution of the flight action is obtained through the fully connected layer.

Benefits of technology

It improves the accuracy and efficiency of flight action recognition, enhances the fault tolerance of flight parameter data, and can better handle complex timing data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120046033A_ABST
    Figure CN120046033A_ABST
Patent Text Reader

Abstract

The invention provides a flight action recognition method and system based on an encoder, aims to provide a flight action recognition method adopting a neural network structure, and solves the problems that an existing method is low in recognition accuracy, small in recognition range, limited in flight parameter data quality, incapable of fully utilizing flight parameters and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of flight action recognition, and particularly to an encoder-based flight action recognition method and system. Background Art

[0002] With the development of artificial intelligence, more and more researchers are engaged in the task of flight action recognition, moving towards the goal of achieving objective and efficient flight action recognition. Flight action recognition is a process of automatically identifying and classifying various actions and postures during flight. This technology has wide applications in the fields of aerospace, military, civil aviation, and flight simulation, and has a positive impact on improving flight safety, efficiency, and performance. The automation and efficiency of flight action recognition can also provide an objective and convenient evaluation for pilot assessment. Evaluate the performance of pilots in different scenarios. This helps to cultivate pilots' adaptability, decision-making ability, and operation skills, and improve their response ability in complex environments.

[0003] Pilot training includes flight simulation training and actual aircraft training. The evaluation method for training is generally based on expert experience for flight action recognition and assessment. Flight actions are generally judged based on flight parameter data, mainly including parameters such as flight speed, flight altitude, acceleration, heading angle, pitch angle, and bank angle. The amount of flight parameter data is very large, and various parameters of the aircraft need to be recorded under different flight states and environmental conditions, and the data quality is often not high. Traditional flight action recognition methods are mainly based on expert systems, extracting flight parameter change features from flight parameter data according to expert experience to establish a rule base, and using computer reasoning to identify actions. This method is difficult to identify complex actions and cannot make full use of flight parameter data, resulting in problems such as low accuracy and low efficiency. Another method based on Bayesian networks uses probabilistic reasoning to identify actions. However, this method requires a large amount of data to train the network, and the selection of network structure and parameters is also relatively complex. Summary of the Invention

[0004] In view of the technical problems existing in the above background art, the present invention proposes an encoder-based flight action recognition method and system, and the technical solutions adopted are as follows:

[0005] An encoder-based flight action recognition method, the method comprising:

[0006] S1: Obtain original flight data and perform data preprocessing on the original flight data;

[0007] S2: Set the time step and fixed window width, divide the preprocessed original flight data into several data samples according to the time step and fixed window width, and set an encoder for each data sample to encode the original flight data;

[0008] S3: Input the encoded data samples into the multi-head attention layer, calculate the similarity between each vector in the data samples, and output the similarity index. At the same time, perform a residual connection on the similarity index output by the multi-head attention layer;

[0009] S4: Take the similarity index output in S3 as the input, send it to the feed-forward network layer, and perform a linear mapping on the input. At the same time, combine the encoded original flight data obtained in S2 to output new data samples. At the same time, perform a residual connection on the output of the feed-forward network layer. Take the output of the feed-forward network layer as the input and send it to S2, and repeat this process multiple times;

[0010] S5: Input the data samples after multiple iterations into the fully connected layer to obtain the probability distribution of the data samples.

[0011] Preferably, the data preprocessing in S1 includes: denoising and smoothing the original flight data, filling in missing values, and removing outliers, and converting the original flight parameter data into data suitable for the recognition model.

[0012] Preferably, S2 includes:

[0013] S21: Set the interval of the data samples in the time dimension, set the number of data points included in each sample as the window width value, select a starting point from the preprocessed original data as the starting time step of the first sample, and move forward one time step each time at the time step interval starting from the starting point, and intercept the data within the window width as a sub-sample;

[0014] S22: Repeat the process of S21 until the entire data set is covered, generate several sub-samples, and set an encoder for each sub-sample.

[0015] Preferably, S3 includes:

[0016] Take the encoded samples as the input and input them into the multi-head attention layer. At the same time, calculate the similarity between each vector in the data samples and output the similarity index. At the same time, perform a residual connection on the similarity index output by the multi-head attention layer. And the similarity index is obtained through the following formula:

[0017]

[0018] Among them, a and b represent two vectors to be compared in the data sample, and H(a) and H(b) represent the information entropy of vectors a and b. Moreover, H(a) and H(b) are obtained through the following calculation formula:

[0019]

[0020] Among them, P i (a) and P i (b) represent the normalized probability values of the i-th elements in vectors a and b.

[0021] Preferably, the S4 includes:

[0022] S41: Send the similarity index into the feed-forward network. The feed-forward network includes a layer of linear fully connected layers, which perform a linear mapping on the input. Among them, the linear fully connected layer maps the input to a higher-dimensional space and fuses it with the encoded original flight data obtained in S2 to obtain a new data sample;

[0023] S42: Establish a residual connection between the new data sample obtained in S41 and the encoder in S2. Take the new data sample obtained in S41 as the input and send it to the encoder in step S2 for encoding.

[0024] Preferably, the S5 includes:

[0025] S51: Obtain the data samples after multiple iterations and preprocess the iterated data samples to ensure that the format and type of the data samples match the input requirements of the fully connected layer;

[0026] S52: Pass through the fully connected layer, perform a linear transformation on the preprocessed data samples, and input the output of the linear transformation into the Softmax layer to convert the output into a probability distribution. Use the probability distribution output by Softmax as the final output probability, and each element represents the prediction probability of the data sample on each flight action.

[0027] An encoder-based flight action recognition system, characterized in that the system includes:

[0028] Data collection system: Obtain the original flight data and perform data preprocessing on the original flight data;

[0029] Data encoding system: Set the time step and the fixed window width, divide the preprocessed original flight data into several data samples according to the time step and the fixed window width, and set an encoder for each data sample to encode the original flight data;

[0030] Multi-Head Attention Layer: Input the encoded data samples into the multi-head attention layer, calculate the similarity between each vector in the data samples, and output the similarity index. At the same time, perform a residual connection on the similarity index output by the multi-head attention layer;

[0031] Feed-Forward Network Layer: Take the similarity index output by S3 as the input, send it to the feed-forward network layer, perform a linear mapping on the input, and combine it with the encoded original flight data obtained from S2 to output new data samples. At the same time, perform a residual connection on the output of the feed-forward network layer. Take the output of the feed-forward network layer as the input and send it to S2, and repeat this process multiple times;

[0032] Fully Connected Layer: Input the data samples after multiple iterations into the fully connected layer to obtain the probability distribution of the data samples.

[0033] Preferably, the data encoding system includes:

[0034] Time Step and Window Width Setting System: Set the interval of data samples in the time dimension, set the number of data points included in each sample as the window width value, select a starting point from the preprocessed original data as the starting time step of the first sample, and take the data within the window width as a sub-sample at intervals of the time step. Starting from the starting point, move forward one time step each time;

[0035] Repetition System: Repeat cyclically until the entire dataset is covered, generate several sub-samples, and set an encoder for each sub-sample.

[0036] Preferably, the multi-head attention layer includes:

[0037] Linear Calculation Layer: Send the similarity index into the feed-forward network. The feed-forward network includes a linear fully connected layer, which performs a linear mapping on the input. The linear fully connected layer maps the input to a higher-dimensional space and fuses it with the encoded original flight data obtained from S2 to obtain new data samples;

[0038] Iteration System: Establish a residual connection between the new data samples obtained from S41 and the encoder in S2, and take the new data samples obtained from S41 as the input and send them to the encoder in step S2 for encoding.

[0039] Preferably, the fully connected layer includes:

[0040] Data Preprocessing System: Obtain the data samples after multiple iterations and preprocess the iterated data samples to ensure that the format and type of the data samples match the input requirements of the fully connected layer;

[0041] Probability output system: Through a fully connected layer, perform a linear transformation on the preprocessed data samples, and input the output of the linear transformation into the Softmax layer to convert the output into a probability distribution. Use the probability distribution output by Softmax as the final output probability. Each element represents the predicted probability of the data sample for each flight action.

[0042] Advantages of the present invention: 1. The encoder part of the Transformer model is adopted in the present invention. On the basis of maintaining the advantages of the Transformer model, the model training intensity is reduced, and it has more general applicability.

[0043] 2. The neural network architecture adopted in the present invention uses the self-attention mechanism and positional encoding design, which can more carefully process the dependency relationship between time series data and has higher representation ability.

[0044] 3. Compared with traditional methods, the method adopted in the present invention for flight action recognition has higher recognition accuracy, stronger plasticity, and higher fault tolerance for flight parameter data.

[0045] 4. The neural network architecture adopted in the present invention can handle general time series data classification problems and can take into account the flight parameter data of various different aircraft models. Description of the Drawings

[0046] Figure 1 A flight action recognition method based on an encoder according to the present invention. Detailed Embodiments

[0047] The following describes the preferred embodiments of the present invention with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are only for explaining and illustrating the present invention and are not used to limit the present invention.

[0048] An embodiment of the present invention, a flight action recognition method based on an encoder, the method comprising:

[0049] S1: Obtain the original flight data and perform data preprocessing on the original flight data;

[0050] S2: Set the time step and fixed window width, divide the preprocessed original flight data into several data samples according to the time step and fixed window width, and set an encoder for each data sample to encode the original flight data;

[0051] S3: Input the encoded data samples into the multi-head attention layer, calculate the similarity between each vector in the data samples, and output the similarity index. At the same time, perform a residual connection on the similarity index output by the multi-head attention layer;

[0052] S4: Use the similarity index output in S3 as the input, send it to the feed-forward network layer, perform a linear mapping on the input, and at the same time combine the encoded original flight data obtained in S2 to output a new data sample. At the same time, perform a residual connection on the output of the feed-forward network layer, use the output of the feed-forward network layer as the input, and send it to S2, repeating this process multiple times;

[0053] S5: Input the data sample after multiple iterations into the fully connected layer to obtain the probability distribution of the data sample.

[0054] The working principle and effects of the above technical solution are as follows: First, obtain the original flight data, which includes information such as sensor readings, GPS data, speed, and attitude. Preprocess the data, including denoising, filling missing values. Set the time step and fixed window width, which can determine the length and granularity of each data sample. Divide the preprocessed flight data into several data samples according to these parameters. Set an encoder for each data sample to extract features from the data sample. Input the encoded data sample into the multi-head attention layer in the model. The multi-head attention mechanism can calculate the similarity between input vectors through multiple different "heads", output the similarity index, that is, the attention score, and at the same time apply a residual connection during the calculation of the output to stabilize the learning process and improve the gradient flow of the model. Input the similarity index obtained in step 3 into the feed-forward network layer for further processing. At the same time, combine the encoded original data in step 2 to generate a new data sample through linear mapping. The output data sample undergoes a residual connection and is iterated multiple times, enabling the network to gradually extract higher-level features at each layer. After multiple iterations, the final data sample obtained is sent into the fully connected layer. The final model output is the probability distribution of each flight action, indicating the possibility that the input data is in different flight actions. By using an encoder to encode the original flight data, important features in the data can be effectively extracted, reducing the dimension of the data, thereby improving the efficiency of classification and recognition. The multi-head attention mechanism can capture the similarity between data from different perspectives, is suitable for processing complex time-series data, and can identify complex patterns and relationships hidden in the data. The residual connection provides a shortcut path in the network to help solve the problem of gradient disappearance that may occur in deep networks, thereby enhancing the stability of the model. By setting the time step and window width, the model can dynamically adapt to input data samples of different lengths and flexibly process flight actions at different stages. Through the iterative process between the feed-forward network layer and the encoder, the model can gradually refine the understanding and extraction of features, improving the classification accuracy.

[0055] In an embodiment of the present invention, the data preprocessing in S1 includes: performing denoising and smoothing, filling missing values, and deleting outlier processing on the original flight data, and converting the original flight parameter data into data suitable for the recognition model.

[0056] The working principle and effects of the above technical solution are as follows: estimate and fill in the missing data to ensure data integrity, reduce the problem of unstable model training caused by data missing, detect and delete outliers that significantly deviate from the normal range, and identify outliers by the IQR (Interquartile Range) method. This method can eliminate data bias caused by outliers and improve the representativeness of the dataset and model performance. Convert the original flight parameter data into a data format suitable for the recognition model. The converted data is more suitable for the input requirements of the model in terms of numerical range and format, improving the model convergence speed and accuracy.

[0057] In one embodiment of the present invention, S2 includes:

[0058] S21: Set the interval of data samples in the time dimension, set the number of data points included in each sample as the window width value, select a starting point from the preprocessed original data as the starting time step of the first sample, and take the data within the window width as a subsample by moving forward one time step each time at the interval of the time step starting from the starting point;

[0059] S22: Repeat the process of S21 until the entire dataset is covered, generate several subsamples, and set an encoder for each subsample.

[0060] The working principle and effects of the above technical solution are as follows: By means of the time step, overlap between subsamples is allowed, in this way, more dynamic changes can be captured, and the model can capture short-term features and long-term dependencies. The window width determines the length of each subsample, enabling each sample to focus on local information, and the encoder extracts and represents features from the data on this basis. According to different task requirements, the time step and window width can be adjusted, thereby adjusting the fineness of data segmentation, making the method very flexible in adapting to data characteristics. By means of the sliding window, information loss is minimized as much as possible to ensure that the model can make full use of the information of the entire dataset during training.

[0061] In one embodiment of the present invention, S3 includes:

[0062] Take the encoded samples as input and input them into the multi-head attention layer, calculate the similarity between each vector in the data samples at the same time, and output the similarity index. At the same time, perform residual connection on the similarity index output by the multi-head attention layer. And the similarity index is obtained through the following formula:

[0063]

[0064] Among them, a and b represent two vectors to be compared in the data sample, and H(a) and H(b) represent the information entropy of vector a and vector b. Moreover, H(a) and H(b) are obtained through the following calculation formula:

[0065]

[0066] Among them, P i (a) and P i (b) represent the normalized probability values of the i-th element in vector a and vector b.

[0067] Since multiple iterations may cause the similarity index to be repeatedly transmitted between S2 and the feedforward network layer, thus causing numerical instability or convergence problems. Based on this problem, an evaluation system is set up to evaluate the stability of the similarity during the iteration process. Moreover, the evaluation result is obtained through the following formula:

[0068]

[0069] Among them, N represents the number of iterations, U i+1 represents the similarity index obtained at the (i + 1)-th iteration, and U i represents the similarity index obtained at the i-th iteration;

[0070] When S is close to 0, it means that the similarity index is very stable during the iteration process and hardly changes. At this time, the iteration can be stopped;

[0071] When 0.1 < S < 0.5, it means that the similarity index has certain fluctuations during the iteration process, but the fluctuation range is not large. At this time, the iteration should continue to observe the change trend of the similarity index. If the fluctuation gradually decreases and tends to be stable, the iteration can be considered to stop. If the fluctuation persists or increases, the algorithm structure needs to be optimized, and the number of iterations does not exceed 6 at most;

[0072] When 0.5 ≤ S < 1, it means that the similarity index has large fluctuations during the iteration process. At this time, the encoded data sample should be cleaned again.

[0073] The working principle and effects of the above technical solution are as follows: This formula uses a calculation method of multiplying the cosine value of the included angle and the difference in information entropy to calculate the similarity between each pair of vectors in the data sample. Among them, the cosine value of the included angle is an index to measure the similarity of the directions of two vectors. In the vector space, the smaller the included angle between two vectors, that is, the larger the cosine value, the more similar they are. In the scenario of the present invention, the cosine value of the included angle can capture the similarity in direction of the vectors. Information entropy is an index to measure the uncertainty of data. In information theory, the higher the entropy, the greater the uncertainty of the data, and vice versa. Multiplying the cosine value of the included angle and the difference in information entropy can comprehensively consider the performance of two data objects in terms of direction similarity and information content difference. This combination helps to capture more comprehensive similarity information. Even if two vectors are very similar in direction (large cosine value of the included angle), but if there is a large difference in their information content (large difference in information entropy), the final similarity score may be low. On the contrary, if two vectors are similar in direction and have similar information content, the final similarity score will be high. Through formula calculation, the stability of the algorithm during the iterative process can be quantified. This helps us to more accurately understand the performance of the algorithm and avoid judging the convergence and stability of the algorithm solely based on experience or intuition. When the algorithm reaches a stable state during the iterative process, it can be judged and the iteration can be stopped in advance through the formula. This helps to save computing resources and improve the efficiency of the algorithm.

[0074] In an embodiment of the present invention, the S4 includes:

[0075] S41: Send the similarity index into the feedforward network. The feedforward network includes a linear fully connected layer, which performs a linear mapping on the input. Among them, the linear fully connected layer maps the input to a higher-dimensional space and fuses it with the encoded original flight data obtained in S2 to obtain a new data sample;

[0076] S42: Establish a residual connection between the new data sample obtained in S41 and the encoder in S2, and use the new data sample obtained in S41 as the input and send it to the encoder in step S2 for encoding.

[0077] The working principle and effects of the above technical solution are as follows: The feedforward network layer performs a linear mapping on the input, that is, each input node is connected to each output node, and each connection has a weight. By weighted summation and adding a bias, the value of the output node is obtained. The role of the linear fully connected layer is to map the input to a higher-dimensional space to better capture the features of the data. The residual connection is a commonly used network structure in deep learning. It adds the input and output of the network, enabling the network to better learn the residual information and improve the performance and training effect of the network. In step S42, the new data sample obtained in step S41 is used as the input and sent to the encoder in step S2 for encoding. At the same time, through the residual connection, the new data sample is added to the output of the encoder. In this way, the encoder not only receives the new data sample as the input but also receives the residual information from the previous layer. In a deep network, the vanishing gradient is a common problem. The residual connection helps the gradient flow better during backpropagation by introducing direct cross-layer connections, thus alleviating the problem of vanishing gradients. Since the residual connection allows the network to directly learn the difference between the input and output (i.e., the residual), this helps the network converge to the optimal solution faster. Through the residual connection, the network can more easily learn the deep feature representation of the data, thereby improving the performance of the model.

[0078] In one embodiment of the present invention, S5 includes:

[0079] S51: Obtain the data samples after multiple iterations, and preprocess the iterated data samples to ensure that the format and type of the data samples match the input requirements of the fully connected layer;

[0080] S52: Pass through the fully connected layer, perform a linear transformation on the preprocessed data samples, and input the output of the linear transformation into the Softmax layer to convert the output into a probability distribution. Use the probability distribution output by Softmax as the final output probability, and each element represents the predicted probability of the data sample for each flight action.

[0081] The working principle and effects of the above technical solution are as follows: The data preprocessing system obtains data samples after multiple iterations, and performs data cleaning on the iterated data samples. The preprocessed data samples are input into the fully connected layer for linear transformation. The purpose of this linear transformation is to map the data samples into a new feature space for better classification or regression tasks. The Softmax layer is usually used as the output layer for multi-classification problems. It converts the output of the fully connected layer into a probability distribution, such that the value of each output node is between 0 and 1, and the sum of the values of all output nodes is 1. The Softmax function obtains the prediction probability for each category by calculating the ratio of the exponent of each category to the sum of the exponents of all categories. In this way, each data sample is assigned a prediction probability distribution over all flight actions. Finally, the probability distribution output by the Softmax layer is used as the final output probability. This probability distribution represents the prediction probability of the data sample over each flight action and can be used for subsequent decision-making or classification tasks. The Softmax layer converts the output of the fully connected layer into a probability distribution, which means that the output of each category can be interpreted as the probability that this category is the correct category. This probability interpretability makes the output of the model more intuitive and easier to understand. The Softmax function obtains the prediction probability by calculating the ratio of the exponent of each category to the sum of the exponents of all categories. This method has good stability when dealing with numerical values. Due to the property of the exponential function of rapid growth, even if there are large differences between the original output values, after being processed by the Softmax function, these differences will be smoothly mapped into the probability range between 0 and 1, thus avoiding problems of numerical overflow or underflow.

[0082] An embodiment of the present invention, a flight action recognition system based on an encoder, characterized in that the system includes:

[0083] A data collection system: obtaining original flight data and performing data preprocessing on the original flight data;

[0084] A data encoding system: setting a time step and a fixed window width, dividing the preprocessed original flight data into a number of data samples according to the time step and the fixed window width, and setting an encoder for each data sample to encode the original flight data;

[0085] A multi-head attention layer: inputting the encoded data samples into the multi-head attention layer, calculating the similarity between each vector in the data samples, and outputting a similarity index, and at the same time performing a residual connection on the similarity index output by the multi-head attention layer;

[0086] Feed-forward network layer: Take the similarity index of the output of S3 as the input, send it to the feed-forward network layer, perform a linear mapping on the input, and combine it with the encoded original flight data obtained from S2 to output a new data sample. At the same time, perform a residual connection on the output of the feed-forward network layer, take the output of the feed-forward network layer as the input, and send it to S2, and repeat this process multiple times;

[0087] Fully connected layer: Input the data samples after multiple iterations into the fully connected layer to obtain the probability distribution of the data samples.

[0088] The working principle and effects of the above technical solution are as follows: First, obtain the original flight data, which includes information such as sensor readings, GPS data, speed, and attitude. Preprocess the data, including denoising and filling in missing values. Set the time step and fixed window width, which can determine the length and granularity of each data sample. Divide the preprocessed flight data into several data samples according to these parameters. Set an encoder for each data sample to extract features from the data sample. Input the encoded data sample into the multi-head attention layer in the model. The multi-head attention mechanism can calculate the similarity between input vectors through multiple different "heads" and output the similarity index, that is, the attention score. At the same time, a residual connection is applied during the calculation of the output to stabilize the learning process and improve the gradient flow of the model. Input the similarity index obtained in step 3 into the feed-forward network layer for further processing. At the same time, combine the encoded original data in step 2 to generate a new data sample through linear mapping. The output data sample undergoes a residual connection and is iterated multiple times, enabling the network to gradually extract higher-level features at each layer. After multiple iterations, the final data sample obtained is sent into the fully connected layer. The final model output is the probability distribution of each flight action, indicating the likelihood that the input data is in different flight actions. By using an encoder to encode the original flight data, important features in the data can be effectively extracted, reducing the dimension of the data, thereby improving the efficiency of classification and recognition. The multi-head attention mechanism can capture the similarity between data from different perspectives, is suitable for processing complex time-series data, and can identify complex patterns and relationships hidden in the data. The residual connection provides a shortcut path in the network to help solve the problem of gradient disappearance that may occur in deep networks, thereby enhancing the stability of the model. By setting the time step and window width, the model can dynamically adapt to input data samples of different lengths and flexibly process flight actions at different stages. Through the iterative process between the feed-forward network layer and the encoder, the model can gradually refine the understanding and extraction of features, improving the classification accuracy.

[0089] In an embodiment of the present invention, the data encoding system includes:

[0090] Time step and window width setting system: Set the interval of data samples in the time dimension, set the number of data points included in each sample as the window width value, select a starting point from the preprocessed original data as the starting time step of the first sample, and at intervals of the time step, starting from the starting point, move forward one time step each time, and intercept the data within the window width as a subsample;

[0091] Repeating system: Loop and repeat until the entire data set is covered, generate a number of subsamples, and set an encoder for each subsample.

[0092] The working principle and effect of the above technical solution are as follows: Through the time step, overlap between subsamples is allowed. In this way, more dynamic changes can be captured, and the model can capture short-term features and long-term dependencies. The window width determines the length of each subsample, enabling each sample to focus on local information. Based on this, the encoder extracts and represents features from the data. According to different task requirements, the time step and window width can be adjusted, thereby adjusting the fineness of data segmentation, making the method very flexible in adapting to data characteristics. By means of a sliding window, information loss is minimized as much as possible to ensure that the model can fully utilize the information of the entire data set during training.

[0093] In an embodiment of the present invention, the multi-head attention layer includes:

[0094] Linear calculation layer: Send the similarity index into the feed-forward network. The feed-forward network includes a linear fully-connected layer, which performs a linear mapping on the input. Among them, the linear fully-connected layer maps the input to a higher-dimensional space and fuses it with the encoded original flight data obtained by S2 to obtain a new data sample;

[0095] Iterative system: Establish a residual connection between the new data sample obtained by S41 and the encoder in S2, take the new data sample obtained by S41 as the input, and send it to the encoder in step S2 for encoding.

[0096] The working principle and effects of the above technical solution are as follows: The feedforward network layer performs a linear mapping on the input, that is, each input node is connected to each output node, and each connection has a weight. By weighted summation and adding a bias, the value of the output node is obtained. The role of the linear fully connected layer is to map the input to a higher-dimensional space to better capture the features of the data. Residual connection is a commonly used network structure in deep learning. It adds the input and output of the network, enabling the network to better learn the residual information and improve the performance and training effect of the network. In step S42, the new data sample obtained in step S41 is used as the input and sent to the encoder in step S2 for encoding. At the same time, through the residual connection, the new data sample is added to the output of the encoder. In this way, the encoder not only receives the new data sample as the input but also receives the residual information from the previous layer. In a deep network, the vanishing gradient is a common problem. Residual connection helps the gradient flow better during backpropagation by introducing direct cross-layer connections, thus alleviating the problem of vanishing gradient. Since residual connection allows the network to directly learn the difference between the input and output (i.e., the residual), this helps the network converge to the optimal solution faster. Through residual connection, the network can more easily learn the deep feature representation of the data, thereby improving the performance of the model.

[0097] In one embodiment of the present invention, the fully connected layer includes:

[0098] Data preprocessing system: Obtain the data samples after multiple iterations and preprocess the iterated data samples to ensure that the format and type of the data samples match the input requirements of the fully connected layer;

[0099] Probability output system: Through the fully connected layer, perform a linear transformation on the preprocessed data samples, and input the output of the linear transformation into the Softmax layer to convert the output into a probability distribution, and use the probability distribution output by Softmax as the final output probability. Each element represents the predicted probability of the data sample on each flight action.

[0100] The working principle and effects of the above technical solution are as follows: The data preprocessing system obtains the data samples after multiple iterations, and performs data cleaning on the iterated data samples. The preprocessed data samples are input into the fully connected layer for linear transformation. The purpose of this linear transformation is to map the data samples to a new feature space for better classification or regression tasks. The Softmax layer is usually used as the output layer for multi-classification problems. It converts the output of the fully connected layer into a probability distribution, such that the value of each output node is between 0 and 1, and the sum of the values of all output nodes is 1. The Softmax function obtains the predicted probability for each class by calculating the ratio of the exponential of each class to the sum of the exponentials of all classes. In this way, each data sample is assigned a predicted probability distribution over all flight actions. Finally, the probability distribution output by the Softmax layer is used as the final output probability. This probability distribution represents the predicted probability of the data sample over various flight actions and can be used for subsequent decision-making or classification tasks. The Softmax layer converts the output of the fully connected layer into a probability distribution, which means that the output of each class can be interpreted as the probability that this class is the correct class. This probabilistic interpretability makes the output of the model more intuitive and easy to understand. The Softmax function obtains the predicted probability by calculating the ratio of the exponential of each class to the sum of the exponentials of all classes. This method has good stability when dealing with numerical values. Due to the property of the exponential function of rapid growth, even if there are large differences between the original output values, after being processed by the Softmax function, these differences will be smoothly mapped to the probability range between 0 and 1, thus avoiding the problems of numerical overflow or underflow.

[0101] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these changes and modifications.

Claims

1. A flight action recognition method based on an encoder, characterized in that: The method comprises: S1: acquiring original flight data, and performing data preprocessing on the original flight data; S2: setting a time step and a fixed window width, dividing the preprocessed original flight data into a number of data samples according to the time step and the fixed window width, and setting an encoder for each data sample to encode the original flight data; S3: Input the encoded data sample into the multi-head attention layer, calculate the similarity between each vector in the data sample, and output the similarity index. At the same time, perform residual connection on the similarity index output by the multi-head attention layer. S4: Take the similarity index of the output of S3 as input, send it to the feedforward network layer, perform linear mapping on the input, and combine it with the encoded original flight data obtained by S2, output a new data sample, and perform residual connection on the output of the feedforward network layer, take the output of the feedforward network layer as input, and send it to S2, and repeat this process multiple times; S5: Input the data samples after multiple iterations into the fully connected layer to obtain the probability distribution of the data samples.

2. The method for identifying flight actions based on an encoder according to claim 1, characterized in that: The data preprocessing in S1 includes: performing denoising and smoothing, gap filling and outlier deletion processing on the original flight data, and converting the original flight parameter data into data that is suitable for the recognition model.

3. The method for identifying flight actions based on an encoder according to claim 1, characterized in that: The S2 includes: S21: Set the interval of data samples in the time dimension, set the number of data points contained in each sample, and use it as the window width value. Select a starting point from the preprocessed raw data as the starting time step of the first sample. Use the time step as the interval, start from the starting point, move forward one time step each time, and intercept the data within the window width as a sub-sample. S22: Repeat the process of S21 until the entire data set is covered, generate several sub-samples, and set an encoder for each sub-sample.

4. The method for identifying flight actions based on an encoder according to claim 1, characterized in that: The S3 includes: The encoded sample is used as input to the multi-head attention layer, and the similarity between each vector in the data sample is calculated and the similarity index is output. At the same time, the similarity index output by the multi-head attention layer is residually connected, and the similarity index is obtained by the following formula: Where a and b represent two vectors to be compared in the data sample, H(a) and H(b) represent the information entropy of vector a and vector b, and H(a) and H(b) are obtained by the following calculation formula: Among them, P i (a) and P i (b) represents the normalized probability value of the i-th element in vector a and vector b.

5. The method for identifying flight actions based on an encoder according to claim 1, characterized in that: The S4 includes: S41: sending the similarity index to a feedforward network, wherein the feedforward network includes a linear fully connected layer, which performs linear mapping on the input, wherein the linear fully connected layer maps the input to a higher dimensional space and fuses it with the encoded original flight data obtained in S2 to obtain a new data sample; S42: Establish a residual connection between the new data sample obtained in S41 and the encoder in S2, and send the new data sample obtained in S41 as input to the encoder in step S2 for encoding.

6. The method for identifying flight actions based on an encoder according to claim 1, characterized in that: The S5 includes: S51: obtaining data samples after multiple iterations, and preprocessing the data samples after the iterations to ensure that the format and type of the data samples match the input requirements of the fully connected layer; S52: Through the fully connected layer, the preprocessed data samples are linearly transformed, and the output of the linear transformation is input to Softmax to convert the output into a probability distribution. The probability distribution output by Softmax is used as the final output probability, and each element represents the predicted probability of the data sample for each flight action.

7. A flight action recognition system based on an encoder, characterized in that: The system comprises: Data collection system: acquiring raw flight data and performing data preprocessing on the raw flight data; Data encoding system: set the time step and fixed window width, divide the pre-processed raw flight data into several data samples according to the time step and fixed window width, and set an encoder for each data sample to encode the raw flight data; Multi-head attention layer: Input the encoded data samples into the multi-head attention layer, calculate the similarity between each vector in the data sample, and output the similarity index. At the same time, perform residual connection on the similarity index output by the multi-head attention layer; Feedforward network layer: The similarity index of the output of S3 is used as input and sent to the feedforward network layer, and the input is linearly mapped, and the encoded original flight data obtained by S2 is combined to output new data samples. At the same time, the output of the feedforward network layer is residually connected, and the output of the feedforward network layer is used as input and sent to S2, and this is repeated many times; Fully connected layer: Input the data samples after multiple iterations into the fully connected layer to obtain the probability distribution of the data samples.

8. The encoder-based flight action recognition system according to claim 7, characterized in that: The data encoding system comprises: Time step and window width setting system: set the interval of data samples in the time dimension, set the number of data points contained in each sample and use it as the window width value, select a starting point from the preprocessed raw data as the starting time step of the first sample, use the time step as the interval, start from the starting point, move forward one time step each time, and intercept the data within the window width as a sub-sample; Repeating system: The loop repeats until the entire dataset is covered, generating several subsamples and setting an encoder for each subsample.

9. The encoder-based flight action recognition system according to claim 7, characterized in that: The multi-head attention layer includes: Linear calculation layer: The similarity index is sent to the feedforward network, which includes a linear fully connected layer to perform linear mapping on the input. The linear fully connected layer maps the input to a higher dimensional space and fuses it with the encoded original flight data obtained by S2 to obtain a new data sample. Iterative system: Establish a residual connection between the new data sample obtained in S41 and the encoder in S2, and send the new data sample obtained in S41 as input to the encoder in step S2 for encoding.

10. The encoder-based flight action recognition system according to claim 7, characterized in that: The fully connected layer includes: Data preprocessing system: obtains data samples after multiple iterations and preprocesses the data samples after iterations to ensure that the format and type of the data samples match the input requirements of the fully connected layer; Probability output system: Through the fully connected layer, the preprocessed data samples are linearly transformed, and the output of the linear transformation is input into Softmax to convert the output into a probability distribution. The probability distribution output by Softmax is used as the final output probability. Each element represents the predicted probability of the data sample for each flight action.

Citation Information

Patent Citations

  • Maneuvering target tracking method and device based on Unet framework and multiple attention mechanisms

    CN116879879A

  • Intention-based aircraft trajectory prediction method

    CN117349594A

  • Transformer based video coding

    US20240267543A1

  • Image processing method and apparatus, and device and storage medium

    WO2022116104A1