Online Recognition Method of Target Tactical Intention Based on Deep Learning in a Simulated Environment

By adopting deep learning methods in an air combat simulation environment, combining situation and sensor information, a target tactical intention recognition model is established, and the problems of slow recognition speed and insufficient accuracy in traditional methods are solved, achieving more efficient tactical intention recognition.

CN115204286BActive Publication Date: 2025-07-11NORTHWESTERN POLYTECHNICAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210810491.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-11
Publication Date
2025-07-11
Estimated Expiration
2042-07-11

AI Technical Summary

Technical Problem

It is difficult for the existing technology to accurately identify the target's tactical intentions in air combat simulation, especially in complex battlefield environments with high dynamics and strong games. Traditional methods cannot effectively utilize the situation information of our own and enemy, resulting in slow recognition speed, insufficient accuracy and poor generalization ability.

Method used

The online recognition method of target tactical intentions based on deep learning is adopted. By obtaining real-time battlefield information in the simulated environment, including the situation characteristics and sensor status information of one's and the opponent's side, a target tactical intention space and feature description model are established, and feature extraction and classification recognition are used for cascaded convolutional layers, bidirectional long and short-term memory neural networks and self-attention mechanisms.

Benefits of technology

It improves the reliability and practicality of target tactical intention recognition, can identify target tactical intentions faster and more accurately, adapt to the complexity of modern air combat simulation scenarios, and has higher recognition speed and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115204286B_ABST
    Figure CN115204286B_ABST
Patent Text Reader

Abstract

The present invention relates to an online recognition method for target tactical intentions based on deep learning in a simulated environment, including: Step 1: Obtain real-time battlefield information in the simulated environment, where the battlefield information includes the situation feature information of both the friendly side and the opposing side, as well as the sensor status information; Step 2: Perform various normalization processes on the battlefield information, and input the battlefield information after various normalization processes into the trained target tactical intention recognition model to obtain the recognition result of the target tactical intention. The online recognition method for target tactical intentions based on deep learning in the simulated environment of the present invention proposes the concept of target tactical behavior intention recognition for the problem that it is difficult to recognize the target intention from the task level according to the existing situation information. It recognizes the target tactical behavior from the perspective of air combat confrontation, and its result is more practical. Compared with the traditional task intention, the tactical intention feature is more obvious, and its recognition result is more reliable and practical.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer simulation and artificial intelligence, and particularly relates to an online recognition method for target tactical intention based on deep learning in a simulation environment. Background Art

[0002] An air combat simulation system is to conduct a detailed and realistic simulation of the entire combat process of a fighter jet by means of computer simulation. In order to effectively improve the authenticity of the user experience and the ease of operation of the combat game and simulation system, it is necessary to simulate and design the combat game and simulation system from the perspective of actual air combat. More importantly, it is the tactical simulation and its convenient interactive design, so as to improve the user's operation level in the combat game and simulation system while restoring the authenticity of air combat. Accurately and online identifying the target intention during the confrontation process between the air combat simulation and the target can create a tactical advantage for the fighter jet through in-depth situation awareness, which is the key to seizing air superiority and defeating the enemy, and also the basis for realizing intelligent auxiliary decision-making.

[0003] Currently, the main methods for air combat intention recognition are expert systems, Bayesian, evidence theory, and fuzzy inference. They mainly establish the mapping relationship between intention and situation characteristics by summarizing and inducing expert experience or real air combat cases; or establish an air combat knowledge graph through expert experience, and analyze the input situation characteristic information through the idea of template matching to match the corresponding intention; or obtain the Pearson correlation coefficient matrix through fuzzy mathematics theory, stratify the intention types, and conduct hierarchical recognition according to the single-step situation characteristic information. With the continuous development of the machine learning field, people have begun to apply algorithms such as deep neural networks to intention recognition. The main methods include using support vector machines (SVM), neural networks, etc. to cluster features. Such methods do not need to establish a recognition process model, but use learning algorithms to establish an implicit function mapping relationship from situation information to tactical intention.

[0004] The research on the prior art mainly focuses on mission intentions. However, in modern information-based air combat, the complex battlefield environment information with high dynamics and intense game makes it difficult to identify the target intention from the mission level based on the existing situation information. Considering the dynamic game between the two sides in the air combat process, it is difficult to accurately identify the tactical intention only based on the enemy's motion state information. Most of the research only considers the enemy's information, lacks the introduction of the motion state information of our side for comprehensive judgment, and ignores the working state information of the sensor. The target tactical intention reflects a series of changes in tactical behaviors, which are intuitively manifested as the situation changes in the time-space domain, including track information and sensor working state information, etc. It has the characteristics of continuity and dynamics. Therefore, it is impossible to accurately identify the target intention only based on the information at the current moment. The RNN network plays a significant role in processing time series features, but the LSTM network only considers unidirectional time information, ignores the bidirectional characteristics of the time series, and the deep neural network has the characteristics of slow convergence speed and slow recognition speed, which cannot meet the timeliness. At the same time, there are also problems such as insufficient online recognition accuracy, small samples, and poor generalization ability. Summary of the Invention

[0005] In order to solve the above problems existing in the prior art, the present invention provides an online recognition method for target tactical intention based on deep learning in a simulated environment. The technical problems to be solved by the present invention are realized through the following technical solutions:

[0006] The present invention provides an online recognition method for target tactical intention based on deep learning in a simulated environment, including:

[0007] Step 1: Obtain the real-time battlefield information in the simulated environment, where the battlefield information includes the situation feature information of both the friendly side and the opposing side and the sensor state information;

[0008] Step 2: Perform various normalization processes on the battlefield information, and input the battlefield information after various normalization processes into the trained target tactical intention recognition model to obtain the recognition result of the target tactical intention.

[0009] In an embodiment of the present invention, the situation feature information of both sides includes: the altitude of the opposing target, the altitude of the friendly fighter, the speed of the opposing target, the speed of the friendly fighter, the relative altitude, the relative distance, the target entry angle, and the target azimuth angle;

[0010] The sensor state information includes: the air combat ability of the opposing target, the air combat ability of the friendly fighter, the interference state of the friendly fighter, and the recognition result of the radar state of the opposing target.

[0011] In one embodiment of the present invention, various normalization processes are performed on the battlefield information, including: normalizing three features, namely the target entry angle, the target azimuth angle, and the relative height in the battlefield information in a normalization manner from -1 to 1, and normalizing the remaining features in the battlefield information in a normalization manner from 0 to 1.

[0012] In one embodiment of the present invention, the training process of the target tactical intention recognition model includes:

[0013] S1: Establish a target tactical intention space and a feature description model;

[0014] S2: Obtain an intention sample data set, and preprocess the intention sample data set to obtain a sample set;

[0015] S3: Construct a target tactical intention recognition network, which includes a cascaded input layer, a feature extraction module, and a classification and recognition module. Among them, the feature extraction module includes a cascaded convolutional layer, a bidirectional long short-term memory neural network layer, and a self-attention mechanism layer;

[0016] S4: Divide the sample set into a training set and a validation set, input the training set and the validation set into the target tactical intention recognition network to train and optimize its network parameters to obtain the optimal network parameters, and obtain the trained target tactical intention recognition model according to the optimal network parameters.

[0017] In one embodiment of the present invention, the S1 includes:

[0018] S11: Establish a target tactical intention space, which includes five intention types: attack, defense, detection, interference, and escape;

[0019] S12: According to the target track and sensor working state information, establish a feature description model, and the feature description model is:

[0020]

[0021] In the formula, represents the feature information of the i-th sample at time t, represents the first-dimensional feature;

[0022] The first-dimensional feature to the twelfth-dimensional feature respectively represent the height of the opposing target, the height of one's own fighter plane, the speed of the opposing target, the speed of one's own fighter plane, the relative height, the relative distance, the target entry angle, the target azimuth angle, the air combat ability of the opposing target, the air combat ability of one's own fighter plane, the interference state of one's own fighter plane, and the recognition result of the radar state of the opposing target;

[0023] Among them, the ranges of the target entry angle and the target azimuth angle are [-π, π].

[0024] In an embodiment of the present invention, the S2 includes:

[0025] S21: Obtain the intention sample data set through an air combat confrontation simulation platform. The intention sample data set includes multiple samples and corresponding intention type labels, and each sample is represented by a feature description model;

[0026] S22: Remove the samples in the intention sample data set with a time series length less than N according to the preset time series length N to obtain a cleaned sample data set;

[0027] S23: Sample the samples in the cleaned sample data set to obtain a sampled sample data set. The sampling method is:

[0028]

[0029] k = n i / N;

[0030]

[0031] Among them, the i-th sample before sampling is represents the time series length after the i-th sample data is intercepted, n i represents the time series length of the i-th sample data, k represents the sampling interval, and S icut represents the i-th sample after sampling;

[0032] S23: Encode the interference state of the friendly fighter aircraft and the recognition result of the adversarial target radar state in the sampled sample data set, and perform one-hot encoding on the intention type label corresponding to each sample;

[0033] S24: Perform multiple normalization processes on the encoded sampled sample data set to obtain a sample set. The multiple normalization processes include normalization from 0 to 1 and normalization from -1 to 1. Among them,

[0034] The normalization process from 0 to 1 is:

[0035]

[0036] The normalization process from -1 to 1 is:

[0037]

[0038] In the formula, represents the eigenvalue of the n-th dimension at the t-th moment of the i-th sample after normalization, Represents the original feature value of the nth dimension of the ith sample at time t, maxs n Represents the maximum value of the nth-dimensional feature of the sample, min s n Represents the minimum value of the nth-dimensional feature of the sample.

[0039] In an embodiment of the present invention, the convolutional layer uses one-dimensional temporal convolution, and its mathematical model is:

[0040]

[0041] Among them, M j Is the jth convolutional region, H i Is the convolution element contained in the region, W ij Is the weight matrix corresponding to the convolution kernel, b j Represents the output corresponding bias, H j Is the jth feature map of the output, f(·) represents the activation function, and the LeakyReLU activation function is used. The mathematical model of this activation function is:

[0042] a i (j) = f(H j ) = max(0, H j ) + leak * min(0, H j );

[0043] In the formula, a i (j) is the activation value of H j , H j Represents the convolution output value, and leak is an adjustable constant value.

[0044] In an embodiment of the present invention, the bidirectional long short-term memory neural network layer includes a forgetting gate, a memory gate, and an output gate, and its mathematical model is:

[0045]

[0046]

[0047]

[0048] In the formula, Represents the splicing function, x t Represents the current input, Represents the hidden layer state obtained by the forward LSTM, Represents the hidden layer state obtained by the backward LSTM, and L represents the time series length.

[0049] In an embodiment of the present invention, the self-attention mechanism of the self-attention mechanism layer is:

[0050]

[0051]

[0052] In the formula, represents the input of the self-attention layer, att ∈ R N×Dv represents the output under the attention distribution, D K is the dimension representing Q and K, W Q W K W V are the mapping weight matrices that the self-attention layer needs to learn and train.

[0053] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0054] 1. The method for online recognition of target tactical intentions based on deep learning in the simulation environment of the present invention proposes the concept of target tactical behavior intention recognition for the problem that it is difficult to recognize target intentions from the task level according to the existing situation information. It recognizes target tactical behaviors from the perspective of air combat confrontation, and the result is more practical. Compared with traditional task intentions, the tactical intention features are more obvious, and its recognition result is more reliable and practical.

[0055] 2. The method for online recognition of target tactical intentions based on deep learning in the simulation environment of the present invention establishes a new target tactical intention space and feature description model for modern air combat simulation scenarios. On the basis of introducing the situation of the opposing side and the own side, it comprehensively considers the state characteristics of sensors, and constructs a target tactical intention recognition model to recognize target tactical intentions. Compared with traditional recognition methods, the recognition method of the present invention has higher reliability and faster recognition speed.

[0056] The above description is only an overview of the technical solution of the present invention. In order to be able to understand the technical means of the present invention more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the present invention more obvious and understandable, the following specific preferred embodiments are given and described in detail in conjunction with the accompanying drawings as follows. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] Figure 1 is a schematic diagram of a method for online recognition of target tactical intentions based on deep learning in a simulation environment provided by an embodiment of the present invention;

[0058] Figure 2 is a flowchart of a method for online recognition of target tactical intentions based on deep learning in a simulation environment provided by an embodiment of the present invention;

[0059] Figure 3 is a schematic diagram of the definition of the target motion coordinate system and characteristic parameters provided by an embodiment of the present invention;

[0060] Figure 4 It is a schematic diagram of one-hot encoding for tactical intention types provided by an embodiment of the present invention;

[0061] Figure 5 It is a schematic diagram of the target tactical intention recognition network structure provided by an embodiment of the present invention;

[0062] Figure 6 It is an internal structure diagram of the BiLSTM cell unit provided by an embodiment of the present invention;

[0063] Figure 7 It is a curve graph of the training results of the target tactical intention recognition model provided by an embodiment of the present invention;

[0064] Figure 8 It is an online recognition result graph of the target tactical intention recognition model provided by an embodiment of the present invention;

[0065] Figure 9 It is a comparison graph of the target tactical intention recognition results under different algorithms provided by an embodiment of the present invention. Detailed implementation manners

[0066] In order to further elaborate on the technical means and effects adopted by the present invention to achieve the predetermined invention purpose, the following, in combination with the accompanying drawings and specific implementation manners, details a method for online recognition of target tactical intentions based on deep learning in a simulated environment proposed according to the present invention.

[0067] The foregoing and other technical contents, features, and effects of the present invention can be clearly presented in the following detailed description in conjunction with the accompanying drawings. Through the description of the specific implementation manners, a more in-depth and specific understanding of the technical means and effects adopted by the present invention to achieve the predetermined purpose can be obtained. However, the attached drawings are only for reference and illustration purposes and are not used to limit the technical solutions of the present invention.

[0068] Embodiment 1

[0069] Due to the increasing complexity of the battlefield environment, the dynamic game of air combat confrontation makes the situation information have a certain degree of deception, making it difficult to effectively identify through the existing situation information. Therefore, it is necessary to identify the target intention from the tactical behavior level. Compared with traditional task intentions, tactical intention features are more obvious, and its recognition results are more reliable and practical. For this reason, this embodiment provides a method for online recognition of target tactical intentions based on deep learning in a simulated environment. Please refer to Figure 1 and Figure 2 , Figure 1 It is a schematic diagram of a method for online recognition of target tactical intentions based on deep learning in a simulated environment provided by an embodiment of the present invention; Figure 2It is a flowchart of an online recognition method for target tactical intentions based on deep learning in a simulation environment provided by an embodiment of the present invention. As shown in the figure, the online recognition method for target tactical intentions in this embodiment includes:

[0070] Step 1: Obtain real-time battlefield information in the simulation environment;

[0071] Among them, the battlefield information includes the situation characteristic information of both the friendly side and the adversarial side and the sensor status information.

[0072] In this embodiment, the situation characteristic information of both sides includes: the height of the adversarial target, the height of the friendly fighter, the speed of the adversarial target, the speed of the friendly fighter, the relative height, the relative distance, the target entry angle, and the target azimuth angle. The sensor status information includes: the air combat ability of the adversarial target, the air combat ability of the friendly fighter, the interference status of the friendly fighter, and the recognition result of the radar status of the adversarial target.

[0073] Step 2: Perform various normalization processes on the battlefield information, and input the battlefield information after various normalization processes into the trained target tactical intention recognition model to obtain the recognition result of the target tactical intention.

[0074] In this embodiment, the target tactical intentions include five intention types: attack, defense, detection, interference, and escape.

[0075] First, the training process of the target tactical intention recognition model in this embodiment is described. Specifically, the training steps of the target tactical intention recognition model include:

[0076] S1: Establish a target tactical intention space and a feature description model;

[0077] Specifically, S1 includes:

[0078] S11: Establish a target tactical intention space;

[0079] In this embodiment, referring to the thinking logic and reasoning mode of pilots for judging the situation in air combat, the target tactical intentions are designed into five types: attack, defense, detection, interference, and escape. That is, the target tactical intention space includes five intention types: attack, defense, detection, interference, and escape.

[0080] S12: Establish a feature description model according to the target track and sensor working status information;

[0081] During the air combat process, continuous multi-moment situation information of both the friendly side and the opposing side can be obtained through on-board sensors, including state information and track information, which intuitively reflects the changes of both the friendly side and the opposing side in the time domain and space domain. When executing different intentions, the situation changes accordingly, and its change pattern is determined by the intention category. Therefore, reasonable and effective feature information is needed to describe the target tactical intention.

[0082] In this embodiment, from the two perspectives of the target track and the sensor working state information, a feature model description is established for the target tactical intention space. In air combat, the change of the target track is reflected in a series of motion parameters such as coordinate azimuth, speed, yaw angle, pitch angle, roll angle, angle of attack, track inclination angle, and sideslip angle. However, some of these parameters are difficult to accurately obtain through on-board sensors, some parameters are coupled with each other, and some parameters have no practical significance for intention recognition and low correlation. The sensor working information refers to the working state of the friendly sensors and the recognition results of the working states of the opposing sensors (mainly including the recognition of the opposing radar state and the opposing target type), which is specifically reflected in the air combat capabilities of the target and the friendly fighter and the sensor working state.

[0083] At the same time, in order to minimize the feature dimension as much as possible to avoid information redundancy and improve the algorithm recognition efficiency, eight parameters, namely the altitude of the opposing target, the altitude of the friendly fighter, the speed of the opposing target, the speed of the friendly fighter, the relative distance, the relative altitude, the target entry angle, and the target azimuth angle, are selected to describe the spatial situation between the opposing target and the friendly fighter. Four parameters, namely the air combat capabilities of the opposing target, the air combat capabilities of the friendly fighter, the interference state of the friendly fighter, and the radar state of the opposing target, are selected to describe the sensor state information. The target tactical intention is described by the above 12 kinds of feature information.

[0084] In this embodiment, in addition to the opposing situation information, the friendly feature information is also selected from the above 12 kinds of feature information. The main purpose is to consider the process of the confrontation game between the two sides and the deceptive behavior of the opposing side. Therefore, considering the friendly feature information can comprehensively grasp the battlefield situation and the advantages and disadvantages of the situation of both sides, so as to conduct more accurate recognition.

[0085] Specifically, the expression of the feature description model is:

[0086]

[0087] In the formula, represents the feature information of the i-th sample at the t-th moment, Represents the first - dimensional feature; the first - dimensional feature to the twelfth - dimensional feature respectively represent the height of the adversarial target, the height of one's own fighter, the speed of the adversarial target, the speed of one's own fighter, the relative height, the relative distance, the target approach angle, the target azimuth angle, the air - combat ability of the adversarial target, the air - combat ability of one's own fighter, the interference state of one's own fighter, and the recognition result of the radar state of the adversarial target.

[0088] Please refer to the schematic diagram of the target motion coordinate system and the definition of feature parameters as shown in Figure 3 to explain the definitions of the target approach angle and the target azimuth angle. The expression of the target azimuth angle q f and the target approach angle q j is:

[0089]

[0090] In the formula, Ox g y g z g is the geographic coordinate system, O u x hu y hu z hu is the track coordinate system of the adversarial target, O t x ht y ht z ht is the track coordinate system of one's own fighter. D represents the distance vector from one's own fighter to the adversarial target. Among them, q f is positive when rotating clockwise from v u to D, and q j is positive when rotating counter - clockwise from v u to D. v u and v t respectively represent the speed vectors of one's own fighter and the adversarial target. q f ∈[-π,π], q j ∈[-π,π]. (D x , D y , D z ), (v ux , v uy , v uz ) and (v tx , v ty , v tz ) respectively represent the projections of D, v u and v t on the three axes of the Ox g y g z g geographic coordinate system. D x =x t -x u , D y =y t-y u , D z = z t -z u , where [x t , y t , z t and [x u , y u , z u represent the spatial position coordinates of the adversarial target and the own fighter in Ox g y g z g space, respectively.

[0091] S2: Obtain an intention sample data set, and preprocess the intention sample data set to obtain a sample set;

[0092] Specifically, S2 includes:

[0093] S21: Obtain an intention sample data set through an air combat confrontation simulation platform. The intention sample data set includes multiple samples and corresponding intention type labels, and each sample is represented by a feature description model;

[0094] In this embodiment, multiple air combat confrontation samples corresponding to each intention type are obtained by using an intention sample generation model.

[0095] Due to the problem of complex calculation amount caused by the too long sample length, which affects the network training efficiency and convergence, when constructing the sample set, it is necessary to sample the samples and select an appropriate time series length for analysis. However, due to the inconsistent time series sample lengths caused by different sample confrontation end times, and some sample lengths are too short and do not meet the sampling conditions, it is necessary to perform data cleaning on the intention sample data set to ensure that all sample lengths meet the sampling requirements.

[0096] S22: According to the preset time series length N, remove the samples in the intention sample data set with a time series length less than N to obtain a cleaned sample data set;

[0097] S23: Sample the samples in the cleaned sample data set to obtain a sampled sample data set. The sampling method is:

[0098]

[0099] k = n i / N(4);

[0100]

[0101] Among them, the i-th sample before sampling is represents the time series length after intercepting the i-th sample data, n irepresents the time series length of the i-th sample data, k represents the sampling interval, and S icut represents the i-th sample after sampling;

[0102] In this embodiment, through the above sampling process, the situation features can be maximally retained, so as to completely reflect the intention behavior information for the network to learn.

[0103] S23: Encode the recognition results of the interference state of our own fighter jets and the radar state of the opposing target in the sampled sample dataset, and perform one-hot encoding on the intention type label corresponding to each sample;

[0104] Specifically, these two features, the recognition results of the interference state of our own fighter jets and the radar state of the opposing target, are non-numerical data and need to be converted into an encodable form that can be recognized for training. Among them, the interference state of our own fighter jets is encoded as 0 and 1, indicating no interference and being interfered respectively; the recognition result codes of the radar state of the opposing target, 0 and 1, indicate that the radar of the opposing target is turned off and turned on respectively.

[0105] Since it involves a multi-classification problem and considering the impact of label values on the network, it is necessary to perform one-hot encoding on the intention type label corresponding to the sample. The specific operation is as Figure 4 shown in the schematic diagram of one-hot encoding of tactical intention types. By setting different bits of the binary number to 1 and the rest to 0, it can cooperate with the cross-entropy loss function to train more effectively.

[0106] S24: Perform multiple normalization processes on the encoded sampled sample dataset to obtain a sample set. The multiple normalization processes include normalization from 0 to 1 and normalization from -1 to 1.

[0107] Since the value ranges of the target entry angle and the target azimuth angle are defined, and the positive and negative values of these values reflect the feature of the directionality of speed, and better intention recognition can be achieved through the comprehensive judgment of the speed directions of both sides, therefore, the normalization method from -1 to 1 is used to normalize the target entry angle, target azimuth angle, and relative height of the samples in the encoded sampled sample dataset to retain the original features of the data to the greatest extent.

[0108] The normalization method from 0 to 1 is used to normalize the target height, the height of our own fighter jets, the target speed, the speed of our own fighter jets, the relative distance, the air combat ability of the opposing target, the air combat ability of our own fighter jets, the recognition results of the interference state of our own fighter jets, and the radar state of the opposing target in the encoded sampled sample dataset.

[0109] Specifically, the normalization process from 0 to 1 is as follows:

[0110]

[0111] The normalization process from -1 to 1 is as follows:

[0112]

[0113] In the formula, represents the eigenvalue of the nth dimension at the t-th moment of the i-th sample after normalization, represents the original eigenvalue of the nth dimension at the t-th moment of the i-th sample, maxs n represents the maximum value of the nth-dimensional feature of the sample, min s n represents the minimum value of the nth-dimensional feature of the sample.

[0114] S3: Construct a target tactical intention recognition network;

[0115] As Figure 5 shown in the schematic diagram of the target tactical intention recognition network structure, the target tactical intention recognition network in this embodiment includes a cascaded input layer, a feature extraction module, and a classification and recognition module. Among them, the feature extraction module includes a cascaded convolutional layer, a bidirectional long short-term memory neural network layer, and a self-attention mechanism layer.

[0116] In this embodiment, the input layer is used to obtain information of the input sample.

[0117] Specifically, in order to extract more effective features to make the network easier to train and improve the recognition effect, the convolutional layer uses one-dimensional temporal convolution. By setting the convolutional kernel size and number, the input information is feature-extracted. Each convolutional kernel obtains a feature map with a length of L, and all feature maps are concatenated to form the final feature representation map. Among them, the mathematical model of the convolutional layer is:

[0118]

[0119] Among them, M j is the j-th convolutional region, H i is the convolutional elements included in the region, W ij is the weight matrix corresponding to the convolutional kernel, b j represents the corresponding bias of the output, H j is the j-th feature map of the output, f(·) represents the activation function, and the LeakyReLU activation function is used to avoid the problem of update sawtooth during network training, which leads to a slow training speed. The mathematical model of this activation function is:

[0120] a i (j) = f(H j ) = max(0, H j ) + leak * min(0, H j ) (9);

[0121] In the formula, a i(j) is H j The activation value, H j Represents the convolution output value, and leak is an adjustable constant value.

[0122] Since the LSTM network effectively solves the long-term dependency problem through the design of three gates and can extract the information contained in a longer sequence, the output of the structure at the current moment is only related to the present and the past, while ignoring the consideration of the future. Therefore, in this embodiment, BiLSTM (bidirectional long short-term memory neural network layer) is used.

[0123] BiLSTM adds a backward layer to LSTM. The forward layer retains and extracts sequence information from the past to the future, and the backward layer is from the future to the past, aiming to mine more potential information. Figure 6 As shown in the figure, it includes a forget gate, a memory gate and an output gate. The BiLSTM network combines the information extracted from the forward layer and the backward layer as the final output, thereby reflecting more comprehensive situation feature information for network training and learning.

[0124] Specifically, the mathematical model of the bidirectional long short-term memory neural network layer is:

[0125]

[0126]

[0127]

[0128] In the formula, represents the concatenation function, x t Indicates the current input. represents the hidden layer state obtained by the forward LSTM, Represents the hidden layer state obtained by backward LSTM, and L represents the time series length.

[0129] Furthermore, the self-attention mechanism of the self-attention mechanism layer is:

[0130]

[0131]

[0132] In the formula, Represents the input of the self-attention layer, att∈R N×Dv represents the output under the attention distribution, D K To represent the dimensions of Q and K, W Q , W K , W V It is the mapping weight matrix that needs to be learned and trained for the self-attention layer.

[0133] For the intention recognition problem, its purpose is to extract key features. In order to avoid interference from other information on the classification results, in this embodiment, the self-attention mechanism can effectively reduce the dimension input to the subsequent sorfmax classification layer, improve the convergence speed, and further improve the recognition accuracy.

[0134] Furthermore, the classification and recognition module of this embodiment includes a flatten layer and a sorftmax classification layer. The output of the bidirectional long short-term memory neural network layer is one-dimensionalized by the flatten layer for multi-dimensional information, and then input into the sorftmax classification layer to obtain the target tactical intention recognition result. The mathematical model of the sorftmax classification layer is:

[0135] y = softmax(w L Y + b L ) (15);

[0136] In the formula, w L , b L are the weight matrix and corresponding bias to be trained by the softmax classification layer, Y is the output obtained by att through the flatten layer, and y is the output target tactical intention recognition result.

[0137] S4: Divide the sample set into a training set and a validation set, input the training set and the validation set into the target tactical intention recognition network to train and optimize its network parameters, obtain the optimal network parameters, and obtain the trained target tactical intention recognition model according to the optimal network parameters.

[0138] Specifically, divide the sample set into a training set and a validation set according to a ratio of 7:3, set basic parameters such as the optimizer, the number of training times, and the batch size, and then start training the target tactical intention recognition network. Set the monitoring function. When the loss of the training set no longer decreases, stop training and save the model. Then, perform parameter tuning according to the experimental results of the validation set. By the method of controlling variables, loop the above operation steps until the optimal parameters are obtained.

[0139] In this embodiment, through the above parameter tuning steps, the network parameters are finally determined as follows: the time series length is 50, the convolution kernel size is 1×1, the number of convolution kernels is 12, the number of BiLSTM neurons is 256, the number of attention is 256, the optimizer is Adam, the loss function is cross-entropy loss, the learning rate is 0.001, the batch size is 32, and the number of iterations is 50.

[0140] Further, before inputting the battlefield information obtained in Step 1 into the trained target tactical intention recognition model for recognition, various normalization processes need to be performed on the battlefield information, including: normalizing the three features of the target entry angle, target azimuth angle, and relative height in the battlefield information using a normalization method from -1 to 1, and normalizing the remaining features in the battlefield information using a normalization method from 0 to 1. For the specific normalization process, refer to the normalization process of the samples, which will not be elaborated here.

[0141] It should be noted that the real-time battlefield information in the simulation environment can obtain the situation feature information of both sides at each moment through data links and other means and save it. When the length of the saved moment information is greater than 50, the currently extracted situation features and the information of the previous 49 moment points are jointly formed into a feature sequence. After normalization processing, it is input into the trained target tactical intention recognition model to obtain the recognition probability of each target tactical intention type corresponding to the current moment, and the one with the highest probability is taken as the target tactical intention type of the opposing target.

[0142] In the simulation environment of this embodiment, the online recognition method of target tactical intention based on deep learning proposes the concept of target tactical behavior intention recognition to address the problem of difficultly recognizing target intentions from the task level based on existing situation information. It recognizes target tactical behaviors from the perspective of air combat confrontation, and its results are more practically significant. Compared with traditional task intentions, the tactical intention features are more obvious, and its recognition results are more reliable and practical.

[0143] In the simulation environment of this embodiment, the online recognition method of target tactical intention based on deep learning establishes a new target tactical intention space and feature description model for modern air combat simulation scenarios. On the basis of introducing the situation of the opposing side and our own side, it comprehensively considers the state features of sensors and constructs a target tactical intention recognition model to recognize target tactical intentions. Compared with traditional recognition methods, the recognition method of the present invention has higher reliability and faster recognition speed.

[0144] Secondly, the target tactical intention recognition model of this embodiment extracts features in the time dimension through one-dimensional convolution, effectively mines the potential features of the time series in combination with the BiLSTM network, and finally introduces the self-attention mechanism to reduce the feature dimension. Selecting typical features can avoid the interference of irrelevant information and has higher recognition accuracy.

[0145] Embodiment 2

[0146] This embodiment illustrates the effect of the online recognition method of target tactical intention based on deep learning in the simulation environment of Embodiment 1 through simulation experiments.

[0147] Simulation Experiment 1

[0148] Taking modern air combat as the research background, a dataset of target tactical intentions is obtained through a confrontation simulation platform. Then, experts revise the sample labels, followed by data cleaning to eliminate abnormal samples. Finally, each sample is sampled, and 50 frames of sequential information (each frame contains 12-dimensional feature information) are extracted at equal length to fully reflect the intention situation characteristics. The final dataset has a sample size of 9256, among which the samples with attack intention account for 17.9%, the samples with defense intention account for 18.0%, the samples with detection intention account for 23.1%, the samples with escape intention account for 17.8%, and the samples with interference intention account for 23.2%. The training set and test set are divided according to the ratio of 7:3, so the number of samples in the test set is 2777.

[0149] The experimental simulation environment is constructed based on the Keras framework using the Python language, and the computer configuration is Windows 10, i7-10700 CPU @ 2.90 GHz.

[0150] The final parameters of the target tactical intention recognition model are shown in Table 1:

[0151] Table 1 Model Parameters

[0152]

[0153] Set the monitoring function to stop training when the loss of the validation set no longer decreases for multiple generations to prevent overfitting problems. The training results of the final network model are as Figure 7 The training result curve graph of the target tactical intention recognition model, where Figure (a) is the loss curve and Figure (b) is the accuracy curve.

[0154] From the above two figures, it can be seen that after combining multiple algorithms, the training curves of the validation set and the training set are consistent, and the final results are approximately the same, indicating that both the loss value and the accuracy have reached convergence. The recognition results of the test set are shown in Table 2, and the recognition accuracy is 94.42%, presented in the form of a confusion matrix.

[0155] Table 2 Intention Recognition Confusion Matrix

[0156]

[0157]

[0158] To better evaluate the model performance, according to the results of the confusion matrix, three evaluation indicators, namely precision, recall, and F1-score, can be calculated. Precision measures the accuracy of recognition, recall measures the coverage rate also known as the recall rate, and the F1-score is the harmonic mean of the two indicators. The results are shown in Table 3.

[0159] Table 3 Evaluation Indicators

[0160]

[0161] At the same time, the trained model is used for real-time online recognition in air combat for testing. Some confrontation cases are selected, and the recognition results are as Figure 8 shown. Among them, (a) and (b) are the schematic diagram and recognition result diagram of Example 1, (c) and (d) are the schematic diagram and recognition result diagram of Example 2, and (e) and (f) are the schematic diagram and recognition result diagram of Example 1.

[0162] It can be seen from (a) and (b) that when recognizing a single escape intention, the effect is good, and there is no misjudgment of other intentions from the output recognition probability. It is known that the target in Example 2 first executes a defense intention and then executes a detection intention. However, the defense postures presented by both sides at this time also conform to certain offensive rules. Since the boundary between defense and attack is relatively blurred, during some periods of the confrontation process, the algorithm is difficult to accurately identify the offensive and defensive intentions. It can be seen from Figure (d) that during the period from 125s to 150s, the recognition probabilities of attack and defense are very close, which conforms to the actual situation. It is known that the target in Example 3 first executes an interference intention and then executes an escape intention. It can be seen from Figure (f) that due to the obvious interference and escape trends, there is only a certain degree of ambiguity in the transition stage of intention switching. For the three randomly selected confrontation samples, good results are obtained in online recognition, and the misclassification in the intention transition stage and the misclassification caused by offensive and defensive games also conform to the actual situation, further verifying the reliability of the method of the present invention.

[0163] Simulation Experiment 2

[0164] Taking modern air combat as the research background, a target tactical intention data set is obtained through an adversarial simulation platform. Then, experts correct the sample labels, and then perform data cleaning to eliminate singular samples. Finally, each sample is sampled, and 50 frames of time series information (each frame contains 12-dimensional feature information) are extracted at equal length to fully reflect the intention situation characteristics. The final data set sample size is 9256, among which the attack intention samples account for 17.9%, the defense intention samples account for 18.0%, the detection intention samples account for 23.1%, the escape intention samples account for 17.8%, and the interference intention samples account for 23.2%. The training set and the test set are divided according to 7:3, so the number of test set samples is 2777.

[0165] The simulation environment is built based on the keras framework using the python language, and the computer configuration is window10, i7-10700CPU@2.90GHz.

[0166] To verify the superiority of the method of the present invention, its effects are compared with other methods, including a stack autoencoder SAE intelligent recognition model, a deep neural network model (DBP) based on the ReLU function and the Adam optimizer, and a multi-classification model constructed using a support vector machine SVM. The algorithm parameters of the three models are shown in Table 4, and the comparison results are as Figure 9 shown.

[0167] Table 4 Algorithm model parameters of different algorithms

[0168]

[0169] As can be seen from the figure, the generalization ability of the models obtained by training the SVM algorithm and the DBP algorithm is weak. In particular, due to too few network parameters in the DBP network, overfitting is likely to occur, resulting in the recognition rate of the training set being much higher than that of the validation set and the test set. And due to the diversity and complexity of the features contained in the time series samples, it is difficult for the SVM algorithm to find a suitable hyperplane in the high-dimensional space to effectively distinguish the target intention. Therefore, the generalization ability of the trained model is also relatively weak. The algorithm of this article is significantly better than other algorithms. The SAE autoencoder model has a certain advantage over the other two algorithms because the increase in its network parameters and the deepening of the number of layers make the model more complex, but the overall recognition rate is low. This is mainly because it is difficult for the network relying only on the encoder to extract effective and comprehensive time series features. The algorithm of this article can effectively extract the target intention features for recognition by combining multiple algorithms and the attention mechanism. And as can be seen from the figure, the accuracy rate of the test set is stronger than other algorithms, and the difference in the accuracy rates of the validation set and the test set is not large, which is outstanding compared with the other three model algorithms, thus verifying the effectiveness of the method of the present invention.

[0170] It should be noted that in this article, the terms "including", "comprising" or any other variant are intended to cover non-exclusive inclusion, so that an article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed. Without more restrictions, an element defined by the statement "including one..." does not exclude the existence of another identical element in the article or device including the said element. Similar words such as "connected" or "linked" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect.

[0171] The above content is a further detailed description of the present invention in combination with specific preferred embodiments. It cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention pertains, without departing from the concept of the present invention, several simple deductions or substitutions can still be made, which should all be regarded as belonging to the protection scope of the present invention.

Claims

1. An online recognition method for target tactical intention based on deep learning in a simulation environment, characterized in that Including: Step 1: Obtain real-time battlefield information in the simulation environment, where the battlefield information includes the situation feature information of both the friendly side and the opposing side, as well as the sensor status information; Step 2: Perform multiple normalization processes on the battlefield information, and input the battlefield information after multiple normalization processes into the trained target tactical intention recognition model to obtain the recognition result of the target tactical intention; The training process of the target tactical intention recognition model includes: S1: Establish a target tactical intention space and a feature description model; S2: Obtain an intention sample data set, and preprocess the intention sample data set to obtain a sample set; S3: Construct a target tactical intention recognition network, where the target tactical intention recognition network includes a cascaded input layer, a feature extraction module, and a classification and recognition module. Among them, the feature extraction module includes a cascaded convolutional layer, a bidirectional long short-term memory neural network layer, and a self-attention mechanism layer; S4: Divide the sample set into a training set and a validation set, input the training set and the validation set into the target tactical intention recognition network to train and optimize its network parameters to obtain the optimal network parameters, and obtain the trained target tactical intention recognition model according to the optimal network parameters; The S2 includes: S21: Obtain the intention sample data set through an air combat confrontation simulation platform. The intention sample data set includes multiple samples and corresponding intention type labels, and each sample is represented by a feature description model; S22: According to the preset time series length N, remove the samples in the intention sample data set with a time series length less than N to obtain a cleaned sample data set; S23: Sample the samples in the cleaned sample data set to obtain a sampled sample data set. The sampling method is: k = n i / N; Among them, the i-th sample before sampling is represents the time series length after intercepting the i-th sample data, n i represents the time series length of the i-th sample data, k represents the sampling interval, S icut represents the i-th sample after sampling; S23: Encode the interference state of the friendly fighter jets and the recognition result of the opposing target radar state in the samples of the sampled sample data set, and perform one-hot encoding on the intention type label corresponding to each sample; S24: Perform multiple normalization processes on the encoded sampled sample data set to obtain a sample set. The multiple normalization processes include normalization from 0 to 1 and normalization from -1 to 1. Among them, The normalization process from 0 to 1 is: The normalization process from -1 to 1 is: Wherein, represents the eigenvalue of the nth dimension at the t-th moment of the i-th sample after normalization, represents the original eigenvalue of the nth dimension at the t-th moment of the i-th sample, maxs n represents the maximum value of the nth-dimensional feature of the sample, mins n represents the minimum value of the nth-dimensional feature of the sample.

2. The method for online recognition of target tactical intentions based on deep learning in a simulation environment according to claim 1, wherein The situation feature information of both sides includes: the height of the opposing target, the height of the friendly fighter jet, the speed of the opposing target, the speed of the friendly fighter jet, the relative height, the relative distance, the target entry angle, and the target azimuth angle; The sensor status information includes: the air combat ability of the opposing target, the air combat ability of the friendly fighter jet, the interference state of the friendly fighter jet, and the recognition result of the opposing target radar state.

3. The online recognition method for target tactical intentions based on deep learning in a simulation environment according to claim 2, characterized in that, Performing multiple normalization processes on the battlefield information includes: performing normalization processing on the three features of the target entry angle, the target azimuth angle, and the relative height in the battlefield information in a normalization manner from -1 to 1, and performing normalization processing on the remaining features in the battlefield information in a normalization manner from 0 to 1.

4. The online recognition method for target tactical intent based on deep learning in a simulation environment according to claim 1, characterized in that, The S1 includes: S11: Establish a target tactical intention space, where the target tactical intention space includes five intention types: attack, defense, detection, interference, and escape; S12: Establish a feature description model according to the target track and sensor working state information, and the feature description model is as follows: In the formula, represents the feature information of the i-th sample at time t, represents the first-dimensional feature; The first dimension feature to the twelfth dimension feature respectively represent the height of the adversarial target, the height of one's own fighter plane, the speed of the adversarial target, the speed of one's own fighter plane, the relative height, the relative distance, the target entry angle, the target azimuth angle, the air combat ability of the adversarial target, the air combat ability of one's own fighter plane, the interference state of one's own fighter plane, and the recognition result of the radar state of the adversarial target; Among them, the ranges of the target entry angle and the target azimuth angle are [-π, π].

5. The method for online recognition of target tactical intentions based on deep learning in a simulation environment according to claim 1, characterized in that, The convolutional layer uses one-dimensional temporal convolution, and its mathematical model is: Among them, M j is the j-th convolution region, H i is the convolved element included in the region, W ij is the weight matrix corresponding to the convolution kernel, b j represents the corresponding bias of the output, H j is the j-th feature map of the output, f(·) represents the activation function, and the LeakyReLU activation function is adopted. The mathematical model of this activation function is: a i (j) = f(H j ) = max(0, H j ) + leak * min(0, H j ); Where a i (j) is the activation value of H j , H j represents the convolution output value, and leak is an adjustable constant value.

6. The online recognition method for target tactical intention based on deep learning in a simulation environment according to claim 5, wherein The bidirectional long short-term memory neural network layer includes a forgetting gate, a memory gate, and an output gate, and its mathematical model is: In the formula, represents the splicing function, x t represents the current input, represents the hidden layer state obtained by the forward LSTM, represents the hidden layer state obtained by the backward LSTM, and L represents the time series length.

7. The online recognition method for target tactical intent based on deep learning in a simulation environment according to claim 6, characterized in that The self-attention mechanism of the self-attention mechanism layer is: In the formula, represents the input of the self-attention layer, att ∈ R N×Dv represents the output under the attention distribution, D K is the dimension representing Q and K, W Q 、W K 、W V are the mapping weight matrices that the self-attention layer needs to learn and train.