Radar working mode switching intelligent prediction method, system, equipment and medium

By adopting a multi-head attention mechanism and residual GRU network in radar operating mode switching prediction, the problems of high computational complexity and insufficient prediction accuracy in the prior art are solved, and higher prediction accuracy and lower computational complexity are achieved.

CN120146098APending Publication Date: 2025-06-13刘明骞
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510103455.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

While improving the prediction accuracy, the existing radar operating mode switching prediction method has high computational complexity and is difficult to accurately discover factors and conditions with strong correlation with the predicted object, and the hyperparameters are difficult to determine, resulting in large errors.

Method used

Residual gating cyclic unit (RGRU-MHA) network based on multi-head attention is adopted to reasonably allocate the model attention proportion through the multi-head attention mechanism, and residual blocks are added when the features are transmitted by the hidden layer of the GRU network to avoid feature loss and model overfitting or underfitting.

Benefits of technology

It effectively improves the accuracy of model prediction, reduces the computational complexity, and enhances the model's feature mining ability of input data, reducing errors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120146098A_ABST
    Figure CN120146098A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of radar behavior intention prediction, and discloses a radar working mode switching intelligent prediction method, system and device and a medium, and the method comprises the steps: carrying out the sequential coding of a historical radar working mode switching sequence, and carrying out the binary representation, so as to serve as the input of a network constructed in subsequent steps; building a multi-head attention mechanism network, and distributing the same attention degree of the model to different positions of an input sequence; s2, a residual GRU network is built on the basis of S2, a chain connection mode is adopted, the input of a previous neuron is superposed as a residual before information is input into a GRU neuron every time, the information is protected from being damaged in the transmission process, and performance reduction caused by the non-forward effect of the neuron is avoided; and S3, fully training the multi-head attention-based residual GRU network by using the data set obtained in the S1 and adopting an Adam algorithm to obtain a final network model. According to the method, the feature loss during the non-forward effect of the neurons is avoided, the condition of over-fitting or under-fitting of the model is prevented, and the model prediction accuracy is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to, but is not limited to, the technical field of radar behavior intention prediction, and particularly relates to a method, a system, a device and a medium for intelligent prediction of radar working mode switching. Background Technique

[0002] The frequency behavior intention of radar is reflected in its working mode, and the prediction of working mode switching is essentially a prediction of a time series pattern with underlying switching logic. There has been relevant research on the prediction of radar working mode switching in recent years, but the prediction of time series pattern switching has been studied by scholars since the 1980s. In the early days, most of the predictions of time series pattern switching were achieved by mining the correlation between pattern switches and then fitting known functions. Deng Julong, Wang Housheng, Ma Heling and others in the literature (Deng Julong, Wang Housheng, Ma Heling, Liu Guomin, Peng Guozhong. Recursive residual identification prediction model [J]. Journal of Huazhong Institute of Technology, 1982(06): 93-98) completed the prediction of the development trend of marine fishery by using single-sequence recursive fitting. This method has a small amount of calculation and low computational complexity, but the effect is also very poor. In order to improve the prediction accuracy, Zhang Chengjin and Cheng Shu in the literature (Zhang Chengjin, Cheng Shu. Recursive algorithm of Bayesian prediction model [C]. Proceedings of the 1993 Chinese Control and Decision Conference, 1993: 366-370) introduced the Bayesian prediction model, used Kalman filtering, and applied the Sweep transformation of matrices to give several more reliable prediction methods. These algorithms have very good reference value, but there is still a gap from actual application. Intaek Kim and Sung-Rock Lee and others in the literature (Intaek Kim and Sung-Rock Lee. A fuzzy time series prediction method based on consecutive values [C]. 1999 IEEE International Fuzzy Systems, 1999(2): 703-707) used a time series pattern switching prediction method based on a fuzzy rule-based system to handle non-stationary data with long-term mean fluctuations. Then, a method for predicting time series pattern switching by using the correlation between multiple variables was proposed. Kunlin Liu in the literature (Liu K. Recursive parameter method for computing the predicting function of the multivariable ARMAX model [J]. Tsinghua Science and Technology, 2000, 5(1): 115-120) proposed a new method for predicting the function of the Auto Regression and Moving Average model with eXogenous input (ARMAX) model based on the calculation of multivariable correlation. The correlation method did not have good characteristics in the early stage, but it played a very good role in promoting the development of prediction algorithms.Recently, the development of neural networks has provided new methods for mining the correlations between variables, and the results are more remarkable. Lv Siying et al. proposed a short-term prediction model for water bloom based on the Elman neural network in the literature (Lv Siying, Liu Zaiwen, Wang Xiaoyi and Cui Lifeng. Short-term predicting model for water bloom based on Elman neural network[C]. 2008 27th Chinese Control Conference, 2008: 218-221). Shi established an agile manufacturing prediction model and an automatic logistics prediction system using neural networks in the literature (Shi S. Agile Manufacturing Predicting Model Based on Neural Network and Automatic Logistics Predicting System in China[C]. Second International Conference on Intelligent System Design and Engineering Application, 2012: 764-767). Qian Wei predicted the unknown links and possible future link patterns in the network based on the model of complex networks in the literature (Qian Wei. Pattern-based link prediction in complex networks[D]. Nanjing: Southeast University, 2016). Ou Jian, a doctor from the National University of Defense Technology, mined the depth information of signals using artificial neural networks to realize the prediction of radar working mode switching in the literature (Ou Jian. Research on multi-functional radar behavior identification and prediction technology[D]. Changsha: National University of Defense Technology, 2017). However, traditional artificial intelligence methods can only learn the inherent limitations of training data, and how to get rid of this stationary time series attribute for prediction poses a challenge to model construction.André Bauer, Marwin Züfle and others proposed in the literature (A. Bauer, M. Züfle, N. Herbst, S. Kounev and V. Curtef. Telescope: An Automatic Feature Extraction and Transformation Approach for Time Series Forecasting on a Level-Playing Field [C]. 2020 IEEE 36th International Conference on Data Engineering (ICDE), 2020: 1902-1905) to use the Telescope method to extract and transform features from the input time series, and generate an optimized prediction model with these features, greatly improving the applicability of the prediction model. In order to mine more time series characteristics, Qingyang Xu, Qingsong Wen and others explored the prediction method based on periodic time series in the literature (Xu Q, Wen Q and Sun L. Two-Stage Framework for Seasonal Time Series Forecasting [C]. 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2021: 3530-3534), and proposed a prediction model that combines an autoregressive model with a neural network to completely capture the linear and nonlinear features in the time series. The model is divided into two stages. The first stage learns the features beyond the prediction range, and the second stage improves the learning ability of the prediction range. This method greatly improves the accuracy of time series pattern switching prediction. However, there are many predictable time series, including the radar working mode switching sequence, which do not have obvious periodicity. There is a problem of data missing in the time series, which makes it difficult to completely extract the time series features in time series pattern switching prediction, and even impossible to model.Xinyu Chen, Lijun Sun and others proposed a Bayesian temporal factorization framework in the literature (Chen X and Sun L. Bayesian Temporal Factorization for Multidimensional Time Series Prediction[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2022, 44(9): 4659-4673) to solve the modeling problem in the presence of missing values in multidimensional time series, and used low-rank matrices, tensor factorization and vector autoregression to construct model representations of the global and local consistency in large-scale time series data. With the further improvement of neural networks, recurrent networks with better performance are applied to the time series pattern switching prediction model. Wu et al. used recurrent neural networks to model and predict the laser polished surface quality in the literature (Wu H and E.V. Bordatchev. Feasibility Study of the Recurrent Neural Network for Modeling and Predicting Laser Polished Surface Quality[C]. 2021 IEEE Photonics Conference (IPC), 2021: 1-2). Hou Chao proposed a set of pattern description words for subsequent behavior recognition and reasoning based on the long short-term memory neural network, a variant of the recurrent network, through the analysis of the principles of various radar working modes in the literature (Hou Chao. Research on Phased Array Radar Behavior Recognition and Reasoning Technology[D]. Xi'an: Xidian University, 2020), and then established a prediction model. However, this method only targets the timing of radar working mode switching and does not consider the differences in data encoding between different working mode sequences. Xiangqin Hu, Changxin Wu and others established a long short-term memory neural network according to the different lengths and types of task sequences corresponding to different working modes in the literature (Hu X, Wu C, Feng K, et al. LSTM-based Prediction of Multi-function Radar Work Mode[C]. 2023 8th International Conference on Signal and Image Processing (ICSIP), 2023: 682-687), which better explains the behavioral intentions behind the radar working modes.

[0003] Through the above analysis, the problems and defects of the existing technologies are as follows:

[0004] The application and development of neural networks have provided infinite possibilities for prediction and reasoning. However, in the current methods for predicting radar working mode switching, the effects and computational complexity cannot be balanced. Neural networks can improve accuracy, but the complexity also increases accordingly. Moreover, it is impossible to accurately discover the factor conditions that have a strong correlation with the predicted object. In addition, the hyperparameters in neural networks are difficult to determine, resulting in large errors. The research on radar working mode switching prediction methods mainly focuses on exploring the correlation and regularity between historical switching sequences. Precise summarization of the correlation and ensuring that features are not lost or forgotten during transmission are the problems that need to be solved urgently. Summary of the Invention

[0005] In view of the problems existing in the prior art, the present invention provides an intelligent prediction method, system, device and medium for radar working mode switching. The prediction of phased array radar working mode switching is completed by using a Residual Gated Recurrent Unit with Multi-Head Attention (RGRU-MHA) network. Different from the simple Gated Recurrent Unit (GRU) network, in the present invention, for the phased array radar working mode switching sequence, there are often transition working modes, resulting in the situation that adjacent working modes are not necessarily the most strongly correlated. When inputting the working mode switching sequence, the multi-head attention mechanism is adopted to reasonably allocate the attention ratio of the model. And when transmitting features in the hidden layer of the GRU network, a residual block is added to avoid feature loss when the neurons have non-positive effects, and to prevent the model from overfitting or underfitting, effectively improving the model prediction accuracy.

[0006] The present invention is implemented as follows. An intelligent prediction method for radar working mode switching, the method includes:

[0007] S1, encoding the historical radar working mode switching sequence in order and representing it in binary as the input of the network constructed in the subsequent steps;

[0008] S2, building a multi-head attention mechanism network to allocate the same attention degree of the model to different positions of the input sequence;

[0009] S3, building a residual GRU network on the basis of S2, adopting a chain connection method, and adding the input of the previous neuron as a residual before each piece of information is input into the GRU neuron to protect the information from being damaged during transmission and avoid the performance degradation caused by the non-positive effects of neurons;

[0010] S4. Using the data set obtained in S1, the residual GRU network based on multi-head attention is fully trained using the Adam algorithm to obtain the final network model.

[0011] Furthermore, the specific steps of S1 are as follows:

[0012] First, training data is constructed based on the time series of the historical radar working modes obtained by detection. The first T radar working modes are used as input data, and the (T + 1)-th radar working mode is used as the output label to obtain a large number of radar working mode data set samples. And the text information of each radar working mode in the data set is encoded and normalized so that different data vectors represent different radar working modes. Data normalization is to prevent the problem of severe gradient disappearance during backpropagation when the value is too small, and the problems of gradient explosion during backpropagation and loss of the original representation ability due to forced compression of features by the activation function when the input is too large. The data normalization adopts deviation standard normalization, as shown below:

[0013]

[0014] where max and min are the minimum and maximum values after encoding the radar working modes respectively.

[0015] Furthermore, the specific steps of S2 are as follows:

[0016] In the attention mechanism, an attention head includes three parameters: query (q), key (k), and value (v). The semantic correlation score is obtained by calculating the dot product of the query and the key, and then the semantic correlation score is converted into an attention weight through an activation function. Finally, the attention weight is applied to the value to obtain the final attention output.

[0017] The input matrix is X = [x 1 T , x 2 T ,..., x n T T , x i (i = 1, 2,..., n) is the vector of each input unit. Through the linear transformation matrices W Q , W K and W V the query matrix Q = X · W Q = [q 1 T , q 2 T ,..., q n T T , the key matrix K = X · W​​K = [k 1 T , k 2 T ,..., k n T T The sum value matrix V = X · W V = [v 1 T , v 2 T ,..., v n T T , the output of the final single - head attention mechanism is as follows:

[0018]

[0019] Among them, σ is the softmax activation function, is the scaling factor, which can avoid the problem that the value becomes too large after dot product and affects the stability of softmax, and the problem of gradient explosion caused by too large gradient during backpropagation. d k usually takes the dimension of the key vector;

[0020] Although the single - head attention mechanism can, to a certain extent, reasonably allocate the model's attention and help the model focus on important information in the input sequence, in complex feature extraction problems, the result of a single attention allocation cannot help the model completely capture all important information in the input sequence, which will lead to problems such as losing some key information or having an unsatisfactory allocation result; therefore, the multi - head attention MHA mechanism is proposed. Each attention mechanism consists of multiple single - head attention mechanisms. Multiple single - head attention mechanisms process the same input sequence in parallel and independently. Each single - head attention mechanism can adaptively focus on important information at different positions in the input sequence. Finally, the allocation results of each single - head attention mechanism are concatenated and integrated, and then passed through a linear transformation matrix to obtain the final attention allocation result; the multi - head attention mechanism enables the model to simultaneously focus on multiple important information from different positions in the input sequence, preventing the loss of multi - level information and deep features, thereby improving the robustness and generalization ability of the model and further helping subsequent feature extraction; the multi - head attention mechanism is widely used in the field of deep learning, especially in the direction of natural language processing, and has a great performance improvement for models with long - term and short - term dependencies in time - series sequences.

[0021] Furthermore, the specific process of S3 is as follows:

[0022] ​​Like the LSTM network, the GRU network is a variant recurrent network model designed to address the gradient problem caused by the long-distance dependence characteristic in the backpropagation of the RNN. The LSTM network has three gate functions, namely the input gate, the forget gate, and the output gate. The input gate controls how much of the newly added memory is retained, the forget gate controls how much of the previous memory is retained, and the output gate integrates all the retained memories. The GRU network, on the other hand, has a lower complexity and a simpler structure, with only two gate functions, namely the update gate and the reset gate. Compared with the LSTM network, the GRU network simplifies the two gates of the input gate and the forget gate in the LSTM network to an update gate by incorporating the cell state and the hidden state into each gate, and further adjusts the memory retention ratio using the reset gate. This structure enables the GRU network to compress the number of parameters significantly and greatly reduces the computational complexity without significantly affecting the model performance. Even in some tasks with fewer input time steps and severe short-term dependencies, the performance of the GRU network is superior to that of the LSTM network.

[0023] The formula for the forward propagation of the GRU neuron is as follows:

[0024] r t = σ(W r · [h t-1 , x t + b r )

[0025] z t = σ(W z · [h t-1 , x t + b z )

[0026]

[0027] Among them, r t represents the reset gate, z t represents the update gate, h t-1 represents the memory of the previous moment, x t represents the new memory of the current moment, h t represents the output of all memories at the current moment, W r , W z , respectively represent the reset gate weight matrix, the update gate weight matrix, and the candidate set weight matrix, b r , b z and respectively represent the reset gate offset, the update gate offset, and the candidate set offset, tanh represents the hyperbolic tangent function, σ represents the sigmoid function, [h t-1 , x t represents h t-1and x t Splicing is connected, * means bitwise multiplication of vectors;

[0028] In summary, the GRU network model solves the long-distance dependency problem of RNN and the gradient problem in back propagation to a certain extent by introducing the gating mechanism and simplifying the gate function, while reducing the complexity of the algorithm.

[0029] It adopts a multi-layer chain structure, and each unit consists of residual connections and GRU neurons. The input of each unit is the output of the corresponding unit in the previous layer and the output of the previous unit in the same layer.

[0030] Before the vector is input to each hidden layer unit, the residual data is first superimposed and normalized. The residual data is the input information transmitted from the previous unit to this unit. This not only adds more useful information to the input data and enhances the mapping relationship between output and input, but also completely retains the original information input to this neuron in the input of the next neuron. If the data processing effect of this neuron is not positive, it is only necessary to reduce the weight of the non-residual data in the weight training of the next neuron, and the original information is still retained, that is, the non-positive effect can be weakened, which can effectively prevent the degradation problem in the deep neural network training. Degradation means that the deep neural network gradually decreases the loss by increasing the number of network layers, and then reaches saturation. If the number of network layers continues to increase, the loss will increase instead, that is, the neuron has a non-positive effect. The formula for adding residual blocks is as follows:

[0031] i k+1 =Norm(o k +o k-1 )

[0032] Among them, i k+1 is the input of the next neuron k+1, o k is the output of the current neuron k, o k-1 is the output of the previous neuron k-1, Norm(·) is data normalization, and Relu function is generally used.

[0033] Further, the specific process of S4 is as follows:

[0034] Radar working mode switching prediction is mainly based on the residual GRU network model based on multi-head attention. The underlying switching logic of T consecutive radar working mode types is mined. The model extracts and compresses this logical feature and outputs the T+1 type as the prediction result; as shown below:

[0035]

[0036] Specifically, a large number of historical radar working mode type switching sequences are used as prior information. After type coding, every consecutive T radar working mode type data X = {x t-T+1 , x t-T+2 ,..., x t} is used as the input of the model, and the (T + 1)-th radar working mode type is used as the target output to effectively train and learn the model; finally, using the trained model, every time consecutive T radar working mode types are input, the (T + 1)-th type can be accurately output as the prediction result; the multi-head attention mechanism and residual blocks can improve the feature sensitivity of the model and the comprehensiveness of the model's feature mining for the input data;

[0037] The residual GRU network model based on multi-head attention built by S2 and S3 consists of an input layer, a hidden layer, an output layer, and a fully connected layer;

[0038] The input layer is mainly used to receive the input of sequence data and allocate the attention of the model. Before the data is passed to the hidden layer, if the input dimension does not match the hidden layer dimension, the dimension will be matched through linear transformation and processed by the multi-head attention mechanism to improve the model's expression ability and generalization ability for time series data, enabling it to better capture the correlation information before and after time series data, thereby improving the model performance;

[0039] The hidden layer is used to mine the features of sequence data. To improve the model's ability to mine the underlying logic of the sequence, the output layer is mainly used to transfer the feature data mined by the hidden layer to the fully connected layer. The fully connected layer finally compresses and integrates the features to complete the classification function of the model, acting as a "classifier"; the formula of the fully connected layer is as follows:

[0040] y t+1 = S(h·W o + b o )

[0041] Among them, y t+1 represents the probability vector output by the fully connected layer, that is, the data at the (t + 1)-th moment predicted, h represents the output of the output layer, W o represents the weight matrix of the fully connected layer, b o represents the offset of the fully connected layer, and S(·) represents the softmax activation function; then the obtained prediction vector is decoded, and finally the prediction result of the radar working mode is obtained; the number of radar working mode types is k, and the softmax activation function is as follows:

[0042]

[0043] After the network training is completed, it can be used for the prediction of radar working mode switching.

[0044] Another object of the present invention is to provide a system based on the intelligent prediction method for radar working mode switching, the system comprising:

[0045] A data processing module, configured to encode the historical radar working mode switching sequence in order and represent it in binary;

[0046] A multi-head attention module, configured to allocate the attention of the model to different positions of the input sequence;

[0047] A residual GRU module, configured to protect information from being damaged during transmission and avoid performance degradation caused by the non-positive effects of neurons;

[0048] A network training module, configured to fully train the residual GRU network based on multi-head attention by using the Adam algorithm to obtain a final network model.

[0049] Another object of the present invention is to provide a device for the intelligent prediction method of radar working mode switching, comprising: a memory and a processor, the memory stores a computer program, and when the computer program is executed by the processor, the intelligent prediction method of radar working mode switching can be implemented.

[0050] Another object of the present invention is to provide a storage medium for receiving a user input program, and when the computer program stored in the storage medium is executed by a processor, the intelligent prediction method of radar working mode switching can be used to perform the prediction of radar working mode switching of the residual GRU network based on multi-head attention.

[0051] Combined with the above technical solutions and the technical problems solved, the advantages and positive effects of the technical solutions to be protected by the present invention are:

[0052] First, to address the problem that the prediction of radar frequency usage behavior intentions has a strong dependence on the relative positions of features, the system analyzes the relationship between the frequency usage behavior intentions and working modes of phased array radars, and analyzes the basic principles of working modes and the underlying logic of working mode switching. Then, the present invention uses a Residual Gated Recurrent Unit with Multi-Head Attention (RGRU-MHA) network to complete the prediction of phased array radar working mode switching. Different from the simple Gated Recurrent Unit (GRU) network, in the present invention, for the situation where there are often transitional working modes in the phased array radar working mode switching sequence, resulting in the adjacent working modes not necessarily having the strongest correlation, a multi-head attention mechanism is adopted to reasonably allocate the attention ratio of the model when inputting the working mode switching sequence, and a residual block is added when transmitting features in the hidden layer of the GRU network to avoid feature loss in the case of non-positive effects of neurons and prevent overfitting or underfitting of the model, effectively improving the model prediction accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] Figure 1 is a flowchart of an intelligent prediction method for radar working mode switching.

[0054] Figure 2 is a schematic diagram of a single-head attention mechanism.

[0055] Figure 3 is a schematic diagram of a multi-head attention mechanism.

[0056] Figure 4 is a structural diagram of a GRU neuron.

[0057] Figure 5 is a schematic diagram of the hidden layer of a residual GRU network.

[0058] Figure 6 is a flowchart of the prediction of radar working mode switching based on a GRU network with multi-head attention.

[0059] Figure 7 is a residual GRU network model based on a multi-head attention mechanism.

[0060] Figure 8 is a schematic diagram of the structure of a radar working mode switching prediction system based on a residual GRU network with multi-head attention;

[0061] Figure 9 is a schematic diagram of the simulation experiment results of a radar working mode switching prediction system based on a residual GRU network with multi-head attention;

[0062] In the figure: 1. Data processing module; 2. Multi-head attention module; 3. Residual GRU module; 4. Network training module. Detailed implementation manners

[0063] In order to make the objectives, technical solutions and advantages of the present invention more clear and understandable, the present invention will be further described in detail below in conjunction with embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0064] As Figure 1 shown, an embodiment of the present invention provides an intelligent prediction method for radar working mode switching, and the method includes:

[0065] S1. Sequentially encode the historical radar working mode switching sequence and represent it in binary as the input of the network built in the subsequent steps;

[0066] S2. Build a multi-head attention mechanism network to allocate the same attention degree of the model to different positions of the input sequence;

[0067] S3. Build a residual GRU network on the basis of S2, and adopt a chained connection method. Before each piece of information is input into the GRU neuron, the input of the previous neuron is first superimposed as a residual to protect the information from being damaged during the transmission process and avoid performance degradation caused by the non-positive effect of the neuron;

[0068] S4. Use the data set obtained in S1 to fully train the residual GRU network based on multi-head attention by using the Adam algorithm to obtain the final network model.

[0069] An embodiment of the present invention proposes an intelligent prediction method for radar working mode switching based on a multi-head attention mechanism and a residual GRU network, aiming to achieve efficient prediction of future working mode switching through learning and feature extraction of historical data. The working principle thereof will be described in detail below.

[0070] This method first sequentially encodes the historical radar working mode switching sequence, digitally represents different working mode switching behaviors, and further converts them into an input sequence in binary format. Through binary encoding, it is ensured that the input data structure is unified and suitable for neural network processing, and at the same time, information loss caused by different sequence lengths or mode complexities is avoided. This step provides high-quality input data for subsequent network training and prediction.

[0071] Construct a multi - head attention mechanism in the network model to allocate different attentions of the model to each position in the input sequence. The multi - head attention mechanism calculates the correlations between different positions in the sequence through multiple independent attention heads, and can capture both short - term and long - term dependencies simultaneously. This mechanism effectively enhances the model's ability to understand the context of the pattern - switching historical sequence, thereby providing a more accurate feature representation for the prediction task.

[0072] Based on the multi - head attention mechanism, a residual GRU (Gated Recurrent Unit) network is further built. The residual GRU network adopts a chain - connection method. Before each input information enters the GRU neuron, the input of the previous neuron is superimposed as a residual and introduced. This can not only preserve the integrity of historical information during transmission but also avoid performance degradation caused by non - positive effects of neurons (such as vanishing or exploding gradients). In addition, the residual connection enhances the network's learning ability for complex pattern - switching behaviors and improves the accuracy and stability of prediction.

[0073] In the final network training stage, the Adam optimization algorithm is used to fully train the residual GRU network based on multi - head attention. By using the dataset constructed from the historical radar working mode switching sequences, the Adam algorithm adjusts the weights and parameters of the network to minimize the prediction error. The trained network model can efficiently identify the potential patterns in the sequence patterns and output high - precision pattern - switching prediction results.

[0074] In summary, by combining the multi - head attention mechanism and the residual GRU network, this method fully utilizes the time - series characteristics of the radar historical working mode switching data, improves the ability to capture and predict complex pattern - switching behaviors, and provides strong technical support for the intelligence and automation of radar systems.

[0075] The specific steps of S1 are as follows:

[0076] First, construct training data based on the time series of the detected historical radar working modes. Take the first T radar working modes as input data and the (T + 1) - th radar working mode as the output label to obtain numerous radar working mode dataset samples. And encode and normalize the text information of each radar working mode in the dataset so that different data vectors represent different radar working modes. Data normalization is to prevent serious vanishing gradient problems during backpropagation when the value is too small, and gradient explosion problems and the loss of due representation ability caused by forcing the activation function to compress features in the non - linear interval when the input is too large. Data normalization adopts deviation standard normalization as follows:

[0077]

[0078] Among them, max and min are respectively the minimum and maximum values after encoding the radar working mode.

[0079] Specifically, S2 includes:

[0080] In the attention mechanism, an attention head includes three parameters: query (q), key (k), and value (v). The semantic correlation score is obtained by calculating the dot product of the query and the key, and then the semantic correlation score is converted into an attention weight through an activation function. Finally, the attention weight is applied to the value to obtain the final attention output. The single-head attention mechanism is as Figure 2 shown.

[0081] The input matrix is X = [x 1 T , x 2 T ,..., x n T T , x i (i = 1, 2,..., n) is the vector of each input unit. Through the linear transformation matrices W Q , W K and W V the query matrix Q = X · W Q = [q 1 T , q 2 T ,..., q n T T , the key matrix K = X · W K = [k 1 T , k 2 T ,..., k n T T and the value matrix V = X · W V = [v 1 T , v 2 T ,..., v n T T The output of the final single-head attention mechanism is as follows:

[0082]

[0083] Among them, σ is the softmax activation function, ​​​​is the scaling factor, which can avoid the problem that the value becomes too large after the dot product, affecting the stability of softmax, and the problem of gradient explosion caused by too large gradients during backpropagation, d k usually takes the dimension of the key vector;

[0084] Although the single-head attention mechanism can, to a certain extent, reasonably allocate the model's attention and help the model focus on important information in the input sequence, in complex feature extraction problems, the result of a single attention allocation cannot help the model completely capture all the important information in the input sequence, which may lead to the problem of losing some key information or an unsatisfactory allocation result; therefore, the multi-head attention MHA mechanism is proposed. Each attention mechanism consists of multiple single-head attention mechanisms. Multiple single-head attention mechanisms process the same input sequence in parallel and independently. Each single-head attention mechanism can adaptively focus on important information at different positions in the input sequence. Finally, the allocation results of each single-head attention mechanism are concatenated and integrated, and then passed through a linear transformation matrix to obtain the final attention allocation result; the multi-head attention mechanism enables the model to simultaneously focus on multiple important information from different positions in the input sequence, preventing the loss of multi-level information and deep features, thereby improving the robustness and generalization ability of the model and further helping subsequent feature extraction; the multi-head attention mechanism is widely used in the field of deep learning, especially in the direction of natural language processing, and has greatly improved the performance of models with long-term and short-term dependencies on time series sequences. The multi-head attention mechanism is as Figure 3 shown.

[0085] The specific process of the above-mentioned S3 is as follows:

[0086] Like the LSTM network, the GRU network is a variant recurrent network model designed to solve the gradient problem caused by the long-distance dependence characteristic of the RNN during backpropagation; the LSTM network has three gate functions, namely the input gate, the forget gate, and the output gate; the input gate controls how much of the newly added memory is retained, the forget gate controls how much of the previous memory is retained, and the output gate integrates all the retained memories; while the GRU network has a lower complexity and a simpler structure, with only two gate functions, namely the update gate and the reset gate; compared with the LSTM network, the GRU network simplifies the two gates of the input gate and the forget gate in the LSTM network to an update gate on the basis of integrating the cell state and the hidden state in each gate, and further adjusts the memory retention ratio using the reset gate; this structure enables the GRU network to compress the number of parameters a lot and greatly reduces the computational complexity on the basis of not significantly affecting the model performance; even in some tasks with fewer input time steps and severe short-term dependencies, the performance of the GRU network is better than that of the LSTM network; the specific structure of the GRU neuron is as Figure 4 shown.

[0087] The formula for the forward propagation of GRU neurons is as follows:

[0088] r t = σ(W r · [h t-1 , x t + b r )

[0089] z t = σ(W z · [h t-1 , x t + b z )

[0090]

[0091] Among them, r t represents the reset gate, z t represents the update gate, h t-1 represents the memory of the previous moment, x t represents the new memory of the current moment, h t represents the output of all memories at the current moment, W r , W z , respectively represent the reset gate weight matrix, the update gate weight matrix, and the candidate set weight matrix, b r , b z and respectively represent the reset gate offset, the update gate offset, and the candidate set offset, tanh represents the hyperbolic tangent function, σ represents the sigmoid function, [h t-1 , x t represents the concatenation of h t-1 and x t ; * represents the element-wise multiplication of vectors;

[0092] Generally speaking, by introducing the gating mechanism and simplifying the gate function, the GRU network model solves the long-distance dependence problem of RNN and the gradient problem in backpropagation to a certain extent, while reducing the algorithm complexity;

[0093] Adopts a multi-layer chain structure, and each unit consists of a residual connection and a GRU neuron; among them, the input of each unit is the output of the corresponding unit in the previous layer and the output of the previous unit in the same layer; a unit in the hidden layer is as Figure 5 shown.

[0094] Before the vector is input to each hidden layer unit, the residual data is first superimposed and normalized. The residual data is the input information transmitted from the previous unit to this unit. This not only adds more useful information to the input data and enhances the mapping relationship between output and input, but also completely retains the original information input to this neuron in the input of the next neuron. If the data processing effect of this neuron is not positive, it is only necessary to reduce the weight of the non-residual data in the weight training of the next neuron, and the original information is still retained, that is, the non-positive effect can be weakened, which can effectively prevent the degradation problem in the deep neural network training. Degradation means that the deep neural network gradually decreases the loss by increasing the number of network layers, and then reaches saturation. If the number of network layers continues to increase, the loss will increase instead, that is, the neuron has a non-positive effect. The formula for adding residual blocks is as follows:

[0095] i k+1 =Norm(o k +o k-1 )

[0096] Among them, i k+1 is the input of the next neuron k+1, o k is the output of the current neuron k, o k-1 is the output of the previous neuron k-1, Norm(·) is data normalization, and Relu function is generally used.

[0097] The specific process of S4 is as follows:

[0098] Radar working mode switching prediction is mainly based on the residual GRU network model based on multi-head attention. The underlying switching logic of T consecutive radar working mode types is mined. The model extracts and compresses this logical feature and outputs the T+1 type as the prediction result; as shown below:

[0099]

[0100] Specifically, a large number of historical radar working mode type switching sequences are used as prior information. After type coding, each T consecutive radar working mode type data X = {x t-T+1 ,x t-T+2 ,...,x t} as the input of the model, and the T+1th radar working mode type as the target output, so as to effectively train and learn the model; finally, using the trained model, each time T consecutive radar working mode types are input, the T+1th type can be accurately output as the prediction result; the multi-head attention mechanism and residual block can improve the feature sensitivity of the model and the comprehensiveness of the model's feature mining of the input data; the radar working mode switching prediction model based on the GRU network with multi-head attention is as followsFigure 6 as shown

[0101] The residual GRU network model based on multi - head attention built by S2 and S3. The model consists of an input layer, a hidden layer, an output layer, and a fully - connected layer; The residual GRU network model based on multi - head attention is as Figure 7 shown

[0102] The input layer is mainly used to receive the input of sequence data and allocate the attention of the model. Before the data is passed to the hidden layer, if the input dimension does not match the hidden - layer dimension, the dimension will be matched through linear transformation, and processed by the multi - head attention mechanism to improve the model's ability to express and generalize time - series data, enabling it to better capture the correlation information before and after time - series data, thereby improving the model performance;

[0103] The hidden layer is used to mine the features of sequence data. To improve the model's ability to mine the underlying logic of the sequence, the output layer is mainly used to transfer the feature data mined by the hidden layer to the fully - connected layer. The fully - connected layer finally compresses and integrates the features to complete the classification function of the model, acting as a "classifier"; The formula of the fully - connected layer is as follows:

[0104] y t+1 = S(h·W o + b o )

[0105] where, y t+1 represents the probability vector output by the fully - connected layer, that is, the data at the (t + 1)-th moment predicted, h represents the output of the output layer, W o represents the weight matrix of the fully - connected layer, b o represents the offset of the fully - connected layer, S(·) represents the softmax activation function; Then the obtained prediction vector is decoded, and finally the prediction result of the radar working mode is obtained; The number of radar working - mode types is k, and the softmax activation function is as follows:

[0106]

[0107] After the network training is completed, it can be used for the prediction of radar working - mode switching.

[0108] As Figure 8 shown, an embodiment of the present invention provides a system based on the intelligent prediction method for radar working - mode switching. The system includes:

[0109] The data - processing module 1, used to encode the historical radar working - mode switching sequence in order and represent it in binary;

[0110] The multi - head attention module 2, used to allocate the attention of the model to different positions of the input sequence;

[0111] The residual GRU module 3 is used to protect information from being damaged during transmission and avoid performance degradation caused by the non-positive effects of neurons;

[0112] The network training module 4 is used to fully train the residual GRU network based on multi-head attention using the Adam algorithm to obtain the final network model.

[0113] An embodiment of the present invention provides a device for the intelligent prediction method of radar working mode switching, including: a memory and a processor. The memory stores a computer program, and when the computer program is executed by the processor, it can implement the intelligent prediction method of radar working mode switching.

[0114] An embodiment of the present invention provides a storage medium for receiving a user input program. When the computer program stored in the storage medium is executed by a processor, it can perform the intelligent prediction method of radar working mode switching and predict the radar working mode switching based on the residual GRU network with multi-head attention.

[0115] The technical effects of the present invention will be described in detail below in combination with simulation experiments.

[0116] The simulation uses the open-source deep learning framework Pytorch based on the Python language. According to the phased array radar working mode switching logic, a radar working mode switching data set is simulated, which includes six radar working modes: speed search mode, range measurement while searching mode, tracking while searching mode, search plus tracking mode, single target tracking mode, and situation awareness mode. The speed search mode is encoded as 0, the range measurement while searching mode is encoded as 1, the search plus tracking mode is encoded as 2, and so on, and is represented by the binary method. The total number of samples in the data set is 40,000. Randomly select 36,000 samples as the training set to train the model, 2,000 samples as the validation set to supervise the training, and 2,000 samples as the test set to test the training results. Build a residual GRU network model based on multi-head attention. The following are the selected parameters of the model: the training batch size batchsize is 50, the training iteration steps epoch is 50, the number of heads of the multi-head attention mechanism is 4, the input layer dimension is 16, the number of hidden layers is 4, the hidden layer dimension is 32, and the output data dimension of the fully connected layer is (6, 1). The cross-entropy loss function is used as the objective function and the Adam algorithm is used as the training optimizer. The parameter initialization of the Adam optimizer is as follows: the learning rate η is 0.005, the decay coefficient β 1 for the first moment estimation is 0.9, and the decay coefficient β 2 for the second moment estimation is 0.999. The prediction accuracies of the six radar working modes are as Figure 9As shown, after 20 iterations, the accuracy curves of various working modes tend to be stable, and the prediction accuracy can reach over 90%. When initializing the parameters of the fully connected layer, the weight of the speed search mode is set to 1, and other weights are set to 0. Therefore, all initial outputs are in the speed search mode. After iteration, the accuracy of the speed search mode gradually decreases from 100% to about 90%, and the accuracy of other radar working modes rapidly increases from 0% to about 90%. Ultimately, the differences in the prediction accuracies of different radar working modes are mainly due to the different necessities of each working mode in the task, and individual working modes may not necessarily exist in the switching sequence.

[0117] It should be noted that the embodiments of the present invention can be implemented by hardware, software, or a combination of software and hardware. The hardware part can be implemented using dedicated logic; the software part can be stored in a memory and executed by a suitable instruction execution system, such as a microprocessor or dedicated designed hardware. Those of ordinary skill in the art can understand that the above devices and methods can be implemented using computer-executable instructions and / or included in processor control code, such as provided on a carrier medium such as a disk, CD, or DVD-ROM, a programmable memory such as read-only memory (firmware), or a data carrier such as an optical or electronic signal carrier. The devices and modules of the present invention can be implemented by hardware circuits of programmable hardware devices such as very large scale integrated circuits or gate arrays, semiconductors such as logic chips, transistors, etc., or programmable logic devices such as field programmable gate arrays, or can be implemented by software executed by various types of processors, or can be implemented by a combination of the above hardware circuits and software, such as firmware.

[0118] The above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be covered by the protection scope of the present invention.

Claims

1. A radar working mode switching intelligent prediction method, characterized in that: The method includes: S1, encode the historical radar working mode switching sequence in sequence and represent it in binary to serve as the input of the network built in the subsequent steps; S2, build a multi-head attention mechanism network to distribute the model's attention to different positions in the input sequence; S3, builds a residual GRU network based on S2, using a chain connection method. Each time information is input into a GRU neuron, the input of the previous neuron is superimposed as a residual to protect the information from being destroyed during transmission and to avoid performance degradation caused by non-positive effects of neurons; S4, using the data set obtained in S1, uses the Adam algorithm to fully train the residual GRU network based on multi-head attention to obtain the final network model.

2. The radar working mode switching intelligent prediction method according to claim 1 is characterized in that: The S1 specifically includes: First, the training data is constructed based on the time series of historical radar working modes obtained by detection. The first T radar working modes are used as input data, and the T+1th radar working mode is used as the output label to obtain a large number of radar working mode data set samples; and the text information of each radar working mode in the data set is encoded and normalized, so that different data vectors represent different radar working modes; data normalization is to prevent the serious gradient vanishing problem in back propagation when the input is too small, and the gradient explosion problem in back propagation when the input is too large, and the activation function will forcibly compress the features in the nonlinear interval, thereby losing its due representation ability; data normalization uses deviation standard normalization, as shown below: Among them, max and min are the minimum and maximum values ​​after the radar working mode is encoded.

3. The radar operating mode switching intelligent prediction method according to claim 1, characterized in that: The S2 specifically includes: In the attention mechanism, an attention head includes three parameters: query (query, q), key (key, k) and value (value, v). The semantic relevance score is obtained by calculating the dot product of the query and the key, and then the semantic relevance score is converted into an attention weight through an activation function. Finally, the attention weight is applied to the value to obtain the final attention output; The input matrix is ​​X = [x1 T ,x2 T ,...,x n T ] T , x i (i=1,2,...,n) is the vector of each input unit, through the linear change matrix W Q , W K and W V Get the query matrix Q = X·W Q =[q1 T ,q2 T ,…,q n T ] T , bond matrix K = X·W K =[k1 T ,k2 T ,...,k n T ] T Sum matrix V = X·W V =[v1 T ,v2 T ,...,v n T ] T , the final output of the single-head attention mechanism is as follows: Among them, σ is the softmax activation function, is a scaling factor, which can avoid the problem of too large a value after the dot product affecting the stability of softmax and the problem of gradient explosion caused by too large a gradient during back propagation. k Usually the dimension of the key vector is taken; Although the single-headed attention mechanism can reasonably allocate the model's attention to a certain extent and help the model focus on important information in the input sequence, in complex feature extraction problems, a single attention allocation result cannot help the model fully capture all the important information in the input sequence, which will lead to the loss of some key information or the allocation result is too unsatisfactory. Therefore, the multi-headed attention MHA mechanism is proposed. Each attention mechanism is composed of multiple single-headed attention mechanisms. Multiple single-headed attention mechanisms process the same input sequence in parallel and independently. Each single-headed attention mechanism can adaptively focus on important information at different positions in the input sequence. Finally, the allocation results of each single-headed attention mechanism are spliced ​​and integrated, and then the final attention allocation result is obtained after passing through the linear transformation matrix. The multi-headed attention mechanism enables the model to simultaneously pay attention to multiple important information from different positions in the input sequence, preventing the loss of multi-level information and deep features, thereby improving the robustness and generalization ability of the model, and further helping subsequent feature extraction. The multi-headed attention mechanism is widely used in the field of deep learning, especially in the direction of natural language processing, and has greatly improved the performance of models with long-term and short-term dependencies on time series.

4. The radar operating mode switching intelligent prediction method according to claim 1, characterized in that: The specific process of S3 is as follows: GRU network, like LSTM network, is a variant recurrent network model for solving the gradient problem caused by the long-distance dependency characteristics of RNN in back propagation. LSTM network has three gate functions, namely input gate, forget gate and output gate. The input gate controls how much newly added memory is retained, the forget gate controls how much previous memory is retained, and the output gate integrates all retained memories. GRU network has lower complexity and simpler structure, with only two gate functions, namely update gate and reset gate. Compared with LSTM network, GRU network is based on the integration of cell state and hidden state in each gate, simplifies the input gate and forget gate in LSTM network into update gate, and uses reset gate to further adjust the memory retention ratio. This structure enables GRU network to compress the number of parameters a lot, greatly reducing the computational complexity without significantly affecting the model performance. Even in some tasks with less input time and severe short-term dependency, the performance of GRU network is better than that of LSTM network. The formula for the forward propagation of a GRU neuron is as follows: r t =σ(W r ·[h t-1 ,x t ]+b r ) z t =σ(W z ·[h t-1 ,x t ]+b z ) Among them, r t represents the reset gate, z t represents the update gate, h t-1 Indicates the memory of the previous moment, x t Indicates the new memory at the current moment, h t Represents the output of all memories at the current moment, W r , W z , Represents the reset gate weight matrix, update gate weight matrix and candidate set weight matrix, respectively, r , b z and They represent the reset gate offset, update gate offset, and candidate set offset, respectively. Tanh represents the hyperbolic tangent function, σ represents the sigmoid function, [h t-1 ,x t ] means h t-1 and x t Splicing is connected, * means bitwise multiplication of vectors; In summary, the GRU network model solves the long-distance dependency problem of RNN and the gradient problem in back propagation to a certain extent by introducing the gating mechanism and simplifying the gate function, while reducing the complexity of the algorithm. It adopts a multi-layer chain structure, and each unit consists of residual connections and GRU neurons. The input of each unit is the output of the corresponding unit in the previous layer and the output of the previous unit in the same layer. Before the vector is input to each hidden layer unit, the residual data is first superimposed and normalized. The residual data is the input information transmitted from the previous unit to this unit. This not only adds more useful information to the input data and enhances the mapping relationship between output and input, but also completely retains the original information input to this neuron in the input of the next neuron. If the data processing effect of this neuron is not positive, it is only necessary to reduce the weight of the non-residual data in the weight training of the next neuron, and the original information is still retained, that is, the non-positive effect can be weakened, which can effectively prevent the degradation problem in the deep neural network training. Degradation means that the deep neural network gradually decreases the loss by increasing the number of network layers, and then reaches saturation. If the number of network layers continues to increase, the loss will increase instead, that is, the neuron has a non-positive effect. The formula for adding residual blocks is as follows: i k+1 =Norm(o k +o k-1 ) Among them, i k+1 is the input of the next neuron k+1, o k is the output of the current neuron k, o k-1 is the output of the previous neuron k-1, Norm(·) is data normalization, and Relu function is generally used.

5. The radar working mode switching intelligent prediction method according to claim 2 is characterized in that: The specific process of S4 is as follows: Radar working mode switching prediction is mainly based on the residual GRU network model based on multi-head attention. The underlying switching logic of T consecutive radar working mode types is mined. The model extracts and compresses this logical feature and outputs the T+1 type as the prediction result; as shown below: Specifically, a large number of historical radar working mode type switching sequences are used as prior information. After type coding, each T consecutive radar working mode type data X = {x t-T+1 ,x t-T+2 ,...,x t } as the input of the model, and the T+1th radar working mode type as the target output, so as to effectively train and learn the model; finally, using the trained model, each time inputting T consecutive radar working mode types, the T+1th type can be accurately output as the prediction result; the multi-head attention mechanism and residual block can improve the feature sensitivity of the model and the comprehensiveness of the model's feature mining of the input data; The residual GRU network model based on multi-head attention built by S2 and S3 consists of an input layer, a hidden layer, an output layer, and a fully connected layer; The input layer is mainly used to receive the input of sequence data and allocate model attention. Before the data is passed to the hidden layer, if the input dimension does not match the hidden layer dimension, the dimension will be matched through linear change and processed through the multi-head attention mechanism to improve the model's expression and generalization capabilities for time series data, so that it can better capture the correlation information before and after the time series data, thereby improving the model performance; The hidden layer is used to mine the features of sequence data. In order to improve the model's ability to mine the underlying logic of the sequence, the output layer is mainly used to pass the feature data mined by the hidden layer to the fully connected layer. The fully connected layer finally compresses and integrates the features to complete the classification function of the model and plays the role of a "classifier". The formula of the fully connected layer is as follows: y t+1 =S(h·W o +b o ) Among them, y t+1 represents the probability vector of the output of the fully connected layer, that is, the predicted data at the t+1th moment, h represents the output of the output layer, and W o represents the weight matrix of the fully connected layer, b o represents the offset of the fully connected layer, S(·) represents the softmax activation function; the obtained prediction vector is decoded, and the prediction result of the radar working mode is finally obtained; the number of radar working mode types is k, and the softmax activation function is as follows: After the network training is completed, it can be used to predict the radar working mode switching.

6. A system based on the radar working mode switching intelligent prediction method according to any one of claims 1 to 5, characterized in that: The system includes: A data processing module is used to encode the historical radar working mode switching sequence in sequence and express it in binary; Multi-head attention module, used to distribute the model's attention to different positions in the input sequence; The residual GRU module is used to protect information from being destroyed during transmission and to avoid performance degradation caused by non-positive effects of neurons; The network training module is used to fully train the multi-head attention-based residual GRU network using the Adam algorithm to obtain the final network model.

7. The device of the radar working mode switching intelligent prediction method according to any one of claims 1 to 5, characterized in that: include: A memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the radar working mode switching intelligent prediction method according to any one of claims 1 to 5 can be implemented.

8. A storage medium for receiving a user input program, characterized in that: When the computer program stored in the storage medium is executed by the processor, it can perform radar working mode switching prediction based on the residual GRU network with multi-head attention based on the radar working mode switching intelligent prediction method according to any one of claims 1 to 5.

Citation Information

Cited By

  • Network switching device and control method

    CN121098702A

  • Underground water level simulation method based on hysteresis response characteristics of multi-aquifer system

    CN121389767A