A waveform adaptive selection method based on deep reinforcement learning
Through the waveform adaptive selection method of deep reinforcement learning, a discrete system state space model and entropy state reward mechanism are constructed, which solves the adaptability and stability problems of traditional radar waveform selection methods in complex environments, and realizes the rapid and accurate waveform parameter selection of radar systems under different tasks.
Patent Information
- Application Number
- CN202211277900.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-19
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2042-10-19
AI Technical Summary
Traditional radar waveform selection methods have poor adaptability under complex interference conditions and high computational complexity. The DQN-based method has the problem of unstable state estimation in radar positioning and tracking scenarios.
A waveform adaptive selection method based on deep reinforcement learning is adopted. By constructing a discrete system state space model, using the volumetric Kalman filter and DQN network, and combining the entropy state reward mechanism, the waveform parameter selection is optimized.
It improves the radar system's adaptability and waveform parameter selection accuracy in complex environments, reduces the motion space exploration time in the initial learning phase, and enhances its rapid adaptability under different tasks.
Smart Images

Figure CN115561723B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of radar waveform selection, relates to a radar waveform adaptive selection method, and further relates to a waveform adaptive selection method based on deep reinforcement learning. Background Art
[0002] The statements in this section merely provide background information related to the present disclosure and may not constitute prior art.
[0003] Adaptive selection of radar transmit waveform parameters is one of the important research contents in the field of radar waveform selection and is widely used in fields such as military reconnaissance. The purpose of adaptive selection of radar waveform parameters is to enable the radar to perform typical radar tasks in a complex, competitive, and dynamic radar environment, meet the actual application requirements of modern radar high-precision detection, tracking, and surveillance, and give the radar adaptive capabilities.
[0004] The cognitive radar system is a closed-loop circuit consisting of a cognitive controller, cognitive sensors, and the environment. The cognitive sensors feed back the received environmental information to the cognitive controller, which then selects the transmitted waveform parameters based on the feedback information, thereby continuously improving the system's estimation and detection performance. Therefore, waveform selection is the most important issue in cognitive radar systems.
[0005] Traditional radar waveform selection methods include methods that build a small waveform library based on pre-set application scenarios to fix radar transmit waveform parameters, and methods that learn the relationship between radar transmit waveform parameters and tracking and positioning accuracy based on genetic algorithms, neural networks, etc. However, although the method of fixing radar transmit waveform parameters provides a small waveform library for application scenarios, the radar faces the problem of limited application under complex interference conditions. For an unknown radar task, the waveform parameter agility algorithm based on the criterion function needs to search in the waveform library, which has a relatively high computational complexity. The waveform selection method based on neural networks requires obtaining sufficient effective training samples, which is also relatively difficult for unknown radar tasks. Obtaining a set of new application scenarios based on genetic algorithms and neural network algorithms has problems such as poor adaptability, high computational complexity, and long processing time. Therefore, in recent years, radar waveform selection has attempted to adopt methods based on reinforcement learning. Reinforcement learning can gradually adapt to the environment by interacting with the environment and utilizing the feedback information of the environment in a dynamic and complex environment. When selecting waveforms, the feedback information of the environment can be used as the input of reinforcement learning, and decisions can be made based on the feedback to select the corresponding transmit waveform parameters.
[0006] Deep reinforcement learning outperforms traditional algorithms in radar waveform selection by modeling radar simulation environments and adaptively selecting waveforms from a radar waveform library. DQN (Deep Q-learning Network), a representative deep reinforcement learning network, learns the relationship between action selection, the system's states before and after the action, and the corresponding rewards, aiming to create a correspondence between actions and rewards in the current state. However, in radar positioning and tracking scenarios, due to the Markov nature of radar observations, interference noise, and other issues, the direct use of DQN leads to large fluctuations in the value network in discrete space and unstable state estimation. Summary of the Invention
[0007] The purpose of the present invention is to address the above-mentioned problems and provide a waveform adaptive selection method based on deep reinforcement learning, so that the radar system has environmental perception capabilities, can adaptively, dynamically and stably adjust the radar waveform, and achieve reasonable resource allocation, thereby solving the above-mentioned problems.
[0008] The technical solutions of the present invention are as follows:
[0009] A waveform adaptive selection method based on deep reinforcement learning, comprising:
[0010] Step S1: Model the simulation scenario by constructing a discrete system state space model. Based on the simulation scenario, initialize the discrete system state space model and randomly initialize the weights of the DQN network.
[0011] Step S2: Construct the target entropy state by minimizing the Cramer-Rao lower bound of the target parameter estimate, and construct the true entropy reward of the waveform parameter selection action by the true entropy state before and after.
[0012] Step S3: using a cubature Kalman filter to predict the next moment target state of the action with different waveform parameters selected at the current moment;
[0013] Step S4: constructing a feed-forward loop path based on the predicted entropy state to form a predicted entropy reward;
[0014] Step S5: Design an action reward function for the multi-node scenario, use the DQN network to approximate the expected value of rewards for different actions in different states, and obtain the final waveform parameter to select the action.
[0015] Furthermore, the step S1 includes:
[0016] The state space model of the discrete system is defined as follows:
[0017]
[0018] in:
[0019] f (.) is the vector conversion function;
[0020] h (.) is another vector transformation function that maps the target from the state space to the observation space;
[0021] Indicates that the system is k The state of the moment;
[0022] Indicates k The observed value at the time;
[0023] represents an additional processing noise as a driving force for updating the state;
[0024] is the additional observation noise;
[0025] The system equations are defined as follows:
[0026]
[0027] in;
[0028] is the relative distance between the radar and the target
[0029] The speed of the target;
[0030] is the target acceleration;
[0031] The system covariance noise is defined as follows:
[0032]
[0033] in:
[0034] ;
[0035] The Cramer-Rao lower bound on the observation noise covariance is defined as follows:
[0036]
[0037] in:
[0038] ;
[0039] is the signal-to-noise ratio, ;
[0040] ;
[0041] .
[0042] Furthermore, the step S2 includes:
[0043] Step S21: The cognitive sensor obtains the current k The observation vector at time instant;
[0044] Step S22: The current K The Cramer-Rao lower bound of the moment-by-moment observation noise covariance is used as the input to the cubature Kalman filter to obtain the estimated error covariance matrix prediction value of the target state at the next moment and the estimated error covariance matrix ;
[0045] Step S23: Use the estimated error covariance matrix to predict the value and the estimated error covariance matrix Perform posterior probability estimation on the target state to obtain the predicted entropy state of the target parameter estimation and the true entropy state of the target parameter estimation;
[0046] Step S24: True entropy state Feedback is given to the cognitive controller to form feedback information.
[0047] Step S25: The cognitive controller receives the real entropy state information fed back by the cognitive sensor;
[0048] Step S26: Obtaining a real entropy reward from the entropy state information at the previous and next moments;
[0049] Step S27: storing a set of target states, waveform parameter selection actions, and obtained real entropy rewards;
[0050] Step S28: Select transmission waveform parameters.
[0051] Furthermore, the step S23 includes:
[0052] The entropy state is a measure of the posterior probability estimate distribution The uncertainty measure of the target state is measured by Shannon entropy, which is in the form of:
[0053]
[0054] By treating the target action as the observation noise covariance The estimated error covariance matrix is obtained by inputting , the estimated error covariance matrix Perform Shannon entropy measurement to obtain the current K The target true entropy state at the moment .
[0055] Furthermore, the step S26 includes:
[0056] Among them, the real entropy reward for a single node to execute an action is:
[0057]
[0058] in:
[0059] for k The actual entropy state at time -1;
[0060] for k The true entropy state at the moment.
[0061] Furthermore, the step S4 includes:
[0062] Step S41: Obtain N groups of waveform parameter action selections with high reward expectations from the DQN value approximation network;
[0063] Step S42: Take the N sets of waveform parameter actions as the input of the volumetric Kalman filter in step S22 to obtain the N sets of estimated error covariance matrix prediction values of the target state at the next moment ;
[0064] Step S43: N groups k Estimated error covariance matrix predicted value at time +1 Passed to the cognitive controller in step S3 to obtain the predicted entropy state of action selection ;
[0065] Step S44: reward N groups of prediction entropy As k +1 The basis for selecting the emission waveform parameters at time;
[0066] Step S45: Store N groups of target states, waveform parameter selection actions, and obtained predicted entropy rewards.
[0067] Furthermore, the step S43 includes:
[0068] The predicted entropy reward in the planning phase is:
[0069]
[0070] in:
[0071] for k The predicted entropy state at the moment;
[0072] fork The entropy state at a given moment.
[0073] Furthermore, the step S5 includes:
[0074] Step S51: Each node obtains the entropy reward value from its own perspective using step S2;
[0075] Step S52: impose a severe penalty on the action selection when the total bandwidth exceeds the limit of 30M bandwidth;
[0076] Step S53: Store the target state before and after the multi-node perspective, the selected waveform parameter action, and the predicted / real entropy reward obtained using step S23 / step S26.
[0077] Furthermore, the step S52 includes:
[0078] Each node competes for bandwidth for its own observation accuracy, and imposes severe penalties on waveform parameter selection actions that exceed the total bandwidth limit;
[0079]
[0080] in:
[0081] C is the penalty constant and ;
[0082] is the sum of the entropy rewards of each node;
[0083] The entropy reward corresponding to a waveform parameter selection action.
[0084] Furthermore, it also includes:
[0085] The DQN network that performs value approximation on parameter action selection is trained using the following steps:
[0086] (a) Randomly initialize or load a pre-trained model to initialize the DQN network parameters;
[0087] Initialize the waveform library parameters, experience replay memory library, number of iterations, and use random weights to adjust the weights of the value function network Initialize and initialize the weights of the target network ,make ;
[0088] During the training process, the pre-loaded network parameters were fine-tuned based on the network parameters. The learning rate was set to 0.001, the batch size was set to 32, and the Adam optimizer was used. After 200 generations of training, the network converged and the network parameters were obtained.
[0089] (b) using the target state before and after the multi-node perspectives recorded in step S53, the selected waveform parameter action, and the obtained real / predicted entropy reward as input for DQN network learning and training;
[0090] (c) Iterate steps S3-S5 to train the network.
[0091] Compared with the existing technology, the beneficial effects of the present invention are:
[0092] 1. A waveform adaptive selection method based on deep reinforcement learning. When performing DQN value approximation, it not only uses the estimated error covariance to obtain the relationship between the action and selection of the learning stage decision in the current state, but also obtains the predicted value of its estimated error oblique variance by trying multiple groups of waveform selection actions in the current state. It measures the uncertainty of the predicted value of the estimated target state under the planned action selection, overcoming the characteristics of slow action space exploration and low production speed in the early stage of training. Compared with existing technologies, it improves the rate of network learning and fitting.
[0093] 2. A waveform adaptive selection method based on deep reinforcement learning can quickly adjust and fit for different tasks. By adding and modifying reward functions and designing system equations, it can be easily transferred to other tasks. It enables the radar to interact with environmental information and make waveform parameter decisions based on environmental feedback, thereby improving the accuracy of waveform parameter selection. BRIEF DESCRIPTION OF THE DRAWINGS
[0094] Figure 1 This is a flowchart of a waveform adaptive selection method based on deep reinforcement learning;
[0095] Figure 2 is the state space model of the discrete system;
[0096] Figure 3 It is a single-node vertical nonlinear motion scene graph;
[0097] Figure 4 It is a single-node horizontal nonlinear motion scene graph;
[0098] Figure 5 It is a multi-node horizontal nonlinear motion scene graph;
[0099] Figure 6 Schematic diagram of single-node vertical nonlinear motion distance tracking error of the present invention and other methods;
[0100] Figure 7 Schematic diagram of single-node horizontal nonlinear motion distance tracking error compared with other methods. DETAILED DESCRIPTION
[0101] It should be noted that relational terms such as "first" and "second" are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus comprising the element.
[0102] The features and performance of the present invention are further described in detail below with reference to the embodiments.
[0103] Example 1
[0104] See also Figure 1-7 , a waveform adaptive selection method based on deep reinforcement learning, comprising the following steps:
[0105] Step S1: Model the simulation scenario by constructing a discrete system state space model. Based on this simulation scenario, initialize the discrete system state space model and randomly initialize the weights of the DQN network. It should be noted that the entire subsequent process is based on this discrete system state space simulation scenario; and initializing the network weights and state space is performed on this simulation scenario.
[0106] The discrete system state space model is a state transition model for nonlinear moving objects. The state changes of the target in the entire system method follow the laws of this discrete state space. The observed values and estimated values of the observed target can be obtained through the discrete state space model using a cubic Kalman filter.
[0107] Initializing network weights means randomly initializing network parameters for DQN using pytorch, or loading pre-trained weights as the initial network weights;
[0108] Initializing the state space model means giving the target an initial imprecise observation state, and then accurately positioning it by selecting the emission waveform;
[0109] Step S2: construct the target entropy state by minimizing the Cramer-Rao lower bound of the target parameter estimate, and construct the real entropy reward of the waveform parameter selection action by the real entropy state before and after the moment; that is, step S2 can be understood as constructing a cognitive sensor and a cognitive controller;
[0110] Step S3: using a cubature Kalman filter to predict the next moment target state of the action with different waveform parameters selected at the current moment;
[0111] Step S4: Construct the predicted entropy reward to form a feedforward loop according to the predicted entropy state; It should be noted that in order to alleviate the problem of high volatility in the DQN network, the volumetric Kalman filter is used K The entropy state of the estimated error covariance matrix prediction value at time +1 is K The entropy state of the moment-by-moment estimation error covariance matrix constitutes the predicted entropy reward in the planning phase. Based on this, multiple sets of training data are generated to overcome the difficulty of exploring the large action space and improve the robustness of the algorithm.
[0112] Step S5: Design an action reward function for the multi-node scenario, use the DQN network to approximate the expected value of rewards for different actions in different states, and obtain the final waveform parameter to select the action.
[0113] In this embodiment, specifically, step S1 includes:
[0114] The state space model of the discrete system extended by the state space model and the entropy state model is defined as follows:
[0115]
[0116] in:
[0117] f (.) is the vector conversion function;
[0118] h (.) is another vector transformation function that maps the target from the state space to the observation space;
[0119] Indicates that the system is k The state of the moment;
[0120] Indicates k The observed value at the time;
[0121] represents an additional processing noise as a driving force for updating the state;
[0122] is the additional observation noise;
[0123] The system equations are defined as follows:
[0124]
[0125] in;
[0126] is the relative distance between the radar and the target
[0127] The speed of the target;
[0128] is the target acceleration;
[0129] The noise is assumed to be Gaussian white noise, and the system covariance noise is defined as follows:
[0130]
[0131] in:
[0132] ;
[0133] The Cramer-Rao lower bound on the observation noise covariance is defined as follows:
[0134]
[0135] in:
[0136] ;
[0137] is the signal-to-noise ratio, ;
[0138] ;
[0139] .
[0140] As attached Figure 5 As shown in the figure, taking four nodes as an example, the frequency modulation slope b and pulse width λ are selected as the actions for selecting radar waveform parameters. The [λ, b] selected by the four nodes together constitute a selectable action. All selectable actions constitute an action space, and the size of the action space is 4096.
[0141]
[0142] The parameters, b, and λ, which constitute the transmitted waveform, are the pulse width and frequency modulation slope, which directly affect radar positioning and the observation effect obtained from the above formula. The observation effect can be influenced by choosing different frequency modulation slopes b and pulse widths λ.
[0143] The multi-node waveform library has 2 different pulse widths, 4 different frequency modulation slopes, and a total of 8 different waveform parameter pair combinations. The action space size is 4096. The four nodes select their own frequency modulation slope and pulse width. The bandwidth of a single node can reach up to 14.4MHz, and the total bandwidth of the four nodes is limited to 30MHz.
[0144] The target's horizontal distances relative to the four nodes are 5494m, 6029m, 7536m, and 8597m, its altitude is 6427m, its horizontal speed is 150m / s, its acceleration is 5, and it performs a nonlinear motion. In the initial states obtained by the four radar nodes, the horizontal distances are 5500m, 6000m, 7500m, and 8600m.
[0145] The entire observation process takes 10 seconds, the time interval between each selection of the transmitted waveform is 0.1 seconds, the carrier frequency is 10.4e9, the bandwidth limit is 30 MHz, and the initial covariance matrix is diag([1e6,1e4,1e4]).
[0146] Except for the initial estimated state, all other parameters of each node remain consistent, and the first moment is randomly selected to transmit the waveform parameter action.
[0147] In this embodiment, specifically, step S2 includes:
[0148] Building a cognitive perceptron:
[0149] Step S21: Get the current k The observation vector at the moment; preferably, the K The time refers to the current time, and the observation vector refers to the observed target state;
[0150] Step S22: The current K Cramer-Rao lower bound on the covariance of moment-to-moment observation noise R As the input of the cubature Kalman filter, the estimated error covariance matrix prediction value of the target state at the next moment is obtained and the estimated error covariance matrix ;
[0151] Step S23: Use the estimated error covariance matrix to predict the value and the estimated error covariance matrix Perform posterior probability estimation on the target state to obtain the predicted entropy state of the target parameter estimation and the true entropy state of the target parameter estimation;
[0152] It should be noted that the entropy state is a measure of the posterior probability estimate distribution The uncertainty measure of the target state is measured by Shannon entropy, which is in the form of:
[0153]
[0154] By treating the target action as the observation noise covariance The estimated error covariance matrix is obtained by inputting , the estimated error covariance matrix Perform Shannon entropy measurement to obtain the currentK The target true entropy state at the moment ;
[0155] Step S24: True entropy state Feedback is given to the cognitive controller to form feedback information.
[0156] Building a Cognitive Controller:
[0157] Step S25: receiving the real entropy state information fed back by the cognitive sensor;
[0158] Step S26: Obtaining a real entropy reward from the entropy state information at the previous and next moments;
[0159] Among them, the real entropy reward for a single node to execute an action is:
[0160]
[0161] in:
[0162] for k The actual entropy state at time -1;
[0163] for k The true entropy state of the moment;
[0164] The entropy state is a measure of the uncertainty of the estimated target state. Measuring the uncertainty of the target state before and after the action constitutes a short-term reward. The larger the reward value, the higher the target accuracy observed at the current moment.
[0165] The entropy state of the estimated error covariance matrix is used to construct the true entropy reward corresponding to the action, which is used as the reward function in the learning phase.
[0166] Step S27: storing a set of target states, waveform parameter selection actions, and obtained real entropy rewards;
[0167] Step S28: Select the transmission waveform parameters. Preferably, the DQN value approximation network gives the action selection with the maximum expected value in the current state.
[0168] In this embodiment, specifically, step S4 includes:
[0169] Step S41: Obtaining N sets of waveform parameter actions with high expected rewards from the DQN value approximation network; preferably, the results output by the DQN value approximation network are sorted from largest to smallest, and the top N results are selected as the obtained N sets of waveform parameter actions with high expected rewards;
[0170] Step S42: Take the N sets of waveform parameter actions as the input of the volumetric Kalman filter in step S22 to obtain the N sets of estimated error covariance matrix prediction values of the target state at the next moment ;
[0171] Step S43: N groups k Estimated error covariance matrix predicted value at time +1 Passed to the cognitive controller constructed in step S3 to obtain the predicted entropy state of action selection ;
[0172] The predicted entropy reward in the planning phase is:
[0173]
[0174] in:
[0175] for k The predicted entropy state at the moment;
[0176] for k The entropy state of the moment;
[0177] The entropy state of the predicted value of the estimated error covariance matrix is used to form the predicted entropy reward corresponding to the action, which is used as the reward function for the feedforward in the planning stage;
[0178] Step S44: reward N groups of prediction entropy As k +1 The basis for selecting the emission waveform parameters at time;
[0179] Step S45: Store N groups of target states, waveform parameter selection actions, and obtained predicted entropy rewards.
[0180] In this embodiment, specifically, step S5 includes:
[0181] Step S51: Each node obtains the entropy reward value from its own perspective using step S2;
[0182] Step S52: impose a severe penalty on the action selection when the total bandwidth exceeds the limit of 30M bandwidth;
[0183] Each node competes for bandwidth for its own observation accuracy, and imposes severe penalties on waveform parameter selection actions that exceed the total bandwidth limit;
[0184]
[0185] in:
[0186] C is the penalty constant and ;
[0187] is the sum of the entropy rewards of each node;
[0188] The entropy reward corresponding to a waveform parameter selection action;
[0189] For example, when returning rewards, the bandwidth of the selected waveform parameters is judged. If the total bandwidth is greater than the limit of 30MHz, a penalty of 100 entropy reward value corresponding to the waveform selection action is given, that is, C=100;
[0190] Step S53: Store the target state before and after the multi-node perspective, the selected waveform parameter action, and the predicted / real entropy reward obtained by step S23 / step S26;
[0191] In this embodiment, the DQN network that performs value approximation on parameter action selection is trained using the following steps:
[0192] (a) Randomly initialize or load a pre-trained model to initialize the DQN network parameters;
[0193] Initialize the waveform library parameters, experience replay memory library, number of iterations, and use random weights to adjust the weights of the value function network Initialize and initialize the weights of the target network ,make ;
[0194] During the training process, the pre-loaded network parameters are fine-tuned based on the network parameters. The learning rate is set to 0.001, the batch size is set to 32, and the Adam optimizer is used. 200 generations of training are performed. The network converges and the network parameters of the network are obtained.
[0195] (b) using the target state before and after the multi-node perspectives recorded in step S53, the selected waveform parameter action, and the obtained real / predicted entropy reward as input for DQN network learning and training;
[0196] (c) Iterate steps S3-S5 to train the network.
[0197] The root mean square error (EA-RMSE) between the estimated target position and the true target position is used as an indicator of tracking accuracy:
[0198]
[0199] in:
[0200] is the real location;
[0201] To estimate the location.
[0202] In the embodiment of the present invention, multi-node bandwidth allocation uses the overall average bandwidth of the nodes as the quality of the bandwidth allocation result, and the time required to obtain the selected waveform parameters of the 10s tracking process is used as the training time of the method.
[0203] The following is a further explanation of the technical effects of the present invention in conjunction with simulation experiments:
[0204] 1. Simulation conditions
[0205] The overall simulation process is as follows Figure 2 As shown;
[0206] The simulation platform is: Apple Sillicon M1 CPU with a main frequency of 3.20GHz, 8.0GB of memory, macOS 12.3 operating system, Gym environment, PyTorch deep learning platform, and Python 3.8 language implementation.
[0207] Simulation content
[0208] Simulation 1: Using the present invention and the traditional waveform selection method, Figure 3 The target is tracked in a single-node vertical nonlinear motion simulation scenario. The distance tracking results are as follows: Figure 6 As shown, compared with the traditional method, the tracking result of the present invention is more stable.
[0209] Simulation 2: Using the present invention and the traditional waveform selection method, Figure 4 The target is tracked in a single-node horizontal nonlinear motion simulation scenario. The distance tracking results are as follows: Figure 7 As shown, the adaptive selection of waveform parameters can obtain the optimal / suboptimal waveform parameter selection.
[0210] Simulation 3: Using the present invention and the traditional waveform selection method, the nonlinear motion of a single node target is simulated. Figure 5 Tracking of horizontal nonlinear moving targets and adaptive bandwidth allocation in multi-node simulation scenarios.
[0211] 3. Simulation Results Analysis
[0212] As can be seen from Tables 1 and 2, in the simulation results, the traditional waveform selection method requires a lot of time in processing, and the DQN-based deep reinforcement learning method can quickly adapt to task requirements; as can be seen from Table 3, compared with the existing technology, the simulation results can make better use of bandwidth resources under limited bandwidth constraints.
[0213] Table 1 Single node vertical tracking results
[0214]
[0215] Table 2 Single node horizontal tracking results
[0216]
[0217] Table 3 Multi-node horizontal tracking bandwidth allocation results
[0218]
[0219] The simulation results were compared with the actual target position. Compared with the existing method, the simulation results of the present invention can achieve the desired tracking effect with acceptable time efficiency. The selected waveform parameters can meet the task requirements while taking into account both accuracy and efficiency. As shown in Table 3, the node-adaptive bandwidth allocation of the present method can more effectively utilize bandwidth resources than the traditional method within the 30MHz bandwidth limit, and can achieve higher tracking accuracy.
[0220] This paper uses an action selection strategy based on expected value, minimizing the Cramer-Rao lower bound of the estimated target state as a component of the target entropy state to form an entropy-based action reward. Through the feedforward link, this paper enables the network to achieve better results under different task requirements.
[0221] The above-described embodiments merely represent specific implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of protection of the present application. It should be noted that a person skilled in the art would be able to make numerous variations and improvements without departing from the technical concept of the present application, and all such variations and improvements fall within the scope of protection of the present application.
[0222] This background section is provided to generally present the context of the invention, and the work of the presently named inventors, the work to the extent described in this background section, and aspects of the description in this section that did not constitute prior art at the time of filing are neither explicitly nor implicitly admitted to be prior art to the present invention.
Claims
1. A waveform adaptive selection method based on deep reinforcement learning, characterized in that: include: Step S1: Model the simulation scenario by constructing a discrete system state space model. Based on the simulation scenario, initialize the discrete system state space model and randomly initialize the weights of the DQN network. Step S2: Construct the target entropy state by minimizing the Cramer-Rao lower bound of the target parameter estimate, and construct the real entropy reward of the waveform parameter selection action by the real entropy state before and after the moment; that is, construct the cognitive sensor and cognitive controller; Step S3: using a cubature Kalman filter to predict the next moment target state of the action with different waveform parameters selected at the current moment; Step S4: constructing a feed-forward loop path based on the predicted entropy state to form a predicted entropy reward; Step S5: Design an action reward function for the multi-node scenario, use the DQN network to approximate the expected value of rewards for different actions under different states, and obtain the final waveform parameter to select the action; The step S4 comprises: Step S41: Obtain N groups of waveform parameter action selections with high reward expectations from the DQN value approximation network; Step S42: Take the N sets of waveform parameter actions as the input of the volumetric Kalman filter in step S22 to obtain the N sets of estimated error covariance matrix prediction values of the target state at the next moment ; Step S43: N groups k Estimated error covariance matrix predicted value at time +1 Passed to the cognitive controller in step S2 to obtain the predicted entropy state of action selection , the predicted entropy reward in the planning stage is: in: for k The predicted entropy state at the moment; for k The entropy state of the moment; Step S44: reward N groups of prediction entropy As k +1 The basis for selecting the emission waveform parameters at time; Step S45: Store N groups of target states, waveform parameter selection actions, and obtained predicted entropy rewards.
2. The waveform adaptive selection method based on deep reinforcement learning according to claim 1, characterized in that: The step S1 comprises: The state space model of the discrete system is defined as follows: in: f (.) is the vector conversion function; h (.) is another vector transformation function that maps the target from the state space to the observation space; Indicates that the system is k The state of the moment; Indicates k The observed value at the time; represents an additional processing noise as a driving force for updating the state; is the additional observation noise; express k -1 waveform parameters; express k The true entropy state of the target parameter estimate at each moment; express k The estimated error covariance matrix of the target state at time +1; The system equations are defined as follows: in; is the relative distance between the radar and the target The speed of the target; is the target acceleration; The system covariance noise is defined as follows: in: ; The Cramer-Rao lower bound on the observation noise covariance is defined as follows: in: ; is the signal-to-noise ratio, ; ; 。 3. The waveform adaptive selection method based on deep reinforcement learning according to claim 2, characterized in that: The step S2 includes: Step S21: The cognitive sensor obtains the current k The observation vector at time instant; Step S22: The current K The Cramer-Rao lower bound of the moment-by-moment observation noise covariance is used as the input to the cubature Kalman filter to obtain the estimated error covariance matrix prediction value of the target state at the next moment and the estimated error covariance matrix ; Step S23: Use the estimated error covariance matrix to predict the value and the estimated error covariance matrix The posterior probability of the target state is estimated to obtain the predicted entropy state of the target parameter estimate and the current The true entropy state of the target parameter estimate at the moment ; Step S24: True entropy state Feedback to the cognitive controller to form feedback information; Step S25: The cognitive controller receives the real entropy state information fed back by the cognitive sensor; Step S26: Obtaining a real entropy reward from the entropy state information at the previous and next moments; Step S27: storing a set of target states, waveform parameter selection actions, and obtained real entropy rewards; Step S28: Select transmission waveform parameters.
4. The waveform adaptive selection method based on deep reinforcement learning according to claim 3, characterized in that: The step S23 includes: The entropy state is a measure of the posterior probability estimate distribution The uncertainty measure of the target state is measured by Shannon entropy, which is in the form of: By treating the target action as the observation noise covariance The estimated error covariance matrix is obtained by inputting , the estimated error covariance matrix Perform Shannon entropy measurement to obtain the current K The target true entropy state at the moment .
5. The waveform adaptive selection method based on deep reinforcement learning according to claim 4, characterized in that: Step S26, include: Among them, the real entropy reward for a single node to execute an action is: in: for k The actual entropy state at time -1; for k The true entropy state at a given moment.
6. The waveform adaptive selection method based on deep reinforcement learning according to claim 5, characterized in that: The step S5 comprises: Step S51: Each node obtains the entropy reward value from its own perspective using step S2; Step S52: impose a severe penalty on the action selection when the total bandwidth exceeds the limit of 30M bandwidth; Step S53: Store the target state before and after, the selected waveform parameter action, the obtained predicted entropy reward and the actual entropy reward from the perspective of multiple nodes.
7. The waveform adaptive selection method based on deep reinforcement learning according to claim 6, characterized in that: The step S52 includes: Each node competes for bandwidth for its own observation accuracy, and imposes severe penalties on waveform parameter selection actions that exceed the total bandwidth limit; in: C is the penalty constant and ; is the sum of the entropy rewards of each node; The entropy reward corresponding to a waveform parameter selection action.
8. The waveform adaptive selection method based on deep reinforcement learning according to claim 7, characterized in that: Also includes: The DQN network that performs value approximation on parameter action selection is trained using the following steps: (a) Randomly initialize or load a pre-trained model to initialize the DQN network parameters; Initialize the waveform library parameters, experience replay memory library, number of iterations, and use random weights to adjust the weights of the value function network Initialize and initialize the weights of the target network ,make ; During the training process, the pre-loaded network parameters were fine-tuned based on the network parameters. The learning rate was set to 0.001, the batch size was set to 32, and the Adam optimizer was used. After 200 generations of training, the network converged and the network parameters were obtained. (b) using the target state before and after the multi-node perspectives recorded in step S53, the selected waveform parameter action, and the obtained real / predicted entropy reward as input for DQN network learning and training; (c) Iterate steps S3-S5 to train the network.
Citation Information
Patent Citations
Broadband wireless communication autonomous frequency selection method and system based on deep reinforcement learning
CN111726217A
Waveform agility radar radiation source identification method based on deep reinforcement learning
CN113158886A