An online vehicle-mounted radar anti-jamming waveform design method based on double deep Q network
By adopting an online anti-interference waveform design method for vehicle-mounted radar based on dual deep Q-networks, the problem of decreased perception performance of vehicle-mounted radar in complex electromagnetic environments is solved, online updates and environmental adaptation are realized, and the robustness and perception performance of the radar are improved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- JILIN UNIVERSITY
- Filing Date
- 2022-10-18
- Publication Date
- 2026-07-21
AI Technical Summary
Existing vehicle radar anti-jamming technologies rely on specific interference environment parameters, have poor robustness when the environment changes abruptly, lack online update capabilities, and thus lead to a decline or failure of perception performance.
An online anti-jamming waveform design method for vehicle-mounted radar based on dual-deep Q-network is adopted. By estimating the signal-to-interference ratio, a Markov decision process is constructed and a dual-deep Q-network model is trained to achieve online adaptive adjustment under electromagnetic interference environment and output the optimal anti-jamming waveform.
It improves the perception performance of vehicle-mounted radar in complex electromagnetic environments, enables online updates, reduces the demand for communication hardware, and enhances environmental adaptability.
Smart Images

Figure CN115718422B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a method for designing anti-interference waveforms for vehicle-mounted radar. This method can be used in fields such as vehicle tracking and positioning, autonomous navigation, and advanced driver assistance systems (ADAS), and belongs to the field of radar anti-interference. Background Technology
[0002] With the development of information technology and the improvement of people's living standards, automobiles have gradually become necessities in daily life, and at the same time, vehicle safety performance has become a primary concern. However, the surge in the number of cars and vehicles equipped with millimeter-wave radar in urban traffic systems has led to severe radar interference problems, causing a decline or even failure in parameter perception performance. As a key perception device for autonomous driving, the accurate perception of position, speed, and azimuth angle by vehicle-mounted radar can not only effectively improve the safety performance of autonomous driving, but also provide effective support for subsequent autonomous driving fusion perception and planning decision-making processes. Therefore, researching anti-interference technologies under complex interference conditions such as the same modulation mode, similar waveform patterns, and overlapping waveform bandwidth is of great significance.
[0003] In recent years, numerous interference mitigation technologies have emerged. Based on whether the interference mitigation occurs at the receiving or transmitting end, they can be broadly categorized into two types: interference cancellation and interference avoidance. Interference cancellation technologies are generally used at the receiving end, performing relevant signal processing in the time, frequency, or time-frequency domains to reduce or eliminate interference. For example, parameter estimation techniques are used to reconstruct the interference signal at the receiving end, and then the received signal is used to subtract the anti-interference signal to achieve anti-interference. Interference avoidance technologies are generally processed at the transmitting end, avoiding interference through coordinated design in the time, frequency, time-frequency, and spatial domains. For example, frequency bands are divided equally according to parameter sensing resolution requirements, and different vehicles use different frequency band radar signals to fundamentally avoid interference. Most of these methods rely on specific interference environment parameters, achieving anti-interference only in limited situations. They are highly dependent on the environment and hardware, and these methods may fail when the environment changes abruptly. They lack online update capabilities, leading to instability and poor robustness of existing technologies. Summary of the Invention
[0004] The purpose of this invention is to address the problems that most existing interference processing technologies rely on specific interference environment parameters, can only achieve anti-interference for vehicle radar in limited situations, have high dependence on environment and hardware, fail when the environment changes abruptly, lack online update performance, and result in instability and poor robustness of existing interference processing technologies. Therefore, this invention proposes an online anti-interference waveform design method for vehicle radar based on dual deep Q networks.
[0005] The specific process of the online vehicle radar anti-jamming waveform design method based on dual-depth Q-network is as follows:
[0006] Step 1: Estimate the signal-to-interference ratio (SIR) of the vehicle-mounted radar received signal;
[0007] Step 2: Based on the signal-to-interference ratio (SINR) from Step 1, construct the Markov decision process for vehicle-mounted radar;
[0008] Step 3: Based on the Markov decision process of the vehicle radar in Step 2, construct and train a dual-deep Q-network model.
[0009] Step 4: Online adaptive adjustment of parameters of the dual-depth Q-network model under abrupt electromagnetic interference environment to obtain the dual-depth Q-network model under abrupt electromagnetic interference environment.
[0010] Step 5: Obtain the state space of the vehicle radar, input the dual-depth Q-network model under electromagnetic interference under abrupt change conditions obtained in Step 4, and output the optimal action corresponding to the optimal action value function value, which is the anti-interference waveform of the vehicle radar.
[0011] The beneficial effects of this invention are as follows:
[0012] This invention addresses the interference problem of vehicle-mounted radar in complex electromagnetic environments with identical modulation schemes, similar waveform patterns, and overlapping waveform bandwidths. It proposes an online anti-interference waveform design method for vehicle-mounted radar based on Double Deep Q Network (DDQN). This method solves the problem of severe inaccuracy or failure of vehicle-mounted radar parameters such as position and velocity caused by complex electromagnetic interference. At the same time, it realizes online updating of the model by using the model's action space constraint and exploration strategy adjustment method, thereby improving the model's environmental adaptability.
[0013] (1) The online vehicle radar anti-interference waveform design method based on dual deep Q network described in this invention uses deep reinforcement learning algorithm to design radar parameter waveforms, which overcomes the dependence of traditional methods on the environment, solves the problem that the perception performance of vehicle radar is seriously reduced or even fails due to complex electromagnetic interference environment in urban roads, realizes the online update capability under sudden electromagnetic environment, and effectively improves the environmental perception performance under complex interference conditions.
[0014] (2) The online vehicle radar anti-interference waveform design method based on dual-depth Q network described in this invention can implement effective anti-interference waveform design without the need for external environmental communication to provide additional information, which effectively reduces the demand for communication hardware equipment. Attached Figure Description
[0015] Figure 1 This is a principle block diagram of an online vehicle radar anti-interference waveform design method based on dual-depth Q-network as described in this invention.
[0016] Figure 2 The figure shows the target detection simulation results of the online vehicle radar anti-interference waveform design method based on dual-depth Q network described in this invention under an interference-free environment.
[0017] Figure 3 The figure shows the target detection simulation results of the online vehicle radar anti-interference waveform design method based on dual-depth Q network described in this invention under complex electromagnetic interference environment.
[0018] Figure 4 This is a simulation diagram comparing the anti-interference success rate of the online vehicle radar anti-interference waveform design method based on dual-depth Q network described in this invention with that of traditional methods under stable environmental conditions with 50% overlap in signal bandwidth.
[0019] Figure 5 This is a simulation diagram comparing the anti-interference success rate of the online vehicle radar anti-interference waveform design method based on dual-depth Q network described in this invention with that of traditional methods under the condition of 50% signal bandwidth overlap and sudden environmental changes.
[0020] Figure 6 This is a simulation diagram comparing the anti-interference success rate of the online vehicle radar anti-interference waveform design method based on dual-depth Q network described in this invention with that of traditional methods under the condition of 100% signal bandwidth overlap and stable environment.
[0021] Figure 7 This is a simulation diagram comparing the anti-interference success rate of the online vehicle radar anti-interference waveform design method based on dual-depth Q network described in this invention with that of traditional methods under the condition of 100% signal bandwidth overlap and sudden environmental changes. Detailed Implementation
[0022] Specific Implementation Method 1: The specific process of this implementation method for online vehicle-mounted radar anti-interference waveform design based on dual-depth Q-network is as follows:
[0023] Step 1: Estimate the signal-to-interference ratio (SIR) of the vehicle-mounted radar received signal;
[0024] Step 2: Based on the signal-to-interference ratio (SINR) from Step 1, construct the Markov decision process for vehicle-mounted radar;
[0025] Step 3: Based on the Markov decision process of the vehicle radar in Step 2, construct and train a dual-deep Q-network model.
[0026] Step 4: The parameters of the dual-deep Q-network model (intelligent anti-interference network based on deep reinforcement learning) are adaptively adjusted online under the sudden electromagnetic interference environment to obtain the dual-deep Q-network model under the sudden electromagnetic interference environment.
[0027] Step 5: Obtain the state space of the vehicle radar, input the dual-depth Q-network model under electromagnetic interference under abrupt change conditions obtained in Step 4, and output the optimal action corresponding to the optimal action value function value, which is the anti-interference waveform of the vehicle radar.
[0028] Specific Implementation Method Two: This implementation method differs from Specific Implementation Method One in that step one involves estimating the signal-to-interference ratio (SIR) of the vehicle-mounted radar received signal; the specific process is as follows:
[0029] Step 11: Provide the power of the target echo signal and interference signal received by the vehicle-mounted radar receiver; the specific process is as follows:
[0030] Let P be the power of the target echo signal received by the vehicle-mounted radar receiver. r ,have
[0031]
[0032] Let the power of the interference signal received by the vehicle-mounted radar receiver be P. i ,have
[0033]
[0034] Where P t The vector represents the transmit power of the vehicle-mounted radar, G represents the antenna gain, σ represents the radar cross-section of the target, R represents the distance from the target or interference source to the vehicle-mounted radar receiver, and A represents the radar cross-section of the target. e Indicates the effective aperture of the vehicle-mounted radar antenna. λ is the carrier wavelength of the signal transmitted by the vehicle-mounted radar;
[0035] Steps one and two: Mix the target echo signal and interference signal received by the vehicle-mounted radar receiver with the vehicle-mounted radar transmitted signal (original transmitted signal); the specific process is as follows:
[0036] The target echo signal received by the vehicle-mounted radar receiver is mixed with the vehicle-mounted radar transmitted signal (original transmitted signal) to obtain the frequency domain expression. for
[0037]
[0038] Where δ(·) is the impulse function, f 0,v τ is the carrier frequency of the signal transmitted by the vehicle-mounted radar, τ is the time delay of the echo signal, and k is the carrier frequency of the signal transmitted by the vehicle-mounted radar. v The slope of the chirp signal emitted by the vehicle-mounted radar; j is the imaginary unit, j 2 =-1; f is the radar frequency;
[0039] The interference signal received by the vehicle-mounted radar receiver is mixed with the vehicle-mounted radar transmitted signal (original transmitted signal) to obtain a frequency domain expression. for
[0040]
[0041] Where t arr t is the time it takes for the interference signal to enter the radar receiver bandwidth. end K is the time it takes for the interference signal to leave the radar receiver bandwidth. if It is the absolute value of the difference between the slopes of the chirp signal transmitted by the radar transmitter and the chirp signal transmitted by the jamming source. The initial phase is rect(·), which is a rectangular window function with an amplitude of 1.
[0042] Step 13: Estimate the signal-to-interference ratio (SIR) of the vehicle-mounted radar received signal; the specific process is as follows:
[0043] The frequency domain expression of the mixing of the echo signal obtained in steps one and two with the vehicle-mounted radar transmitted signal (original transmitted signal) And the frequency domain expression of the mixing of the interference signal and the vehicle-mounted radar transmitted signal (original transmitted signal). Estimate the signal-to-interference ratio (SIR) of the vehicle-mounted radar signal received at the target detection point; the expression is:
[0044]
[0045] in, P v-v P represents the received power of the target echo signal at the detection point. i-v The interference power introduced by the interference signal at the detection point; Δf is the frequency domain sampling interval; A rectangular function with unit amplitude;
[0046] Considering |exp[j(2πf 0,v τ-πk v τ 2 )| and The value is 1. Therefore, the expression for estimating the signal-to-interference ratio can be approximated as follows:
[0047]
[0048] because Combined with the echo signal power P obtained in step one r and interference signal power P i and δ 2 (f) The expression for estimating the signal-to-interference ratio is further simplified to:
[0049]
[0050] The other steps and parameters are the same as in Specific Implementation Method 1.
[0051] Specific Implementation Method Three: This implementation method differs from Specific Implementation Method One or Two in that, in step two, a Markov decision process for the vehicle-mounted radar is constructed based on the signal-to-interference ratio (SINR) of step one; the specific process is as follows:
[0052] Step 2.1: Construct a Markov decision process environment model; the specific process is as follows:
[0053] To address the complex electromagnetic interference environment of vehicle-mounted radar, deep reinforcement learning techniques are employed for waveform parameter selection. This avoids prior requirements regarding the complex electromagnetic environment. Deep reinforcement learning can use Markov decision processes to model the state space s of the vehicle-mounted radar at time t. t Defined as [o t ,o t-1 ,...,o t-L ], where L is the number of observation vectors constituting the state space, o t =[r t ,sir t ,p t ,a t [] represents the observation vector at time t, r t Let sir be the reward function at time t. t Let p be the signal-to-interference ratio (SIR) of the vehicle-mounted radar received signal at time t. t Let Ot be the row vector representing the position of the interference source to be detected in the environment relative to the vehicle radar at time t (assuming there are W interference sources, each with position (x, y) coordinates, I combine the W coordinates into a large row vector so that Ot can be one-dimensional for easier input to the subsequent neural network). Where W represents the number of interference sources to be detected in the environment. Let be the row vector of the position of the W-th interference source to be detected in the environment at time t relative to the vehicle-mounted radar. Let a be the Cartesian coordinate of the i-th interference source to be detected (radars installed on other vehicles) in the environment at time t, with the vehicle-mounted radar as the reference frame, i = 1, 2, ..., W, a t For the action space, representing s t The radar waveform parameters used by the vehicle-mounted radar under the specified conditions;
[0054] Step 22: Based on the information-to-interference ratio from Step 1, construct the reward function and action space of the Markov decision process; the specific process is as follows:
[0055] The reward function of a Markov decision process is
[0056] Where sir0 is the set signal-to-interference ratio (SIR) threshold, sir t Let be the signal-to-interference ratio (SIR) of the signal received by the vehicle-mounted radar at time t;
[0057] The action space set of a Markov decision process can be represented as A = {a t0 ,a t1 ,...,a tM};
[0058] Where M is the number of action spaces. For the first in the action space set Actions (s) t The radar waveform parameters used by the vehicle-mounted radar in this state are also the first... (the slope of the chirp signal emitted by the vehicle-mounted radar in the action space)
[0059] Other steps and parameters are the same as in specific implementation method one or two.
[0060] Specific Implementation Method Four: This implementation method differs from Specific Implementation Methods One to Three in that, in step three, a dual-deep Q-network model (an intelligent anti-interference network model based on deep reinforcement learning) is constructed and trained based on the Markov decision process of the vehicle-mounted radar in step two; the specific process is as follows:
[0061] Step 3: 1. Construct and initialize the Dual-Depth Q-Network (DDQN) network structure;
[0062] Step 3.2: Utilizing the experience gained from evaluating network-controlled vehicle-mounted radar data collection [r] t ,sir t ,p t ,a t ];
[0063] Step 3: Calculate the time difference function using the target network;
[0064] Steps 3 and 4: Construct the time difference error function and update the neural network parameters;
[0065] Step 35: Repeat steps 32 to 34 until convergence, to obtain the training data of the dual-depth Q-network model.
[0066] Steps 32, 33, and 34 are the training process. Step 32 continuously collects the simulated scene data and puts it into the experience pool. Then, steps 33 and 34 continuously construct the loss function and update it.
[0067] The other steps and parameters are the same as those in one of the specific implementation methods one to three.
[0068] Specific Implementation Method Five: This implementation method differs from Specific Implementation Methods One to Four in that, in step three, a dual-depth Q-network (DDQN network) structure is constructed and initialized; the specific process is as follows:
[0069] Step 3.11: Construct a Dual-Depth Q-Network Model (an intelligent anti-interference network model based on deep reinforcement learning) (DDQN network); the specific process is as follows:
[0070] DDQN networks use Q(s) t ,a t ;w) and Q - (s t ,a t ;w - Two neural networks approximate the evaluation network and the target network respectively, which can effectively solve the overestimation problem of traditional Deep Q Network (DQN);
[0071] Intelligent anti-interference networks based on deep reinforcement learning (DDQN networks) include an evaluation network Q(s) t ,a t ;w) and target network Q - (s t ,a t ;w - );
[0072] Evaluate network Q(s) t ,a t The w) consists of one convolutional layer and five fully connected layers.
[0073] The state space s of the vehicle radar at time t t Input to evaluate network Q(s) t ,a t The optimal action value function is output after passing through a convolutional layer (w), a convolutional layer to extract features, and then a five-layer fully connected layer.
[0074] The state space s of the vehicle radar at time t t Input to evaluate network Q(s) t ,a t ;w), determine the input state space s t This determines the corresponding a. t a t There could be 3 corresponding ones, a t1 a t2 a t3 Evaluate the network output Q(s) t ,a t Find 3 Q(s) t ,a t Take the maximum value of Q(s) t ,a t The maximum value is the optimal action value function value, corresponding to a. t The optimal action;
[0075] Target network Q - (s t ,a t ;w - Adopting and evaluating network Q(s) t ,a t The network structure is the same as that of w), but the parameter updates of the target network lag behind the evaluation function.
[0076] The inputs in the evaluation network and the target network t Let be the state space of the vehicle-mounted radar at time t;
[0077] The output a in the evaluation network and the target network t For s t The radar waveform parameters used by the vehicle-mounted radar under the specified conditions;
[0078] w represents the parameters to be learned by the network. - The target network parameters do not need to be learned; the evaluation network's learning parameters w are periodically assigned to the target network parameters w. - That's all;
[0079] The DDQN (Double Deep Q Network) is a double deep Q network;
[0080] Step 3.12: Initialize the dual-deep Q-network model (an intelligent anti-interference network based on deep reinforcement learning); the specific process is as follows:
[0081] The network parameters of the initial dual-depth Q-network model (initializing the evaluation network and the target network together to the same value at the beginning) include the empirical replay library size D, the learning rate lr, ε0 in the decay ε-greedy policy, the reward decay factor γ, the batch size for one update, and the target network Q. - (s t ,a t ;w - The parameter lag update step number threshold C (the evaluation network is updated with each input data, while the target network is not. The target network is updated only once after C training cycles, where C is the number of update steps).
[0082] ε0 is the initial exploration rate; C is the threshold for the number of update steps.
[0083] The other steps and parameters are the same as those in one of the specific implementation methods one to four.
[0084] Specific Implementation Method Six: This implementation method differs from Specific Implementation Methods One to Five in that step three-two employs an evaluation network-controlled vehicle radar to collect experience [r] t ,sirt ,p t ,a t The specific process is as follows:
[0085] In the initial state space s t The waveform parameter a is selected based on the attenuation ε-greedy strategy. t Then the environment provides the reward function r. t (the state space s of the vehicle radar at time t) t Input to evaluate network Q(s) t ,a t ;w), determine the input state space s t This determines the corresponding a. t a t There could be 3 corresponding ones, a t1 a t2 a t3 Evaluate the network output Q(s) t ,a t Find 3 Q(s) t ,a t Take the maximum value of Q(s) t ,a t The maximum value is the optimal action value function value, corresponding to a. t To determine the optimal action, select action a. t Afterwards, the simulation environment collects and transmits the corresponding waveforms, simulates the received echo signal-to-interference ratio, calculates the feedback (which is then fed into the neural network), and the environment jumps to the next state space s. t+1 Combined with the signal-to-interference ratio estimate at time t, sir t And the position row vector p of the interference source to be detected relative to the vehicle radar in the environment at time t. t , integrated into experience [r t ,sir t ,p t ,a t Store the data in the experience pool until the experience pool is full, then randomly sample data of size batch_size for learning.
[0086] The decay ε-greedy strategy can be represented as p(s) t ,a t ), in ε0 is the initial exploration rate, N exp ε represents the total number of exploration steps, k represents the current exploration step number, and ε represents the total number of exploration steps. k This represents the exploration rate corresponding to the k-th exploration.
[0087] The attenuation ε-greedy strategy can effectively balance the exploratory requirements of the early training of the network and the efficiency requirements of the later training.
[0088] The other steps and parameters are the same as those in one of the specific implementation methods one to five.
[0089] Specific Implementation Method Seven: This implementation method differs from Specific Implementation Methods One through Six in that step three-three uses a target network to calculate the time difference function; the specific process is as follows:
[0090] Using the target network Q - (s t ,a t ;w - Calculate the time difference function y t ,
[0091] Where γ represents the discount attenuation factor and · represents multiplication.
[0092] Time difference function y t and evaluation network Q(s) t ,a t Both w and y are essentially approximations of the optimal action value function, but y t Because it contains the true return function value r t Compared to Q(s) t ,a t ;w) is more accurate, while Q(s) is more accurate. t ,a t w) It needs to learn and update the parameter w to continuously approach y. t .
[0093] The other steps and parameters are the same as those in one of the specific implementation methods one to six.
[0094] Specific Implementation Method Eight: This implementation method differs from Specific Implementation Methods One to Seven in that steps three and four involve constructing a time difference error function and updating the neural network parameters; the specific process is as follows:
[0095] Construct the time difference error function e t =Q(s) t ,a t ;w)-y t Using stochastic gradient descent to... Parameters are updated, and the network Q(s) is evaluated when the update step reaches an integer multiple of C (updating for the first time when C is reached, updating for the second time when 2C is reached, etc.). t ,a t The parameter w in (w) is assigned to the target network Q. - (s t ,a t ;w - );
[0096] Where loss(w) represents the loss function of the dual-depth Q-network model; C is the threshold number of update steps.
[0097] The other steps and parameters are the same as those in one of the specific implementation methods one to eight.
[0098] Specific Implementation Method Nine: This implementation method differs from Specific Implementation Methods One through Seven in that, in step four, the parameters of the dual-deep Q-network model (an intelligent anti-interference network based on deep reinforcement learning) under abrupt electromagnetic interference conditions are adaptively adjusted online to obtain the dual-deep Q-network model under abrupt electromagnetic interference conditions; the specific process is as follows:
[0099] Step 41: Action space constraint for the dual-depth Q-network model; the specific process is as follows:
[0100] When the model has finished training, if it faces a sudden change in the electromagnetic environment, the model's generalization performance will drop significantly. The network model can be appropriately adjusted to limit the action space and achieve rapid reconvergence. The specific method is as follows: A represents the action space set before the electromagnetic environment change. The set of n optimal action spaces before the mutation represents the n sets of waveform parameters that were most frequently used in the model during the time interval T before the mutation; after the electromagnetic environment mutates, it means... Actions in a concentrated action sequence can be ignored; only actions need to be considered. The set of actions in the model accelerates model convergence;
[0101] Subtract the optimal set of n action spaces before the electromagnetic environment mutation from the action space set A before the mutation. As a new set of action spaces;
[0102] Step 4.2: Exploration strategy correction for the dual-depth Q-network model; the specific process is as follows:
[0103] To improve the model's early exploration and ensure optimal action selection in later training stages, a decaying ε-greedy strategy is used in the new action space set based on a dual-deep Q-network model for action selection. k The current exploration step count is reset to zero, and k is restarted from zero to ensure the exploratory nature of the decaying ε-greedy strategy. Substitute In, p(s) t ,a t Take the values of two cases. a t(The actions are in the new action space set); that is, select the optimal waveform parameters. When the model training is stable, the ε-greedy strategy tends to be a greedy strategy, and the probability of strategy exploration will continue to decrease. If the electromagnetic interference environment changes abruptly, the generalization ability of the model will drop significantly. At this time, it is necessary to reset the decaying ε-greedy strategy so that the model can explore within the limited space with a high probability when the environment changes abruptly, so as to achieve rapid reconvergence of the model.
[0104] Repeat steps 32, 33, and 34 to obtain the dual-depth Q-network model under electromagnetic interference conditions with abrupt change.
[0105] The other steps and parameters are the same as those in one of the specific implementation methods one to eight.
[0106] Example 1:
[0107] This invention addresses the problem of vehicle-mounted radar interference in complex electromagnetic environments with identical modulation methods, similar waveform patterns, and overlapping waveform bandwidths. It constructs an urban highway section where four interfering vehicles (i.e., four interference sources) exist within the radar's effective range. These interfering vehicles undergo various dynamic motions relative to the radar, such as constant speed, deceleration, and acceleration. Through deep reinforcement learning, the invention enables intelligent selection of radar waveform parameters and conducts online controllable anti-interference design.
[0108] See Figure 1 The simulation experiment steps of the vehicle radar anti-interference waveform design method based on dual deep Q-network described in this invention, using MATLAB and Python simulation software, are as follows:
[0109] 1. Estimate the signal-to-interference ratio (SIR) of the vehicle-mounted radar received signal.
[0110] 1) Give the power of the target echo signal and interference signal received by the vehicle-mounted radar receiver.
[0111] Let P be the power of the target echo signal received by the vehicle-mounted radar receiver. r ,have
[0112]
[0113] Let the power of the interference signal received by the vehicle-mounted radar receiver be P. i ,have
[0114]
[0115] Where P t G represents the transmit power of the vehicle-mounted radar, σ represents the antenna gain, σ represents the radar cross-section of the target, and R represents the distance from the target or interference source to the vehicle-mounted radar receiver. λ represents the effective aperture of the vehicle-mounted radar antenna, and λ is the carrier wavelength of the vehicle-mounted radar transmitted signal.
[0116] In the simulation experiment, P t The value is taken as 10dBm, the antenna gain G is 12dBi, and σ is taken as the typical value of the automotive radar cross-section of 100m². 2 The carrier wavelength λ is 3.947 mm, and the distance R between the target and the vehicle radar fluctuates between 50 m and 100 m.
[0117] 2) The echo signal and the interference signal are mixed with the original transmitted signal respectively.
[0118] The target echo signal received by the vehicle-mounted radar receiver is mixed with the original transmitted signal to obtain the frequency domain expression. for
[0119]
[0120] Where δ(·) is the impulse function, f 0,v τ is the carrier frequency of the signal transmitted by the vehicle-mounted radar, τ is the time delay of the echo signal, and k is the carrier frequency of the signal transmitted by the vehicle-mounted radar. v The slope of the chirp signal emitted by the vehicle-mounted radar;
[0121] The interference signal received by the vehicle-mounted radar receiver is mixed with the original transmitted signal to obtain a frequency domain expression. for
[0122]
[0123] Where t arr t is the effective time for the interference signal to enter the radar receiver bandwidth. end K is the effective time for the interference signal to leave the radar receiver bandwidth. if It is the absolute value of the difference in slope between the chirp signals emitted by the radar transmitter and the jamming source. The initial phase is rect(·), which is a rectangular window function with an amplitude of 1.
[0124] In the simulation experiment, the carrier frequency f of the vehicle-mounted radar transmitted signal 0,v Set to 76GHz, chirp slope k v Set to {5.3×10 12 Hz / s, 6.3×10 12 Hz / s,...,21.3×10 12 Any element in the set {Hz / s} with an interval of 1.0 × 10⁻⁶ 12 Hz / s, initial phase Set to a random distribution of 0-2π, with the received radar echo signal as the starting point, t arr Set to 0-20us, and t arr<t end Assuming the radar transmitter and the jamming source use different chirp slopes, then K if The value is {5.3 × 10}. 12 Hz / s, 6.3×10 12 Hz / s,...,21.3×10 12 The difference between any two distinct slopes in Hz / s;
[0125] 3) Obtain the signal-to-interference ratio (SIR) of the radar received signal.
[0126] Combining the frequency domain expression of the echo signal obtained in step 2) with the original transmitted signal And the frequency domain expression for the mixing of the interference signal and the original transmitted signal. The signal-to-interference ratio (SIR) of the radar received signal at the target detection point is estimated as follows:
[0127]
[0128] in, P v-v P represents the received power of the target echo signal at the detection point. i-v The interference power introduced by the interference signal at the detection point is considered as |exp[j(2πf 0,v τ-πk v τ 2 )| and The value is 1. Where Δf is the frequency domain sampling interval. Since it is a rectangular function with unit amplitude, the signal-to-interference ratio estimation expression can be approximated as:
[0129]
[0130] because Combined with the echo signal power P obtained in step 1), r and interference signal power P i and δ 2 (f) The expression for estimating the signal-to-interference ratio is further simplified to:
[0131]
[0132] In the simulation experiment, Δf was set to 135 kHz, and the slope fluctuation range of the chirp signal emitted by the vehicle-mounted radar was 5.3 × 10⁻⁶. 12 Hz / s~21.3×10 13 Hz / s, with an interval of 1.0×10 12 Hz / s, K ifIt is the absolute value of the difference between the slopes of the chirp signals emitted by the radar transmitter and the jamming source, and its value varies with the simulation environment.
[0133] 2. Constructing a Markov Decision Process for Vehicle-Mounted Radar
[0134] 1) Constructing a Markov decision process environment model
[0135] To address the complex electromagnetic interference environment of vehicle-mounted radar, deep reinforcement learning techniques are employed for waveform parameter selection. This avoids prior requirements regarding the complex electromagnetic environment. Deep reinforcement learning can use Markov decision processes for modeling, with the state space s at time t... t Defined as [o t ,o t-1 ,...,o t-L ], where L is the number of observation vectors constituting the state space, o t =[r t ,sir t ,p t ,a t [r] is the observation vector at time t. t Let sir be the reward function at time t. t p is the signal-to-interference ratio estimate at time t. t Let be the row vector of the target to be detected in the environment at time t relative to the vehicle-mounted radar. Where W represents the number of interference sources to be detected in the environment. Let a be the Cartesian coordinates of the i-th target in the environment at time t, with the vehicle-mounted radar as the reference frame. t The waveform parameters used by the vehicle-mounted radar at time t;
[0136] In the simulation experiment, W=4, L=7. Let be the Cartesian coordinates of the i-th target in the environment at time t, initialized to [1m, 50m, -2m, 70m, 0m, 80m, -0.5m, 90m];
[0137] 2) Construct the reward function and action space of the Markov decision process.
[0138] The reward function of a Markov decision process is sir0 is the set signal-to-interference ratio (SIR) threshold, and the action space set can be represented as A = {k}. v0 ,k v1 ,...,k vM}, where M is the number of action spaces, k vi , i = 1, ..., M is the slope of the chirp signal emitted by the i-th action space vehicle radar;
[0139] In the simulation experiment, sir0 is 15dB, M = 17, and the action space set A = {5.3 × 10 12 Hz / s, 6.3×10 12 Hz / s,...,21.3×10 12 Hz / s}, with an interval of 1.0×10 12 Hz / s;
[0140] 3. Construct and train an intelligent anti-interference network model based on deep reinforcement learning.
[0141] 1) Construct and initialize the DDQN network structure.
[0142] (1) The DDQN network adopts Q(s) t ,a t ;w) and Q - (s t ,a t ;w - Two neural networks approximate the evaluation network and the target network respectively, effectively solving the overestimation problem of traditional Deep Q Networks (DQNs). The target network Q... - (s t ,a t ;w - Adopting and evaluating network Q(s) t ,a t The network structure is the same as that of the target network, but the parameter updates of the target network lag behind the evaluation function; the evaluation network Q(s) has the same structure. t ,a t w) Employing convolutional and fully connected networks to process state s t Convolution is performed to further extract features, followed by several fully connected network layers to output the optimal action value function; the input parameters s in the evaluation network and the target network are then evaluated. t Let a be the state of the vehicle-mounted radar at time t. t For s t The radar waveform parameters used by the vehicle-mounted radar under the given conditions, where w is the parameter to be learned by the evaluation network. - The target network parameter does not need to be learned; the parameter w is periodically assigned to w. - That's all;
[0143] In the simulation experiment, the Double-DQN network structure adopts a one-layer convolutional network and a five-layer fully connected network. The convolutional kernel size is 2×2, the convolutional stride is 1, zero padding is used, and the number of convolutional channels is 32. After flattening, it outputs the action value function of each state through five fully connected layers.
[0144] (2) Initialized network parameters include the experience replay library size D, learning rate lr, ε0 in the decay ε-greedy policy, reward decay factor γ, batch size for each update, and target network Q. - (s t ,a t ;w - The parameter lag update pace C;
[0145] In the simulation experiment, the experience replay library size D was 5000, the learning rate lr was set to 0.0001, the ε0 in the decay ε-greedy policy was set to 0.2, the reward decay factor γ was set to 0.9, the batch size for each update was set to 32, and the target network Q... - (s t ,a t ;w - The parameter lag update step C is set to 5; the neural network parameter w is initialized with a Gaussian random function with a mean of 0 and a variance of 0.01.
[0146] 2) Utilize evaluation network control vehicle radar data collection experience
[0147] First, evaluate the network Q(s) t ,a t w) Random initialization, in the initial state s t The waveform parameter a is selected based on the attenuation ε-greedy strategy. t Then the environment provides the reward function r. t The environment transitions to the next state. t+1 Combined with the signal-to-interference ratio estimate at time t, sir t and position vector p t , integrated into experience [r t ,sir t ,p t ,a t The experience pool is filled until it is full, after which random sampling of batch size ε is performed for learning. The decaying ε-greedy strategy can be expressed as follows: in ε0 is the initial exploration rate, N exp denoted by , where k represents the current number of exploration steps. The decay ε-greedy strategy can effectively balance the exploratory requirements of the early training of the network with the efficiency requirements of the later training.
[0148] In the simulation experiment, ε0 is taken as 0.2, N exp The value is 500000;
[0149] 3) Calculate the time difference function using the target network.
[0150] Using the target network Q - (s t ,a t ;w - Calculate the time difference function y t , Time difference function y t and evaluation network Q(s) t ,a t Both w and y are essentially approximations of the optimal action value function, but y t Because it contains the true return function value r t Compared to Q(s) t ,a t ;w) is more accurate, while Q(s) is more accurate. t ,a t w) It needs to learn and update the parameter w to continuously approach y. t ;
[0151] 4) Construct the time difference error function and update the neural network parameters
[0152] Construct the time difference error function e t =Q(s) t ,a t ;w)-y t Using stochastic gradient descent to... Perform parameter updates, and evaluate the network Q(s) when the update step reaches an integer multiple of C. t ,a t The parameter w in (w) is assigned to the target network Q. - (s t ,a t ;w - );
[0153] 4. Online adaptive adjustment of model parameters under sudden electromagnetic interference environment
[0154] 1) Model motion space constraints
[0155] When the model has finished training, if it faces a sudden change in the electromagnetic environment, the model's generalization performance will drop significantly. The network model can be appropriately adjusted to limit the action space and achieve rapid reconvergence. The specific method is as follows: A represents the action space set before the electromagnetic environment change. The set of n optimal action spaces before the mutation represents the n sets of waveform parameters that were most frequently used in the model during the time interval T before the mutation. After the electromagnetic environment mutates, it means... Actions in a concentrated action sequence can be ignored; only actions need to be considered. The set of actions in the model accelerates model convergence;
[0156] In the simulation experiment, n is set to 5;
[0157] 2) Model exploration strategy correction
[0158] To improve the model's early exploration and ensure optimal action selection in the later stages of training, the model uses a continuously decaying ε-greedy strategy for action selection, i.e., selecting the optimal waveform parameters. When the model training stabilizes, the ε-greedy strategy tends to become a greedy strategy, and the probability of strategy exploration will continuously decrease. If the electromagnetic interference environment changes abruptly, the model's generalization ability will drop significantly. At this time, it is necessary to reset the decaying ε-greedy strategy so that the model can explore within a limited space with a high probability when the environment changes abruptly, thereby achieving rapid reconvergence of the model.
[0159] The simulation results of the target detection method of the present invention in an interference-free environment are as follows: Figure 2 As shown, the signal-to-noise ratio in the simulation experiment is 30dB. Considering four target vehicles (interference vehicles), in an interference-free environment, it is assumed that the radar waves emitted by the interference sources cannot interfere with the vehicle-mounted radar and only act as the target. Taking the vehicle-mounted radar as the reference point, the vector composed of their initial radial distance and radial velocity relative to the reference point is represented as [50m, -30m / s, 80m, 30m / s, 110m, -20m / s, 140m, -40m / s]. The four interference vehicles move at constant speed, constant acceleration, constant deceleration, and constant speed respectively, with the acceleration vector represented as [0m / s²]. 2 5m / s 2 -5m / s 2 0m / s 2 The actual instantaneous distance and velocity after 0.5 seconds can be expressed as [35m, -30m / s, 64.375m, 32.5m / s, 100.625m, -17.5m / s, 120m, -40m / s]. Using the conventional 2-DFFT transform method for target detection of the jamming vehicle, the estimated instantaneous distance and velocity after 0.5 seconds using the method of this invention are [34.75m, -32.4m / s, 64.25m, 32.3m / s, 100.5m, -17.4m / s, 120.25m, -40.2m / s]. Simulation results of target detection using the method of this invention under complex electromagnetic interference environments are as follows: Figure 3 As shown, in the simulation experiment, the signal-to-noise ratio was 30dB, the signal-to-interference ratio was 15dB, and the estimated values were [34.5m, -32.3m / s, 64.0m, 32.6m / s, 101.0m, -17.6m / s, 120.5m, -40.4m / s]. The target detection simulation results show that the method of the present invention has good anti-interference performance.
[0160] Simulation results comparing the anti-interference success rates of the proposed method and the traditional method under stable environmental conditions with 50% signal bandwidth overlap are as follows: Figure 4 As shown, Method 1 is the method of this invention, Method 2 is the traditional waveform parameter selection method based on DDQN, and Method 3 is the traditional random waveform emission method. The success rate is defined as the ratio of the number of correct waveform selections in one training cycle to the total number of selections. From the simulation graphs, we can see that under steady-state conditions, the methods of this invention and Method 2 have similar performance indicators, both outperforming Method 3. Simulation results under sudden environmental changes are shown below. Figure 5 As shown, when the number of training segments equals 280, the interference environment suddenly changes. In the experiment, the waveform parameters of the interference sources are randomly changed, and the chirp slope of the four interference sources is changed from the original {5.3×10}. 12 Hz / s, 8.3×10 12 Hz / s, 14.3×10 12 Hz / s, 16.3×10 12 The Hz / s value randomly varies to {5.3 × 10 Hz / s}. 12 Hz / s, 6.3×10 12 Hz / s, 7.3×10 12 Hz / s, 8.3×10 12 At Hz / s, due to the limited generalization ability of the model, the success rate of waveform parameter selection in Method 1 and Method 2 decreases. However, the method of this invention uses action space limitation and exploration strategy reset method, which enables the performance of the model to rebound quickly and reach a level close to that of the stable environment. Method 2 does not take any measures, the model performance drops significantly, and it takes a long time to converge. Moreover, the performance of the model at convergence is not as good as that at the stable environment. Method 3 uses a random strategy, which is independent of whether there are sudden changes in the environment, and the performance does not change significantly.
[0161] Simulation results comparing the anti-interference success rates of the proposed method and the traditional method under stable environmental conditions, with 100% signal bandwidth overlap, are as follows: Figure 6 As shown. Simulation results under sudden environmental changes are as follows. Figure 7 As shown, due to the complete bandwidth overlap, the interference effect increases, causing some waveform parameters that are available under the 50% bandwidth overlap condition to become unusable. This results in a significant decrease in the performance of Method 1, Method 2, and Method 3, but their trends are consistent with those under the 50% bandwidth overlap condition.
[0162] Although the method of this invention, which employs action domain limitation and exploration strategy resetting, performs comparably to method 2 under stable environmental conditions, it achieves rapid reconvergence and real-time updates when facing abrupt electromagnetic changes, thanks to its effective action domain limitation and high exploration rate resetting algorithm. This demonstrates the advantages of this method over methods 2 and 3. In summary, this embodiment proves the effectiveness and reliability of the proposed dual-depth Q-network-based vehicle radar anti-interference waveform design method.
[0163] This invention may have other embodiments. Without departing from the spirit and essence of this invention, those skilled in the art can make various corresponding changes and modifications according to this invention, but these corresponding changes and modifications should all fall within the protection scope of the appended claims.
Claims
1. A method for designing anti-interference waveforms for online vehicle-mounted radar based on dual-depth Q-networks, characterized in that: The specific process of the method is as follows: Step 1: Estimate the signal-to-interference ratio (SIR) of the vehicle-mounted radar received signal; Step 2: Based on the signal-to-interference ratio (SINR) from Step 1, construct the Markov decision process for vehicle-mounted radar; Step 3: Based on the Markov decision process of the vehicle radar in Step 2, construct and train a dual-deep Q-network model. Step 4: Online adaptive adjustment of parameters of the dual-depth Q-network model under abrupt electromagnetic interference environment to obtain the dual-depth Q-network model under abrupt electromagnetic interference environment. Step 5: Obtain the state space of the vehicle radar, input the dual-depth Q-network model under electromagnetic interference under abrupt change conditions obtained in Step 4, and output the optimal action corresponding to the optimal action value function value, which is the anti-interference waveform of the vehicle radar. The step one involves estimating the signal-to-interference ratio (SIR) of the vehicle-mounted radar received signal; the specific process is as follows: Step 11: Provide the power of the target echo signal and interference signal received by the vehicle-mounted radar receiver; the specific process is as follows: Let the power of the target echo signal received by the vehicle-mounted radar receiver be... ,have Let the power of the interference signal received by the vehicle-mounted radar receiver be... ,have in Indicates the transmission power of the vehicle-mounted radar. Indicates antenna gain. The radar cross-section of the target, This indicates the distance from the target or interference source to the vehicle-mounted radar receiver. Indicates the effective aperture of the vehicle-mounted radar antenna. , The carrier wavelength for the signal transmitted by the vehicle-mounted radar; Steps one and two: Mix the target echo signal and interference signal received by the vehicle-mounted radar receiver with the vehicle-mounted radar transmitted signal respectively; the specific process is as follows: The target echo signal received by the vehicle-mounted radar receiver is mixed with the transmitted signal of the vehicle-mounted radar to obtain the frequency domain expression. for in Let be the impulse function. It is the carrier frequency of the signal transmitted by the vehicle-mounted radar. It is the time delay of the echo signal. The slope of the chirp signal emitted by the vehicle-mounted radar; The imaginary unit, ; For radar frequency; The interference signal received by the vehicle-mounted radar receiver is mixed with the transmitted signal of the vehicle-mounted radar to obtain the frequency domain expression. for in This refers to the time it takes for the interference signal to enter the radar receiver bandwidth. The time it takes for the interference signal to leave the radar receiver bandwidth. It is the absolute value of the difference between the slopes of the chirp signal transmitted by the radar transmitter and the chirp signal transmitted by the jamming source. For the initial phase, It is a rectangular window function with an amplitude of 1; Step 13: Estimate the signal-to-interference ratio (SIR) of the vehicle-mounted radar received signal; the specific process is as follows: The frequency domain expression of the mixing of the echo signal obtained in steps one and two with the vehicle-mounted radar transmission signal. And the frequency domain expression of the mixing of the interference signal and the vehicle radar transmitted signal. Estimate the signal-to-interference ratio (SIR) of the vehicle-mounted radar signal received at the target detection point; the expression is: in, , The received power of the target echo signal at the detection point. The interference power introduced by the interference signal at the detection point; The frequency domain sampling interval; A rectangular function with unit amplitude; Considering and The value is 1. Therefore, the expression for the signal-to-interference ratio estimation can be approximated as: because Combined with the echo signal power obtained step by step and interference signal power as well as The expression for the signal-to-interference ratio (SINR) estimation is further simplified to: 。 2. The online vehicle-mounted radar anti-interference waveform design method based on dual-depth Q-network according to claim 1, characterized in that: In step two, based on the signal-to-interference ratio (SIR) of step one, a Markov decision process for the vehicle-mounted radar is constructed; the specific process is as follows: Step 2.1: Construct a Markov decision process environment model; the specific process is as follows: exist The state space of the vehicle radar at all times Defined as ,in The number of observation vectors that make up the state space. for The observation vector at time t, for The reward function at time step, for The signal-to-interference ratio (SIR) of the vehicle-mounted radar receiving signals at all times. for The row vector of the position of the interference source to be detected relative to the vehicle-mounted radar in the environment at any given time. ,in The number of interference sources to be detected in the environment. for The row vector of the position of the W-th interference source to be detected relative to the vehicle-mounted radar in the environment at any given time. for The first reference frame in the real-time environment is the vehicle-mounted radar. The Cartesian coordinates of the interference source to be detected. , For action space, representing The radar waveform parameters used by the vehicle-mounted radar under the specified conditions; Step 22: Based on the information-to-interference ratio from Step 1, construct the reward function and action space of the Markov decision process; the specific process is as follows: The reward function of a Markov decision process is ; in For the set signal-to-interference ratio threshold, for The signal-to-interference ratio (SIR) of the vehicle-mounted radar receiving signals at all times; The action space set of a Markov decision process can be represented as ; in The number of action spaces, For the first in the action space set One action, .
3. The online vehicle-mounted radar anti-interference waveform design method based on dual-depth Q-network according to claim 2, characterized in that: In step three, a dual-depth Q-network model is constructed and trained based on the Markov decision process of the vehicle-mounted radar in step two; the specific process is as follows: Step 3:
1. Construct and initialize a dual-depth Q-network structure; Step 3.2: Utilize the evaluation network to control the vehicle-mounted radar to collect experience; Step 3: Calculate the time difference function using the target network; Steps 3 and 4: Construct the time difference error function and update the neural network parameters; Step 35: Repeat steps 32 to 34 until convergence, and obtain the trained dual-depth Q-network model.
4. The online vehicle-mounted radar anti-interference waveform design method based on dual-depth Q-network according to claim 3, characterized in that: In step 3.1, a dual-depth Q-network structure is constructed and initialized; the specific process is as follows: Step 3.11: Construct a dual-depth Q-network model; the specific process is as follows: Intelligent anti-interference networks based on deep reinforcement learning include evaluation networks. and target network ; Evaluation Network It consists of one convolutional layer and five fully connected layers. The state space of the vehicle radar at all times Input evaluation network Features are extracted through a convolutional layer, and then the optimal action value function is output through five fully connected layers. Target Network Adoption and evaluation of networks Same network structure; Inputs in the evaluation network and the target network for The state space where the vehicle-mounted radar is located at all times; Outputs in the evaluation network and the target network for The radar waveform parameters used by the vehicle-mounted radar under the specified conditions; To evaluate the parameters to be learned in the network, The target network parameters do not need to be learned; the network parameters to be learned will be evaluated periodically. Assign values to target network parameters That's all; Step 3.12: Initialize the dual-depth Q-network model; the specific process is as follows: The network parameters of the initialized dual-depth Q-network model include the size of the empirical replay library. Learning rate ,attenuation In strategy Return decay factor One update learning size Target network The threshold C for the number of parameter lag update steps; C represents the initial exploration rate; C is the threshold for the number of update steps.
5. The online vehicle-mounted radar anti-interference waveform design method based on dual-depth Q-network according to claim 4, characterized in that: Step 3.2 involves evaluating the experience collected by the vehicle-mounted radar controlled by the network; the specific process is as follows: In the initial state space According to attenuation Strategy for selecting waveform parameters Then the environment provides the reward function. The environment jumps to the next state space. , combined Information-to-interference ratio estimation at time point and The positional line vector of the interference source to be detected relative to the vehicle radar in the environment at any given time. Integrate into experience Store the experience in the experience pool until it is full, then proceed. Random sampling of size for learning; attenuation The strategy can be represented as , ; in , This is the initial exploration rate. This represents the total number of steps explored. Indicates the current number of exploration steps. This represents the exploration rate corresponding to the k-th exploration.
6. The online vehicle-mounted radar anti-interference waveform design method based on dual-depth Q-network according to claim 5, characterized in that: In step 3, the target network is used to calculate the time difference function; The specific process is as follows: Using target network Calculate the time difference function , ; in Indicates the discount decay factor. It represents multiplication.
7. The online vehicle-mounted radar anti-interference waveform design method based on dual-depth Q-network according to claim 6, characterized in that: In steps three and four, the time difference error function is constructed and the neural network parameters are updated; the specific process is as follows: Constructing the time difference error function Using stochastic gradient descent to... Perform parameter updates, and when the update pace reaches... When the value is an integer multiple, the network will be evaluated. Parameters in Assigned to the target network ; in denoted by , represents the loss function of the dual-depth Q-network model; C is the threshold number of update steps.
8. The online vehicle-mounted radar anti-interference waveform design method based on dual-depth Q-network according to claim 7, characterized in that: In step four, the parameters of the dual-depth Q-network model are adaptively adjusted online under the abrupt electromagnetic interference environment to obtain the dual-depth Q-network model under the abrupt electromagnetic interference environment; the specific process is as follows: Step 41: Action space constraint for the dual-depth Q-network model; the specific process is as follows: This represents the set of action spaces prior to a sudden change in the electromagnetic environment. Indicates the optimal state before the mutation. A set of action spaces; The action space set before the electromagnetic environment change Subtract the optimal value before the mutation A set of action spaces , as a new set of action spaces; Step 4.2: Exploration strategy correction for the dual-depth Q-network model; the specific process is as follows: Using attenuation based on a dual-depth Q-network model The strategy selects actions from the new action space set. Set the current exploration steps to zero, and let Re-count from zero to ensure decay. Exploratory nature of the strategy; Repeat steps 32, 33, and 34 to obtain the dual-depth Q-network model under electromagnetic interference conditions with abrupt change.