Multi-transmit multi-receive radar networking cooperative anti-jamming method based on nested reinforcement learning
Patent Information
- Application Number
- CN202510367698.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2026-08-28
- Estimated Expiration
- 2045-03-26
AI Technical Summary
[0005]本发明实施例提供了一种基于嵌套强化学习的多发多收雷达组网协同抗干扰方法,可以解决当前的多雷达系统的抗干扰方法适应性和协同能力较差的问题
[0018]本发明实施例与现有技术相比存在的有益效果是:由于本发明提供的方法是基于多智能体强化学习算法对频点决策模型进行训练的,相较于单雷达系统,其在应对复杂干扰环境中展现出显著优势;并且本发明通过资源互补和策略协同,多个雷达站可以形成更加有效的抗干扰网络,提高系统内各收发站的协同性能。
Smart Images

Figure CN120178173B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of radar technology, specifically relating to a multi-transmitter, multi-receiver radar networking collaborative anti-jamming method based on nested reinforcement learning. Background Technology
[0002] With the development of jamming technology, in order to effectively counter intelligent jammers with learning capabilities and autonomous strategy adjustments, radar must fully perceive and analyze the external jamming environment and combine multi-dimensional information such as time, frequency, and space to construct flexible anti-jamming strategies. Through adaptive learning, radar can adjust its operating mode in real time and dynamically optimize anti-jamming measures, thereby maintaining a competitive advantage in complex and ever-changing environments. Compared to traditional anti-jamming methods that rely on fixed rules or expert systems, intelligent adaptive anti-jamming strategies can continuously optimize their own decisions based on the jammer's behavior patterns, giving radar greater adaptability and robustness in long-term confrontations.
[0003] However, in scenarios where a single radar counters multiple intelligent jammers, the radar often faces severe resource constraints and insufficient decision-making freedom. A single radar is not only limited in its sensing range, power allocation, and strategy adjustment, but its individual anti-jamming strategies often fail to achieve ideal results when facing multiple jammers operating in concert. Therefore, employing a multi-radar system for coordinated anti-jamming becomes a feasible solution. A multi-radar system consists of multiple geographically dispersed radar stations, and through information sharing, resource complementarity, and strategy coordination, the overall anti-jamming capability of the system can be significantly enhanced.
[0004] However, current anti-jamming methods for multi-radar systems cannot adapt to complex and ever-changing real-world scenarios, and the coordination capabilities between radars are poor. Summary of the Invention
[0005] This invention provides a multi-receiver radar network collaborative anti-jamming method based on nested reinforcement learning, which can solve the problem of poor adaptability and collaborative ability of current multi-radar system anti-jamming methods.
[0006] In a first aspect, embodiments of the present invention provide a collaborative anti-jamming method for multi-transmitter / multi-receiver radar networks based on nested reinforcement learning, the method comprising:
[0007] The first and second state information of the multiple-transmitter multiple-receiver radar system at the current moment are obtained, wherein the multiple-transmitter multiple-receiver radar system includes multiple transmitting stations and multiple receiving stations;
[0008] The first state information of the multiple-transmitter multiple-receiver radar system at the current moment is input into the trained radar transceiver station decision model to obtain the transmission and reception strategy of the multiple-transmitter multiple-receiver radar system. The transmission and reception strategy is used to indicate the transmitting station of the multiple-transmitter multiple-receiver radar system at the current moment for transmitting signals.
[0009] The second state information of the multiple-transmitter multiple-receiver radar system at the current moment and the transmit / receive strategy are input into the trained radar frequency decision model to obtain the frequency strategy of the multiple-transmitter multiple-receiver radar system.
[0010] The frequency point strategy is used to indicate the frequency point of the transmitted signal. The trained radar transceiver station decision model and the trained frequency point decision model are obtained by nested reinforcement learning based on single agent reinforcement learning algorithm and multi agent reinforcement learning algorithm, respectively.
[0011] The multiple transmit / receive radar system transmits signals according to an anti-jamming strategy, wherein the anti-jamming strategy consists of the transmit / receive strategy and the frequency point strategy.
[0012] Secondly, embodiments of the present invention provide a multi-transmitter, multi-receiver radar network collaborative anti-jamming device based on nested reinforcement learning, comprising:
[0013] An acquisition unit is used to acquire first and second state information of a multiple-transmitter multiple-receiver radar system, wherein the multiple-transmitter multiple-receiver radar system includes multiple transmitting stations and multiple receiving stations.
[0014] The first decision unit is used to input the first state information of the multiple-transmitter multiple-receiver radar system at the current moment into the trained radar transceiver station decision model to obtain the transmission and reception strategy of the multiple-transmitter multiple-receiver radar system. The transmission and reception strategy is used to indicate the receiving station for receiving signals and the transmitting station for transmitting signals of the multiple-transmitter multiple-receiver radar system at the current moment.
[0015] The second decision unit is used to input the second state information of the multiple-transmitter multiple-receiver radar system at the current moment and the transmit / receive strategy into the trained radar frequency decision model to obtain the frequency strategy of the multiple-transmitter multiple-receiver radar system.
[0016] The frequency point strategy is used to indicate the frequency point of the transmitted signal. The trained radar transceiver station decision model and the trained frequency point decision model are obtained by nested reinforcement learning based on single agent reinforcement learning algorithm and multi agent reinforcement learning algorithm, respectively.
[0017] A control unit is configured to control the transmission of signals by the multiple-transmitter multiple-receiver radar system according to an anti-jamming strategy, wherein the anti-jamming strategy consists of the transmit / receive strategy and the frequency point strategy.
[0018] The beneficial effects of the embodiments of the present invention compared with the prior art are as follows: Since the method provided by the present invention is based on training the frequency point decision model using a multi-agent reinforcement learning algorithm, it shows significant advantages in dealing with complex interference environments compared with a single radar system; and through resource complementarity and strategy coordination, multiple radar stations can form a more effective anti-interference network, improving the collaborative performance of each transceiver station in the system. Attached Figure Description
[0019] Figure 1 This is a schematic diagram of the structure of a multi-transmitter, multi-receiver radar network collaborative anti-jamming device based on nested reinforcement learning, provided in an embodiment of the present invention.
[0020] Figure 2 A flowchart illustrating the implementation of a training method for a radar transceiver station decision model and a frequency point decision model provided in an embodiment of the present invention;
[0021] Figure 3 A flowchart illustrating the implementation of a method for training a frequency point decision model based on the transmit / receive strategy at time t, as provided in an embodiment of the present invention.
[0022] Figure 4 A flowchart illustrating the implementation of a multi-transmitter, multi-receiver radar networking collaborative anti-jamming method based on nested reinforcement learning, provided for an embodiment of the present invention;
[0023] Figures 5a-5f The diagram shown is a schematic of the curve of the detection probability at the end of each CPI as a function of the number of training iterations, obtained from the simulation experiment of this invention.
[0024] Figure 6 This is a schematic diagram illustrating how the radar transceiver station strategy changes during training, obtained from the simulation experiment of this invention.
[0025] Figure 7a and Figure 7b This is a frequency point diagram obtained from the simulation experiment of this invention. Detailed Implementation
[0026] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of the invention. However, those skilled in the art will understand that the invention can be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods are omitted so as not to obscure the description of the invention with unnecessary detail.
[0027] It should be understood that, when used in this specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.
[0028] It should also be understood that the term "and / or" as used in this specification and the appended claims refers to any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0029] As used in this specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if [described condition or event] is detected" may be interpreted, depending on the context, as meaning "once determined," "in response to determination," "once [described condition or event] is detected," or "in response to detection of [described condition or event]."
[0030] Furthermore, in the description of this invention and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0031] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of the invention include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.
[0032] The present invention will be further described in detail below with reference to specific embodiments, but the implementation of the present invention is not limited thereto.
[0033] The multi-transmitter, multi-receiver radar networking collaborative anti-jamming method based on nested reinforcement learning provided in this invention can be applied to electronic devices such as mobile terminals, personal laptops, and supercomputers. This invention does not impose any restrictions on the specific type of electronic device.
[0034] Example 1
[0035] Figure 1The diagram shown is a structural schematic of a multi-transmitter, multi-receiver radar network cooperative anti-jamming device based on nested reinforcement learning, provided by an embodiment of the present invention. As an example and not a limitation, the device 100 may include an acquisition unit 110, a first decision unit 120, a second decision unit 130, and a control unit 140.
[0036] For example, the acquisition unit 110 can acquire the first state information and the second state information of the multi-receiver radar system at the current moment. The first decision unit 120 can input the first information at the current moment into the trained radar transceiver station decision model to obtain the transmission and reception strategy of the multi-receiver radar system; wherein, the transmission and reception strategy is used to indicate the transmitting station of the multi-receiver radar system for transmitting signals at the current moment; the second decision unit 130 inputs the second state information and the transmission and reception strategy of the multi-receiver radar system at the current moment into the trained radar frequency point decision model to obtain the frequency point strategy of the multi-receiver radar system; wherein, the frequency point strategy is used to indicate the frequency point of the transmitted signal, and the trained radar transceiver station decision model and the trained frequency point decision model are obtained by nested reinforcement learning based on single-agent reinforcement learning algorithm and multi-agent reinforcement learning algorithm, respectively; the control unit 140 controls the transmission of signals of the multi-receiver radar system according to the anti-jamming strategy, wherein, the anti-jamming strategy consists of the transmission and reception strategy and the frequency point strategy.
[0037] Since the device provided by this invention is based on a multi-agent reinforcement learning algorithm to train the frequency decision model, it shows significant advantages in dealing with complex interference environments compared to a single radar system. Furthermore, through resource complementarity and strategy coordination, multiple radar stations can form a more effective anti-jamming network, improving the collaborative performance of each transceiver station within the system.
[0038] Example 2
[0039] As an example, since the radar transceiver decision model is used to determine the transmitting stations that are currently active, the action space of the radar transceiver decision model can include the identifiers of each transmitting station in the radar system, such as the transmitting station number. The radar transceiver decision model can select the identifiers of multiple transmitting stations from the action space as transmit / receive strategies to enable the transmitting stations indicated by these identifiers to transmit signals.
[0040] For example, the action space of a radar transceiver station decision model can be represented as:
[0041]
[0042] in, Let I be the transmit / receive strategy at time T, where I is the number of active transmitting stations and H is the total number of radar transmitting stations.
[0043] In one example, the state space of the radar transceiver decision model can be composed of the real-time position information (e.g., real-time two-dimensional coordinates) of each radar in the radar system. The first state information can then include the real-time position information of each radar in the radar system. The transceiver strategy generated by the radar transceiver decision model can be used to indicate which transmitting station is currently active.
[0044] For example, the state space of a radar transceiver station decision model can be represented as:
[0045]
[0046] in, x represents the first state information at time T. h y h Let H and Y be the coordinates of the h-th radar at time T, respectively, in the x-axis and y-axis directions, where h is a positive integer less than or equal to H.
[0047] Specifically, the motion state of each radar in the radar system can be updated using the following dynamic model:
[0048]
[0049] in, Let be the coordinates of the h-th radar at time T+1 in the x-axis and y-axis directions, respectively. Let v be the coordinates of the h-th radar at time T along the x-axis and y-axis, respectively. h Let θ be the velocity of the h-th radar. h Let θ be the motion direction angle of the h-th radar, and ΔT be the time step.
[0050] Figure 2 The diagram illustrates a training method for a radar transceiver station decision model and a frequency point decision model provided in an embodiment of the present invention. As an example and not a limitation, the method may include steps S201-S207, which are described below.
[0051] S201, Obtain the first state information of the multiple transmit / receive radar system at time t.
[0052] For example, time t can be started from the beginning of the i-th round of training.
[0053] For example, the first state information at time t used during training can be sampled from the state space of the radar transceiver decision model.
[0054] S202, input the first state information at time t into the radar transceiver station decision model after the (i-1)th round of training to obtain the first training experience at time t; and put the first training experience at time t into the first experience playback pool.
[0055] In some embodiments, the radar transceiver decision model can be an actor-critic architecture, including a first policy network and a first value network.
[0056] In one possible implementation, the first training experience may include the transmit / receive strategy at time t (i.e., the action the radar system will perform at time t), the first reward at time t, and the first state information at time t and time t′.
[0057] For example, time t′ is the next time after time t. For instance, if time t′ is 3:05:49 AM on February 14, 2025, then time t′ could be 3:05:50 AM on February 14, 2025.
[0058] In one example, the first policy network can select an action from the action space as the transmit / receive policy at time t based on the first state information at time t.
[0059] For example, the transmit / receive strategy at time t can be expressed as:
[0060]
[0061] in, Let θ be the transmit / receive strategy at time t. j For the first policy network π s The current latest network parameters (here, the network parameters after the (i-1)th round of training), This represents the first state information at time t.
[0062] In one example, the first reward at time t. The joint detection probability Pd of the radar system at time t can be given. total ,Right now
[0063] S203, according to the transmit / receive strategy at time t, the frequency decision model after the (i-1)th round of training is trained in the i-th round to obtain the frequency decision model after the i-th round of training.
[0064] For example, a frequency decision model can be trained under the transmit / receive strategy at time t, so that the frequency decision network can learn from the experience in this scenario.
[0065] S204, determine whether the number of first training experiences in the first experience replay pool is greater than the first preset experience threshold.
[0066] In one example, if the number of first training experiences in the first experience replay pool is not greater than the first preset experience threshold, then t = t + 1 can be set to continue accumulating experience from step S201.
[0067] Specifically, reinforcement learning involves two time concepts: rounds and time steps. A round is a larger time scale representing the complete process of a task, containing multiple steps. A time step is a small unit within a round, representing each interaction between the agent and the environment. In radar systems, the focus is often on a coherent processing interval (CPI), which contains multiple pulses. If each pulse within a CPI is understood as a time step in reinforcement learning, then one CPI corresponds to one round in reinforcement learning. Furthermore, in reinforcement learning, agent training involves multiple rounds, which is analogous to multiple CPIs in radar. Therefore, in a radar transceiver decision model, one CPI corresponds to one time step, i.e., the time difference T1 between time t+1 and time t equals one CPI; in a frequency decision model, one CPI corresponds to one round.
[0068] In another example, if the number of first training experiences in the first experience replay pool is greater than the first preset experience threshold, then the following step S205 can be performed.
[0069] S205, sample the first training experience data from the first experience replay pool, and based on the first update model, update the network parameters of the radar transceiver station decision model after the (i-1)th round of training according to the sampled first training experience data, to obtain the radar transceiver station decision model after the i-th round of training.
[0070] In one example, the radar transceiver station decision model can update the number of network books based on the Proximal Policy Optimization (PPO) algorithm. The first update model can then satisfy the following formula:
[0071]
[0072] Among them, L PPO (θ j Let θ be the loss function of the first policy network. j The network parameters are updated for the first strategy network. Expressing expectations, For the first policy network with parameters θ j Based on the first state information at time t The generated transmit / receive policy indicates that the first policy network is in state 1. Choose action The probability of; This indicates that the first policy network has parameters of... According to The generated send and receive strategy, For the network parameters before the first strategy network update, A t For generalized dominance estimation, clip(·) denotes the cutoff function, and ∈ is the clipping threshold; Let θ be the loss function of the first value network. m The updated network parameters are for the first value network. Indicates the expectation. The parameter is θ m The first value network in state The value obtained below This is the first cumulative return.
[0073] in:
[0074]
[0075] Where γ is the discount factor. This is the first reward at time t+j.
[0076] S206, Determine whether the training stop requirements are met.
[0077] In one example, if the training stopping condition is not met, then i = i + 1 can be set to proceed to the next round of training.
[0078] For example, the training stopping condition can be determined based on the size of the training rounds and the performance of the two models after the i-th round of training.
[0079] In another example, if the training stopping condition is met, step S207 can be performed.
[0080] S207, output the radar transceiver station decision model after the i-th round of training and the frequency point decision model after the i-th round of training as the trained radar transceiver station decision model and the trained frequency point decision model, respectively.
[0081] By nesting reinforcement learning between the two decision models, the frequency decision model can learn from experience under different transmit / receive strategies and environments, and the radar transceiver station decision model can learn from experience under different environments. This enables the two models to formulate coordinated strategies for each transceiver station in a multi-transmitter / multi-receiver radar system in complex real-world environments, thereby improving the radar system's anti-jamming capability.
[0082] Based on Example 2, the present invention also provides Example 3.
[0083] As an example, the radar's transmitted signal can be a frequency-hopping sequence, and the transmitted signal contains several (e.g., K) sub-pulses within a single pulse. Since the frequency decision model is used to determine the frequency of the transmitted signal, the action space of the frequency decision model can be composed of the frequency number of each sub-pulse. Users can pre-set multiple (e.g., N) operating frequencies for the radar system and assign them numbers. When making a decision, the frequency decision model can select K operating frequency numbers from its action space to set the operating frequencies corresponding to these numbers as the frequencies of the corresponding sub-pulses.
[0084] For example, the action space of the frequency point decision model can be represented as:
[0085]
[0086] in, For the frequency point strategy at time k, 0≤a i ≤N-1 indicates that the carrier frequency of the i-th sub-pulse is K is the number of sub-pulses, and N is the total number of operating frequency points.
[0087] Optionally, dividing a pulse into multiple sub-pulses can give the radar higher degrees of freedom; however, as the radar's frequency and the number of sub-pulses increase, using a discrete action space can lead to an excessively large dimension of the action space, resulting in decreased training efficiency and increased computational costs. Therefore, when the frequency and the number of sub-pulses are large, a continuous action space discretization method can be used to convert the continuous frequency information of the K sub-pulses output by the frequency decision model into discrete frequency information.
[0088] Specifically, the expression for discrete frequency point information can be:
[0089]
[0090] Among them, a d Discrete action representation of radar frequency, a c The continuous motion representation of the radar frequency point is given by tanh, which is the hyperbolic tangent function, and floor, which is the floor function.
[0091] In some embodiments, the state space of the frequency point decision model may consist of the observation information (e.g., the observation matrix) of the radar system, and the second state information may include the observation information of the radar system.
[0092] For example, the state space of the frequency point decision model can be represented as:
[0093]
[0094] in, This represents the second state information of the radar system at time k. Let m be the observation matrix obtained from the transmission signal of the h′-th transmitting station activated at time km, where m is a positive integer less than or equal to M and h′ is a positive integer less than or equal to I.
[0095] For example, the observation matrix of the radar system at time k can be represented as:
[0096]
[0097] in, Let represent the observation matrix of radar system at time k, indicating the observation of radar y at time k.
[0098] by For example, it satisfies:
[0099]
[0100] in, Let a be the radar information at the Nth frequency point of the Kth subpulse. K Indicates the frequency at which the radar transmits subpulse K; α q Let q be the sidelobe attenuation coefficient of the q-th jammer. When the q-th jammer interferes with the radar y as a main lobe jammer, its value is 1. When it interferes with the sidelobe jammer, it is attenuated according to the antenna pattern coefficient. q is a positive integer less than or equal to Q. Let q be the interference intensity of the q-th jammer on the N-th frequency point of the radar at the K-th subpulse.
[0101] In one possible implementation, the jamming action of the jammer on the radar system can be represented by a jamming matrix, which can represent the jamming intensity of the jammer at different frequencies at different times.
[0102] For example, the interference matrix can be represented as:
[0103]
[0104] in, Let be the interference matrix of the q-th jammer at time k. Let be the interference intensity of the q-th jammer at the n-th frequency point on the radar's i′-th subpulse, typically taking a value between [0,1]. The time indicates that the q-th jammer is currently in the signal receiving state.
[0105] In one example, jamming strategies can be mainly divided into two types: frequency-targeting jamming and intermittent sampling jamming. Frequency-targeting jamming refers to the jammer precisely interfering with the radar frequency, usually by adjusting the frequency of the jamming signal to be close to or coincide with the radar's operating frequency. Intermittent sampling jamming is a type of decoy jamming that samples at specific times and uses the sampled signal as the jamming signal. Its typical operating mode is "receive x, transmit y," where "receive x" means the jammer will sample the signal within x sub-pulse periods, and "transmit y" means the jammer will transmit the jamming signal within y sub-pulse periods.
[0106] For example, necessary operational isolation can be implemented between a receiving jammer and a transmitting jammer operating on the same platform. Specifically, high transmit / receive isolation can be used to achieve intermittent observation. For instance, the transmitting jamming signal can be interrupted so that the receiver can effectively intercept the signal.
[0107] It should be understood that the state space and action space of the frequency point decision model should be constructed together with the state space and action space of the radar transceiver station decision model before the first training.
[0108] Figure 3 The diagram illustrates a flowchart of a method for training a frequency decision model based on a transmit / receive strategy at time t, as provided in an embodiment of the present invention. This method is illustrative and not limiting; it can be a specific possible implementation of step S203 described above. The method may include steps S301-S304, which are described below.
[0109] S301, acquire the second state information of the multiple transmit / receive radar system at time k.
[0110] For example, k can be accumulated starting from time t.
[0111] S302, input the second state information at time k and the transmit / receive strategy at time t into the frequency decision model after training for the (i-1)th time to obtain the second training experience at time k.
[0112] In some embodiments, the frequency decision model can be an actor-critic architecture, including a second policy network and a second value network.
[0113] In one possible implementation, the second training experience may include the frequency policy at time k, the second reward at time k, and the second state information at times k and k′.
[0114] Similarly, time k′ is the next time after time k.
[0115] In one example, the first policy network can select an action as the frequency policy at time k from the action space of the frequency decision network, based on the second state information at time k and the transmit / receive policy at time t.
[0116] For example, the frequency point strategy at time k can be expressed as:
[0117]
[0118] in, Let θ be the frequency strategy at time t. i For the second policy network π f The latest network parameters at the current moment (here, the network parameters after the (i-1)th round of training), This represents the first state information at time k.
[0119] In one example, the second reward at time k. Let be the non-coverage rate of the radar system at time k, which satisfies the following formula:
[0120]
[0121] in, For the second reward at time k, num uncovered K represents the number of undisturbed sub-pulses in a pulse, where K is the number of sub-pulses in a pulse.
[0122] S303, determine whether the number of second training experiences in the second experience replay pool is greater than the second preset experience threshold.
[0123] In one example, if the number of second training experiences in the second experience replay pool is not greater than the second preset experience threshold, then k = k + 1 can be set to continue accumulating the second training experience from step S301.
[0124] For example, the time difference T2 between time k+1 and time k is less than T1.
[0125] In another instance, if the number of second training experiences in the second experience replay pool is greater than the second preset experience threshold, then step S304 can be performed.
[0126] S304. Sample the second training experience data from the second experience replay pool, and based on the second update model, update the network parameters of the frequency point decision model after the (i-1)th round of training according to the sampled second training experience data to obtain the frequency point decision model after the i-th round of training.
[0127] In one example, the frequency point decision model can update network parameters based on the Multi-Agent Proximal Policy Optimization (MAPPO) algorithm. The second update model can then satisfy the following formula:
[0128]
[0129] Among them, L MAPPO (θ i Let θ be the loss function of the second policy network. i The network parameters are updated for the second strategy network. Expressing expectations, For the second policy network with parameters θ i Based on the second state information at time k The generated transmit / receive policy indicates that the second policy network is in state 1. Choose action The probability of; This indicates that the second policy network has parameters of According to The generated send and receive strategy, For the network parameters before the second strategy network update, A k For generalized dominance estimation, clip(·) denotes the cutoff function, and ∈ is the clipping threshold; Let θ be the loss function of the second value network. n The network parameters are updated for the second value network. Indicates the expectation. The parameter is θ n The second value network in state The value obtained below This is the second cumulative return.
[0130] For example, a truncation function can restrict the result to the interval [1-∈, 1+∈], which determines the tolerance for ratio changes with each policy update.
[0131] For example, Generalized Advantage Estimation (GAE) can measure the advantage in a given state. Take action Compared to state The degree of improvement of the baseline value. The generalized state estimate of the frequency point decision model can be calculated using the following formula:
[0132]
[0133] Where γ is the discount factor, λ is a parameter of GAE used to balance the trade-off between bias and variance; δ k Time difference error is the difference between the agent's policy and the actual reward in the current state. It measures the error between the performance of the current policy in the current state and the expected reward. It is the state at time k. The value estimate, Based on the state at time k′ Calculate the discount for future rewards.
[0134] Similarly, the second cumulative reward at time k It can satisfy: This is the second reward at time k+j′.
[0135] Since the method provided by this invention is based on training the frequency decision model using a multi-agent reinforcement learning algorithm, it exhibits significant advantages in dealing with complex interference environments compared to a single radar system. Furthermore, through resource complementarity and strategy coordination, multiple radar stations can form a more effective anti-jamming network, improving the collaborative performance of each transceiver station within the system.
[0136] Furthermore, this invention constructs the state space of the radar transceiver station decision model based on the radar networking system in motion, taking into account the impact of radar position changes on radar-jamming countermeasures, which can improve the anti-jamming performance of the decision.
[0137] Furthermore, this invention is based on a multi-agent reinforcement learning algorithm framework, which enables decision-making when multiple radars and multiple jammers do not fully know each other's information. Multi-agent reinforcement learning is a solution for sequential decision-making with missing prior information.
[0138] Based on Examples 2 and 3, the present invention also provides Example 4.
[0139] Figure 4 The diagram illustrates a flowchart of a multi-transmitter, multi-receiver radar network cooperative anti-jamming method based on nested reinforcement learning, provided by an embodiment of the present invention. As an example and not a limitation, the method may include steps S401-S404, which are described below.
[0140] S401, acquire the first and second state information of the multiple-transmitter multiple-receiver radar system at the current moment.
[0141] For example, a multiple-transmitter multiple-receiver radar system may include a transmitting station and multiple receiving stations.
[0142] S402, input the first state information of the multiple-transmitter multiple-receiver radar system at the current moment into the trained radar transceiver station decision model to obtain the transmission and reception strategy of the multiple-transmitter multiple-receiver radar system.
[0143] For example, the transmit / receive strategy is used to indicate the transmitting station of the radar system for transmitting signals at the current moment.
[0144] S403: Input the second state information and transmit / receive strategy of the multi-transmitter multi-receiver radar system at the current moment into the trained radar frequency decision model to obtain the frequency strategy of the multi-transmitter multi-receiver radar system.
[0145] For example, a frequency point strategy is used to indicate the frequency point of the transmitted signal.
[0146] For example, the trained radar transceiver station decision model and the trained frequency point decision model can be trained according to the methods described in Embodiments 2 and 3 above.
[0147] S404 controls the transmission signals of the multiple-transmitter, multiple-receiver radar system according to anti-jamming strategies.
[0148] For example, the anti-jamming strategy of a radar system may include a transceiver strategy and a frequency strategy.
[0149] Since the method provided by this invention is based on training two decision models using a multi-agent reinforcement learning algorithm, it exhibits significant advantages in dealing with complex interference environments compared to a single radar system. Furthermore, through resource complementarity and strategy synergy, multiple radar stations can form a more effective anti-jamming network, improving the collaborative performance of each transceiver station within the system.
[0150] To better illustrate the beneficial effects of the present invention, the following simulation experiment was conducted.
[0151] Simulation conditions
[0152] The experiment involved a total of four radars in a two-transmit, four-receive configuration, with ten available frequency points. Each pulse contained four sub-pulses, and each CPI (Continuous Pipeline Intake) consisted of 32 pulses. The radar power was 30 kW, the transmit antenna gain was 30 dB, the receive antenna gain was 30 dB, and the target RCS was 15 m. 2 The radar receiver bandwidth is 25MHz; the total number of jammers is 4, and the main lobe drying ratio is 30dB.
[0153] The jammer and target positions are fixed at [100km, 100km]. The positions, speeds, and orientation angles of the four radars are shown in Table 1.
[0154] Table 1
[0155] Radar 0 [110,0] 300 96° Radar 1 [10,0] 800 48° Radar 2 [0,0] 1000 45° Radar 3 [0,150] 500 -27°
[0156] Specifically, the operational space of a radar transceiver station is finite and discrete, and coded actions can be obtained by encoding combinations of radar numbers. Its operational space is as follows:
[0157] a s ∈{0,1,2,3,4,5}
[0158] Wherein, the a s The transceiver station decision action of the radar at time t is represented by the numbering of the four radars as radar 0, radar 1, radar 2 and radar 3 respectively. The mapping relationship between the transceiver station decision action and the radar transmitting station number is shown in Table 2.
[0159] Table 2
[0160] Radar number [0,1] [0,2] [0,3] [1,2] [1,3] [2,3]
[0161] Simulation Experiment Content
[0162] Under simulation conditions, training was performed every 32 CPIs, with each CPI being 32 pulses long, for a total of 1024 pulses. This corresponded to 1024 data points sampled by MAPPO and 32 data points sampled by PPO. Each pulse was considered a time step, and the model parameters were saved after 300,000 time steps. The simulation results are as follows:
[0163] Figures 5a to 5f The figure shows the curve of the detection probability at the end of each CPI as a function of the number of training iterations, obtained from the simulation experiment. Both jammers use intermittent sampling jamming, but their transmit and receive times differ. Figure 5a The two jammers in China transmit and receive at the same times: one transmit and one receive, and the other transmit and receive. Figure 5b They are respectively one send and one receive, and one send and two receive. Figure 5c They are respectively receiving one and sending two, receiving two and sending four. Figure 5d They are respectively one send and one receive, and one send and three receive. Figure 5e They are respectively one send and one receive, and one send and four receive. Figure 5f The configurations are one receiver and four transmitters, respectively. As shown in Figure 5, with the increase in the number of training iterations, the detection probability roughly converges to between 0.9 and 1, indicating that the radar network's collaborative anti-jamming capability is relatively good.
[0164] Figure 6 The diagram illustrates how radar transceiver strategies change during training. The x-axis represents the number of training steps, and the y-axis represents the transceiver's decision-making actions. From Figure 6It can be seen that as the number of training iterations increases, the transceiver decision network tends to select the transmitting station with the largest angular difference for the jammer. For the radars, since radar 1 and radar 2 are relatively close, if the main lobe width of the jammer can completely cover them, one jammer can simultaneously and strongly jam both radars. When two jammers are jamming radar 0, their sidelobes have negligible interference with radar 3, indicating that the output of the radar transceiver strategy network is reasonable.
[0165] Figure 7a and Figure 7b These are frequency point maps obtained from training radar transmitter station 1 (corresponding to radar 0, with one receiver and four transmitters) and radar transmitter station 2 (corresponding to radar 3, with one receiver and two transmitters). In the graphs, the x-axis represents the sub-pulse index, the y-axis represents the radar transmitter frequency point index, the light blue line represents the jammer signal, the green line represents the unjammed radar signal, and the yellow line represents the jammed radar signal. From... Figure 7a As can be seen, the radar has completely learned the jammer's jamming strategy. When the jammer intercepts the signal, the radar transmits radio frequency point 9, and when the jammer jams, the radar transmits radio frequency point 1. Figure 7b As can be seen, all frequencies successfully avoided interference, while Figure 7a Of the 20 pulses, only one frequency was interfered with, indicating that the frequency decision network can successfully learn the jammer's transmission mode and avoid interference.
[0166] Therefore, the method provided by this invention can obtain an effective radar anti-jamming strategy and improve the anti-jamming performance of the radar system. At the same time, the method also considers the changes of moving radar in the spatial dimension when making decisions, further improving the anti-jamming performance of the radar system.
[0167] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.
Claims
1. A collaborative anti-jamming method for multi-transmitter / multi-receiver radar networks based on nested reinforcement learning, comprising: The first and second state information of the multiple-transmitter multiple-receiver radar system at the current moment are obtained, wherein the multiple-transmitter multiple-receiver radar system includes multiple transmitting stations and multiple receiving stations; The first state information of the multiple-transmitter multiple-receiver radar system at the current moment is input into the trained radar transceiver station decision model to obtain the transceiver strategy of the multiple-transmitter multiple-receiver radar system. The transceiver strategy is used to indicate the receiving station for receiving signals and the transmitting station for transmitting signals of the multiple-transmitter multiple-receiver radar system at the current moment. The second state information of the multiple-transmitter multiple-receiver radar system at the current moment and the transmit / receive strategy are input into the trained radar frequency decision model to obtain the frequency strategy of the multiple-transmitter multiple-receiver radar system. The frequency point strategy is used to indicate the frequency point of the transmitted signal. The trained radar transceiver station decision model and the trained frequency point decision model are obtained by nested reinforcement learning based on single agent reinforcement learning algorithm and multi agent reinforcement learning algorithm, respectively. The transmit / receive strategy and the frequency point strategy constitute an anti-jamming strategy, and the multiple transmit / receive radar system is controlled to transmit signals according to the anti-jamming strategy. The training methods for the radar transceiver station decision model and the frequency point decision model include: Obtain the first [unclear] of the multiple-send multiple-receiver radar system t First state information at any given moment; The first t The first state information at time i is input into the radar transceiver decision model after the (i-1)th round of training to obtain the first state information. t The first training experience at the moment; and the first t The first training experience at time step is placed into the first experience replay pool, wherein the first... t The first training experience at any given moment includes: t Time-based send and receive strategies, the first t The first reward of the moment, the first t Time and the The first state information at time t, the first The time is the first t The moment after the moment; According to the first t The transmit / receive strategy at each moment is used to train the frequency decision model after the (i-1)th round of training in the i-th round, resulting in the frequency decision model after the i-th round of training. Determine whether the number of first training experiences in the first experience replay pool is greater than the first preset experience threshold; If the number of the first training experiences is not greater than the first preset experience threshold, then let t=t +1, obtain the first [number] of the multi-transmitter multi-receiver radar system. t The first state information at time step [1] is used to continue accumulating the first training experience, wherein the [1]th [time step]... t +1 time and the first t The time difference is equal to the coherent processing interval of the multiple-transmitter multiple-receiver radar system. If the number of the first training experience items is greater than the first preset experience threshold, then the first training experience data is sampled from the first experience replay pool, and based on the first update model, the network parameters of the radar transceiver station decision model after the (i-1)th round of training are updated according to the sampled first training experience data to obtain the radar transceiver station decision model after the i-th round of training. Determine whether the training stop requirement is met. If it is met, output the radar transceiver station decision model after the i-th round of training and the frequency point decision model trained in the i-th round as the trained radar transceiver station decision model and the trained frequency point decision model, respectively.
2. The method according to claim 1, characterized in that, The first t The first reward at any given moment is the first time the multi-receiver radar system... t The probability of joint detection at any given time.
3. The method according to claim 1, characterized in that, The radar transceiver station decision model includes a first strategy network and a first value network; Wherein, the first update model satisfies the following formula: in, Let the loss function of the first policy network be . The network parameters are the updated network parameters after the first strategy network. Expressing expectations, For the first policy network with parameters as According to the first t First state information at time 1 The generated transmit / receive policy also indicates that the first policy network is in state 1. Choose action The probability of; This indicates that the first policy network has parameters of... According to The generated send and receive strategy, The network parameters before the first strategy network update are: For the generalized advantage estimation of the first-policy network, This represents the truncation function. This is the clipping threshold; Let be the loss function of the first value network. The updated network parameters for the first value network. Indicates the expectation. The parameter is The first value network in state The value obtained below This is the first cumulative return.
4. The method according to claim 1, characterized in that, According to the first t The transmit / receive strategy at each time step is used to train the frequency decision model after the (i-1)th training round in the i-th round, resulting in the frequency decision model after the i-th training round, which includes: Obtain the first [unclear] of the multiple-send multiple-receiver radar system k Second state information at time; The first k The second state information at time and the first t The transmit / receive strategy at time i is input into the frequency decision model after the (i-1)th round of training to obtain the i-th time frequency decision strategy. k The second training experience at time point, wherein the first k The second training experience at that moment includes: k Frequency strategy at time, the first k The second reward of the moment, the first k Time and the The second state information at time t, the first The time is the first k The moment after the moment; Determine whether the number of second training experiences in the second experience replay pool is greater than the second preset experience threshold; If the number of the second training experience points is not greater than the second preset experience threshold, then let k=k +1, the first k +1 time and the first k The time difference between the two points is less than the coherent processing interval; If the number of second training experiences is greater than the second preset experience threshold, then second training experience data is sampled from the second experience replay pool, and based on the second update model, the network parameters of the frequency point decision model after the (i-1)th round of training are updated according to the sampled second training experience data to obtain the frequency point decision model after the i-th round of training.
5. The method according to claim 4, characterized in that, The first k The second reward at time t satisfies the following formula: in, For the first k The second reward of the moment This represents the number of undisturbed sub-pulses within a single pulse. This represents the number of neutron pulses in a single pulse.
6. The method according to claim 4, characterized in that, The frequency point decision model includes a second policy network and a second value network; The second update model satisfies the following formula: in, Let the loss function be the second policy network. The network parameters are updated according to the second strategy network. Expressing expectations, For the second policy network with parameters According to the first k Second state information at time 1 The generated send / receive policy also indicates that the second policy network is in state 1. Choose action The probability of; This indicates that the second policy network has parameters of... According to The generated send and receive strategy, The network parameters before the second strategy network update are: For generalized advantage estimation of the second-policy network, This represents the truncation function. This is the clipping threshold; Let be the loss function of the second value network. The updated network parameters for the second value network. Indicates the expectation. The parameter is The second value network in state The value obtained below This is the second cumulative return.
7. The method according to any one of claims 1-6, characterized in that, The first status information includes the real-time location information of each radar in the multiple-transmitter-multiple-receiver radar system.
8. The method according to any one of claims 1-6, characterized in that, The second status information includes observation information from the multiple-transmitter, multiple-receiver radar system.
9. A multi-transmitter, multi-receiver radar network cooperative anti-jamming device based on nested reinforcement learning, comprising: An acquisition unit is used to acquire first and second state information of a multiple-transmitter multiple-receiver radar system, wherein the multiple-transmitter multiple-receiver radar system includes multiple transmitting stations and multiple receiving stations. The first decision unit is used to input the first state information of the multiple-transmitter multiple-receiver radar system at the current moment into the trained radar transceiver station decision model to obtain the transmission and reception strategy of the multiple-transmitter multiple-receiver radar system. The transmission and reception strategy is used to indicate the receiving station for receiving signals and the transmitting station for transmitting signals of the multiple-transmitter multiple-receiver radar system at the current moment. The second decision unit is used to input the second state information of the multiple-transmitter multiple-receiver radar system at the current moment and the transmit / receive strategy into the trained radar frequency decision model to obtain the frequency strategy of the multiple-transmitter multiple-receiver radar system. The frequency point strategy is used to indicate the frequency point of the transmitted signal. The trained radar transceiver station decision model and the trained frequency point decision model are obtained by nested reinforcement learning based on single agent reinforcement learning algorithm and multi agent reinforcement learning algorithm, respectively. A control unit is configured to combine the transceiver strategy and the frequency point strategy to form an anti-jamming strategy, and control the multiple transmit / receive radar system to transmit signals according to the anti-jamming strategy. The training methods for the radar transceiver station decision model and the frequency point decision model include: Obtain the first [unclear] of the multiple-send multiple-receiver radar system t First state information at any given moment; The first t The first state information at time i is input into the radar transceiver decision model after the (i-1)th round of training to obtain the first state information. t The first training experience at the moment; and the first t The first training experience at time step is placed into the first experience replay pool, wherein the first... t The first training experience at any given moment includes: t Time-based send and receive strategies, the first t The first reward of the moment, the first t Time and the The first state information at time t, the first The time is the first t The moment after the moment; According to the first t The transmit / receive strategy at each time step is used to train the frequency decision model after the (i-1)th training round to obtain the frequency decision model after the i-th training round. Determine whether the number of first training experiences in the first experience replay pool is greater than the first preset experience threshold; If the number of the first training experiences is not greater than the first preset experience threshold, then let t=t +1, obtain the first [number] of the multi-transmitter multi-receiver radar system. t The first state information at time step [1] is used to continue accumulating the first training experience, wherein the [1]th [time step]... t +1 time and the first t The time difference is equal to the coherent processing interval of the multiple-transmitter multiple-receiver radar system. If the number of the first training experience items is greater than the first preset experience threshold, then the first training experience data is sampled from the first experience replay pool, and based on the first update model, the network parameters of the radar transceiver station decision model after the (i-1)th round of training are updated according to the sampled first training experience data to obtain the radar transceiver station decision model after the i-th round of training. Determine whether the training stop requirement is met. If it is met, output the radar transceiver station decision model after the i-th round of training and the frequency point decision model trained in the i-th round as the trained radar transceiver station decision model and the trained frequency point decision model, respectively.
Citation Information
Patent Citations
Frequency agility multi-radar cooperative anti-interference method based on reinforcement learning
CN116125397A
Realtime electronic countermeasure optimization
US20220163627A1