Multi-transmitting multi-receiving radar networking cooperative anti-interference method based on nested reinforcement learning

By applying nested reinforcement learning algorithms in multi-radar systems to generate anti-jamming strategies, the problem of insufficient anti-jamming adaptability and coordination capabilities of multi-radar systems in complex scenarios is solved, and the anti-jamming performance of the system is significantly improved.

CN120178173AActive Publication Date: 2025-06-20XIDIAN UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510367698.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-06-20
Estimated Expiration
2045-03-26

AI Technical Summary

Technical Problem

When facing complex and changing practical scenarios, the anti-interference method has poor adaptability and coordination capabilities.

Method used

The multi-send and multi-receive radar networking collaborative anti-jamming method is adopted based on nested reinforcement learning. By obtaining the status information of the multi-radar system, inputting it into the trained decision model, generating a transmitting and receiving strategy and frequency point strategy, and then controlling the signal transmitted by the radar system to achieve anti-jamming.

Benefits of technology

It significantly improves the anti-interference capability of the multi-radar system in complex interference environments, enhances the overall anti-interference performance of the system and the coordinated performance of each transceiver and receiver station.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120178173A_ABST
    Figure CN120178173A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-transmitting and multi-receiving radar networking cooperative anti-interference method based on nested reinforcement learning. The method comprises the following steps: acquiring first state information and second state information of a multi-transmitting and multi-receiving radar system at the current moment; inputting the first state information of the multi-input multi-output radar system at the current moment into the trained radar transmit-receive station decision model to obtain a transmit-receive strategy of the multi-input multi-output radar system; inputting the second state information of the multi-input multi-output radar system at the current moment and the receiving and transmitting strategy into the trained radar frequency point decision model to obtain a frequency point strategy of the multi-input multi-output radar system; the trained radar transceiver station decision model and the trained frequency point decision model are obtained by performing nested reinforcement learning based on a single-agent reinforcement learning algorithm and a multi-agent reinforcement learning algorithm; and controlling the multiple-input multiple-output radar system to transmit signals according to the anti-interference strategy. The method can adapt to complex scenes and improve the cooperative anti-interference capability among radars.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of radar technology, and in particular relates to a multi-transmit and multi-receive radar networking collaborative anti-interference method based on nested reinforcement learning. Background Art

[0002] With the development of jamming technology, in order to effectively counter intelligent jammers with learning capabilities and autonomous adjustment strategies, radars must fully perceive and analyze the external jamming environment, and build flexible anti-jamming strategies by combining multi-dimensional information such as time, frequency, and space. Through adaptive learning, radars can adjust their working modes in real time and dynamically optimize anti-jamming measures, thereby maintaining a competitive advantage in a complex and changing environment. Compared with traditional anti-jamming methods that rely on fixed rules or expert systems, intelligent adaptive anti-jamming strategies can continuously optimize their own decisions based on the behavior patterns of jammers, making radars more adaptable and robust in long-term confrontations.

[0003] However, in the scenario of a single radar against multiple intelligent jammers, the radar often faces serious problems of limited resources and insufficient decision-making freedom. A single radar is not only limited in perception range, power allocation and strategy adjustment, but also in the face of multiple jammers operating in coordination, its single anti-interference strategy often fails to achieve the desired effect. Therefore, the use of a multi-radar system for coordinated anti-interference becomes a feasible solution. The multi-radar system consists of multiple geographically dispersed radar stations, which can significantly enhance the overall anti-interference capability of the system through information sharing, resource complementarity and strategy coordination.

[0004] However, the current anti-interference methods of multi-radar systems cannot adapt to complex and changeable actual scenarios, and the coordination ability between radars is poor. Summary of the invention

[0005] The embodiment of the present invention provides a multi-transmitter and multi-receiver radar network collaborative anti-interference method based on nested reinforcement learning, which can solve the problem of poor adaptability and coordination ability of the current anti-interference method of the multi-radar system.

[0006] In a first aspect, an embodiment of the present invention provides a multi-transmit and multi-receive radar network collaborative anti-interference method based on nested reinforcement learning, the method comprising:

[0007] Acquire first state information and second state information of a multi-transmitter and multi-receiver radar system at a current moment, wherein the multi-transmitter and multi-receiver radar system includes a plurality of transmitting stations and a plurality of receiving stations;

[0008] Input the first state information of the multi-transmitter and multi-receiver radar system at the current moment into the trained radar transceiver station decision model to obtain the transceiver strategy of the multi-transmitter and multi-receiver radar system, where the transceiver strategy is used to indicate the transmitter station for transmitting signals by the multi-transmitter and multi-receiver radar system at the current moment;

[0009] Input the second state information of the multi-transmitter and multi-receiver radar system at the current moment and the transceiver strategy into the trained radar frequency point decision model to obtain the frequency point strategy of the multi-transmitter and multi-receiver radar system;

[0010] Among them, the frequency point strategy is used to indicate the frequency point of the transmitted signal, and the trained radar transceiver station decision model and the trained frequency point decision model are obtained by nested reinforcement learning based on the single-agent reinforcement learning algorithm and the multi-agent reinforcement learning algorithm respectively;

[0011] Control the multi-transmitter and multi-receiver radar system to transmit signals according to the anti-jamming strategy, where the anti-jamming strategy is composed of the transceiver strategy and the frequency point strategy.

[0012] In a second aspect, an embodiment of the present invention provides a multi-transmitter and multi-receiver radar networking collaborative anti-jamming device based on nested reinforcement learning, including:

[0013] An acquisition unit, where the acquisition unit is used to acquire the first state information and the second state information of the multi-transmitter and multi-receiver radar system, where the multi-transmitter and multi-receiver radar system includes a plurality of transmitter stations and a plurality of receiver stations;

[0014] A first decision-making unit, where the first decision-making unit is used to input the first state information of the multi-transmitter and multi-receiver radar system at the current moment into the trained radar transceiver station decision model to obtain the transceiver strategy of the multi-transmitter and multi-receiver radar system, where the transceiver strategy is used to indicate the receiver station for receiving signals and the transmitter station for transmitting signals by the multi-transmitter and multi-receiver radar system at the current moment;

[0015] A second decision-making unit, where the second decision-making unit is used to input the second state information of the multi-transmitter and multi-receiver radar system at the current moment and the transceiver strategy into the trained radar frequency point decision model to obtain the frequency point strategy of the multi-transmitter and multi-receiver radar system;

[0016] Among them, the frequency point strategy is used to indicate the frequency point of the transmitted signal, and the trained radar transceiver station decision model and the trained frequency point decision model are obtained by nested reinforcement learning based on the single-agent reinforcement learning algorithm and the multi-agent reinforcement learning algorithm respectively;

[0017] A control unit, which is used to control the signal transmission of the multiple-input multiple-output radar system according to an anti-interference strategy, where the anti-interference strategy is composed of the transceiver strategy and the frequency point strategy.

[0018] The beneficial effects of the embodiments of the present invention compared with the prior art are as follows: Since the method provided by the present invention trains the frequency point decision model based on the multi-agent reinforcement learning algorithm, compared with a single radar system, it shows significant advantages in dealing with complex interference environments; and through resource complementarity and strategy coordination, multiple radar stations can form a more effective anti-interference network, improving the collaborative performance of each transceiver station in the system. Description of the Drawings

[0019] Figure 1 It is a schematic structural diagram of a multiple-input multiple-output radar networking collaborative anti-interference device provided by an embodiment of the present invention;

[0020] Figure 2 It is a flowchart of the implementation of a training method for a radar transceiver station decision model and a frequency point decision model provided by an embodiment of the present invention;

[0021] Figure 3 It is a flowchart of the implementation of a method for training a frequency point decision model according to the transceiver strategy at the t-th moment provided by an embodiment of the present invention;

[0022] Figure 4 It is a flowchart of the implementation of a multiple-input multiple-output radar networking collaborative anti-interference method based on nested reinforcement learning provided by an embodiment of the present invention;

[0023] Figures 5a - 5f It shows a schematic diagram of the curve of the detection probability varying with the number of training times at the end of each CPI obtained from the simulation experiment of the present invention;

[0024] Figure 6 It is a schematic diagram of the strategy of the radar transceiver station varying with training obtained from the simulation experiment of the present invention;

[0025] Figure 7a and Figure 7b It is a frequency point diagram obtained from the simulation experiment of the present invention. Detailed Embodiments

[0026] In the following description, specific details such as specific system structures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present invention. However, those skilled in the art should clearly understand that the present invention can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present invention.

[0027] It should be understood that, as used in the specification of the present invention and the appended claims, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not preclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.

[0028] It should also be understood that the term "and / or" as used in the specification of the present invention and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0029] As used in the specification of the present invention and the appended claims, the term "if" can be interpreted, depending on the context, as "when", "once", "in response to determining", or "in response to detecting". Similarly, the phrases "if determined" or "if [the described condition or event] is detected" can be interpreted, depending on the context, as meaning "once determined", "in response to determining", "once [the described condition or event] is detected", or "in response to detecting [the described condition or event]".

[0030] In addition, in the description of the specification of the present invention and the appended claims, the terms "first", "second", "third", etc. are only used for differentiating descriptions and cannot be construed as indicating or implying relative importance.

[0031] Reference to "one embodiment" or "some embodiments" or the like described in the specification of the present invention means that a specific feature, structure, or characteristic described in connection with that embodiment is included in one or more embodiments of the present invention. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification do not necessarily all refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "comprising", "including", "having", and their variants all mean "including but not limited to", unless otherwise specifically emphasized in other ways.

[0032] The following further describes the present invention in detail with reference to specific embodiments, but the embodiments of the present invention are not limited thereto.

[0033] The method for collaborative anti-jamming of multi-transmitter multi-receiver radar networking based on nested reinforcement learning provided by the embodiments of the present invention can be applied to electronic devices such as mobile terminals, personal laptop computers, supercomputers, etc. The embodiments of the present invention do not impose any restrictions on the specific types of electronic devices.

[0034] Embodiment 1

[0035] Figure 1The following is a schematic structural diagram of a multi-transmitter and multi-receiver radar networking collaborative anti-jamming device provided by an embodiment of the present invention. By way of example and not limitation, device 100 may include an acquisition unit 110, a first decision-making unit 120, a second decision-making unit 130, and a control unit 140.

[0036] Exemplarily, the acquisition unit 110 may acquire first state information and second state information of the multi-transmitter and multi-receiver radar system at the current moment. The first decision-making unit 120 may input the first information at the current moment into the trained radar transceiver station decision-making model to obtain a transceiver strategy for the multi-transmitter and multi-receiver radar system; wherein, the transceiver strategy is used to indicate the transmitter station for transmitting signals by the multi-transmitter and multi-receiver radar system at the current moment; the second decision-making unit 130 inputs the second state information and the transceiver strategy of the multi-transmitter and multi-receiver radar system at the current moment into the trained radar frequency point decision-making model to obtain a frequency point strategy for the multi-transmitter and multi-receiver radar system; wherein, the frequency point strategy is used to indicate the frequency point of the transmitted signal, and the trained radar transceiver station decision-making model and the trained radar frequency point decision-making model are respectively obtained through nested reinforcement learning based on a single-agent reinforcement learning algorithm and a multi-agent reinforcement learning algorithm; the control unit 140 controls the multi-transmitter and multi-receiver radar system to transmit signals according to the anti-jamming strategy, wherein the anti-jamming strategy is composed of the transceiver strategy and the frequency point strategy.

[0037] Since the device provided by the present invention trains the frequency point decision-making model based on the multi-agent reinforcement learning algorithm, compared with a single radar system, it shows significant advantages in dealing with complex interference environments; and through resource complementarity and strategy coordination, multiple radar stations can form a more effective anti-jamming network, improving the collaborative performance of each transceiver station in the system.

[0038] Embodiment 2

[0039] As an example, since the radar transceiver station decision-making model is used to determine the transmitter station enabled at the current moment, the action space of the radar transceiver station decision-making model may include the identifiers of each transmitter station in the radar system, such as the transmitter station number. The radar transceiver station decision-making model may select the identifiers of multiple transmitter stations from the action space as the transceiver strategy to enable the transmitter stations indicated by these identifiers to transmit signals.

[0040] Exemplarily, the action space of the radar transceiver station decision-making model may be expressed as:

[0041]

[0042] wherein, is the transceiver strategy at the T-th moment, I is the number of enabled transmitter stations, and H is the total number of radar transmitter stations.

[0043] In one example, the state space of the radar transceiver station decision model can be composed of the real-time position information (such as real-time two-dimensional coordinates) of each radar in the radar system. Then, the first state information can include the real-time position information of each radar in the radar system. The transceiver strategy generated by the radar transceiver station decision model can be used to indicate the transmitting station enabled at the current moment.

[0044] Exemplarily, the state space of the radar transceiver station decision model can be expressed as:

[0045]

[0046] Wherein, is the first state information at the T-th moment, x h , y h are respectively the coordinate values of the h-th radar at the T-th moment in the x-axis direction and the y-axis direction, and h is a positive integer less than or equal to H.

[0047] Specifically, the motion state of each radar in the radar system can be updated through the following dynamic model:

[0048]

[0049] Wherein, are respectively the coordinate values of the h-th radar at the (T + 1)-th moment in the x-axis direction and the y-axis direction, are respectively the coordinate values of the h-th radar at the T-th moment in the x-axis direction and the y-axis direction, v h is the speed of the h-th radar, θ h is the motion direction angle of the h-th radar, and ΔT is the time step.

[0050] Figure 2 The figure shows a flowchart of the implementation of a training method for a radar transceiver station decision model and a frequency point decision model provided by an embodiment of the present invention. By way of example and not limitation, the method may include steps S201 - S207, which will be described below for each step.

[0051] S201, obtain the first state information of the multiple-input multiple-output radar system at the t-th moment.

[0052] Exemplarily, the t-th moment can be superimposed starting from the starting moment of the i-th round of training.

[0053] Exemplarily, the first state information at the t-th moment used during training can be sampled from the state space of the radar transceiver station decision model.

[0054] S202. Input the first state information at the t-th moment into the radar transceiver decision model after the (i - 1)-th round of training to obtain the first training experience at the t-th moment; and put the first training experience at the t-th moment into the first experience replay pool.

[0055] In some embodiments, the radar transceiver decision model can be an actor-critic architecture, including a first policy network and a first value network.

[0056] In a possible implementation, the first training experience can include the transceiver strategy at the t-th moment (i.e., the action to be executed by the radar system at the t-th moment), the first reward at the t-th moment, and the first state information at the t-th moment and the t'-th moment.

[0057] Exemplarily, the t'-th moment is the next moment of the t-th moment. For example, if the t'-th moment is 3:05:49 am on February 14, 2025, then the t'-th moment can be 3:05:50 am on February 14, 2025.

[0058] In an example, the first policy network can select an action from the action space as the transceiver strategy at the t-th moment according to the first state information at the t-th moment.

[0059] Exemplarily, the transceiver strategy at the t-th moment can be expressed as:

[0060]

[0061] where, is the transceiver strategy at the t-th moment, θ j is the current latest network parameter of the first policy network π s (here it is the network parameter after the (i - 1)-th round of training), represents the first state information at the t-th moment.

[0062] In an example, the first reward at the t-th moment can be the joint detection probability Pd total of the radar system at the t-th moment, that is

[0063] S203. Perform the i-th round of training on the frequency point decision model after the (i - 1)-th round of training according to the transceiver strategy at the t-th moment to obtain the frequency point decision model after the i-th round of training.

[0064] Exemplarily, the frequency point decision model can be trained under the transceiver strategy at the t-th moment so that the frequency point decision network can learn the experience in this scenario.

[0065] S204. Determine whether the number of the first training experiences in the first experience replay pool is greater than the first preset experience threshold.

[0066] In one example, if the number of first training experiences in the first experience replay pool is not greater than the first preset experience threshold, then t can be set to t + 1, and the experience accumulation can continue from step S201.

[0067] Specifically, there are two time concepts in reinforcement learning, namely episode and time step. An episode is a relatively large time scale representing the complete process of a task. It contains multiple steps. A time step is a small unit within an episode, representing each interaction between the agent and the environment. In the radar operating mode, the situation of a coherent processing interval (CPI) is often concerned. A CPI contains multiple pulses. If each pulse within the CPI is understood as a time step in reinforcement learning, then a CPI can correspond to an episode in reinforcement learning. And in reinforcement learning, the training of the agent is multi-episode, so it is equivalent to multiple CPIs in the radar. Therefore, in the radar transceiver station decision model, a CPI can correspond to a time step, that is, the time difference T1 between the (t + 1)-th moment and the t-th moment is equal to a CPI; in the frequency point decision model, a CPI can correspond to an episode.

[0068] In another example, if the number of first training experiences in the first experience replay pool is greater than the first preset experience threshold, then the following step S205 can be performed.

[0069] S205. Sample the first training experience data from the first experience replay pool, and based on the first update model, update the network parameters of the radar transceiver station decision model after the (i - 1)-th round of training according to the sampled first training experience data, to obtain the radar transceiver station decision model after the i-th round of training.

[0070] In one example, if the radar transceiver station decision model can update the network parameters based on the Proximal Policy Optimization (PPO) algorithm, then the first update model can satisfy the following formula:

[0071]

[0072] where L PPO (θ j ) is the loss function of the first policy network, θ j is the network parameter after the update of the first policy network, denotes expectation, is the transceiver policy generated by the first policy network according to the first state information at the t-th moment j when the parameter is θ, and represents the probability that the first policy network selects the action when the state is ; denotes the transceiver policy generated by the first policy network when the parameter is at and is the network parameter before the update of the first policy network, A t is the generalized advantage estimation, clip(·) represents the truncation function, and ∈ is the clipping threshold; is the loss function of the first value network, and θ m is the network parameter after the update of the first value network, represents taking the expectation, denotes the first value network with parameter θ m obtained at state and is the first cumulative return.

[0073] Among them:

[0074]

[0075] Among them, γ is the discount factor, is the first reward at the (t + j)-th moment.

[0076] S206, determine whether the training stop requirement is satisfied.

[0077] In one example, if the training stop condition is not satisfied, then i = i + 1 can be set to perform the next round of training.

[0078] Exemplarily, it can be determined whether the training stop condition is satisfied according to the size of the training round and the performance of the two models after the i-th round of training.

[0079] In another example, if the training stop condition is satisfied, then step S207 can be performed.

[0080] S207, output the radar transceiver station decision model after the i-th round of training and the frequency point decision model trained in the i-th round as the trained radar transceiver station decision model and the trained frequency point decision model respectively.

[0081] By performing nested reinforcement learning on the two decision models, the frequency point decision model can learn the experience under different transceiver policies and different environments, and the radar transceiver station decision model can learn the experience under different environments; thus enabling these two models to make the collaborative strategy of each transceiver station in the multi-transmit and multi-receive radar system in a complex actual environment, and improving the anti-interference ability of the radar system.

[0082] On the basis of Embodiment 2, the present invention also provides Embodiment 3.

[0083] As an example, the transmitted signal of the radar can be a frequency hopping sequence, and the transmitted signal contains several (e.g., K) sub-pulses within one pulse. Since the frequency point decision model is used to determine the frequency point of the transmitted signal, the action space of the frequency point decision model can be composed of the frequency point numbers of each sub-pulse. The user can pre-set multiple (e.g., N) operating frequency points for the radar system and number them. When making a decision, the frequency point decision model can select K operating frequency point numbers from its action space to set the operating frequency points corresponding to these numbers as the frequency points of the corresponding sub-pulses.

[0084] Exemplarily, the action space of the frequency point decision model can be expressed as:

[0085]

[0086] Wherein, is the frequency point strategy at the k-th moment, 0 ≤ a i ≤ N - 1, indicating that the carrier frequency of the i-th sub-pulse is K is the number of sub-pulses, and N is the total number of operating frequency points.

[0087] Optionally, dividing one pulse into multiple sub-pulses can make the radar have higher degrees of freedom; however, with the increase in the frequency points and the number of sub-pulses of the radar, when using a discrete action space, the dimension of the action space will be too large, resulting in a decrease in training efficiency and an increase in computational cost. Therefore, when the frequency points and the number of sub-pulses are large, the method of discretizing the continuous action space can be used to convert the continuous frequency point information of the K sub-pulses output by the frequency point decision model into discrete frequency point information.

[0088] Specifically, the expression of the discrete frequency point information can be:

[0089]

[0090] Wherein, a d represents the discrete action representation of the radar frequency point, a c represents the continuous action representation of the radar frequency point, tanh is the hyperbolic tangent function, and floor is the floor function.

[0091] In some embodiments, the state space of the frequency point decision model can be composed of the observation information of the radar system (e.g., the observation matrix), then the second state information can include the observation information of the radar system.

[0092] Exemplarily, the state space of the frequency point decision model can be expressed as:

[0093]

[0094] Wherein, represents the second state information of the radar system at the k-th moment, is the observation matrix obtained from the transmission signal of the h'-th transmitting station enabled at the (k - m)th moment, where m is a positive integer less than or equal to M, and h' is a positive integer less than or equal to I.

[0095] Exemplarily, the observation matrix of the radar system at the k-th moment can be expressed as:

[0096]

[0097] where, is the observation of the radar y at the k-th moment represented by the observation matrix of the radar system at the k-th moment.

[0098] Taking as an example, it satisfies:

[0099]

[0100] where, is the radar information at the K-th sub-pulse and the N-th frequency point of the radar, and a K represents the frequency point transmitted by the radar in sub-pulse K; α q is the sidelobe attenuation coefficient of the q-th jammer. When the q-th jammer is a main lobe jammer to the radar y, its value is 1. When it is a sidelobe jammer, it is attenuated according to the antenna pattern coefficient. q is a positive integer less than or equal to Q; is the interference intensity of the q-th jammer on the N-th frequency point at the K-th sub-pulse of the radar.

[0101] In a possible implementation, the interference action of the jammer on the radar system can be represented by an interference matrix, and the interference matrix can represent the interference intensity of the jammer acting on each frequency point at different times.

[0102] Exemplarily, the interference matrix can be expressed as:

[0103]

[0104] where, is the interference matrix of the q-th jammer at the k-th moment. is the interference intensity of the q-th jammer on the n-th frequency point at the i'-th sub-pulse of the radar, and usually takes values in [0, 1]. When it means that the q-th jammer is in the state of receiving signals at this time.

[0105] In one example, the jamming strategies of the jammer can be mainly divided into two types: frequency-point jamming and intermittent sampling jamming. Frequency-point jamming refers to the jammer performing precise noise jamming on the radar frequency. Usually, by adjusting the frequency of the jamming signal, it is made to be close to or coincide with the operating frequency of the radar. Intermittent sampling jamming is a false target jamming that samples at specific moments and uses the sampled signal as the jamming signal. Generally, its operating mode is to receive for x sub-pulse periods and transmit for y sub-pulse periods. Receiving for x means that the jammer will select to sample the signal within x sub-pulse periods, and transmitting for y means that the jammer will transmit the jamming signal within y sub-pulse periods.

[0106] Exemplarily, necessary working isolation can be carried out between the receiving jammer and the transmitting jammer working on the same platform. Specifically, high transceiver isolation can be used to achieve intermittent observation. For example, interrupting the transmission of the jamming signal so that the receiver can effectively intercept the signal.

[0107] It should be understood that the state space and action space of the frequency-point decision model should be constructed together with the state space and action space of the radar transceiver station decision model before the first training.

[0108] Figure 3 The figure shows a flowchart of an implementation of a method for training a frequency-point decision model according to the transceiver strategy at the t-th moment provided by an embodiment of the present invention. By way of example and not limitation, this method can be a specific possible implementation manner of the above step S203. This method can include steps S301 - S304, and each step will be described below.

[0109] S301, obtain the second state information of the multiple-input multiple-output radar system at the k-th moment.

[0110] Exemplarily, k can be incremented starting from the t-th moment.

[0111] S302, input the second state information at the k-th moment and the transceiver strategy at the t-th moment into the frequency-point decision model after the (i - 1)-th training to obtain the second training experience at the k-th moment.

[0112] In some embodiments, the frequency-point decision model can be an actor-critic architecture, including a second policy network and a second value network.

[0113] In a possible implementation manner, the second training experience can include the frequency-point strategy at the k-th moment, the second reward at the k-th moment, and the second state information at the k-th moment and the k'-th moment.

[0114] Similarly, the k'-th moment is the next moment of the k-th moment.

[0115] In one example, the first policy network can select an action as the frequency point policy at the k-th moment from the action space of the frequency point decision network according to the second state information at the k-th moment and the transceiver policy at the t-th moment.

[0116] Exemplarily, the frequency point policy at the k-th moment can be expressed as:

[0117]

[0118] where is the frequency point policy at the t-th moment, and θ i is the latest network parameter of the second policy network π f at the current moment (here it is the network parameter after the (i - 1)-th round of training), represents the first state information at the k-th moment.

[0119] In one example, the second reward at the k-th moment can be the uncovered rate of the radar system at the k-th moment, which can satisfy the following formula:

[0120]

[0121] where is the second reward at the k-th moment, num uncovered is the number of sub-pulses not interfered in one pulse, and K is the number of sub-pulses in one pulse.

[0122] S303, Determine whether the number of the second training experiences in the second experience replay pool is greater than the second preset experience threshold.

[0123] In one example, if the number of the second training experiences in the second experience replay pool is not greater than the second preset experience threshold, then k = k + 1 can be set, and the accumulation of the second training experiences continues from step S301.

[0124] Exemplarily, the time difference T2 between the (k + 1)-th moment and the k-th moment is less than T1.

[0125] In another example, if the number of the second training experiences in the second experience replay pool is greater than the second preset experience threshold, then step S304 can be performed.

[0126] S304, Sample the second training experience data from the second experience replay pool, and based on the second update model, update the network parameters of the frequency point decision model after the (i - 1)-th round of training according to the sampled second training experience data to obtain the frequency point decision model after the i-th round of training.

[0127] In one example, the frequency point decision model can update network parameters based on the Multi-Agent Proximal Policy Optimization (MAPPO) algorithm, and the second update model can satisfy the following formula:

[0128]

[0129] Where L MAPPO (θ i ) is the loss function of the second policy network, and θ i is the network parameter after the update of the second policy network. represents the expectation. is the transceiver policy generated by the second policy network according to the second state information at the k-th moment when the parameter is θ i , and represents the probability that the second policy network selects the action when the state is ; represents the transceiver policy generated by the second policy network according to when the parameter is . is the network parameter before the update of the second policy network, A is the Generalized Advantage Estimation, and clip(·) represents the clipping function, and ∈ is the clipping threshold; k is the loss function of the second value network, and θ is the network parameter after the update of the second value network. n represents the expectation. represents the value obtained by the second value network with the parameter θ in the state n , and is the second cumulative return.

[0130] Exemplarily, the clipping function can limit the result to the interval [1 - ∈, 1 + ∈], which determines the tolerance of the ratio change during each policy update.

[0131] Exemplarily, the Generalized Advantage Estimation (GAE) can measure the improvement degree of taking the action in the state compared to the baseline value in the state . The generalized state estimation of the frequency point decision model can be calculated by the following formula:

[0132]

[0133] Among them, γ is the discount factor, and λ is the parameter of GAE, which is used to balance the trade-off between bias and variance; δ k is the temporal difference error, which is the gap between the agent's policy and the actual reward in the current state, and measures the error between the performance of the current policy in the current state and the expectation. is the state at the k-th moment value estimate of is based on the state at the k'-th moment calculated discounted future reward.

[0134] Similarly, the second cumulative return at the k-th moment can satisfy: is the second reward at the k + j'-th moment.

[0135] Since the method provided by the present invention trains the frequency point decision model based on the multi-agent reinforcement learning algorithm, compared with the single radar system, it shows significant advantages in dealing with complex interference environments; and through resource complementarity and policy coordination, multiple radar stations can form a more effective anti-interference network, improving the coordination performance of each transceiver station in the system.

[0136] Furthermore, the present invention constructs the state space of the radar transceiver station decision model based on the radar networking system in the motion state, considering the influence of the radar position change on the radar-jammer confrontation, and can improve the anti-interference performance of the decision-making.

[0137] Moreover, based on the multi-agent reinforcement learning algorithm framework, the present invention can make decisions when multiple radars and multiple jammers do not fully understand each other's information. Multi-agent reinforcement learning is a solution to sequential decision-making with missing prior information.

[0138] Based on Embodiment 2 and Embodiment 3, the present invention also provides Embodiment 4.

[0139] Figure 4 The following shows the implementation flowchart of a multi-transmitter multi-receiver radar networking cooperative anti-interference method based on nested reinforcement learning provided by an embodiment of the present invention. By way of example and not limitation, the method may include steps S401 - S404, and the following is an explanation of each step.

[0140] S401, obtain the first state information and the second state information of the multi-transmitter multi-receiver radar system at the current moment.

[0141] Exemplarily, the multi-transmitter multi-receiver radar system may include a transmitter station and multiple receiver stations.

[0142] S402. Input the first state information of the multi-transmitter and multi-receiver radar system at the current moment into the trained radar transceiver station decision model to obtain the transceiver strategy of the multi-transmitter and multi-receiver radar system.

[0143] Exemplarily, the transceiver strategy is used to indicate the transmitter station for transmitting signals by the radar system at the current moment.

[0144] S403. Input the second state information and the transceiver strategy of the multi-transmitter and multi-receiver radar system at the current moment into the trained radar frequency point decision model to obtain the frequency point strategy of the multi-transmitter and multi-receiver radar system.

[0145] Exemplarily, the frequency point strategy is used to indicate the frequency point of the transmitted signal.

[0146] Exemplarily, the trained radar transceiver station decision model and the trained frequency point decision model can be trained according to the methods described in Embodiment 2 and Embodiment 3 above.

[0147] S404. Control the multi-transmitter and multi-receiver radar system to transmit signals according to the anti-jamming strategy.

[0148] Exemplarily, the anti-jamming strategy of the radar system may include the transceiver strategy and the frequency point strategy.

[0149] Since the method provided by the present invention trains two decision models based on the multi-agent reinforcement learning algorithm, compared with a single radar system, it shows significant advantages in dealing with complex interference environments; and through resource complementarity and strategy coordination, multiple radar stations can form a more effective anti-jamming network, improving the coordination performance of each transceiver station in the system.

[0150] To better illustrate the beneficial effects of the present invention, the following simulation experiments were carried out.

[0151] Simulation conditions

[0152] In the experiment, the total number of radars is set to 4, the radar configuration is two transmitters and four receivers, the number of available radar frequency points is 10, one pulse contains 4 sub-pulses, and one CPI contains 32 pulses; the radar power is 30 kw, the transmitting antenna gain is 30 db, the receiving antenna gain is 30 db, and the target RCS is 15 m 2 , the radar receiver bandwidth is 25 Mh; the total number of jammers is 4, and the main lobe dry ratio is 30 db.

[0153] The positions of the jammers and the target are fixed at [100 km, 100 km]. The positions, speeds, and orientation angles of the four radars are shown in Table 1:

[0154] Table 1

[0155] Radar number Coordinates (km) Speed (m / s) Orientation angle Radar 0 [110,0] 300 96° Radar 1 [10,0] 800 48° Radar 2 [0,0] 1000 45° Radar 3 [0,150] 500 -27°

[0156] Specifically, the action space of the radar transceiver station is finite and discrete, and the encoded action can be obtained by encoding the combination of radar numbers. Its action space is as follows:

[0157] a s ∈{0, 1, 2, 3, 4, 5}

[0158] Among them, the a s represents the transceiver station decision action of the radar at time t. The four radars are numbered as Radar 0, Radar 1, Radar 2, and Radar 3 respectively. The mapping relationship between the transceiver station decision action and the radar transmitter station number is shown in Table 2.

[0159] Table 2

[0160] Action 0 1 2 3 4 5 Transmitting radar number [0,1] [0,2] [0,3] [1,2] [1,3] [2,3]

[0161] Contents of the simulation experiment

[0162] Under the simulation conditions, training is carried out once every 32 CPIs. Each CPI has a length of 32 and a total of 1024 pulses, corresponding to MAPPO sampling 1024 pieces of data and PPO sampling 32 pieces of data. Each pulse is used as a time step. When 300,000 time steps are reached, the model parameters are saved, and the simulation results are as follows:

[0163] Figures 5a to 5f The curve shows the detection probability at the end of each CPI obtained from the simulation experiment varying with the number of training times. The jamming types of the two jammers are both intermittent sampling jamming, but the transceiver times are different. Among them Figure 5a the transceiver times of the two jammers are receive-transmit-receive-transmit and receive-transmit-receive-transmit respectively, Figure 5b are receive-transmit-receive-transmit and receive-transmit-receive-two respectively, Figure 5c are receive-transmit-receive-two and receive-two-transmit-four respectively, Figure 5d are receive-transmit-receive-transmit and receive-transmit-receive-three respectively, Figure 5e are receive-transmit-receive-transmit and receive-transmit-receive-four respectively, Figure 5f are receive-transmit-receive-four and receive-transmit-receive-four respectively. It can be seen from Figure 5 that as the number of training times increases, the detection probability generally converges between 0.9 and 1, indicating that the radar network collaborative anti-jamming ability is good.

[0164] Figure 6 The figure shows the schematic diagram of the radar transceiver station strategy changing with training. The x-axis is the number of training steps, and the y-axis is the transceiver station decision action. From Figure 6It can be seen that as the number of training times increases, the transceiver decision network tends to select the transmitting stations with the largest angular difference for the jammer. For the radar, since Radar 1 and Radar 2 are relatively close, if the main lobe width of the jammer can completely cover them, one jammer can strongly jam these two radars simultaneously. When two jammers jam Radar 0, their side lobes have little interference on Radar 3. It can be seen that the result output by the radar transceiver strategy network is reasonable.

[0165] Figure 7a and Figure 7b are the frequency point maps obtained by training Radar Transmitting Station 1 (corresponding to Radar 0, the jammer receives one and transmits to four) and Radar Transmitting Station 2 (corresponding to Radar 3, the jammer receives one and transmits to two), respectively. In the figure, the x-axis is the sub-pulse index, the y-axis is the radar transmitting frequency point index, the light blue line represents the jammer signal, the green line represents the radar signal not jammed, and the yellow line represents the jammed radar signal. It can be seen from Figure 7a that the radar has fully learned the jamming strategy of the jammer. When the jammer intercepts the signal, the radar transmits at frequency point 9, and when the jammer jams, the radar transmits at frequency point 1. Figure 7b It can be seen that all frequency points have successfully avoided jamming. And among the 20 pulses in Figure 7a , only one frequency point is jammed, indicating that the frequency point decision network can successfully learn the transmitting mode of the jammer and avoid jamming.

[0166] Therefore, through the method provided by the present invention, an effective radar anti-jamming strategy can be obtained, and the anti-jamming performance of the radar system can be improved. At the same time, the change of the moving radar in the spatial dimension is also considered during decision-making, further improving the anti-jamming performance of the radar system.

[0167] In the above embodiments, the descriptions of each embodiment have their own focuses. For the parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

Claims

1. A collaborative anti-interference method for multiple-transmit and multiple-receive radar networking based on nested reinforcement learning, characterized in that: include: Acquire first state information and second state information of a multi-transmitter and multi-receiver radar system at a current moment, wherein the multi-transmitter and multi-receiver radar system includes a plurality of transmitting stations and a plurality of receiving stations; Inputting the first state information of the multi-transmit multi-receive radar system at the current moment into the trained radar transceiver station decision model to obtain the transceiver strategy of the multi-transmit multi-receive radar system, wherein the transceiver strategy is used to indicate the transmitting station used by the multi-transmit multi-receive radar system to transmit signals at the current moment; Inputting the second state information of the multiple-transmit multiple-receive radar system at the current moment and the transmission and reception strategy into the trained radar frequency decision model to obtain the frequency strategy of the multiple-transmit multiple-receive radar system; The frequency strategy is used to indicate the frequency of the transmitted signal, and the trained radar transceiver station decision model and the trained frequency decision model are obtained by nested reinforcement learning based on a single-agent reinforcement learning algorithm and a multi-agent reinforcement learning algorithm respectively; The multi-transmit multi-receive radar system is controlled to transmit signals according to an anti-interference strategy, wherein the anti-interference strategy is composed of the transceiver strategy and the frequency strategy.

2. The method according to claim 1, characterized in that The training method of the radar transceiver station decision model and the frequency point decision model includes: Acquire first state information of the multiple-transmit multiple-receive radar system at time t; Input the first state information at the tth moment into the radar transceiver station decision model after the i-1th round of training to obtain the first training experience at the tth moment; and put the first training experience at the tth moment into the first experience replay pool, wherein the first training experience at the tth moment includes: the transceiver strategy at the tth moment, the first reward at the tth moment, the first state information at the tth moment and the t′th moment, and the t′th moment is the next moment of the tth moment; Performing an i-th round of training on the frequency decision model after the i-1-th round of training according to the receiving and sending strategy at the t-th moment, to obtain the frequency decision model after the i-th round of training; Determine whether the number of first training experiences in the first experience replay pool is greater than a first preset experience threshold; If the number of first training experiences in the first experience replay pool is not greater than the first preset experience threshold, set t=t+1, obtain the first state information of the multi-transmitter and multi-receiver radar system at the tth moment, and continue to accumulate the first training experience, wherein the time difference between the t+1th moment and the tth moment is equal to the coherent processing interval of the multi-transmitter and multi-receiver radar system; If the number of first training experiences in the first experience replay pool is greater than the first preset experience threshold, first training experience data is sampled from the first experience replay pool, and based on the first update model, the network parameters of the radar transceiver station decision model after the i-1th round of training are updated according to the sampled first training experience data to obtain the radar transceiver station decision model after the i-th round of training; Determine whether the training stop requirement is met, and if so, output the radar transceiver station decision model after the i-th round of training and the frequency point decision model trained in the i-th round as the trained radar transceiver station decision model and the trained frequency point decision model, respectively.

3. The method according to claim 2, characterized in that The first reward at the t-th moment is the joint detection probability of the multi-transmitter and multi-receiver radar system at the t-th moment.

4. The method according to claim 2, characterized in that: The radar transceiver station decision model includes a first strategy network and a first value network; The first update model satisfies the following formula: Among them, L PPO (θ j ) is the loss function of the first strategy network, θ j is the updated network parameter of the first strategy network, Express expectations, For the first policy network, the parameter is θ j According to the first state information at time t The generated sending and receiving strategy also indicates that the first strategy network is in state Select action probability; It means that the first policy network has parameters Time Based The generated sending and receiving strategies, is the network parameter before the first strategy network is updated, A t is the generalized advantage estimate of the first policy network, clip(·) represents the truncation function, ∈ is the clipping threshold; is the loss function of the first value network, θ m is the updated network parameter of the first value network, Expressing hope, Denote the parameter as θ m The first value network in the state The value obtained is The first cumulative return.

5. The method according to claim 2, characterized in that: The step of performing the i-th round of training on the frequency decision model after the i-1-th round of training according to the receiving and sending strategy at the t-th moment to obtain the frequency decision model after the i-th round of training includes: Acquire second state information of the multiple-transmit multiple-receive radar system at a kth moment; Inputting the second state information at the kth moment and the sending and receiving strategy at the tth moment into the frequency decision model after the i-1th round of training, obtaining the second training experience at the kth moment, wherein the second training experience at the kth moment includes: the frequency strategy at the kth moment, the second reward at the kth moment, and the second state information at the kth moment and the k′th moment, wherein the k′th moment is the next moment of the kth moment; Determine whether the number of second training experiences in the second experience replay pool is greater than a second preset experience threshold; If the number of second training experiences in the second experience replay pool is not greater than the second preset experience threshold, set k=k+1, and the time difference between the k+1th moment and the kth moment is less than the coherent processing interval; If the number of second training experiences in the second experience replay pool is greater than the second preset experience threshold, second training experience data is sampled from the second experience replay pool, and based on the second update model, the network parameters of the frequency decision model after the i-1th round of training are updated according to the sampled second training experience data to obtain the frequency decision model after the i-1th round of training.

6. The method according to claim 5, characterized in that The second reward at the kth moment satisfies the following formula: Among them, r k f is the second reward at the kth moment, num uncovered is the number of undisturbed sub-pulses in a pulse, and K is the number of sub-pulses in a pulse.

7. The method according to claim 5, characterized in that The frequency decision model includes a second strategy network and a second value network; Wherein, the second update model satisfies the following formula: Among them, L MAPPO (θ i ) is the loss function of the second policy network, θ i is the updated network parameter of the second strategy network, Express expectations, For the second policy network, the parameter is θ i According to the second state information at the kth moment The generated sending and receiving strategy also indicates that the second strategy network is in state Select action probability; It means that the second strategy network has parameters Time Based The generated sending and receiving strategies, is the network parameter before the second strategy network update, A k is the generalized advantage estimate of the second policy network, clip(·) represents the truncation function, ∈ is the clipping threshold; is the loss function of the second value network, θ n is the updated network parameter of the second value network, Expressing hope, Denote the parameter as θ n The second value network in the state The value obtained is is the second cumulative return.

8. The method according to any one of claims 1 to 7, characterized in that: The first status information includes real-time position information of each radar in the multiple-transmit multiple-receive radar system.

9. The method according to any one of claims 1 to 7, characterized in that: The second state information includes observation information of the MTMR radar system.

10. A multi-transmit and multi-receive radar network collaborative anti-interference device based on nested reinforcement learning, characterized in that: include: An acquisition unit, the acquisition unit is used to acquire first state information and second state information of a multi-transmitter and multi-receiver radar system, wherein the multi-transmitter and multi-receiver radar system includes a plurality of transmitting stations and a plurality of receiving stations; a first decision unit, the first decision unit being used to input first state information of the multiple-transmit multiple-receive radar system at a current moment into a trained radar transceiver station decision model to obtain a transceiver strategy of the multiple-transmit multiple-receive radar system, wherein the transceiver strategy is used to indicate a receiving station for receiving signals and a transmitting station for transmitting signals of the multiple-transmit multiple-receive radar system at a current moment; A second decision unit, the second decision unit is used to input the second state information of the multi-transmit multi-receive radar system at a current moment and the transceive strategy into the trained radar frequency decision model to obtain the frequency strategy of the multi-transmit multi-receive radar system; The frequency strategy is used to indicate the frequency of the transmitted signal, and the trained radar transceiver station decision model and the trained frequency decision model are obtained by nested reinforcement learning based on a single-agent reinforcement learning algorithm and a multi-agent reinforcement learning algorithm respectively; A control unit, wherein the control unit is used to control the multiple-transmit multiple-receive radar system to transmit signals according to an anti-interference strategy, wherein the anti-interference strategy is composed of the transmit-receive strategy and the frequency strategy.

Citation Information

Patent Citations

  • Frequency agility multi-radar cooperative anti-interference method based on reinforcement learning

    CN116125397A

  • Radar rapid anti-interference strategy acquisition method based on model independent element learning

    CN118035748A

  • Realtime electronic countermeasure optimization

    US20220163627A1