Electronic countermeasure offline decision-making method, system, device and storage medium

By constructing an adversary agent in radar jamming decision-making and alternately training the jammer agent, the non-optimality problems of traditional radar jamming decision-making methods and the high cost of online learning are solved, and efficient and safe jamming decision-making is achieved in complex electromagnetic environments.

CN119830989BActive Publication Date: 2025-09-16UNIV OF SCI & TECH OF CHINA +1

Patent Information

Application Number
CN202411893456.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-20
Publication Date
2025-09-16
Estimated Expiration
2044-12-20

AI Technical Summary

Technical Problem

In existing technologies, radar jamming decision-making methods rely on experience or templates, which cannot guarantee that the selected jamming pattern is optimal. In addition, online reinforcement learning is costly and poses security risks in complex environments. There is no solution for applying offline reinforcement learning to radar jamming decision-making.

Method used

The jammer is modeled as a jammer agent, and the opponent agent is constructed through historical game sequences. The jammer and opponent agents are trained alternately, and the offline adversarial learning method is used to optimize the jamming strategy, reduce the number of interactions with the real environment, and improve the effectiveness and security of decision-making.

Benefits of technology

In complex electromagnetic environments, it provides efficient and safe radar jamming decision-making, ensures the performance lower bound of the jamming strategy in real environments, reduces the number of interactions with the environment, and improves the accuracy and adaptability of jamming decisions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119830989B_ABST
    Figure CN119830989B_ABST
Patent Text Reader

Abstract

The present invention discloses an electronic countermeasure offline decision-making method, system, device and storage medium. Due to the fact that in a complex electromagnetic environment, the radar state information obtained by the jammer is partially observable and has deviations, the jamming strategy obtained by the general decision-making method is often difficult to guarantee performance when deployed in a real environment due to its limitations. Therefore, the solution provided by the present invention models the electromagnetic environment into an opponent intelligent agent, and the jammer intelligent agent and the opponent intelligent agent are interactively trained to obtain the final jammer intelligent agent. This can better make radar jamming decisions, making the jammer's decision-making in a complex electromagnetic environment more efficient and safe.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of offline decision-making technology in electronic countermeasure scenarios, and in particular to an electronic countermeasure offline decision-making method, system, device and storage medium. Background Art

[0002] Radar jamming decision-making is a crucial component of radar countermeasures, and the jamming strategy output by the decision-making process directly impacts the effectiveness of the jamming. In practical applications, although jammers can obtain radar operating modes through reconnaissance and other means, multiple jamming patterns can be selected for the same radar operating mode. Traditional jamming decision-making methods rely on experience or templates to select jamming patterns, which cannot guarantee the optimal jamming pattern. Deep reinforcement learning, on the other hand, is a method that combines both perception and decision-making capabilities. Through deep reinforcement learning modeling, jammers can continuously learn from their experience during the radar countermeasures process, dynamically selecting and optimizing the optimal jamming pattern. This approach not only improves the accuracy and effectiveness of jamming decisions but also enhances the jamming system's adaptability in complex environments.

[0003] Online reinforcement learning models, which collect samples and learn policies through online trial-and-error interactions with the environment, are an effective approach for solving sequential perception-based decision-making problems. However, this online, interactive, active learning paradigm can lead to high costs and safety issues when collecting samples in complex real-world environments. Offline reinforcement learning, a data-driven reinforcement learning paradigm, emphasizes learning policies from static sample datasets without exploratory interaction with the environment, offering a viable solution for deploying reinforcement learning algorithms in real-world environments. However, no relevant solutions have been found for applying offline reinforcement learning to radar jamming decision-making. Summary of the Invention

[0004] The purpose of the present invention is to provide an electronic countermeasure offline decision-making method, system, device and storage medium, which can improve the effectiveness of the interference strategy and make the interference party's decision-making in a complex electromagnetic environment efficient and safe.

[0005] The purpose of the present invention is achieved through the following technical solutions:

[0006] An electronic countermeasure offline decision-making method, comprising:

[0007] Model the interferer as an interferer agent;

[0008] Take the given historical game sequence as the original dataset and use it to construct the opponent agent;

[0009] Alternately training the interfering agent and the opponent agent, the steps comprising: in a current iteration, using the opponent agent to perform data extrapolation, adding the generated data to an extended data set, and using the original data set and the extended data set to train the interfering agent, wherein the training objective is to maximize the interference value function when the interfering agent and the opponent agent compete; training the opponent agent according to an offline adversarial learning method combined with training the interfering agent trained in the current iteration, wherein the training objective is to minimize the interference value function when the interfering agent and the opponent agent compete; and continuously iterating to obtain a final interfering agent;

[0010] The final jammer agent is deployed in a real environment for radar jamming decision-making.

[0011] An electronic countermeasure offline decision-making system, comprising:

[0012] an interfering party agent modeling unit, configured to model the interfering party as an interfering party agent;

[0013] The opponent agent construction unit is used to take the given historical game sequence as the original data set and use it to construct the opponent agent;

[0014] An agent alternating training unit is configured to alternately train an interfering agent and an opponent agent, the steps comprising: in a current iteration, using the opponent agent to perform data extrapolation, adding the generated data to an extended data set, and using the original data set and the extended data set to train the interfering agent, with the training objective being to maximize the interference value function when the interfering agent and the opponent agent confront each other; training the opponent agent according to an offline adversarial learning method combined with training of the interfering agent trained in the current iteration, with the training objective being to minimize the interference value function when the interfering agent and the opponent agent confront each other; and continuously iterating to obtain the final interfering agent;

[0015] The deployment application unit is used to deploy the final jammer agent in a real environment to make radar jamming decisions.

[0016] A processing device comprising: one or more processors; a memory for storing one or more programs;

[0017] When the one or more programs are executed by the one or more processors, the one or more processors implement the aforementioned method.

[0018] A readable storage medium stores a computer program, which implements the aforementioned method when the computer program is executed by a processor.

[0019] It can be seen from the technical solution provided by the present invention that in a complex electromagnetic environment, the radar status information obtained by the jammer is partially significant and biased. When the jamming strategy obtained by the general decision-making method is deployed in a real environment, due to its limitations, the performance is often difficult to guarantee. Therefore, the solution provided by the present invention, by modeling the electromagnetic environment into an opponent intelligent agent, the jammer intelligent agent and the opponent intelligent agent interactively train to obtain the final jammer intelligent agent, can better make radar jamming decisions, making the jammer's decision-making in a complex electromagnetic environment efficient and safe. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0021] Figure 1 A flowchart of an electronic countermeasure offline decision-making method provided by an embodiment of the present invention;

[0022] Figure 2 A flowchart of the agent training provided by an embodiment of the present invention;

[0023] Figure 3 A schematic diagram of an electronic countermeasure offline decision-making system provided by an embodiment of the present invention;

[0024] Figure 4 A schematic diagram of a processing device provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0025] The following is a clear and complete description of the technical solutions in the embodiments of the present invention, in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0026] First, the following terms may be used in this article:

[0027] The terms "include," "comprises," "contains," "has," or other similar expressions should be interpreted as non-exclusive. For example, "including certain technical features (such as raw materials, components, ingredients, carriers, dosage forms, materials, dimensions, parts, components, mechanisms, devices, steps, procedures, methods, reaction conditions, processing conditions, parameters, algorithms, signals, data, products, or manufactured articles, etc.) should be interpreted as including not only the technical features explicitly listed, but also other technical features known in the art that are not explicitly listed.

[0028] The term "consisting of" excludes any technical features not explicitly listed. If used in a claim, this term renders the claim closed, excluding any technical features other than those explicitly listed, except for conventional impurities associated with them. If this term appears only in a clause of a claim, it limits only the elements explicitly listed in that clause; elements listed in other clauses are not excluded from the claim as a whole.

[0029] The following describes in detail the electronic countermeasure offline decision-making method, system, device, and storage medium provided by the present invention. Any information not described in detail in the embodiments of the present invention is prior art known to those skilled in the art. For any conditions not specified in the embodiments of the present invention, the procedures were performed according to conventional conditions in the art or the conditions recommended by the manufacturer. For any reagents or instruments used in the embodiments of the present invention without manufacturer identification, they are all commercially available conventional products.

[0030] Example 1

[0031] The embodiment of the present invention provides an electronic countermeasure offline decision-making method, such as Figure 1 As shown, it mainly includes the following steps:

[0032] Step 1: Model the interferer as an interferer agent.

[0033] Step 2: Take the given historical game sequence as the original dataset and use it to construct the opponent agent.

[0034] Preferably, an adversary agent can be constructed based on the original data set through maximum likelihood estimation.

[0035] Preferably, the opponent agent represents the transfer function of the electromagnetic environment in the electronic countermeasure (training process), and its optimization goal is to minimize the benefit of the interfering agent (ie, the interference value function).

[0036] Step 3: Alternately train the interfering agent and the opponent agent.

[0037] In summary, the alternating training steps of the embodiment of the present invention include: (1) in the current iteration, using the opponent intelligent agent to perform data extrapolation and adding the generated data to the extended data set; (2) using the original data set and the extended data set to train the interfering party intelligent agent; (3) training the opponent intelligent agent according to the offline adversarial learning method in combination with the training of the interfering party intelligent agent after the current iteration training; (4) continuously iterating the above (1) to (3) to obtain the final interfering party intelligent agent.

[0038] Step 4: deploy the final jammer agent in a real environment for radar jamming decision-making.

[0039] Preferably, the final interfering agent is deployed in a real environment, and its strategy can be analyzed for performance to verify the effectiveness of the strategy's lower bound. This approach provides a priori performance results for deploying the strategy in a real environment, making the deployment of the interfering agent safe and efficient in complex electromagnetic environments.

[0040] In order to more clearly demonstrate the technical solution and technical effects provided by the present invention, the method provided by the embodiment of the present invention is described in detail below with reference to specific embodiments.

[0041] 1. Introduction to the principles of the plan.

[0042] In actual electronic warfare scenarios, the jammer's understanding of the electromagnetic environment is incomplete and biased, making it difficult to guarantee the performance of jamming decisions in real environments. To address this issue, the present invention constructs an adversary agent through a historical game sequence under incomplete electromagnetic environment cognition. The jammer's agent alternates training with the constructed adversary agent without interacting with the real environment. The result is a conservative environmental model and a performance-guaranteed jamming strategy, enabling the jammer's decision-making in complex electromagnetic environments to be efficient and secure.

[0043] Specifically, in an electronic countermeasure scenario, the interfering party is usually modeled as an intelligent agent, and the electromagnetic environment is regarded as the environment in reinforcement learning. At the same time, this invention introduces the idea of ​​robust adversarial reinforcement learning (RL) and introduces a new intelligent agent, which represents the transfer function of the electromagnetic environment. The purpose is to make the interfering party's benefits as small as possible (that is, minimize the benefits of the interfering party's intelligent agent). The optimization goal set here is:

[0044]

[0045] Where π is the jammer agent’s strategy (i.e., the jamming decision to be solved at the end), M is the electromagnetic environment in which the jammer agent and the opponent agent (jammer and radar) are located. is the strategy of the opponent agent; Π, The strategy of the interfering agent π and the strategy of the opponent agent are A collection of Indicates that the strategy of the interfering agent is π and the strategy of the opponent agent is Interference value function of the interfering agent when the electromagnetic environment is M.

[0046] Combining this idea with offline reinforcement learning, the radar jamming decision problem of the jammer agent can be modeled as: In the case of , find a strategy π that satisfies the following constraints:

[0047]

[0048] in, is the opponent agent, is the interference value function when the interfering agent with strategy π competes with the opponent agent, is a set of state transition functions, expressed as:

[0049]

[0050] In the above formula, represents expectation, s represents state, and a represents action. represents the state transition law of the opponent agent, ξ represents the degree of variational distance closeness, since and Both represent the state and reward at the next moment after taking action a in state s. Since the possible state and reward at the next moment are uncertain, they are represented by dots; TV(P1,P2) is the variational distance between distributions P1 and P2. represents the maximum likelihood estimate of the MDP given by the historical game sequence. The above constraint formula defines a pessimistic electromagnetic environment model, because the purpose of this pessimistic electromagnetic environment model is to minimize the gain of the interferer.

[0051] In the embodiment of the present invention, the above optimization objective formula is equivalent to the above constraint formula, mainly because the present invention uses the strategy Abstracted as the corresponding agent (i.e., the opponent agent)

[0052] Those skilled in the art will understand that the historical game sequence is a set of four-tuples {state, action, next moment state, reward}.

[0053] The performance of the strategy derived from this pessimistic electromagnetic environment model is analyzed: Let T represent the state transition function in the real environment. It can be proved that for any strategy π of the interfering agent, there is a 1-σ probability:

[0054]

[0055] Among them, σ is a set value (a value close to 0), which means that the above formula is very likely to be true. represents the interference value function when the interfering agent with strategy π competes with the opponent agent, This represents the interference value function when an interfering agent with strategy π is deployed in a real environment against a real adversary. The above formula indicates that the performance of any interfering strategy in the real environment is at least as good as the performance of the strategy in the pessimistic model. In other words, the performance of strategy π in the pessimistic model is the lower bound of the performance of the strategy in the real environment.

[0056] 2. Detailed introduction of the plan.

[0057] Figure 2 The training process is shown; it mainly includes:

[0058] 1. Use historical game sequences as original data sets And normalize it.

[0059] 2. Using the original data set, a simple method (such as maximum likelihood estimation) is used to give the initial radar state transfer function (i.e. the opponent agent ).

[0060] 3. Alternately train the agent’s strategy and the opponent’s agent.

[0061] Each iteration consists of the following steps:

[0062] (1) Through the opponent agent Perform K-step extrapolation to generate new data and add it to the extended sequence set (i.e., extended dataset).

[0063] (2) Train the interfering agent. The training goal is to maximize the interference value function when the interfering agent with strategy π competes with the opponent agent. The training data comes from Use reinforcement learning to update the interfering agent's policy π and Q network

[0064] (3) Train the opponent agent. The training goal is to minimize the interference value function when the interference agent with strategy π competes with the opponent agent. An offline adversarial learning method is used to update the opponent intelligent agent. Specifically, according to the offline adversarial learning method, the policy gradient is calculated in combination with the strategy of the interfering intelligent agent after the current iterative training, and the policy gradient is used to guide the training of the opponent intelligent agent.

[0065] In the embodiment of the present invention, the update of strategy π can adopt conventional reinforcement learning methods (for example, actor-critic algorithm), so it will not be described in detail. The key issue is how to update the strategy π from a given historical game sequence. Starting from, an opponent intelligent agent is constructed to represent the transfer function of the electromagnetic environment. To represent the opponent agent, where is the probability of obtaining reward r and transitioning to state s′ after executing (s, a). The present invention represents the adversary agent through an integration of neural networks, each of which generates a Gaussian distribution about the next state and reward: It means that after executing (s, a), we get reward r and transfer to the Gaussian distribution of state s′, u φ (s,a) is the mean corresponding to (s,a), ∑ φ The standard deviation corresponding to (s,a) is Indicates that the mean is u φ (s,a), the standard deviation is ∑ φ (s, a) is a Gaussian distribution, where (s, a) means taking action a when in state s. The action of the opponent agent is the next state and interference benefit detected by the jammer from the electromagnetic environment. Under different action strategies, the electromagnetic environment state information and interference benefit obtained by the jammer are different. The goal of the opponent agent is to make the jammer's benefit (i.e. ) is as small as possible. In order to achieve the above goals, the present invention uses the policy gradient method. For the opponent agent, the actual work to be done is to minimize the value function of the policy π (that is, the value function of the policy π mentioned above). ), so this value function can in turn be used as an indicator to guide the training of the opponent agent. The policy gradient can be calculated as:

[0066]

[0067] in, For expectations, Indicates that state s obeys The distribution of represents the access distribution of state s under strategy π, that is, the steady-state distribution or weighted distribution of the state under a given strategy π, which depends on the transition probability distribution and strategy π; a~π means that the action selection strategy is π, Indicates that the state s′ and reward r at the next moment come from the Gaussian distribution of the opponent agent γ is the setting coefficient; and The meaning of is the same, here we use the neural network model to represent the opponent agent, φ is the neural network model parameter, that is, Represents the interference value function when the interfering agent with strategy π confronts the neural network model, represents the gradient with respect to φ.

[0068] In addition, for the opponent agent, the training loss function also needs to consider being close to the maximum likelihood estimate on the dataset. The training loss function can be expressed as:

[0069]

[0070] in, It is related to the maximum likelihood estimate given by the historical game sequence, λ is the set coefficient, Indicates that (s,a,r,s′) comes from the original dataset (i.e. historical game sequence).

[0071] Through iterative alternating training steps, we can finally get a trained interference agent, which includes a strategy π and Q network The strategy π is a performance-guaranteed strategy, and the Q network The corresponding value function is a pessimistic value function, which is the lower bound of the value function in the real environment and can reduce the number of interactions during the training process. Specifically, during online training, the jammer (corresponding to the jammer agent) has to interact with the radar to gain experience. The present invention uses historical data to construct an opponent agent, and then the radar interacts with this opponent agent, reducing the number of interactions between the jammer and the radar. The interaction between the two only occurs during the data collection phase and provides a valuable reference for the deployment of the strategy in the real environment. The strategy π is deployed to the environment for operation, and its performance is analyzed. The analysis results are compared with the performance in the pessimistic model to verify the effectiveness of the lower bound of the strategy performance. This method provides a priori performance results for the deployment of the strategy in the real environment, making the deployment of the reinforcement learning algorithm in a complex electromagnetic environment safe and efficient.

[0072] Through the description of the above embodiments, those skilled in the art will clearly understand that the above embodiments can be implemented through software or by using software plus a necessary general-purpose hardware platform. Based on this understanding, the technical solutions of the above embodiments can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, a USB flash drive, a mobile hard disk, etc.) and includes a number of instructions for causing a computer device (such as a personal computer, a server, or a network device) to execute the methods described in the various embodiments of the present invention.

[0073] Example 2

[0074] The present invention also provides an electronic countermeasure offline decision-making system, which is mainly used to implement the method provided in the above embodiment, such as Figure 3 As shown, the system mainly includes:

[0075] an interfering party agent modeling unit, configured to model the interfering party as an interfering party agent;

[0076] The opponent agent construction unit is used to take the given historical game sequence as the original data set and use it to construct the opponent agent;

[0077] An agent alternating training unit is configured to alternately train an interfering agent and an opponent agent, the steps comprising: in a current iteration, using the opponent agent to perform data extrapolation, adding the generated data to an extended data set, and using the original data set and the extended data set to train the interfering agent, with the training objective being to maximize the interference value function when the interfering agent and the opponent agent confront each other; training the opponent agent according to an offline adversarial learning method combined with training of the interfering agent trained in the current iteration, with the training objective being to minimize the interference value function when the interfering agent and the opponent agent confront each other; and continuously iterating to obtain the final interfering agent;

[0078] The deployment application unit is used to deploy the final jammer agent in a real environment to make radar jamming decisions.

[0079] Considering that the relevant technical details involved in the system have been introduced in detail in the previous embodiments, they will not be repeated here.

[0080] Those skilled in the art will clearly understand that for the convenience and brevity of description, only the division of the above-mentioned functional modules is used as an example. In actual applications, the above-mentioned functions can be assigned to different functional modules as needed, that is, the internal structure of the system can be divided into different functional modules to complete all or part of the functions described above.

[0081] Example 3

[0082] The present invention also provides a processing device, such as Figure 4 As shown, it mainly includes: one or more processors; a memory for storing one or more programs; wherein, when the one or more programs are executed by the one or more processors, the one or more processors implement the method provided by the aforementioned embodiment.

[0083] Furthermore, the processing device further includes at least one input device and at least one output device; in the processing device, the processor, memory, input device, and output device are connected via a bus.

[0084] In the embodiment of the present invention, the specific types of the memory, input device, and output device are not limited; for example:

[0085] The input device can be a touch screen, image acquisition device, physical button or mouse;

[0086] The output device may be a display terminal;

[0087] The memory may be a random access memory (RAM) or a non-volatile memory, such as a disk memory.

[0088] Example 4

[0089] The present invention also provides a readable storage medium storing a computer program, which implements the method provided in the above embodiment when the computer program is executed by a processor.

[0090] In the embodiments of the present invention, the computer-readable storage medium may be provided in the aforementioned processing device, for example, as a memory in the processing device. Alternatively, the computer-readable storage medium may be a USB flash drive, a removable hard drive, a read-only memory (ROM), a magnetic disk, or an optical disk, among other media capable of storing program code.

[0091] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by any person skilled in the art within the technical scope disclosed in the present invention should be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims. The information disclosed in the background technology section of this article is only intended to deepen the understanding of the overall background technology of the present invention, and should not be regarded as an admission or any form of implication that the information constitutes prior art already known to those skilled in the art.

Claims

1. An electronic countermeasure offline decision-making method, characterized in that: include: Model the interferer as an interferer agent; Take the given historical game sequence as the original dataset and use it to construct the opponent agent; Alternately training the interfering agent and the opponent agent, the steps comprising: in a current iteration, using the opponent agent to perform data extrapolation, adding the generated data to an extended data set, and using the original data set and the extended data set to train the interfering agent, wherein the training objective is to maximize the interference value function when the interfering agent and the opponent agent compete; training the opponent agent according to an offline adversarial learning method combined with training the interfering agent trained in the current iteration, wherein the training objective is to minimize the interference value function when the interfering agent and the opponent agent compete; and continuously iterating to obtain a final interfering agent; Deploying the final jammer agent in a real environment for radar jamming decision-making; The radar jamming decision-making problem of the jammer agent is modeled as: given a historical game sequence, find a strategy π that satisfies the following constraints: Among them, π is the strategy of the interfering agent, Π is the set of strategies π of the interfering agent, is the transfer function of the opponent agent, representing the electromagnetic environment during training; is the set of state transition functions, is the interference value function when the interfering agent with strategy π competes with the opponent agent; The training of the opponent agent according to the offline adversarial learning method in combination with the interfering agent after the current iterative training includes: The offline adversarial learning method is combined with the strategy of the interfering agent after the current iterative training to calculate the strategy gradient, and the strategy gradient is used to guide the training of the opponent agent; Among them, the calculation method of the policy gradient is expressed as: in, For expectations, Indicates that state s obeys The distribution of represents the access distribution of state s under strategy π, a~π represents the action selection strategy is π, represents the interference value function when the interfering agent with strategy π confronts the neural network model, where the neural network model represents the opponent agent, φ is the neural network model parameter, Indicates the state s at the next moment ′ and the reward r comes from the Gaussian distribution of the opponent agent It indicates that after executing (s, a), the reward r is obtained and the state s′ is transferred to the Gaussian distribution. (s, a) indicates that action a is taken in state s; γ is the set coefficient, represents the gradient with respect to φ.

2. The electronic countermeasure offline decision-making method according to claim 1, characterized in that: The constructing of the opponent intelligent agent includes: constructing the opponent intelligent agent through maximum likelihood estimation based on the original data set.

3. The electronic countermeasure offline decision-making method according to claim 1, characterized in that: The training loss function of the opponent agent is expressed as: in, is the training loss function of the opponent agent, λ is the setting coefficient, represents the Gaussian distribution of obtaining reward r and transferring to state s′ after executing (s, a), represents (s,a,r,s ′ ) from the original dataset 4. The electronic countermeasure offline decision-making method according to claim 1, characterized in that: Also includes: The final interfering agent is deployed in a real environment, and the performance of the strategy of the final interfering agent is analyzed to verify the effectiveness of the lower bound of the strategy performance.

5. The electronic countermeasure offline decision-making method according to claim 1, characterized in that: When analyzing the performance of the final interfering agent's strategy, we use T to represent the state transition function in the real environment. For any strategy, there is a 1-σ probability: Among them, σ is a set value to make the above formula valid. Represents the interference agent and the opponent agent with strategy π Interference value function during confrontation, opponent agent represents the transfer function of the electromagnetic environment during training, is the set of state transition functions, The interference value function represents the interference value of the interfering agent with strategy π when deployed in a real environment against an actual opponent.

6. An electronic countermeasure offline decision-making system, characterized in that: The method for implementing any one of claims 1 to 5 comprises: an interfering party agent modeling unit, configured to model the interfering party as an interfering party agent; The opponent agent construction unit is used to take the given historical game sequence as the original data set and use it to construct the opponent agent; An agent alternating training unit is configured to alternately train an interfering agent and an opponent agent, the steps comprising: in a current iteration, using the opponent agent to perform data extrapolation, adding the generated data to an extended data set, and using the original data set and the extended data set to train the interfering agent, with the training objective being to maximize the interference value function when the interfering agent and the opponent agent confront each other; training the opponent agent according to an offline adversarial learning method combined with training of the interfering agent trained in the current iteration, with the training objective being to minimize the interference value function when the interfering agent and the opponent agent confront each other; and continuously iterating to obtain the final interfering agent; The deployment application unit is used to deploy the final jammer agent in a real environment to make radar jamming decisions.

7. A processing device, characterized in that: include: one or more processors; a memory for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 5.

8. A readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.

Citation Information

Patent Citations

  • DQN-based radar confrontation intelligent decision-making method

    CN113378466A

  • Performance boundary analysis method, system and equipment of agent game strategy and medium

    CN117313786A

Cited By

  • Intelligent electromagnetic environment sensing and anti-interference system

    CN121637031A