Electronic countermeasures two-layer modeling and decision optimization method, system, equipment and medium

By modeling both sides of the electronic countermeasure as intelligent agents, using Markov decision process and game theory to establish a time scale model, and optimizing the strategy network of radar and jammer, the problem of limited improvement in countermeasure performance in existing technologies is solved, and signal-level countermeasure performance is improved.

CN119805377BActive Publication Date: 2025-09-19UNIV OF SCI & TECH OF CHINA +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510008840.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-03
Publication Date
2025-09-19
Estimated Expiration
2045-01-03

AI Technical Summary

Technical Problem

The performance improvement of existing electronic countermeasure scenarios is limited, and most of them remain at the functional level, failing to be refined to the signal level, and the confrontation often only involves one side.

Method used

The two electronic countermeasure parties are modeled as radar agents and jammer agents respectively. A time scale model is established based on Markov decision process and game theory. The strategy network and value network are optimized through multi-agent reinforcement learning algorithm to achieve the equilibrium strategy of radar and jammer.

Benefits of technology

It improves the performance of electronic countermeasures, enhances the anti-interference capability of radar and the interference capability of jammer, and solves the Markov problem of signal-level parameters under the frequency of function-level parameter changes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119805377B_ABST
    Figure CN119805377B_ABST
Patent Text Reader

Abstract

The present invention discloses a two-layer modeling and decision optimization method, system, device and medium for electronic countermeasures. For electronic countermeasure scenarios, Markov decision process modeling is performed on both sides of the electronic countermeasure at different time scales to solve the problem that signal-level parameters in the electronic countermeasure process do not have Markov properties at the time scale of the function-level parameter change frequency. At the same time, decision optimization of anti-interference methods and interference styles can be performed based on the model to improve countermeasure performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of electronic countermeasure technology, and in particular to an electronic countermeasure double-layer modeling and decision optimization method, system, equipment and medium. Background Art

[0002] The Markov decision process is a time-dependent mathematical framework consisting of four elements: state, action, state transition function, and reward function. In a Markov decision process, there is typically an agent that performs an action. A Markov decision process is a continuous interaction between the agent and its environment. At a given moment, the agent selects an action based on its current state and submits it to the environment. The environment then receives a reward value and the next state based on the state transition function and reward function, which it then feeds back to the agent. The function that determines the agent's action selection based on the current state is called a policy. Generally speaking, the agent's goal is to maximize the cumulative reward, so it needs to update its policy based on the feedback from the environment during each interaction, ultimately arriving at an optimal policy.

[0003] At present, there are many electronic warfare scenarios based on Markov decision process modeling, but most of them remain at the functional level, have not been refined to the signal level, and often only involve unilateral confrontation. Therefore, the confrontation performance needs to be improved. Summary of the Invention

[0004] The purpose of the present invention is to provide an electronic countermeasure dual-layer modeling and decision optimization method, system, equipment and medium, which can perform electronic countermeasure dual-layer modeling and decision optimization from the functional level and signal level, and improve countermeasure performance.

[0005] The purpose of the present invention is achieved through the following technical solutions:

[0006] An electronic countermeasure two-layer modeling and decision optimization method, comprising:

[0007] The two sides of the electronic warfare are modeled as corresponding intelligent agents, the radar side is modeled as the radar agent, and the jammer side is modeled as the jammer agent;

[0008] According to the different time intervals of the confrontation, a corresponding time scale model is established. The electronic confrontation process of the radar agent and the jammer agent is modeled based on the Markov decision process and game theory, and the radar agent and the jammer agent are optimized. The optimization process includes: the radar agent and the jammer agent each make action decisions and execute them according to the current state based on the Markov decision process, and then obtain the corresponding reward value and new state based on game theory. The radar agent and the jammer agent update the strategy network and the value network based on their respective reward values ​​and new states; wherein, the strategy network and the value network are both used in the action decision process, the strategy network is used to output the probability distribution of the action, and the value network is used to output the value of the action; after the optimization is completed, the equilibrium strategy of the corresponding time scale model is obtained;

[0009] The equilibrium strategy of the corresponding time scale model is used to guide the one-to-one radar confrontation decision-making process.

[0010] An electronic countermeasure two-layer modeling and decision optimization system, comprising:

[0011] The agent modeling unit is used to model the two electronic countermeasure parties as corresponding agents, the radar party is modeled as a radar agent, and the jammer party is modeled as a jammer agent;

[0012] The electronic countermeasures two-layer modeling and decision optimization unit is used to establish a corresponding time scale model according to the time interval of the confrontation, model the electronic countermeasure process of the radar agent and the jammer agent based on the Markov decision process and game theory, and optimize the radar agent and the jammer agent. The optimization process includes: the radar agent and the jammer agent each make action decisions and execute them according to the current state based on the Markov decision process, and then obtain the corresponding reward value and new state based on game theory. The radar agent and the jammer agent update the strategy network and value network based on their respective reward values ​​and new states; wherein, the strategy network and value network are both used in the action decision process, the strategy network is used to output the probability distribution of the action, and the value network is used to output the value of the action; after the optimization is completed, the equilibrium strategy of the corresponding time scale model is obtained;

[0013] The decision-making unit is used to guide the one-to-one radar confrontation decision-making process using the equilibrium strategy of the corresponding time scale model.

[0014] A processing device comprising: one or more processors; a memory for storing one or more programs;

[0015] When the one or more programs are executed by the one or more processors, the one or more processors implement the aforementioned method.

[0016] A readable storage medium stores a computer program, which implements the aforementioned method when the computer program is executed by a processor.

[0017] It can be seen from the technical solution provided by the present invention that, for electronic countermeasure scenarios, Markov decision process modeling is performed on both sides of the electronic countermeasure at different time scales to solve the problem that the signal-level parameters in the electronic countermeasure process do not have Markov properties at the time scale of the function-level parameter change frequency. At the same time, the decision optimization of the anti-interference mode and interference style can be carried out according to the model to improve the countermeasure performance (that is, to improve the anti-interference capability of the radar and the interference capability of the jammer). BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0019] Figure 1 A flowchart of a two-layer modeling and decision optimization method for electronic countermeasures provided by an embodiment of the present invention;

[0020] Figure 2 A flowchart of a search task provided by an embodiment of the present invention;

[0021] Figure 3 A flowchart of a confirmation task provided by an embodiment of the present invention;

[0022] Figure 4 A flowchart of a tracking task provided by an embodiment of the present invention;

[0023] Figure 5 A flowchart of a lost tracking task provided by an embodiment of the present invention;

[0024] Figure 6 A schematic diagram of the switching logic of a radar agent provided by an embodiment of the present invention;

[0025] Figure 7 A schematic diagram of a large-time-scale model solving and optimization process provided by an embodiment of the present invention;

[0026] Figure 8 A schematic diagram of a small time scale model solving and optimization process provided by an embodiment of the present invention;

[0027] Figure 9 A schematic diagram of an electronic countermeasures dual-layer modeling and decision optimization system provided by an embodiment of the present invention;

[0028] Figure 10A schematic diagram of a processing device provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0029] The following is a clear and complete description of the technical solutions in the embodiments of the present invention, in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0030] First, the following terms may be used in this article:

[0031] The terms "include," "comprises," "contains," "has," or other similar expressions should be interpreted as non-exclusive. For example, "including certain technical features (such as raw materials, components, ingredients, carriers, dosage forms, materials, dimensions, parts, components, mechanisms, devices, steps, procedures, methods, reaction conditions, processing conditions, parameters, algorithms, signals, data, products, or manufactured articles, etc.) should be interpreted as including not only the technical features explicitly listed, but also other technical features known in the art that are not explicitly listed.

[0032] The term "consisting of" excludes any technical features not explicitly listed. If used in a claim, this term renders the claim closed, excluding any technical features other than those explicitly listed, except for conventional impurities associated with them. If this term appears only in a clause of a claim, it limits only the elements explicitly listed in that clause; elements listed in other clauses are not excluded from the claim as a whole.

[0033] Unless otherwise specified or limited, the terms "mounted," "connected," "connect," and "fixed" should be interpreted broadly. For example, they can refer to fixed, detachable, or integral connections; mechanical or electrical connections; direct or indirect connections through an intermediary; and internal communication between two components. Those skilled in the art will understand the specific meanings of the above terms in this document based on specific circumstances.

[0034] The following describes in detail the electronic countermeasure dual-layer modeling and decision optimization method, system, device, and medium provided by the present invention. Any information not described in detail in the embodiments of the present invention represents prior art known to those skilled in the art. Where specific conditions are not specified in the embodiments of the present invention, the procedures were performed in accordance with conventional conditions in the art or the conditions recommended by the manufacturer. Instruments used in the embodiments of the present invention, where the manufacturer is not specified, are all commercially available conventional products.

[0035] Example 1

[0036] The embodiment of the present invention provides a two-layer modeling and decision optimization method for electronic countermeasures, such as Figure 1 As shown, it mainly includes the following steps:

[0037] Step 1: Model both sides of the electronic confrontation as corresponding intelligent agents.

[0038] In the embodiment of the present invention, the electronic countermeasures include: a radar and a jammer. The radar is modeled as a radar agent, and the jammer is modeled as a jammer agent.

[0039] Step 2: Establish corresponding time scale models according to different time intervals of confrontation, model the electronic confrontation process of radar agent and jammer agent based on Markov decision process and game theory, and optimize the radar agent and jammer agent.

[0040] The optimization process includes: the radar agent and the jammer agent each make action decisions and execute them according to the current state based on the Markov decision process, and then obtain the corresponding reward value and new state based on game theory. The radar agent and the jammer agent update the policy network and value network based on their respective reward values ​​and new states; wherein, the policy network and value network are both used in the action decision process, the policy network is used to output the probability distribution of the action, and the value network is used to output the value of the action; after the optimization is completed, the equilibrium strategy of the corresponding time scale model is obtained.

[0041] Step 3: Use the equilibrium strategy of the corresponding time scale model to guide the one-to-one radar countermeasure decision-making process, that is, use the equilibrium strategy to decide the actions of the radar and jammer.

[0042] In order to more clearly demonstrate the technical solution and technical effects provided by the present invention, the method provided by the embodiment of the present invention is described in detail below with reference to specific embodiments.

[0043] This paper focuses on the scenario of one-on-one electronic confrontation in air combat. Based on the Markov decision process, the confrontation process model with different perspectives under two time scales is studied and optimized. Specifically, the paper defines two time intervals. One time interval is the time it takes for the radar to execute a scheduling cycle, and the corresponding time scale is called the large time scale model; the other time interval is the radar coherent processing interval t CPI The corresponding time scale is called the small time scale model. The following is a detailed introduction to the two time scale models.

[0044] 1. Large time scale model.

[0045] The parameters of the radar agent are expressed as:

[0046]

[0047] Among them, s R (n), o R (n), a R (n) corresponds to the state, observation, and action of the radar agent in the nth time slot, S R , O R 、A R Correspondingly represents the state space, observation space, and action space, r R is the radar agent’s profit function, W R (n) is the working mode, RL R (n) is the track list; a J (n) is the action of the jammer agent in the nth time slot.

[0048] Those skilled in the art will understand that a time slot is the length of time occupied by a round. Since the time intervals of the large time scale model and the small time scale model are different, the lengths of their time slots are different.

[0049] The observation is the echo signal received by the radar agent, and the echo signal is the echo signal received after executing the beam request. Figures 2 to 5 The figure shows the simulation process of executing beam request. The radar executes a beam request among search, confirmation, tracking, and loss of tracking according to the current working mode and track list. After the beam request is executed, the actions of both the radar and the jammer jointly generate an echo signal, which is received by the radar. After that, the echo signal is processed in the simulation program to obtain the target motion information. x is the coordinate in the spherical coordinate system with the radar as the origin and the front as the polar axis. is the radial velocity of the target.

[0050] The working mode is the scheduling logic of the radar agent, which affects the execution order of the four beam requests: search, confirmation, tracking, and loss of tracking. At the same time, it will switch according to the observation value at the end of the time slot. The switching logic is as follows: Figure 6 As shown, Figure 6 Middle R i (m) is the radius of the jammer in the spherical coordinate system with the radar as the center, β i (m) is the pitch angle of the jammer in the spherical coordinate system with the radar as the center, R1, R2, ..., R4 and β0 are the setting parameters.

[0051] The action is the anti-interference method adopted by the radar agent, including: frequency agility, radio frequency shielding, and frequency modulation slope agility.

[0052] The track list is divided into a temporary track list and a stable track list, which respectively represent the target motion information history of the confirmed target and the tracked target. After each beam request is executed, the track list for the next time slot is maintained.

[0053] The profit function is expressed as:

[0054]

[0055] Among them, -η1, η2, and η3 are the reward values ​​under different circumstances.

[0056] The parameters of the jammer agent are expressed as:

[0057]

[0058] Among them, s J (n), o J (n), a J (n) corresponds to the state, observation, and action of the jammer agent in the nth time slot, S J , O J 、A J Correspondingly represents the state space, observation space, and action space, r J is the reward function of the jammer agent; l J (n),H J (n) is the jammer's own motion information and observation history; a R (n) is the action of the radar agent in the nth time slot;

[0059] The observation is the pulse descriptor of the radar signal intercepted by the jammer agent, that is:

[0060]

[0061] Where T is the transposed symbol, each item represents a pulse description word, and the subscript is the sub-pulse number reached in the nth time slot, K J is the number of sub-pulses intercepted in the nth time slot.

[0062] The jammer's own motion information is consistent with the target information structure obtained by the radar, that is, Considering the target moving model with constant speed, l J The transfer of (n) can be expressed as:

[0063] l J (n) = Fl J (n-1)

[0064] Among them, the matrix T0 is the time slot length of the large time scale model.

[0065] The transfer of observation history can be obtained by the following formula

[0066] H J (n+1)= H J (n)∪o J (n)

[0067] The action is the jamming style taken by the jammer agent, and the jammer agent generates jamming signals based on the observations and actions.

[0068] The reward function is set to the inverse of the radar view model's reward. This is to facilitate solving the model, and is set as a zero-sum game. Those skilled in the art will understand that game theory is a reasonable approach to studying adversarial problems. In games, it is often assumed that all participants can find an optimal strategy. In particular, a game in which the sum of all parties' gains and losses is always zero is called a zero-sum game. It can be mathematically proven that for all zero-sum game problems, an equilibrium solution with mutually optimal strategies can be obtained by solving a linear programming problem.

[0069] The optimization process of solving large time scale models is as follows Figure 7 As shown, input radar simulation platform parameter structure MPARParam (describing some constants of radar, such as power, working airspace range, antenna gain, etc.), training round number e num , exploration rate ε; then, initialize π R (Radar Agent Strategy),π J (Jammer Agent Strategy),Q R (Radar agent action value),Q J (Interference agent action value), V R (radar agent state value), V J (interference machine agent state value),s R ,s J In each time slot of each iteration (Maxtime is an external hyperparameter that limits the maximum time a game can last), the radar agent and the jammer agent each make an action decision based on the current state with a probability of 1-ε based on the Markov decision process, or randomly select an action according to the uniform distribution with a probability of ε. A confrontation simulation is performed for one scheduling cycle to obtain the tracking status of the radar after the confrontation and the state s′ of the radar agent after the transfer. R and the jammer agent state s′ J , thereby obtaining the reward values ​​of the two agents and updating the Q table (a table of long-term rewards corresponding to states and actions, including the aforementioned Q R ,Q J ), strategy (i.e. π R ,π J) and V table (the table of expected long-term returns corresponding to the state corresponds to the aforementioned V R ,V J ), in the relevant formula, α and γ are two hyperparameters, a′ R and a′ J Is a temporary variable in the process of finding the minimum or sum (corresponding to the actions of the two agents), execute Isterminal(s R ,s j ), that is, to determine whether the current state is the terminal state, that is, whether the game between the radar and the jammer has produced a winner; repeat the iteration until the set number of training rounds e is reached num , output the Nash equilibrium strategy.

[0070] Those skilled in the art can understand that the Nash equilibrium strategy is a pair of strategies in which the strategies of the two opposing parties are mutually optimal solutions, that is, both parties use the optimal interference strategy and anti-interference strategy respectively, that is, when one party adopts the Nash equilibrium strategy, the other party does not adopt the Nash equilibrium strategy, and only suboptimal results can be obtained. Therefore, the double-layer electronic countermeasures of the present invention can achieve an improvement in countermeasure performance compared to the unilateral countermeasures of the existing scheme.

[0071] 2. Small time scale model.

[0072] The parameters of the radar agent are expressed as:

[0073] {s r (n)={F(n),f c (n),PRF,Z(n)}∈S r ,a r (n)∈A r ,o r (n)∈O r ,r r (s r (n),a r (n),a j (n)),W R ,a R}

[0074] Among them, s r (n), o r (n), a r (n) corresponds to the state, observation, and action of the radar agent in the nth time slot, S r , O r 、A r Correspondingly represents the state space, observation space, and action space, r r is the radar agent’s profit function, W R with a RThe corresponding fixed working mode and anti-interference action are parameters in the large time scale model and can be directly obtained from the large time scale model. They will directly affect the radar status of the small time scale model, as shown in Table 1 below; a j (n) is the action of the jammer agent in the nth time slot.

[0075] Similar to the large time scale model, the observation is the echo signal received by the radar agent, and the echo signal is processed to obtain the target motion information

[0076] Status r (n)={F(n),f c (n),PRF,Z(n)}, where F(n) = {0,1} is the radio frequency shielding flag. When its value is 0, the signal is processed according to the normal process. When its value is 1, the first sub-pulse observed in the nth time slot is used as the shielding pulse and no signal processing is performed. c (n)={f0,f1,...,f K″} is the carrier frequency, each item is a different carrier frequency point, the subscript is the carrier frequency point number, K″ is the total number of carrier frequency points, and Z(n)={-1,1} is the frequency modulation slope agility indicator.

[0077] The action is a binary vector of 1*Np. When the i-th bit is set to 1, the corresponding anti-interference action is taken to make s r (n) The corresponding parameter changes and issues a r (n) The corresponding sub-pulse signal, s r The parameter variation range of (n) is shown in Table 1.

[0078] Table 1: s r (n) Parameter range

[0079]

[0080] The settings of other working modes and anti-interference actions are omitted in Table 1. The working mode settings can be flexibly adjusted according to the actual scenario, so they are not omitted.

[0081] The sub-pulse signal waveform is expressed as follows:

[0082]

[0083] Among them, A is the pulse amplitude, rect is the rectangular window function, f d is the Doppler frequency, j is the imaginary unit, t is the time variable, B is the bandwidth, t p is the pulse width, and π is the symbol of pi.

[0084] After sending a signal, the radar will receive an echo. The target echo is in the form of:

[0085]

[0086] Where τ(n) is the delay time of the signal return, v(n) is the Doppler frequency, and φ j For interference signal.

[0087] Revenue function r r Defined as the signal-to-interference ratio of the received echo.

[0088] The parameters of the jammer agent are expressed as:

[0089] {s j (n) = {H j (n)}∈S j ,a j (n)∈A j ,o j (n)∈O j ,r j (s j (n),a j (n),a r (n)),a J}

[0090] Among them, s j (n), o j (n), a j (n) corresponds to the state, observation, and action of the jammer agent in the nth time slot, S j , O j 、A j is the corresponding representation of state space, observation space, and action space, r j is the benefit function of the jammer agent, H j (n) is the observation history, a J The interference signal pattern transmitted by the large time scale model is obtained synchronously from the large time scale model and used as a parameter for simulating the signal processing process in the small time scale model. r (n) is the action of the radar agent in the nth time slot.

[0091] The observation is the pulse description word of the radar signal intercepted by the jammer agent. In the case of transmitter-receiver isolation (i.e., the jammer cannot transmit the jamming signal when intercepting the signal), assuming that the duration of each interception by the jammer agent is equal to the repetition interval of the radar pulse, the observation o is defined as j (n)=[o j (n,1),...o j (n,m),...,o j (n,Np)], where oj (n,m) is the observation value of the jammer agent at the mth pulse repetition interval in the nth time slot, defined as:

[0092]

[0093] in, The pulse description word obtained in the nth time slot, that is, the kth pulse description word obtained in the current large time slot j Pulse description word, a j (n,m)=1 means the jammer takes reconnaissance action at the mth pulse repetition interval of the nth time slot, a j (n,m)=0, then according to a j (n) and o j (n,m) transmits an interference signal.

[0094] The state is defined as H j (n) = {o j (1),…,o j (n-1)}, that is, the set of historical observations; the transfer equation can be expressed as

[0095] The actions of the jammer agent are defined as: Where Np is the total number of pulses transmitted by the radar agent in one time slot.

[0096] Similar to the large time scale model, the jammer agent payoff function is defined as the inverse of the radar payoff function.

[0097] The optimization process of solving small time scale models is as follows Figure 8 As shown, input radar simulation platform parameter structure MPARParam (describing some constants of radar, such as power, working airspace range, antenna gain, etc.), anti-interference method a R , interference pattern a J , the number of training rounds e num , exploration rate ε; then, initialize π r (Radar Agent Strategy),π j (Jammer Agent Strategy),Q r (Radar agent action value),Q j (Interference agent action value), V r (radar agent state value), V j (interference machine agent state value),s r ,s jIn each iteration, the radar agent and the jammer agent each make an action decision based on the current state with a probability of 1-ε based on the Markov decision process, or randomly select an action according to the uniform distribution with a probability of ε; the signal-level radar countermeasure process is simulated to obtain the signal-to-interference ratio after signal processing and the radar agent state s′ after the transfer. r and the jammer agent state s′ j , thereby obtaining the reward values ​​of the two agents and updating the Q table (a table of long-term rewards corresponding to states and actions, including the aforementioned Q r ,Q j ), strategy (i.e. π r ,π j ) and V table (the table of expected long-term returns corresponding to the state corresponds to the aforementioned V r ,V j ); Repeat the iteration until the set number of training rounds e is reached num , outputting the Nash equilibrium strategy. Considering that both the small-timescale model and the aforementioned large-timescale model are optimized using the multi-agent reinforcement learning algorithm, the parameters involved have essentially the same meaning. The main difference is that they belong to different timescale models. Therefore, the subscripts are distinguished by uppercase and lowercase letters.

[0098] The above solution provided by the embodiment of the present invention has the following advantages:

[0099] The present invention aims at specific scenarios, starting from the functional level and signal level in the radar confrontation process, and establishing models for both parties based on Markov decision process and game theory and solving optimization. At a large time scale, the definitions of state, action, and benefit in the model are given based on the basic elements in the radar confrontation process, and according to the process of radar signal processing, the real process corresponding to the state transition of the model is explained, completing the large-scale Markov modeling of the radar perspective and the jammer perspective, and the decision optimization of the anti-interference mode and the jamming style can be carried out according to the model. At a small time scale, based on the fact that the antenna transmits and receives, focusing on the game of intermittent observation of the jammer and the agile change of radar waveform parameters, the small-scale Markov modeling of the radar perspective and the jammer perspective is completed, solving the problem that the changes of some parameters in the small-time scale model of radar confrontation do not conform to the Markov property. In terms of model solution optimization, a multi-agent reinforcement learning algorithm is used to perform model solution optimization on a radar confrontation simulation platform.

[0100] To facilitate understanding, the following example uses a one-on-one radar confrontation process in air combat as an example. The main process of this example is as follows:

[0101] 1. Adjust the size of the model state space and action space according to the number of radar working modes, the number of anti-interference actions, and the number of jammer interference patterns.

[0102] 2. Based on the feedback from the radar simulation program, run the small-time-scale model solving algorithm for each anti-interference action-interference pattern pair to obtain the equilibrium strategy of the small-time-scale model.

[0103] 3. Since the benefits of the large-time-scale model depend on the simulation of the signal processing process, and the simulation results of the signal processing process are affected by the equilibrium strategy of the small-time-scale model, the equilibrium strategy of the small-time-scale model is brought into the solution optimization process of the large-time-scale model to obtain the equilibrium strategy of the large-time-scale models of both parties.

[0104] 4. Guide the decision-making process of one-on-one radar confrontation in air combat based on the equilibrium strategy of the large-time-scale model and the equilibrium strategy of the small-time-scale model. Specifically: First, use the large-time-scale model to determine the jamming style or anti-jamming action to be adopted in the current state, and then use the small-time-scale model to determine the intermittent observation behavior in the jamming style or anti-jamming action.

[0105] Through the description of the above embodiments, those skilled in the art will clearly understand that the above embodiments can be implemented through software or by using software plus a necessary general-purpose hardware platform. Based on this understanding, the technical solutions of the above embodiments can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, a USB flash drive, a mobile hard disk, etc.) and includes a number of instructions for causing a computer device (such as a personal computer, a server, or a network device) to execute the methods described in the various embodiments of the present invention.

[0106] Example 2

[0107] The present invention also provides an electronic countermeasure dual-layer modeling and decision optimization system, which is mainly used to implement the method provided in the above embodiment, such as Figure 9 As shown, the system mainly includes:

[0108] The agent modeling unit is used to model the two electronic countermeasure parties as corresponding agents, the radar party is modeled as a radar agent, and the jammer party is modeled as a jammer agent;

[0109] The electronic countermeasures two-layer modeling and decision optimization unit is used to establish a corresponding time scale model according to the time interval of the confrontation, model the electronic countermeasure process of the radar agent and the jammer agent based on the Markov decision process and game theory, and optimize the radar agent and the jammer agent. The optimization process includes: the radar agent and the jammer agent each make action decisions and execute them according to the current state based on the Markov decision process, and then obtain the corresponding reward value and new state based on game theory. The radar agent and the jammer agent update the strategy network and value network based on their respective reward values ​​and new states; wherein, the strategy network and value network are both used in the action decision process, the strategy network is used to output the probability distribution of the action, and the value network is used to output the value of the action; after the optimization is completed, the equilibrium strategy of the corresponding time scale model is obtained;

[0110] The decision-making unit is used to guide the one-to-one radar confrontation decision-making process using the equilibrium strategy of the corresponding time scale model.

[0111] Considering that the main technical details involved in the system have been introduced in detail in the previous embodiments, they will not be repeated here.

[0112] Those skilled in the art will clearly understand that for the convenience and brevity of description, only the division of the above-mentioned functional modules is used as an example. In actual applications, the above-mentioned functions can be assigned to different functional modules as needed, that is, the internal structure of the system can be divided into different functional modules to complete all or part of the functions described above.

[0113] Example 3

[0114] The present invention also provides a processing device, such as Figure 10 As shown, it mainly includes: one or more processors; a memory for storing one or more programs; wherein, when the one or more programs are executed by the one or more processors, the one or more processors implement the method provided by the aforementioned embodiment.

[0115] Furthermore, the processing device further includes at least one input device and at least one output device; in the processing device, the processor, memory, input device, and output device are connected via a bus.

[0116] In the embodiment of the present invention, the specific types of the memory, input device, and output device are not limited; for example:

[0117] The input device can be a touch screen, image acquisition device, physical button or mouse;

[0118] The output device may be a display terminal;

[0119] The memory may be a random access memory (RAM) or a non-volatile memory, such as a disk memory.

[0120] Example 4

[0121] The present invention also provides a readable storage medium storing a computer program, which implements the method provided in the above embodiment when the computer program is executed by a processor.

[0122] In the embodiments of the present invention, the computer-readable storage medium may be provided in the aforementioned processing device, for example, as a memory in the processing device. Alternatively, the computer-readable storage medium may be a USB flash drive, a removable hard drive, a read-only memory (ROM), a magnetic disk, or an optical disk, among other media capable of storing program code.

[0123] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by any person skilled in the art within the technical scope disclosed in the present invention should be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims. The information disclosed in the background technology section of this article is only intended to deepen the understanding of the overall background technology of the present invention, and should not be regarded as an admission or any form of implication that the information constitutes prior art already known to those skilled in the art.

Claims

1. A two-layer modeling and decision optimization method for electronic countermeasures, characterized in that: include: The two sides of the electronic warfare are modeled as corresponding intelligent agents, the radar side is modeled as the radar agent, and the jammer side is modeled as the jammer agent; According to the different time intervals of the confrontation, a corresponding time scale model is established. The electronic confrontation process of the radar agent and the jammer agent is modeled based on the Markov decision process and game theory, and the radar agent and the jammer agent are optimized. The optimization process includes: the radar agent and the jammer agent each make action decisions and execute them according to the current state based on the Markov decision process, and then obtain the corresponding reward value and new state based on game theory. The radar agent and the jammer agent update the strategy network and the value network based on their respective reward values ​​and new states; wherein, the strategy network and the value network are both used in the action decision process, the strategy network is used to output the probability distribution of the action, and the value network is used to output the value of the action; after the optimization is completed, the equilibrium strategy of the corresponding time scale model is obtained; The equilibrium strategy of the corresponding time scale model is used to guide the one-to-one radar confrontation decision-making process.

2. The electronic countermeasure double-layer modeling and decision optimization method according to claim 1 is characterized in that: The establishing of a corresponding time scale model according to different time intervals of the confrontation includes: Two time intervals are defined: one is the time it takes for the radar to execute a scheduling cycle, and the corresponding time scale is called the large time scale model; the other is the coherent processing interval of the radar, and the corresponding time scale is called the small time scale model. For large-time-scale models, during the optimization process, the reward value of the radar agent is determined in combination with the radar's tracking status; for small-time-scale models, during the optimization process, the reward value of the radar agent is determined in combination with the signal-to-interference-and-noise ratio of the received echo signal; in both types of time-scale models, the reward value of the jammer agent is the opposite of the reward value of the radar agent in the corresponding time-scale model.

3. The electronic countermeasure double-layer modeling and decision optimization method according to claim 1 or 2, characterized in that: The radar agent and the jammer agent each make action decisions based on the Markov decision process according to the current state, including: According to the set exploration rate ε, the radar agent and the jammer agent each make an action decision based on the current state with a probability of 1-ε based on the Markov decision process, or randomly select an action according to the uniform distribution with a probability of ε.

4. The electronic countermeasure double-layer modeling and decision optimization method according to claim 2 is characterized in that: In the large time scale model, the parameters of the radar agent are expressed as: Among them, s R (n), o R (n), a R (n) corresponds to the state, observation, and action of the radar agent in the nth time slot, S R , O R 、A R Correspondingly represents the state space, observation space, and action space, r R is the radar agent’s profit function, W R (n) is the working mode, RL R (n) is the track list; a J (n) is the action of the jammer agent in the nth time slot; observation is the target motion information obtained by the radar agent by processing the received echo signal. The echo signal refers to the echo signal received after the beam request is executed; the working mode is the scheduling logic of the radar agent, which affects the execution order of the beam request; the action is the anti-interference method adopted by the radar agent; the track list is divided into a temporary track list and a stable track list, which respectively represent the target motion information history of the confirmed target and the tracked target; The profit function is expressed as: Among them, -η1, η2, and η3 are the reward values ​​under different circumstances.

5. The electronic countermeasure double-layer modeling and decision optimization method according to claim 2 or 4, characterized in that: In the large time scale model, the parameters of the jammer agent are expressed as: Among them, s J (n), o J (n), a J (n) corresponds to the state, observation, and action of the jammer agent in the nth time slot, S J , O J 、A J Correspondingly represents the state space, observation space, and action space, r J is the payoff function of the jammer agent; l J (n),H J (n) is the jammer's own motion information and observation history; a R (n) is the action of the radar agent in the nth time slot; observation is the pulse descriptor of the radar signal intercepted by the jammer agent; action is the jamming style adopted by the jammer agent.

6. The electronic countermeasure double-layer modeling and decision optimization method according to claim 2, characterized in that: In the small time scale model, the parameters of the radar agent are expressed as: {s r (n) = {F(n), f c (n), PRF, Z(n)} ∈ S r , a r (n) ∈ A r , o r (n) ∈ O r , r r (s r (n), a r (n), a j (n)), W R , a R} Wherein, s r (n), o r (n), a r (n) corresponds to the state, observation, and action of the radar agent in the nth time slot, S r , O r 、A r Correspondingly represents the state space, observation space, and action space, r r is the benefit function of the radar agent; W R with a R The corresponding representations are fixed working modes and anti-interference actions, both of which are parameters in the large time scale model and are obtained synchronously from the large time scale model; a j (n) is the action of the jammer agent in the nth time slot; observation is the target motion information obtained by the radar agent by processing the received echo signal; state s r (n)={F(n),f c (n),PRF,Z(n)}, where F(n)={0,1} is the radio frequency shielding flag. When its value is 1, the first sub-pulse observed in the nth time slot is used as the shielding pulse, and no signal processing is performed. c (n)={f0,f1,...,f K″ } is the carrier frequency, each item is a different carrier frequency point, the subscript is the carrier frequency point number, K″ is the total number of carrier frequency points, Z(n)={-1,1} is the frequency modulation slope agility mark; the benefit function r r It is defined as the signal-to-interference ratio of the received echo; The action is a binary vector of 1*Np. When the nth bit is 1, the corresponding anti-interference action is taken to make s r (n) The corresponding parameter changes and issues a r (n) The corresponding sub-pulse signal, the sub-pulse signal waveform is expressed as follows: Among them, A is the pulse amplitude, rect is the rectangular window function, f d is the Doppler frequency, j is the imaginary unit, t is the time variable, B is the bandwidth, t p is the pulse width, and π is the symbol of pi.

7. The electronic countermeasure double-layer modeling and decision optimization method according to claim 2 or 6, characterized in that: In the small time scale model, the parameters of the jammer agent are expressed as: {s j (n)={H j (n)}∈S j ,a j (n)∈A j ,o j (n)∈O j ,r j (s j (n),a j (n),a r (n)),a J } Among them, s j (n), o j (n), a j (n) corresponds to the state, observation, and action of the jammer agent in the nth time slot, S j , O j 、A j is the corresponding representation of state space, observation space, and action space, r j is the benefit function of the jammer agent, H j (n) is the observation history, a J The interference signal pattern delivered to the large time scale model; a r (n) is the action of the radar agent in the nth time slot; the state is defined as H j (n) = {o j (1),…,o j (n-1)}, that is, the set of historical observations; The observation is the pulse description word of the radar signal intercepted by the jammer agent. In the case of transmitter-receiver isolation, if the duration of each interception by the jammer agent is equal to the repetition interval of the radar pulse, then the observation o is defined as j (n)=[o j (n,1),...o j (n,m),...,o j (n,Np)], where o j (n,m) is the observation value of the jammer agent at the mth pulse repetition interval in the nth time slot, defined as: in, The pulse description word obtained in the nth time slot, that is, the kth pulse description word obtained in the current large time slot j Pulse description word, a j (n,m)=1 means the jammer takes reconnaissance action at the mth pulse repetition interval of the nth time slot, a j (n,m)=0, then according to a j (n) and o j (n,m) transmits interference signal; The action of the jammer agent is defined as: a j (n)=[a j (n,1),...,a j (n,m),...,a j (n, Np)], where Np is the total number of pulses transmitted by the radar agent in one time slot.

8. An electronic countermeasures two-layer modeling and decision optimization system, characterized by: include: The agent modeling unit is used to model the two electronic countermeasure parties as corresponding agents, the radar party is modeled as a radar agent, and the jammer party is modeled as a jammer agent; The electronic countermeasures two-layer modeling and decision optimization unit is used to establish a corresponding time scale model according to the time interval of the confrontation, model the electronic countermeasure process of the radar agent and the jammer agent based on the Markov decision process and game theory, and optimize the radar agent and the jammer agent. The optimization process includes: the radar agent and the jammer agent each make action decisions and execute them according to the current state based on the Markov decision process, and then obtain the corresponding reward value and new state based on game theory. The radar agent and the jammer agent update the strategy network and value network based on their respective reward values ​​and new states; wherein, the strategy network and value network are both used in the action decision process, the strategy network is used to output the probability distribution of the action, and the value network is used to output the value of the action; after the optimization is completed, the equilibrium strategy of the corresponding time scale model is obtained; The decision-making unit is used to guide the one-to-one radar confrontation decision-making process using the equilibrium strategy of the corresponding time scale model.

9. A processing device, characterized in that include: one or more processors; a memory for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 7.

10. A readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Q learning interference decision-making method based on hidden Markov model

    CN115062790A

  • Radar space-time-frequency-energy multi-domain combined intelligent active anti-interference method

    CN115932750A