Platform and method for modeling and simulating reinforcement learning environment in electronic confrontation game

By designing a reinforced learning environment modeling platform to simulate radar detection and tracking processes, the problem of insufficient decision-making capabilities in radar confrontation simulation in the existing technology is solved, and the agent's efficient interference decision-making in electronic confrontation is achieved.

CN120235045APending Publication Date: 2025-07-01UNIV OF SCI & TECH OF CHINA +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510364877.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

The existing reinforcement learning environment schemes are unable to achieve a complete closed-loop assessment from interference encounter to confrontation end, resulting in insufficient decision-making capabilities in radar confrontation simulation.

Method used

Design a reinforced learning environment modeling and simulation platform in electronic confrontation game, realize the interaction between the jammer agent and the environment through the simulation stepping module, simulate the radar detection success rate and tracking process, and update the strategy parameters in combination with the Kalman filter.

Benefits of technology

It has improved the actual decision-making ability of the agent in the field of electronic confrontation, and can quickly and accurately select the best interference style, improve the interference success rate, and enhance competitiveness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120235045A_ABST
    Figure CN120235045A_ABST
Patent Text Reader

Abstract

The invention discloses a reinforced learning environment modeling and simulation platform and method in an electronic countermeasure game, which can provide a relatively real training environment for a jammer agent, effectively improve the actual combat decision-making ability of the jammer agent in the field of electronic countermeasure, and particularly, improve the practical combat decision-making ability of the jammer agent in different radar states and interference conditions, and improve the practical combat decision-making ability of the jammer agent in the field of electronic countermeasure. The optimal interference pattern can be quickly and accurately selected, the interference success rate is effectively improved, the competitiveness in an electronic confrontation game is enhanced, and more advantageous decision support is provided for practical application.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of electronic countermeasure game mechanisms and reinforcement learning, and particularly relates to a platform and method for modeling and simulating a reinforcement learning environment in an electronic countermeasure game. Background Art

[0002] Reinforcement learning is a machine learning method that can be divided into value-based methods and policy-based methods. Value-based methods solve reinforcement learning problems by learning the optimal value function, while policy-based methods solve reinforcement learning tasks by directly optimizing the agent's behavior policy function. Reinforcement learning learns how to achieve goals through interaction with the environment. In reinforcement learning, an agent affects the environment by performing actions, and the environment transfers states according to the agent's actions and provides immediate rewards. The goal of the agent is to learn a policy, that is, to select the best action in each state to maximize the long-term cumulative reward. The key elements of reinforcement learning include states, actions, rewards, and policies.

[0003] Current radar countermeasure simulation environment schemes for reinforcement learning usually use radar states, jamming patterns and parameters as the state space and action space, and use the signal-to-interference ratio or target detection probability to evaluate the immediate effect of a single interference as the reward. However, such schemes can only reflect the immediate effect of a single interference and cannot achieve a complete closed-loop evaluation from the interference encounter to the end of the confrontation based on the continuous tracking process of the radar on the target track. Therefore, the decision-making ability still needs to be improved.

[0004] In view of this, the present invention is specifically proposed. Summary of the Invention

[0005] The purpose of the present invention is to provide a platform and method for modeling and simulating a reinforcement learning environment in an electronic countermeasure game, so as to conduct a highly realistic game simulation for the reinforcement learning technology in the field of electronic countermeasures and improve the decision-making ability of the agent in a complex environment.

[0006] The purpose of the present invention is achieved through the following technical solutions:

[0007] A platform for modeling and simulating a reinforcement learning environment in an electronic countermeasure game includes: a simulation stepping module for realizing the interaction between the jammer agent and the environment. In each interaction, the input of the simulation stepping module is the jamming pattern decided by the jammer agent according to the environmental state at the current moment, and the output of the simulation stepping module is the environmental state at the next moment and the game result. The jammer agent updates its own policy parameters by combining the output of the simulation stepping module. The working process of the simulation stepping module includes:

[0008] Step 1: Calculate the working mode of the radar and the anti-jamming actions taken according to the environmental state at the previous moment;

[0009] Step 2: Determine the radar detection success rate based on the radar's operating mode, anti-jamming actions, and the jamming pattern of the input jammer agent, and sample the radar detection results. If the detection is successful, proceed to Step 3; otherwise, mark the radar detection success flag as failed, add a 0 at the end of the association vector. If the size of the association vector exceeds the set window size, remove the first element of the vector, increment the track measurement count by 1, and proceed to Step 5. The association vector is an array that records whether the radar has been successfully detected in recent times, where 0 indicates the most recent detection failure and 1 indicates the most recent detection success.

[0010] Step 3: Generate an error based on the average value and standard deviation of the radar detection distance deviation, and superimpose it on the jammer agent's position to obtain a track point.

[0011] Step 4: Determine the number of track points based on the size of the track measurement count, and take different actions in combination with the track points.

[0012] Step 5: If the number of 1s in the association vector is less than or equal to the track termination threshold, reset the beam request type, track, and the Kalman filter involved in the actions taken in Step 4 to their initial states.

[0013] Step 6: Calculate the radar descriptor based on the operating mode and anti-jamming actions, and perform a game termination determination and return the result.

[0014] A method for modeling and simulating a reinforcement learning environment in an electronic countermeasure game, implemented based on the aforementioned platform. The method includes: implementing a simulation stepping module for the interaction between the agent and the environment through the simulation stepping module in the platform. In each interaction, the input of the simulation stepping module is the jamming pattern determined by the jammer agent according to the environmental state at the current moment. The output of the simulation stepping module is the environmental state and the game result at the next moment. The jammer agent updates its own policy parameters in combination with the output of the simulation stepping module. The working process of the simulation stepping module includes:

[0015] Step 1: Calculate the radar's operating mode and the anti-jamming actions taken based on the environmental state at the previous moment.

[0016] Step 2: Determine the radar detection success rate based on the radar's operating mode, anti-jamming actions, and the jamming pattern of the input jammer agent, and sample the radar detection results. If the detection is successful, proceed to Step 3; otherwise, mark the radar detection success flag as failed, add a 0 at the end of the association vector. If the size of the association vector exceeds the set window size, remove the first element of the vector, increment the track measurement count by 1, and proceed to Step 5. The association vector is an array that records whether the radar has been successfully detected in recent times, where 0 indicates the most recent detection failure and 1 indicates the most recent detection success.

[0017] Step 3: Generate an error based on the average value and standard deviation of the radar detection distance deviation, and superimpose it on the jammer agent position to obtain a trace;

[0018] Step 4: Determine the number of track points according to the size of the track measurement number, and take different actions in combination with the said trace;

[0019] Step 5: If the number of 1s in the association vector is less than or equal to the track termination threshold, reset the beam request type, track, and the Kalman filter involved when taking actions in Step 4 to the initial state;

[0020] Step 6: Calculate the radar descriptor according to the working mode and anti-jamming actions, conduct game termination determination and return the result.

[0021] It can be seen from the technical solutions provided by the present invention above that a relatively realistic training environment can be provided for the agent, effectively improving its actual combat decision-making ability in the field of electronic countermeasures. Specifically, in the face of different radar states and jamming situations, it can quickly and accurately select the best jamming pattern, effectively improve the jamming success rate, enhance the competitiveness in the electronic countermeasure game, and provide more advantageous decision-making support for practical applications. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0023] Figure 1 It is a schematic diagram of a platform for reinforcement learning environment modeling and simulation in an electronic countermeasure game provided by an embodiment of the present invention;

[0024] Figure 2 It is a schematic diagram of the radar working process provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0025] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the protection scope of the present invention.

[0026] First, the following explanations will be given to the terms that may be used in this article:

[0027] Descriptions using terms such as "comprising", "including", "containing", "having" or other similar semantics shall be construed as non-exclusive inclusion. For example, including a technical feature element (such as raw materials, components, ingredients, carriers, dosage forms, materials, dimensions, parts, components, mechanisms, devices, steps, processes, methods, reaction conditions, processing conditions, parameters, algorithms, signals, data, products or articles, etc.) shall be construed as not only including the explicitly listed technical feature element, but also including other technical feature elements well-known in the art that are not explicitly listed.

[0028] The term "consisting of" means excluding any technical feature element that is not explicitly listed. If this term is used in a claim, it will make the claim a closed type, so that it does not include technical feature elements other than the explicitly listed ones, except for conventional impurities related thereto. If this term only appears in a sub-clause of a claim, then it only limits the elements explicitly listed in that sub-clause, and the elements recorded in other sub-clauses are not excluded from the overall claim.

[0029] The following will describe in detail a platform and method for modeling and simulating a reinforcement learning environment in an electronic countermeasure game provided by the present invention. Contents not described in detail in the embodiments of the present invention belong to the prior art well-known to those skilled in the art. For those conditions not specified in the embodiments of the present invention, they are carried out according to the conventional conditions in the art or the conditions recommended by the manufacturer. For the instruments used in the embodiments of the present invention that are not specified by the manufacturer, they are all conventional products that can be obtained through commercial purchase.

[0030] Embodiment 1

[0031] The embodiment of the present invention provides a platform for modeling and simulating a reinforcement learning environment in an electronic countermeasure game, as Figure 1 shown, which mainly includes: a simulation step module and a simulation reset module; wherein, the simulation step module is used to realize the interaction between the jammer agent and the environment; the simulation reset module is used to reset various types of information in the environment to the initial values.

[0032] (1) The simulation step module.

[0033] The input of the simulation step module is the jamming pattern determined by the jammer agent according to the environmental state at the current moment. The output of the simulation step module is the environmental state and the game result at the next moment. The jammer agent updates its own policy parameters in combination with the output of the simulation step module. The working process of the simulation step module includes:

[0034] Step 1: Calculate the working mode of the radar and the anti-jamming actions taken according to the environmental state at the previous moment.

[0035] Step 2: Determine the radar detection success rate based on the radar's working mode, anti-jamming actions, and the jamming pattern of the input jammer agent, and sample the radar detection results. If the detection is successful, proceed to Step 3; otherwise, mark the radar detection success flag as failed, add a 0 at the end of the correlation vector. If the size of the correlation vector exceeds the set window size, remove the first element of the vector, increment the track measurement count by 1, and proceed to Step 5.

[0036] In the embodiments of the present invention, the data source of the radar detection success rate is the statistical result of the simulation signal processing process under fixed working mode, anti-jamming actions, and jamming patterns, which can be in the form of a table (array). Using the working mode, anti-jamming actions, and jamming patterns as indexes, look up the corresponding radar detection success rate in the radar detection success rate, and then randomly determine whether the radar detection is successful based on this probability.

[0037] In the embodiments of the present invention, the correlation vector is an array that records whether the radar detection is successful in the recent several times (for example, the recent 8 times). 0 indicates that the most recent detection failed, and 1 indicates that the most recent detection was successful.

[0038] Step 3: Generate an error based on the average value of the radar detection range deviation and the standard deviation of the radar detection range deviation, and superimpose it on the position of the jammer agent to obtain a trace.

[0039] In the embodiments of the present invention, the average value of the radar detection range deviation and the standard deviation of the radar detection range deviation can be statistically obtained through the simulation signal processing process under fixed working mode, anti-jamming actions, and jamming patterns.

[0040] Step 4: Determine the number of track points based on the size of the track measurement count, and take different actions in combination with the said trace.

[0041] In the following sub-steps, the actions taken mainly refer to behaviors such as updating the correlation vector and the number of track points to maintain the track.

[0042] Step 41: If there is no point in the track, directly add the trace to the track, set the beam request type to confirm, add a 1 at the end of the correlation vector. If the size of the correlation vector exceeds the set window size, remove the first element of the vector, and assign the track measurement count as 1.

[0043] Step 42: If there is one point in the track, estimate the speed based on the point trace and the measurement time. If the estimated speed is within the preset speed threshold, add the point trace to the track, add a 1 at the end of the association vector. If the size of the association vector exceeds the set window size, remove the first bit of the vector, and increment the track measurement count by 1. Otherwise (i.e., the estimated speed is not within the preset speed threshold), discard the point trace, mark the radar detection success flag as failed, add a 0 at the end of the association vector. If the size of the association vector exceeds the set window size, remove the first bit of the vector, and increment the track measurement count by 1.

[0044] In the embodiments of the present invention, the measurement time is the time when the corresponding point trace in the track is recorded, and each step in the platform flows for 0.1 s. The estimated speed mainly refers to the relative speed between the two parties (jammer and radar).

[0045] Step 43: If there are two points in the track, estimate the speed based on the point trace and the measurement time. If the estimated speed is within the preset speed threshold, add the point trace to the track, add a 1 at the end of the association vector. If the size of the association vector exceeds the set window size, remove the first bit of the vector, and increment the track measurement count by 1, and establish a Kalman filter. Otherwise (i.e., the estimated speed is not within the preset speed threshold), discard the point trace, mark the radar detection success flag as failed, add a 0 at the end of the association vector. If the size of the association vector exceeds the set window size, remove the first bit of the vector, and increment the track measurement count by 1.

[0046] Step 44: If there are more than two points in the track, predict the target position through the Kalman filter, that is, the position of the jammer agent. If the distance between the point trace and the target position is less than the preset point trace - track association threshold, add the point trace to the track, add a 1 at the end of the association vector. If the size of the association vector exceeds the set window size, remove the first bit of the vector, and increment the track measurement count by 1, and update the Kalman filter. If the number of 1s in the association vector is greater than or equal to the preset confirmation transfer to tracking threshold, set the beam request type to tracking. Otherwise (i.e., the distance between the point trace and the target position is not less than the preset point trace - track association threshold), discard the point trace, mark the radar detection success flag as failed, add a 0 at the end of the association vector. If the size of the association vector exceeds the set window size, remove the first bit of the vector, and increment the track measurement count by 1.

[0047] Step 5: If the number of 1s in the association vector is less than or equal to the track termination threshold, reset the beam request type, the track, and the Kalman filter involved when taking actions in Step 4 to their initial states;

[0048] Step 6: Calculate the radar descriptor according to the working mode and anti - interference actions, perform the game termination determination and return the result.

[0049] In the embodiments of the present invention, for the radar descriptor obtained by statistically simulating the signal processing process in the fixed anti-interference operation and working mode in advance, the platform only needs to use the working mode and anti-interference operation as index values.

[0050] The platform provided by the above solution in the embodiments of the present invention is mainly used to simulate a radar and conduct a self-defense jamming and 1v1 scenario between the radar and an external agent (i.e., a jammer agent). Based on the above platform provided by the embodiments of the present invention, it is possible to conduct a highly realistic game simulation of reinforcement learning technology in the field of electronic countermeasures to improve the decision-making ability of agents in complex environments. The purpose of this platform is to train and evaluate the strategies of agents by simulating electronic countermeasure scenarios to achieve more efficient countermeasure decisions and action selections.

[0051] In order to more clearly show the technical solutions provided by the present invention and the technical effects produced, the following uses specific embodiments to describe in detail the platform provided by the embodiments of the present invention.

[0052] The embodiments of the present invention provide a platform for modeling and simulating a reinforcement learning environment in an electronic countermeasure game. First, the scenario will be introduced below, and then the details of the platform will be introduced.

[0053] 1. Simulation scenario setting

[0054] The present invention sets the simulation scenario as an air combat one-on-one electronic countermeasure, where both sides approach each other at a certain relative speed at a certain distance. Every once in a while, the radar agent sends a radio frequency signal to the jammer agent. The jammer agent can obtain the radar descriptor information corresponding to the radar agent and decide to send a jamming signal among ten jamming patterns. After the radar agent receives the mixed echo signal and jamming signal, it obtains the measured target position and switches its own working mode and task type. If the multiple measurement deviations of the radar agent are less than the threshold in the tracking state, it is determined that the radar agent is locked and the jamming fails. If the jammer agent keeps the radar unlocked and reaches the specified position, it is determined that the jamming is successful.

[0055] 2. Scenario modeling.

[0056] Taking the perspective of the jammer as an example, the following establishes a Markov decision process model, expressed as:

[0057]

[0058] Among them, s(n), o(n), and a(n) are the state, observation, and action of the environment at time n respectively; r(s(n), a(n)) is the reward obtained after performing action a(n) in state s(n); S, O, and A are the corresponding state space, observation space, and action space respectively.

[0059] Those skilled in the art can understand that the basis for the decision-making of the jammer agent is the observation (i.e., the radar description word generated by the platform); when the jammer agent interacts with the environment, the data received is only the observation and the reward, and the state only exists inside the environment and will not be directly provided to the jammer agent.

[0060] The specific definition of the state s(n) is as follows:

[0061] s(n) = [W(n), BR(n), A s (n)]

[0062] Among them, W(n) is the working mode of the radar.

[0063] In this embodiment, the working mode W(n) of the radar is set to be in {0, …, 9}, which respectively represent 10 working modes such as speed search, …, speed search - vertical screen display. BR(n) ∈ {0, 1, 2} is the type of beam request currently executed by the radar, which respectively represent search, confirmation, and tracking. A s (n) is the radar target detection result in recent several moments. For example, A s (n) = [A(n - 3), A(n - 2), A(n - 1), A(n)] is the radar target detection result in the recent four moments. The radar target detection result A(n) ∈ {0, 1} indicates whether the radar detects a target at the nth moment. A value of 1 means a target is detected, and a value of 0 means no target is detected.

[0064] For the transition of the working mode, it can be obtained through the Figure 2 shown radar working process. The transition of the beam request type is as follows:

[0065]

[0066] Among them,

[0067] For the transition of the radar target detection result, the present invention statistically analyzes the results of signal processing process simulation for a given radar state s(n) and the scenario of the jamming pattern, and obtains the radar target detection success rate in the corresponding scenario. Therefore, A(n + 1) can be directly sampled according to the success rate.

[0068] The specific definition of o(n) is as follows:

[0069] o(n) = {c(n), k(n), A(n), W(n)}

[0070] Among them, c(n) is the carrier frequency of the radar signal at the nth moment, and k(n) is the tuning frequency of the radar at the nth moment. These two variables will be calculated from the values in s(n) during the simulation.

[0071] The action a(n) includes multiple action types, which are defined as follows in this embodiment:

[0072] a(n) ∈ {0, …, 9}

[0073] The above formula represents the selection of interference patterns, where 0, …, 9 respectively refer to aiming suppression interference, blocking suppression interference, range pulling interference, intermittent sampling and forwarding interference, swept-frequency suppression interference, multi-tone interference, comb-shaped spectrum interference, range deception interference, dense false target interference, and noise product interference. The simulation platform generates interference signals according to the actions of the jammer agent, which are superimposed on the radar echo. The interference is manifested as the detection success rate and range measurement error of the radar under the influence of interference. The detection success rate and range measurement error are obtained by statistically processing the results of simulating the signal processing process for a given radar state s(n) and the scenario of the interference pattern.

[0074] The specific definition of the reward r is as follows:

[0075]

[0076] Based on the above introduction, the present invention implements a corresponding platform based on Python. The reinforcement learning algorithm can input the actions of the jammer agent to the platform and obtain the environmental state and game result at the next moment. The input variables and output variables of the platform are shown in Tables 1 and 2 respectively.

[0077] Table 1: Input variables of the platform

[0078] Variable Name Variable Meaning Variable Type Variable Range action Interference Pattern int 0-9

[0079] Table 2: Output variables of the platform

[0080] Variable Name Variable Meaning Variable Type Variable Range c1 Carrier Frequency of Sub - Pulse 1 float 6000.0e - 6008.0e6 c2 Carrier Frequency of Sub - Pulse 2 float 6000.0e - 6008.0e6 c3 Carrier Frequency of Sub - Pulse 3 float 6000.0e - 6008.0e6 k Intra - Pulse Modulation Parameter float 1.5e10, - 1.5e10 amp Radar Detection Success Flag bool 0-1 workmode Radar Working Mode int 0-9 state Environmental State int 0-359 done Termination Flag int 0-2

[0081] 3. Platform introduction.

[0082] This platform is mainly equipped with the simulation step module and simulation reset module described above. The simulation step module is used to realize the interaction between the jammer agent and the environment. Considering that the specific working process involved in the simulation step module has been introduced in detail above, it will not be elaborated here.

[0083] The simulation reset module is mainly used to reset various types of information in the environment to the initial values. Specifically: reset the tracks, filters, working modes, anti-jamming modes, simulation time, jammer positions, jammer speeds, and association vectors in the environment to the initial values, and return the initial state. The termination state flag returns false (value 0).

[0084] The platform provided by the present invention can be used for training and evaluation. Relevant examples are provided below for introduction.

[0085] Example 1: Training a Reinforcement Learning Agent

[0086] Declare a platform object in the reinforcement learning algorithm and call the simulation reset module to initialize the platform parameters.

[0087] Call the simulation step module of the platform multiple times to enable the jammer agent to interact with the environment. In each interaction, the jammer agent selects a jamming pattern based on the current environmental state and inputs it into the platform. The platform calculates the environmental state and game result at the next moment according to the internal algorithm and returns them to the jammer agent. The jammer agent updates its policy parameters based on the feedback results.

[0088] Example 2: Evaluating a Reinforcement Learning Agent

[0089] Evaluate the effectiveness of the agent's policy by conducting simulated confrontation experiments on the platform. Confront the trained agent with a radar set with a pre-set policy or manual operation, and record the game results of both sides.

[0090] The above solutions provided by the embodiments of the present invention mainly achieve the following beneficial effects:

[0091] (1) Targeted scenario simulation.

[0092] Focus on the electronic countermeasure scenario between the jammer and the radar, and closely combine the actual scenario to conduct modeling and simulation. It can accurately simulate the complex process of radar target tracking, covering various radar working modes, beam request types, and dynamic changes in target detection results, providing a relatively realistic training environment for the jammer agent and effectively improving its actual combat decision-making ability in the field of electronic countermeasures.

[0093] (2) In-depth signal-level model construction.

[0094] Based on the Markov decision process model, key elements such as state, observation, action, and reward are comprehensively and meticulously defined. Through the simulation and statistics of the radar signal processing process, key data such as the accurate radar target detection success rate are obtained, and the state transition and reward mechanism are determined accordingly.

[0095] (3) Significantly improve the agent's decision-making ability.

[0096] Through training and evaluation on this platform, the jammer agent can continuously learn and optimize its strategy in a simulation environment close to the real scenario. When facing different radar states and jamming situations, it can quickly and accurately select the best jamming pattern, effectively improve the jamming success rate, enhance its competitiveness in the electronic countermeasure game, and provide more advantageous decision-making support for actual electronic countermeasure operations.

[0097] Those skilled in the art can clearly understand that, for the convenience and conciseness of description, only the above division of each functional module is used as an example. In practical applications, the above functions can be allocated to different functional modules as needed, that is, the internal structure of the platform is divided into different functional modules to complete all or part of the functions described above.

[0098] Embodiment 2

[0099] The embodiment of the present invention provides a method for modeling and simulating a reinforcement learning environment in electronic countermeasure games, which is mainly implemented based on the platform provided in the foregoing embodiment. The method includes: implementing a simulation stepping module for the interaction between the agent and the environment through the simulation stepping module in the platform. In each interaction, the input of the simulation stepping module is the interference pattern determined by the jammer agent according to the environmental state at the current moment. The simulation stepping module outputs the environmental state and the game result at the next moment, and the jammer agent updates its own policy parameters in combination with the output of the simulation stepping module. The working process of the simulation stepping module includes:

[0100] Step 1: Calculate the working mode of the radar and the anti-jamming actions taken according to the environmental state at the previous moment.

[0101] Step 2: Determine the radar detection success rate based on the working mode of the radar, the anti-jamming actions, and the interference pattern of the input jammer agent, and sample the radar detection result. If the detection is successful, go to Step 3; otherwise, mark the radar detection success flag as failed, add a 0 at the end of the association vector. If the size of the association vector exceeds the set window size, remove the first element of the vector, increment the track measurement count by 1, and go to Step 5. Among them, the association vector is an array recording whether the radar has detected successfully in the recent several times. 0 indicates that the most recent detection failed, and 1 indicates that the most recent detection was successful.

[0102] Step 15: Generate an error based on the average value and standard deviation of the radar detection distance deviation, and superimpose it on the position of the jammer agent to obtain a trace.

[0103] Step 18: Determine the number of track points according to the size of the track measurement count, and take different actions in combination with the trace.

[0104] Step 21: If the number of 1s in the association vector is less than or equal to the track termination threshold, reset the beam request type, track, and the Kalman filter involved when taking actions in Step 18 to their initial states.

[0105] Step 24: Calculate the radar descriptor according to the working mode and anti-jamming actions, perform game termination determination, and return the result.

[0106] Furthermore, it also includes: resetting various types of information in the environment to their initial values through the simulation reset module.

[0107] As described above, it is only the preferred specific implementation manner of the present invention. However, the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims. The information disclosed in the background art part of this article is only intended to deepen the understanding of the overall background technology of the present invention, and should not be regarded as an admission or imply in any form that this information constitutes the prior art already known to those skilled in the art.

Claims

1. A platform for modeling and simulating reinforcement learning environments in electronic confrontation games, characterized in that: include: A simulation stepping module for realizing the interaction between the jammer agent and the environment. In each interaction, the simulation stepping module inputs the jamming style decided by the jammer agent according to the current environmental state. The simulation stepping module outputs the environmental state and game result at the next moment. The jammer agent updates its own strategy parameters in combination with the output of the simulation stepping module. The working process of the simulation step module includes: Step 1: Calculate the radar's operating mode and the anti-interference actions taken based on the environmental status at the last moment; Step 2: According to the radar working mode and anti-interference action, the interference pattern of the input jammer intelligent agent is used to determine the radar detection success rate and sample the radar detection results; if the detection is successful, proceed to step 3; otherwise, the radar detection success mark is recorded as failure, and a 0 is added to the end of the association vector. If the size of the association vector exceeds the set window size, the vector head is removed, the track measurement number is increased by 1, and the process proceeds to step 5; where the association vector is an array that records whether the radar has successfully detected the last few times, 0 indicates that the most recent detection failed, and 1 indicates that the most recent detection was successful; Step 3: Generate an error based on the average value of the radar detection distance deviation and the standard deviation of the radar detection distance deviation, and superimpose it with the jammer agent position to obtain a point trace; Step 4: Determine the number of track points according to the size of the track measurement number, and take different actions in combination with the track points; Step 5: If the number of 1s in the association vector is less than or equal to the track termination threshold, the beam request type, track, and the Kalman filter involved in taking the action in step 4 are reset to the initial state; Step 6: Calculate the radar description word according to the working mode and anti-interference action, make a game termination judgment and return the result.

2. The platform for modeling and simulating a reinforcement learning environment in an electronic confrontation game according to claim 1, characterized in that: The different actions taken according to the size of the track and in combination with the point track include: If there is no point in the track, the point track is directly added to the track, the beam request type is set to confirm, a 1 is added to the end of the association vector, and if the size of the association vector exceeds the set window size, one bit is removed from the vector head, and the track measurement number is assigned a value of 1.

3. The platform for modeling and simulating a reinforcement learning environment in an electronic confrontation game according to claim 1, characterized in that: The different actions taken according to the size of the track and in combination with the point track include: If there is a point in the track, the speed is estimated by the point track and the measurement time; if the estimated speed is within the preset speed threshold, the point track is added to the track, and a 1 is added to the end of the associated vector. If the size of the associated vector exceeds the set window size, the first bit of the vector is removed, and the track measurement number is increased by 1; otherwise, the point track is discarded, the radar detection success mark is marked as a failure, and a 0 is added to the end of the associated vector. If the size of the associated vector exceeds the set window size, the first bit of the vector is removed, and the track measurement number is increased by 1.

4. The platform for modeling and simulating a reinforcement learning environment in an electronic confrontation game according to claim 1, characterized in that: The different actions taken according to the size of the track and in combination with the point track include: If there are two points in the track, the speed is estimated by the point track and the measurement time. If the estimated speed is within the preset speed threshold, the point track is added to the track, and a 1 is added to the end of the associated vector. If the size of the associated vector exceeds the set window size, the first bit of the vector is removed, the track measurement number is increased by 1, and a Kalman filter is established; otherwise, the point track is discarded, the radar detection success mark is marked as a failure, a 0 is added to the end of the associated vector, and if the size of the associated vector exceeds the set window size, the first bit of the vector is removed, and the track measurement number is increased by 1.

5. The platform for modeling and simulating a reinforcement learning environment in an electronic confrontation game according to claim 1, characterized in that: The different actions taken according to the size of the track and in combination with the point track include: If there are more than two points in the track, the target position is predicted through the Kalman filter. If the distance between the point track and the target position is less than the preset point track-track association threshold, the point track is added to the track, and a 1 is added to the end of the association vector. If the size of the association vector exceeds the set window size, the first bit of the vector is removed, the track measurement number is increased by 1, and the Kalman filter is updated. If the number of 1s in the association vector is greater than or equal to the preset confirmation-to-tracking threshold, the beam request type is set to tracking; otherwise, the point track is discarded, the radar detection success mark is marked as failure, and a 0 is added to the end of the association vector. If the size of the association vector exceeds the set window size, the first bit of the vector is removed, and the track measurement number is increased by 1.

6. The platform for modeling and simulating a reinforcement learning environment in an electronic confrontation game according to claim 1, characterized in that: Also includes: The simulation reset module is used to reset various information in the environment to initial values.

7. A method for modeling and simulating a reinforcement learning environment in an electronic confrontation game, characterized in that: Based on the platform implementation described in any one of claims 1 to 6, the method includes: implementing a simulation stepping module for the interaction between an agent and an environment through a simulation stepping module in the platform, in each interaction, the simulation stepping module inputs the interference pattern decided by the jammer agent according to the environmental state at the current moment, the simulation stepping module outputs the environmental state and the game result at the next moment, and the jammer agent updates its own strategy parameters in combination with the output of the simulation stepping module; the working process of the simulation stepping module includes: Step 1: Calculate the radar's operating mode and the anti-interference actions taken based on the environmental status at the last moment; Step 2: According to the radar working mode and anti-interference action, the interference pattern of the input jammer intelligent agent is used to determine the radar detection success rate and sample the radar detection results; if the detection is successful, proceed to step 3; otherwise, the radar detection success mark is recorded as failure, and a 0 is added to the end of the association vector. If the size of the association vector exceeds the set window size, the vector head is removed, the track measurement number is increased by 1, and the process proceeds to step 5; where the association vector is an array that records whether the radar has successfully detected the last few times, 0 indicates that the most recent detection failed, and 1 indicates that the most recent detection was successful; Step 3: Generate an error based on the average value of the radar detection distance deviation and the standard deviation of the radar detection distance deviation, and superimpose it with the jammer agent position to obtain a point trace; Step 4: Determine the number of track points according to the size of the track measurement number, and take different actions in combination with the track points; Step 5: If the number of 1s in the association vector is less than or equal to the track termination threshold, the beam request type, track, and the Kalman filter involved in taking the action in step 4 are reset to the initial state; Step 6: Calculate the radar description word according to the working mode and anti-interference action, make a game termination judgment and return the result.

8. The method for modeling and simulating a reinforcement learning environment in an electronic confrontation game according to claim 7, characterized in that: Also includes: The simulation reset module is used to reset various types of information in the environment to their initial values.