Methods, systems, equipment, and media for evaluating adversarial effectiveness based on MCTS and self-game theory.
By conducting self-game training in a radar simulator environment, and combining MCTS and deep neural networks to optimize the policy network and value network, the shortcomings of existing radar countermeasure effectiveness evaluation technologies are addressed, achieving high-precision and real-time countermeasure effectiveness evaluation and strategy optimization.
Patent Information
- Application Number
- CN202411927670.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-25
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2044-12-25
AI Technical Summary
Existing radar countermeasures effectiveness assessment technologies have not fully utilized intelligent technologies, and the core technologies are not mature enough, making it difficult to provide effective strategy optimization and accurate assessment in complex confrontation environments.
We employ a method based on Monte Carlo Tree Search (MCTS) and self-game theory. By constructing a radar simulator environment and an interfering agent, we conduct self-game training. We then utilize Monte Carlo Tree Search and deep neural networks to optimize the policy network and value network, generating high-quality training data to optimize the agent's policy and value assessment.
It improves the accuracy and real-time performance of radar countermeasures assessment, and can provide reliable strategy optimization support in complex confrontation environments, adapting to rapid changes and complex feedback from adversaries' strategies.
Smart Images

Figure CN119849310B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of adversarial effectiveness evaluation technology, and in particular to an adversarial effectiveness evaluation method, system, device and medium based on MCTS and self-game. Background Technology
[0002] In recent years, the development of intelligent technologies has provided new solutions for radar countermeasures effectiveness assessment. For example, artificial intelligence has already been introduced into electronic warfare systems, improving the accuracy and efficiency of assessments. Intelligent radar countermeasures systems, through real-time signal processing and decision optimization using deep learning, can effectively cope with the challenges of complex electromagnetic environments.
[0003] Although some research has yielded preliminary results, overall, current interference effectiveness assessment techniques still need improvement. They have not yet fully utilized intelligent technologies, the core technologies are not mature, and research directions are relatively scattered.
[0004] In view of this, the present invention is hereby proposed. Summary of the Invention
[0005] The purpose of this invention is to provide a method, system, device, and medium for evaluating adversarial effectiveness based on MCTS and self-game, which not only improves the accuracy and real-time performance of effectiveness evaluation, but also provides reliable technical support for strategy optimization in complex adversarial environments.
[0006] The objective of this invention is achieved through the following technical solution:
[0007] An adversarial effectiveness evaluation method based on MCTS and self-game theory includes:
[0008] Step 1: Build a radar simulator environment and construct the jamming agent corresponding to the jammer.
[0009] Step 2: Use the interfering agent to engage in self-game with the radar simulator environment, including: the interfering agent makes the optimal action based on the state information obtained from the radar simulator environment and the MCTS, and applies it to the radar simulator environment; it continues to obtain state information from the radar simulator environment and make optimal action decisions; and stores the data generated by the self-game; where MCTS is Monte Carlo Tree Search.
[0010] Step 3: Train the interfering agent using the stored data;
[0011] Step 4: Alternately execute steps 2 and 3 until the set stopping condition is met to obtain the trained interfering agent. Deploy the trained interfering agent in the actual adversarial scenario to output the optimal action and the corresponding state value. Use the state value as the evaluation result of the adversarial effectiveness.
[0012] An adversarial performance evaluation system based on MCTS and self-game theory includes:
[0013] The environment and agent construction unit is used to construct the radar simulator environment and the jamming agent corresponding to the jammer.
[0014] The sample data collection unit based on MCTS and self-game is used to conduct self-games between a disruptive agent and a radar simulator environment. This includes: the disruptive agent determining the optimal action based on state information obtained from the radar simulator environment, combined with MCTS, and applying it to the radar simulator environment; continuing to obtain state information from the radar simulator environment and making action decisions; storing the data generated by the self-game, including the determined optimal action and the actual results of the confrontation; where MCTS is Monte Carlo Tree Search.
[0015] The agent training unit is used to train the interfering agent using stored data.
[0016] The performance evaluation unit controls the alternating operation of the sample data collection unit and the agent training unit based on MCTS and self-game until the set stopping conditions are met, obtains the trained interfering agent, and deploys the trained interfering agent in the actual adversarial scenario to output the optimal action and the corresponding state value. The state value is used as the evaluation result of adversarial performance.
[0017] A processing device includes: one or more processors; and a memory for storing one or more programs;
[0018] When the one or more programs are executed by the one or more processors, the one or more processors implement the aforementioned method.
[0019] A readable storage medium storing a computer program that, when executed by a processor, implements the aforementioned method.
[0020] As can be seen from the technical solution provided by the present invention, by combining Monte Carlo tree search and deep neural networks (agents), a self-game mechanism is innovatively introduced. During the self-game process, high-quality training data is generated by simulating a real adversarial environment, and the policy network and value network of the agent are continuously optimized. The present invention effectively searches and evaluates in a large state space (mainly referring to the large number of radar states in an adversarial environment), and can adapt to the rapid changes and complex feedback of the opponent's strategy in real time. Through multi-round self-game simulation, it not only improves the accuracy and real-time performance of performance evaluation, but also provides reliable technical support for strategy optimization in complex adversarial environments. Attached Figure Description
[0021] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0022] Figure 1 A flowchart of an adversarial effectiveness evaluation method based on MCTS and self-game provided for an embodiment of the present invention;
[0023] Figure 2 A schematic diagram of a self-game process based on MCTS provided for an embodiment of the present invention;
[0024] Figure 3 A schematic diagram of the Monte Carlo tree search process provided in an embodiment of the present invention;
[0025] Figure 4 A schematic diagram illustrating the training of a model based on Monte Carlo tree search and self-game, provided for an embodiment of the present invention;
[0026] Figure 5 A schematic diagram of a simulation experiment provided in an embodiment of the present invention;
[0027] Figure 6 A schematic diagram of an adversarial effectiveness evaluation system based on MCTS and self-game provided in an embodiment of the present invention;
[0028] Figure 7 This is a schematic diagram of a processing device provided in an embodiment of the present invention. Detailed Implementation
[0029] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the protection scope of the present invention.
[0030] First, the following explanations are provided for the terms that may be used in this article:
[0031] The terms “including,” “comprising,” “containing,” “having,” or other similar semantic descriptions should be interpreted as non-exclusive inclusion. For example, “including a technical feature element (such as raw material, component, ingredient, carrier, dosage form, material, size, part, component, mechanism, device, step, process, method, reaction conditions, processing conditions, parameter, algorithm, signal, data, product or article of manufacture, etc.)” should be interpreted as including not only the expressly listed technical feature element, but also other technical feature elements that are not expressly listed and are well-known in the art.
[0032] The term "composed of" excludes any technical features not expressly listed. When used in a claim, it closes the claim to exclude all technical features other than those expressly listed, except for associated conventional impurities. If the term appears only in a clause of a claim, it limits the claim to the elements expressly listed in that clause; elements recited in other clauses are not excluded from the overall claim.
[0033] The following provides a detailed description of the adversarial effectiveness evaluation method, system, device, and medium based on MCTS and self-game theory provided by this invention. Contents not described in detail in the embodiments of this invention belong to prior art known to those skilled in the art. Where specific conditions are not specified in the embodiments of this invention, conventional conditions in the art or conditions recommended by the manufacturer shall apply. Where the manufacturers of the instruments used in the embodiments of this invention are not specified, they are all conventional products that can be purchased commercially.
[0034] Example 1
[0035] like Figure 1 As shown, an adversarial effectiveness evaluation method based on MCTS and self-game mainly includes the following steps:
[0036] Step 1: Build a radar simulator environment and construct the jamming agent corresponding to the jammer.
[0037] Step 2: Use the interfering agent to engage in self-play with the radar simulator environment and store the data generated by the self-play.
[0038] In this embodiment of the invention, the self-game process based on MCTS includes: the interfering agent makes a decision on the optimal action based on the state information obtained from the radar simulator environment and MCTS (Monte Carlo Tree Search), and applies it to the radar simulator environment, and continues to obtain state information from the radar simulator environment and make a decision on the optimal action.
[0039] In this embodiment of the invention, obtaining state information from the radar simulator environment includes: receiving radar pulse signals from the radar simulator environment and extracting pulse-related parameters, combining all the extracted parameters to obtain a pulse description word vector, and using the pulse description word vector as state information.
[0040] Preferably, the MCTS decision to determine the optimal action includes: 1) performing a Monte Carlo tree search based on the state information obtained from the radar simulator environment, recursively selecting the most valuable child node from the root node until a leaf node is reached; wherein, the Monte Carlo tree search is implemented based on a tree model, in which nodes are associated with state information and edges are associated with actions, the root node is the node corresponding to the state information obtained from the radar simulator environment, and each edge stores information including: number of visits, action probability, average value of the action, and total state value; the most valuable child node is evaluated through the information stored in the corresponding edge; 2) determining the available actions based on the state corresponding to the leaf node, and expanding them to the child nodes of the leaf node; 3) evaluating the action probability and state value of each child node through the interfering agent; 4) updating the information stored in the corresponding edges in the search path in reverse; 5) selecting the optimal action using the information stored in the updated edges.
[0041] In this embodiment of the invention, the data generated by the game includes at least: the result of each game (i.e., the actual result of the confrontation), the optimal action decided, etc.
[0042] Step 3: Train the interfering agent using the stored data.
[0043] In this embodiment of the invention, the interference agent is trained using reinforcement learning with data generated from self-games.
[0044] Step 4: Alternately execute steps 2 and 3 until the set stopping condition is met to obtain the trained interfering agent. Deploy the trained interfering agent in the actual adversarial scenario to output the optimal action and the corresponding state value. Use the state value as the evaluation result of the adversarial effectiveness.
[0045] In this embodiment of the invention, a stopping condition can be set according to the actual situation. For example, the maximum number of times steps 2 and 3 are executed alternately can be set as the stopping condition, or the convergence of the loss function can be set as the stopping condition.
[0046] To more clearly demonstrate the technical solution and its effects provided by the present invention, the method provided by the embodiments of the present invention will be described in detail below with reference to specific examples.
[0047] I. Introduction to Electronic Warfare Scenarios
[0048] Consider the dynamic game between the opposing sides. In this scenario, the radar (R) state is no longer discretized. Under the same operating mode, multiple sets of different waveform parameters often exist, and the radar selects different waveform parameters based on target and environmental characteristics. The jammer (J) receives the radar's pulse signal and extracts its parameters, mainly including pulse width (PW), pulse amplitude (PA), pulse repetition interval (PRI), and radio frequency (RF). All parameters are then combined to obtain a pulse description word (PDW) vector. Upon receiving the PDW, the jammer determines its jamming action based on a pre-set algorithm and applies this action as a signal to the radar. The radar adjusts its operating mode and parameters accordingly to dynamically engage in the game with the jammer. The scenario in this invention will be constructed using a real radar simulation platform; therefore, the radar's own state switching, parameter adjustments, and the impact of jamming actions will be implemented by the platform itself. This scenario will be modeled as a dual-view MDP (Markov Decision Process) between the radar and the jammer.
[0049] In the above scenario, the perspective of the interfering party is:
[0050] {s J (t)∈S J ,a J (t)∈A J ,p J (s J (t+1)|s J (t),a J (t),s R (t),a R (t)),r J (s J (t),a J (t),s R (t),a r (t))}
[0051] The radar perspective is:
[0052] {s R (t)∈S R ,a R (t)∈A R ,p R (s R (t+1)|s R (t),a R (t),s J (t),a J (t)),rR (s J (t),a J (t),s R (t),a R (t))}
[0053] Among them, s J S is the PDW vector (representing state information) of the radar signal received by the jamming party after sorting. J For a series of s J The set composed of a J For the interfering party's interference action (i.e., the action decided), A J p is the set of interference actions. J For the state transition of the radar side, r J For the countermeasure effectiveness of the interfering party, t represents time step t; similarly, s R S represents the radar-detected trajectory of the interfering party, including relevant information such as distance, angle, and radial velocity. R For a series of s R The set composed of a R Anti-jamming actions taken by the radar, A R Anti-interference action set, p R For the state transition of the interfering party, r R This enhances the radar's anti-jamming capabilities.
[0054] Taking the perspective of the interfering party as an example, the goal is to find the best possible situation for the interfering party in the current adversarial context. This invention uses a state value function to characterize this. Since the immediate reward r... J (That is, the effectiveness of the interfering party) can only show the quality of the interfering action at the current time step, and cannot reflect the quality of the interfering strategy. Therefore, a state value function is needed to reflect the long-term value of the strategy.
[0055]
[0056] in, For the long-term reward of the interfering agent under policy π, γ k This is a discount factor, with a value range of 0 < γ. k ≤1, the discount factor is used to reduce the importance of future rewards. The larger the value (i.e., the closer to 1), the more emphasis is placed on long-term rewards; the smaller the value, the more emphasis is placed on short-term rewards. The parameter k represents the consideration of long-term discounted rewards, and the influence is smaller for states further away; s represents the state, s(t) = s J This indicates that the state considered in this part is s. J .
[0057] When interference is performed according to a certain strategy, the state value function is expressed as:
[0058]
[0059] Where, p J (s′ J |s J ,a J ,s R ,a R ) indicates that the interfering party is in state s J The following interference action a J And the radar in state s R Take anti-interference action a R The interfering party receives state s′ J The probability of r J (s J ,a J ,s R ,a R ) indicates that the interfering party is in state s J The following interference action a J And the radar in state s R Take anti-interference action a R Then, the countermeasure effectiveness of the interfering party is determined (i.e., even if rewarded), and γ is the discount factor.
[0060] Ultimately, the optimal strategy of an MDP can be intuitively expressed by the following formula for any adversarial situation s. J Next, find a way to convert value functions. Approximate maximization strategy
[0061]
[0062] Similarly, the radar's objective is also to find a strategy that maximizes the radar's value function. Its state value function is:
[0063]
[0064] in, Let p be the state-value function of a radar with strategy π, used to measure the radar's long-term return. R (s′ R |s J ,a J ,s R ,a R ) indicates that the interfering party is in state s J The following interference action a J And the radar in state s R Take anti-interference action a R The radar then received the status s′ R The probability, r R (sJ ,a J ,s R ,a R ) indicates that the interfering party is in state s J The following interference action a J And the radar in state s R Take anti-interference action a R Afterwards, the radar's countermeasure effectiveness was determined.
[0065] Considering that the target sought by the radar side corresponds to the target sought by the jamming side, the meanings of the relevant symbols can be understood together, so they will not be elaborated further.
[0066] II. Self-game process based on MCTS (Monte Carlo Tree Search).
[0067] In a real-world environment, the adversarial effectiveness r J Often unknowable. To address this challenge, this invention employs an adversarial performance evaluation method based on Monte Carlo tree search and self-game theory. This method combines Monte Carlo tree search with deep neural networks for self-learning, with each search involving a series of simulated self-games.
[0068] like Figure 2 As shown, in the self-game phase, the interfering party and the simulated radar engage in a battle. In each state, the interfering party selects the optimal action through Monte Carlo tree search, guided by the current policy and state value. During each battle, all states, selected actions, and the final outcome are recorded. This adversarial data is stored and used to train a deep neural network. The network parameters are updated by minimizing the policy loss and value loss functions, thereby improving its performance in future battles. Monte Carlo tree search is based on a tree model. In the tree model, nodes are associated with states, and edges are associated with actions performed. Each edge is interconnected with the node it points to. It can be understood that after a state associated with a node performs an action associated with an edge, it transitions to a state associated with another node. The information stored in each edge is: N(s,a), P(s,a), Q(s,a), W(s,a), where the symbols represent: number of visits N, average value Q, total state value W, and action probability P, respectively. The search process is the growth process of the tree model: starting from the root node (the node corresponding to the current state), actions are selected to advance to higher states, the tree nodes are expanded, and the utility is evaluated until the search termination condition is met. The final decision is generated from the resulting tree model and is the action with the highest utility value at the root node. The core of MCTS is a combination of optimal finite search and Monte Carlo estimation. It builds a tree model based on random sampling, a process involving repeated iterative searches that terminates when certain constraints (such as computation time, maximum number of iterations, etc.) are met.
[0069] Figure 2 In this context, s represents the state, p represents the state value distribution, t represents the time step, Tmax represents the maximum time step, and z represents the final outcome of the adversarial process, such as whether the interference is successful or not.
[0070] like Figure 3 The diagram illustrates the Monte Carlo tree search process, which mainly consists of four stages, described below:
[0071] (1) Selection
[0072] Starting from the root node, recursively select the most valuable child node until a leaf node (i.e., a node with no children) is found. The selection strategy is based on a balance between the value of the child node, prior probability, and the number of visits. Specifically, the score of the actions associated with each edge is calculated using UCB (Upper Confidence Bound), and the child node pointed to by the edge with the highest score is selected as the most valuable child node. The relevant calculation formula is expressed as follows:
[0073]
[0074] Among them, a K Let K represent the action corresponding to the edge with the highest score, and let Q(s,a), P(s,a), and N(s,a) represent the information stored in the edge, namely the average value of the action, the probability of the action, and the number of visits, respectively. Let (s,a) represent the action a being performed in state s. b N(s,b) represents the total number of visits to all actions in state s, b is the index symbol, referring to any action among all actions; C is a hyperparameter that controls the exploration intensity, which determines that c, which has a larger weight in action selection, tends to explore new actions more actively.
[0075] (2) Expansion.
[0076] The search proceeds according to the selection method described above until a leaf node is encountered. The leaf node is then added to the search tree, and the environment's mechanisms provide the available actions for the current state (i.e., the state corresponding to the leaf node). These available actions are used to expand the child nodes, while simultaneously initializing their information: {N(s,a)=0,Q(s,a)=0,P(s,a)}.
[0077] (3) Evaluation
[0078] The action probability of each child node is evaluated through the policy network in the interfering agent, and the state value of each child node is evaluated through the value network in the interfering agent. The action probability of each child node is then stored in the edge between each child node and the leaf node.
[0079] Figure 3 In the middle, p l ,(v l )=f θ (.) indicates that the state of the leaf node is input to the interfering agent f. θ Obtain the corresponding action probability distribution p l and value v l Action probability distribution p l The state value p containing two child nodes l,0 With p l,1 .
[0080] (4) Backpropagation.
[0081] Near the end of this simulation, W, Q, and N on the search path are updated. The evaluation values are backpropagated back to all nodes in the search path. For each visited node, its visit count is incremented by 1, and its value is updated to the average value of the k simulations already performed.
[0082] Specifically: the search path includes all nodes and their corresponding edges from the root node to the child nodes; the information stored in all edges of this search path is updated using the following formula:
[0083] N(s,a)←N(s,a)+1
[0084] W(s,a)←W(s,a)+v L
[0085]
[0086] Where ← is the assignment operator, N(s,a) is the number of visits, W(s,a) is the total state value, and v L Let W(s,a) be the state value of the child node; Q(s,a) be the average value of the action; and W(s,a) and N(s,a) on the right side of the equation be the updated information.
[0087] The above process is repeated multiple times. After the MCTS search is completed, the `visit_counts` set is returned at the root node. Then, an action is selected based on the number of visits to the edges of the root node. Specifically, it selects the action with the highest number of visits or samples the distribution derived from `visit_counts` to obtain the final action to execute.
[0088] All training samples in this invention are generated through a self-game approach, with the policy function and value function continuously optimized during training. These two functions guide the Monte Carlo Tree Search (MCTS) search, while the MCTS search results interact with the environment to generate training samples, further optimizing the objective function. Through this gradual improvement of the strategy, higher-quality training data can be continuously obtained.
[0089] III. Model training based on Monte Carlo tree search and self-game.
[0090] like Figure 4 The diagram shows an overview of the training of a model based on Monte Carlo tree search and self-game. The upper part shows the training process of optimizing the policy network through cross-entropy loss and optimizing the value network through mean squared error loss. The lower part shows the reasoning process of the interfering party giving actions during electronic warfare.
[0091] In the training process, data samples are first obtained through self-game analysis, as described in Part II above. In this self-game, the root node of a Monte Carlo search tree is constructed based on the current state. The UCB score is calculated using the policy network and value network, and the action with the highest score is selected, continuing until a leaf node is encountered. The leaf node is then expanded, and the node information along the path is updated in reverse based on the output of the value network, with multiple simulations performed. Actions are executed under various states. After each round of adversarial training, the sample data is stored in an experience replay pool, from which samples are used to train reinforcement learning cross-entropy loss for optimization. As the policy and value function improve, the training samples also continuously improve, resulting in a gradually enhanced capability of the trained agent (implemented using a deep neural network model).
[0092] In this embodiment of the invention, the following loss function is used to optimize the policy network and value network in the agent:
[0093] loss = l v (z t ,v t )+l p (π t ,p t )+c||θ|| 2
[0094] Where, c||θ|| 2 Let θ be the square norm (L2 norm) of the interfering agent's parameters, also known as the weight penalty; c is a hyperparameter that controls the strength of regularization; the larger the value, the more strictly the size of parameter θ is restricted. v (z t ,v t )=(z t -v t ) 2 , l v (z t ,v t ) represents the loss function of the value network (using the mean squared error loss function), T is the transpose sign, and v t The state value output by the value network, z t For self-game data, it refers to the actual outcome of the confrontation (i.e., the game result). p (π t ,p t ) represents the loss function of the policy network (using the cross-entropy loss function), π t The strategy obtained from MCTS, i.e., the optimal action decided in the self-game, p t This represents the action probability distribution output by the policy network.
[0095] The second and third parts are executed alternately, allowing the interfering agent and the MCTS-based self-game process to improve together.
[0096] The method provided in this invention is applicable to complex environments of dynamic game theory between opposing strategies. By combining Monte Carlo Tree Search (MCTS) and deep neural networks, this method can effectively search and evaluate adversarial effectiveness within a large state space. During the game process, the interfering agent engages in multiple rounds of adversarial interaction with the radar simulator environment, generating high-quality training data and continuously optimizing the policy network and value network. In the decision-making process, MCTS selects high-value, low-access actions for searching and updates node information (nodes and edges are interconnected) through backpropagation, thereby providing dynamic effectiveness evaluation and optimization schemes. This method not only improves the accuracy and real-time performance of effectiveness evaluation but also provides reliable technical support for strategy optimization in complex adversarial environments.
[0097] IV. Example Explanation.
[0098] like Figure 5 As shown, an example of a simulation experiment is provided. During the simulation, it is mainly used to provide a state description of the radar. In this example, the simulation platform is implemented using MATLAB, while the entire method is implemented using Python, presenting a platform barrier for data interaction. During training, the platform acts as the environment, and the interfering party acts as the agent; data interaction between the two is achieved through Redis. Redis is an open-source in-memory database that stores data in key-value pairs. Because the data is stored in memory, Redis has extremely fast read and write speeds, capable of processing millions of requests per second. Through Redis, the agent can interact with a large amount of data from the environment during training. The main flow of this example is as follows:
[0099] (1) Before the training phase, a radar environment is constructed based on the existing radar knowledge base.
[0100] (2) During the training process, MCTS is used to construct the jamming agent to simulate the confrontation with the radar.
[0101] (3) The interfering agent collects training data into the buffer by playing a game with the constructed radar simulator environment.
[0102] (4) Sample data from the buffer to train the interfering agent, and alternate between steps (3) and (4) until the stopping condition is met.
[0103] (5) After training, in actual confrontation (applied to real environment), the trained interfering agent makes the decision on the optimal action and obtains the state value corresponding to the optimal action. The state value is used as an immediate reward, which is the evaluation result of the confrontation effectiveness.
[0104] Those skilled in the art will understand that the interfering agent comprises a policy network and a value network, wherein the policy network is responsible for making decisions on the optimal action, and the value network is responsible for outputting the corresponding state value.
[0105] In the above example, based on the self-game process and training using MCTS, accurate adversarial effectiveness can be obtained, guiding the jamming process of electronic countermeasures. Furthermore, Figure 5 The document lists several functions involved in the processing. For example, the collect() function below the collector is a sample collection function, which collects sample data generated from the game process. Since the functions involved are all common functions in this field, their meanings can be understood by referring to conventional techniques, so they will not be elaborated further.
[0106] Through the above description of the embodiments, those skilled in the art can clearly understand that the above embodiments can be implemented by software, or by using software plus necessary general-purpose hardware platforms. Based on this understanding, the technical solutions of the above embodiments can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, external hard drive, etc.), including several instructions to cause a computer device (such as a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments of the present invention.
[0107] Example 2
[0108] This invention also provides an adversarial performance evaluation system based on MCTS and self-game theory, which is mainly used to implement the methods provided in the foregoing embodiments, such as... Figure 6 As shown, the system mainly includes:
[0109] The environment and agent construction unit is used to construct the radar simulator environment and the jamming agent corresponding to the jammer.
[0110] The sample data collection unit based on MCTS and self-game is used to conduct self-games between a disruptive agent and a radar simulator environment. This includes: the disruptive agent determining the optimal action based on state information obtained from the radar simulator environment, combined with MCTS, and applying it to the radar simulator environment; continuing to obtain state information from the radar simulator environment and making action decisions; storing the data generated by the self-game, including the determined optimal action and the actual results of the confrontation; where MCTS is Monte Carlo Tree Search.
[0111] The agent training unit is used to train the interfering agent using stored data.
[0112] The performance evaluation unit controls the alternating operation of the sample data collection unit and the agent training unit based on MCTS and self-game until the set stopping conditions are met, obtains the trained interfering agent, and deploys the trained interfering agent in the actual adversarial scenario to output the optimal action and the corresponding state value. The state value is used as the evaluation result of adversarial performance.
[0113] Since the main technical details involved in the system have been described in detail in the previous embodiments, they will not be repeated here.
[0114] Those skilled in the art will understand that, for the sake of convenience and brevity, the above-described division of functional modules is used as an example. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the system can be divided into different functional modules to complete all or part of the functions described above.
[0115] Example 3
[0116] The present invention also provides a processing device, such as Figure 7 As shown, it mainly includes: one or more processors; a memory for storing one or more programs; wherein, when the one or more programs are executed by the one or more processors, the one or more processors implement the method provided in the foregoing embodiments.
[0117] Furthermore, the processing device also includes at least one input device and at least one output device; in the processing device, the processor, memory, input device, and output device are connected via a bus.
[0118] In this embodiment of the invention, the specific types of the memory, input device, and output device are not limited; for example:
[0119] Input devices can be touchscreens, image acquisition devices, physical buttons, or mice, etc.
[0120] The output device can be a display terminal;
[0121] The memory can be random access memory (RAM) or non-volatile memory, such as disk storage.
[0122] Example 4
[0123] The present invention also provides a readable storage medium storing a computer program that, when executed by a processor, implements the method provided in the foregoing embodiments.
[0124] In this embodiment of the invention, the readable storage medium is a computer-readable storage medium and can be disposed in the aforementioned processing device, for example, as a memory in the processing device. Furthermore, the readable storage medium can also be any medium capable of storing program code, such as a USB flash drive, portable hard drive, read-only memory (ROM), magnetic disk, or optical disk.
[0125] The above description is merely a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims. The information disclosed in the background section is intended only to enhance the understanding of the overall background technology of the present invention and should not be construed as an admission or implication in any way that such information constitutes prior art known to those skilled in the art.
Claims
1. A method for evaluating adversarial effectiveness based on MCTS and self-game theory, characterized in that, include: Step 1: Build a radar simulator environment and construct the jamming agent corresponding to the jammer. Step 2: The interfering agent engages in a self-game with the radar simulator environment, including: the interfering agent, based on state information obtained from the radar simulator environment and combined with Monte Carlo Tree Search (MCTS), determines the optimal action and applies it to the radar simulator environment; it continues to obtain state information from the radar simulator environment and make optimal action decisions; the data generated from the self-game is stored; wherein, MCTS is Monte Carlo Tree Search; obtaining state information from the radar simulator environment includes: receiving radar pulse signals from the radar simulator environment and extracting pulse-related parameters, combining all extracted parameters to obtain a pulse description word vector, and using the pulse description word vector as state information; the interfering agent determining the optimal action based on the state information obtained from the radar simulator environment and combined with MCTS includes: based on the state information obtained from the radar simulator environment... The Monte Carlo tree search is performed on the state information obtained from the environment, recursively selecting the most valuable child node from the root node until a leaf node is reached. The Monte Carlo tree search is based on a tree model where nodes are associated with state information and edges are associated with actions. The root node corresponds to the state information obtained from the radar simulator environment. Each edge stores information including: visit count, action probability, average action value, and total state value. The most valuable child node is evaluated using the information stored in its corresponding edge. The available actions corresponding to the current state are combined and expanded to the child nodes of the leaf node. The interference agent evaluates the action probability and state value of each child node. The information stored in the corresponding edges of the search path is updated in reverse. The optimal action is selected using the updated information stored in the edges. Step 3: Train the interfering agent using the stored data; Step 4: Alternately execute steps 2 and 3 until the set stopping condition is met to obtain the trained interfering agent. Deploy the trained interfering agent in the actual adversarial scenario to output the optimal action and the corresponding state value. Use the state value as the evaluation result of the adversarial effectiveness.
2. The adversarial effectiveness evaluation method based on MCTS and self-game as described in claim 1, characterized in that, The recursive selection of the most valuable child node includes: calculating the score of the action associated with each edge using the upper confidence interval algorithm, and selecting the child node pointed to by the edge with the highest score as the most valuable child node; The relevant calculation formula is expressed as follows: Among them, a k For the action corresponding to the edge with the highest score, Q(s,a), P(s,a), and N(s,a) are the information stored in the edge, which are the average value of the action, the probability of the action, and the number of visits, respectively. (s,a) represents the action a being performed in state s; ∑ b N(s,b) represents the total number of visits to all actions in state s, where b is the index symbol, referring to any action among all actions; C is the hyperparameter.
3. The adversarial effectiveness evaluation method based on MCTS and self-game as described in claim 1, characterized in that, The evaluation of the action probability and state value of each child node by the interfering agent includes: The action probability of each child node is evaluated through the policy network in the interfering agent, and the state value of each child node is evaluated through the value network in the interfering agent. Save the action probability of each child node to the edge between each child node and the leaf node.
4. The adversarial effectiveness evaluation method based on MCTS and self-game as described in claim 1, characterized in that, The information stored in the corresponding edges of the reverse update search path includes: The search path includes all nodes and their corresponding edges from the root node to the child nodes; The information stored in all edges of the search path is updated, including: the number of visits, the average value of actions, and the total state value; the update method is represented as follows: N(s,a)←N(s,a)+1 W(s,a)←W(s,a)+v L Where ← is the assignment operator, N(s,a) is the number of visits, W(s,a) is the total state value, and v L Let Q(s,a) be the state value of the child node, Q(s,a) be the average value of the action, and (s,a) represent the action a performed in state s. W(s,a) and N(s,a) on the right side of the equation are the updated information.
5. The adversarial effectiveness evaluation method based on MCTS and self-game as described in claim 1, characterized in that, The method of training the interference agent using stored data includes: The following loss function is used to optimize the policy network and value network in the interfering agent: loss=l v (z t ,v t )+l p (π t ,p t )+c||θ|| 2 Where θ represents the parameters of the interfering agent, c represents the hyperparameter, and l v (z t ,v t )=(z t -v t ) 2 , T is the transpose symbol, l v (z t ,v t Let ) be the loss function of the value network, l p (π t ,p t Let v be the loss function of the policy network. t p is the state value output by the value network. t Let z be the action probability distribution output by the policy network. t and π t All data are generated from self-games, and are presented in the following order: the game result and the optimal action decided.
6. A performance evaluation system for adversarial competition based on MCTS and self-game theory, characterized in that, To implement the method according to any one of claims 1 to 5, comprising: The environment and agent construction unit is used to construct the radar simulator environment and the jamming agent corresponding to the jammer. The sample data collection unit based on MCTS and self-game is used to conduct self-games between a disruptive agent and a radar simulator environment. This includes: the disruptive agent determining the optimal action based on state information obtained from the radar simulator environment, combined with MCTS, and applying it to the radar simulator environment; continuing to obtain state information from the radar simulator environment and making action decisions; storing the data generated by the self-game, including the determined optimal action and the actual results of the confrontation; where MCTS is Monte Carlo Tree Search. The agent training unit is used to train the interfering agent using stored data. The performance evaluation unit controls the alternating operation of the sample data collection unit and the agent training unit based on MCTS and self-game until the set stopping conditions are met, obtains the trained interfering agent, and deploys the trained interfering agent in the actual adversarial scenario to output the optimal action and the corresponding state value. The state value is used as the evaluation result of adversarial performance.
7. A processing device, characterized in that, include: One or more processors; Memory, used to store one or more programs; Wherein, when the one or more programs are executed by the one or more processors, the one or more processors cause the one or more processors to implement the method as described in any one of claims 1 to 5.
8. A readable storage medium storing a computer program, characterized in that, When a computer program is executed by a processor, it implements the method as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Incomplete information game method and system based on reinforcement learning and electronic equipment
CN112926744A
Control system based on multi-unmanned-aerial-vehicle collaborative game confrontation
WO2021174765A1