Radio anti-interference decision-making method based on Monte Carlo nerve virtual self-game
Through a two-layer decision model based on Monte Carlo neural virtual self-game, the real-time and robustness of anti-interference decisions in a complex electromagnetic adversarial environment are solved, and efficient anti-interference strategy selection and resource optimization are achieved.
Patent Information
- Application Number
- CN202510662764.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-22
- Publication Date
- 2025-07-29
AI Technical Summary
When existing radio equipment faces complex electromagnetic countermeasure environments, traditional anti-interference decision-making methods lack real-time and robustness, making it difficult to effectively identify interference strategies and choose appropriate anti-interference measures.
A two-layer decision model based on Monte Carlo neural virtual self-game is adopted, including the outer Monte Carlo neural virtual self-game model and the inner multi-agent deep deterministic strategy gradient model. Through layered optimization and reinforcement learning, combined with the multi-head self-attention mechanism, the selection of anti-interference measures is optimized.
It improves the anti-interference decision-making level of radio equipment in complex electromagnetic countermeasures, improves strategy learning ability and training efficiency, takes into account anti-interference effect and resource limitations, and is suitable for large-scale search spaces.
Smart Images

Figure CN120389819A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of radio anti - interference decision - making, and more particularly to a radio anti - interference decision - making method based on Monte Carlo neural virtual self - game. Background Art
[0002] Electromagnetic interference devices pose a serious threat to communication system devices, radars and other radio equipment. Therefore, the research on the confrontation game between radio equipment and interference devices has become an extremely important part of the development field of communication signal technology. In the traditional anti - interference decision - making process, the information interaction is poor, mainly relying on subjective experience or fixed strategy templates. Due to the complex and changeable electromagnetic confrontation environment, it is very difficult for such methods to fully identify interference strategies or timely select appropriate anti - interference measures. In existing research, game theory, the mathematical theory and method of confrontation behavior, has been introduced into the decision - making system under uncertain conditions to solve the problem of strategy selection. However, such methods highly rely on the ideas of "template matching" or "reasoning and trial - error", and are difficult to be applied to modern electronic confrontation scenarios with high requirements for adaptability and real - time performance. The confrontation between radio equipment and interference devices has a certain duration, rather than a single - round interaction; the behavior of the opponent can only be partially observed, and the anti - interference resources of radio equipment are limited. Developing an intelligent anti - interference decision - making method with high efficiency, strong real - time performance and good robustness has become inevitable. Summary of the Invention
[0003] The present invention aims to solve at least one of the technical problems existing in the prior art. To this end, a first aspect of the present invention proposes a radio anti - interference decision - making method based on Monte Carlo neural virtual self - game, including:
[0004] Hierarchically optimize the action space of the radio equipment, and the action space of the radio equipment is divided into a transform domain subspace and an anti - interference measure subspace;
[0005] Use a two - layer radio anti - interference decision - making model to conduct confrontation between the radio equipment and the jammer;
[0006] Among them, the two - layer radio anti - interference decision - making model includes an outer - layer anti - interference decision - making layer and an inner - layer anti - interference decision - making layer;
[0007] The outer - layer anti - interference decision - making layer includes a Monte Carlo neural virtual self - game model for selecting the corresponding transform domain in the transform domain subspace;
[0008] The inner - layer anti - interference decision - making layer includes a multi - agent deep deterministic policy gradient model for selecting corresponding anti - interference measures according to the corresponding transform domain;
[0009] Find the global optimal solution through the interaction between the Monte Carlo neural virtual self - game model and the multi - agent deep deterministic policy gradient model.
[0010] The above-mentioned radio anti-jamming decision-making method based on Monte Carlo neural virtual self-play is characterized in that: the transform domain subspace includes time domain, frequency domain, space domain and polarization domain; the anti-jamming measure subspace in the time domain includes pulse compression anti-jamming measures, Doppler processing anti-jamming measures and waveform agility anti-jamming measures; the anti-jamming measure subspace in the frequency domain includes frequency agility anti-jamming measures, frequency diversity anti-jamming measures and phase coding anti-jamming measures; the anti-jamming measure subspace in the space domain includes adaptive beamforming anti-jamming measures, sidelobe concealment anti-jamming measures and sidelobe blanking anti-jamming measures; the anti-jamming measure subspace in the polarization domain includes polarization diversity anti-jamming measures, polarization filtering anti-jamming measures and variable polarization anti-jamming measures.
[0011] The above-mentioned radio anti-jamming decision-making method based on Monte Carlo neural virtual self-play is characterized in that: the jamming techniques in the jammer include noise jamming, deception jamming and composite jamming, and the noise jamming includes blocking jamming, aiming jamming and frequency sweep jamming; the deception jamming includes range gate pull-off jamming, velocity gate pull-off jamming, range-velocity synchronous pull-off jamming and intermittent sampling and forwarding jamming; the composite jamming includes frequency sweep and range gate pull-off composite jamming, and aiming and intermittent sampling and forwarding composite jamming.
[0012] The above-mentioned radio anti-jamming decision-making method based on Monte Carlo neural virtual self-play is characterized in that: the Monte Carlo neural virtual self-play model includes a best response network for Monte Carlo tree search and an average policy network for supervised learning; the best response network and the average policy network are trained through a mixed strategy κ self-play, where κ = (1 - ε)Π + εB, that is, in each action, the agent selects the Monte Carlo tree search result based on the best response network Β with probability ε, or selects the supervised learning result of the average policy network Π with probability 1 - ε.
[0013] The above-mentioned radio anti-jamming decision-making method based on Monte Carlo neural virtual self-play is characterized in that: the agent establishes a best response network through reinforcement learning, and the agent establishes an average policy network through supervised learning;
[0014] After the end of each round of adversarial training, (s, π(s), v) is stored in the replay storage area of reinforcement learning to train the best response network Β; (s, π(s)) is stored in the replay storage area of supervised learning to train the average policy network Π; the loss functions L1 of the best response network and L2 of the average policy network are updated, and the Adam optimizer is used for optimization, L1(θ B ) = -Σ t (π(s t ) logp t - (v(s t ) - zt ) 2 ),L2(ψ Π )=-Σ t π(s t )logp t , where the parameters θ and ψ represent the characteristic parameters of the optimal response network and the average policy network respectively, s is the agent state, π(s) is the agent's policy, v is the predicted value of the given state output by the optimal response network, s t is the agent state at time t, p t is the output of the average policy network, z t is the confrontation result of each round and z t ∈[-1, 1].
[0015] The above radio anti-jamming decision-making method based on Monte Carlo neural virtual self-play is characterized in that: the multi-agent deep deterministic policy gradient model includes an Actor network and a Critic network established based on the Actor-Critic framework. The Actor network and the Critic network update the target network parameters through gradient descent, and a multi-head self-attention layer for alleviating the training difficulty of neural networks with a large number of parameters is introduced into the Actor network and the Critic network of the multi-agent deep deterministic policy gradient model.
[0016] The above radio anti-jamming decision-making method based on Monte Carlo neural virtual self-play is characterized in that: regarding a corresponding transform domain selected by the Monte Carlo neural virtual self-play model as a sub-agent, selecting the corresponding anti-jamming measure action, observing the reward function, updating the loss function L and the target function y, and optimizing using the Adam optimizer.
[0017] The above radio anti-jamming decision-making method based on Monte Carlo neural virtual self-play is characterized in that: the reward function where w is the number of anti-jamming measure actions selected by the sub-agent during a single interaction, W is the total number of anti-jamming measure actions, is the maximum value of the anti-jamming decision factor ζ and ρ0 is the evaluation index value when the radio device is not interfered, ρ J is the evaluation index value when the radio device is interfered, ρ AJ is the evaluation index value after the radio device applies a certain anti-jamming measure, ζ min is the minimum value of the anti-jamming decision factor ζ, ζ max is the maximum value of the anti-jamming decision factor ζ, ω1 is the first positive hyperparameter, ω2 is the second positive hyperparameter and ω1 + ω2 = 1;
[0018] It should be noted that this application defines an anti-interference decision factor to analyze the effectiveness of each anti-interference measure in reducing interference. Using this factor as feedback, it is used to design the reward function in the multi-agent training algorithm. This factor quantifies the enhanced anti-interference ability of the radio device when using a certain anti-interference measure compared to not using it. During the decision-making process, the increase in the number of selected anti-interference measures is related to the improvement of the interference suppression performance; however, this will result in the use of more anti-interference resources. To coordinate these two aspects, we implemented a weighting strategy when designing the reward function, aiming to maximize while reducing the number of selected anti-interference measures, and two positive hyperparameters are selected to control the relative importance of these two goals.
[0019] The loss function L = E z,a,r,z' [(Q(z,a)-y) 2 , where Q(z,a) is the centralized evaluation function of the sub-agent. The input of Q(z,a) is the anti-interference measure action a taken by the sub-agent and the environmental information z, and the output of Q(z,a) is the Q-value of the sub-agent. E is the expectation, and z' is the environmental update information;
[0020] The objective function y = r + γQ'(z',a'), where γ is the update coefficient, Q' is the centralized update evaluation function of the sub-agent, and a' is the anti-interference measure action replaced by the sub-agent.
[0021] The above-mentioned radio anti-interference decision method based on Monte Carlo neural virtual self-play is characterized in that: the interaction process between the Monte Carlo neural virtual self-play model and the multi-agent deep deterministic policy gradient model: the Monte Carlo neural virtual self-play model determines the transform domain and guides the actions of the multi-agent deep deterministic policy gradient model. The anti-interference measures determined by the multi-agent deep deterministic policy gradient model directly determine the anti-interference effect of the radio device, determine the selection of the next decision-making action, and the decision result will be fed back to the Monte Carlo neural virtual self-play model.
[0022] The embodiment of the present invention provides a radio anti-interference decision method based on Monte Carlo neural virtual self-play. Compared with the prior art, its beneficial effects are as follows:
[0023] Combining Monte Carlo tree search with neural virtual self-play effectively improves the performance of large-scale zero-sum imperfect information games; adopting a hierarchical deep reinforcement learning structure, reducing the action space dimension through a two-layer hierarchical selection and joint optimization process, and enhancing the training efficiency; specifically designing a reward function to enable the intelligent agent to consider both the anti-interference effect and resource constraints when selecting strategies and actions; introducing a multi-head self-attention mechanism to mitigate the adverse effects of excessive parameter dimensions on the neural network training process and prevent irrelevant intelligent agents from interfering with the intelligent agent making decisions, having good policy learning ability in scenarios with large-scale search spaces, and effectively improving the anti-interference decision-making level of radio equipment in complex electromagnetic confrontation environments.
[0024] The following will, through the accompanying drawings and embodiments, give a more detailed description of the technical solutions of the present invention. Description of the Drawings
[0025] To more clearly illustrate the technical solutions of the present invention, the accompanying drawings required for use in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other accompanying drawings based on these drawings without creative efforts.
[0026] Figure 1 It is a flowchart of the radio anti-interference decision-making method based on Monte Carlo neural virtual self-play provided by an embodiment of the present invention. Specific Embodiments
[0027] The following will, in combination with the accompanying drawings in the embodiments of the present invention, clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.
[0028] This specification provides the method operation steps as described in the embodiments or the flowchart, but based on routine or non-creative labor, it may include more or fewer operation steps. When actually executed by a system or server product, it can be executed in the order shown in the embodiments or the accompanying drawings or in parallel (for example, in an environment of parallel processors or multi-threaded processing).
[0029] Figure 1 It is a flowchart of the radio anti-interference decision-making method based on Monte Carlo neural virtual self-play provided by an embodiment of the present invention, as Figure 1 shown, the method includes:
[0030] Hierarchically optimize the action space of the radio device, where the action space of the radio device is divided into a transform domain subspace and an anti-jamming measure subspace;
[0031] Use a two-layer radio anti-jamming decision-making model to conduct confrontation between the radio device and the jammer;
[0032] Among them, the two-layer radio anti-jamming decision-making model includes an outer anti-jamming decision-making layer and an inner anti-jamming decision-making layer;
[0033] The outer anti-jamming decision-making layer includes a Monte Carlo neural virtual self-play model for selecting the corresponding transform domain in the transform domain subspace;
[0034] The inner anti-jamming decision-making layer includes a multi-agent deep deterministic policy gradient model for selecting corresponding anti-jamming measures according to the corresponding transform domain;
[0035] Find the global optimal solution through the interaction between the Monte Carlo neural virtual self-play model and the multi-agent deep deterministic policy gradient model.
[0036] In this embodiment, the transform domain subspace includes time domain, frequency domain, space domain, and polarization domain; the anti-jamming measure subspace in the time domain includes pulse compression anti-jamming measures, Doppler processing anti-jamming measures, and waveform agility anti-jamming measures; the anti-jamming measure subspace in the frequency domain includes frequency agility anti-jamming measures, frequency diversity anti-jamming measures, and phase coding anti-jamming measures; the anti-jamming measure subspace in the space domain includes adaptive beamforming anti-jamming measures, sidelobe concealment anti-jamming measures, and sidelobe blanking anti-jamming measures; the anti-jamming measure subspace in the polarization domain includes polarization diversity anti-jamming measures, polarization filtering anti-jamming measures, and variable polarization anti-jamming measures.
[0037] In this embodiment, the jamming techniques in the jammer include noise jamming, deception jamming, and composite jamming. The noise jamming includes blocking jamming, aiming jamming, and sweep jamming; the deception jamming includes range gate pull-off jamming, velocity gate pull-off jamming, range-velocity synchronous pull-off jamming, and intermittent sampling and forwarding jamming; the composite jamming includes sweep and range gate pull-off composite jamming, and aiming and intermittent sampling and forwarding composite jamming.
[0038] In this embodiment, the Monte Carlo neural virtual self-play model includes a best response network for Monte Carlo tree search and an average policy network for supervised learning; the best response network and the average policy network are trained through a mixed strategy κ self-play, where κ=(1 - ε)Π + εB, that is, in each action, the agent selects the Monte Carlo tree search result based on the best response network Β with probability ε, or selects the supervised learning result of the average policy network Π with probability 1 - ε.
[0039] It should be noted that in this dual-layer radio anti-jamming decision-making model, the outer layer uses the Monte Carlo neural virtual self-play model to explore and learn the first layer of the action space of radio equipment. This model will determine which transformation domain the radio equipment should use to respond in the confrontation and transmit the decision result to the inner-layer anti-jamming decision-making layer to guide the inner-layer anti-jamming decision-making layer to further explore and make decisions on which specific anti-jamming measures to take in this domain.
[0040] In this embodiment, the agent establishes an optimal response network through reinforcement learning and an average policy network through supervised learning.
[0041] After the end of each round of confrontation training, (s, π(s), v) is stored in the replay storage area of reinforcement learning to train the optimal response network Β; (s, π(s)) is stored in the replay storage area of supervised learning to train the average policy network Π; the loss functions L1 of the optimal response network and L2 of the average policy network are updated and optimized using the Adam optimizer. L1(θ B ) = -∑ t (π(s t ) log p t - (v(s t ) - z t ) 2 ), L2(ψ Π ) = -∑ t π(s t ) log p t , where the parameters θ and ψ represent the feature parameters of the optimal response network and the average policy network respectively, s is the agent state, π(s) is the agent's policy, v is the predicted value of the given state output by the optimal response network, s t is the agent state at time t, p t is the output of the average policy network, z t is the result of each round of confrontation and z t ∈ [-1, 1].
[0042] In this embodiment, the multi-agent deep deterministic policy gradient model includes an Actor network and a Critic network established based on the Actor-Critic framework. The Actor network and the Critic network update the target network parameters through gradient descent, and a multi-head self-attention layer for alleviating the training difficulty of neural networks with a large number of parameters is introduced into the Actor network and the Critic network of the multi-agent deep deterministic policy gradient model.
[0043] It should be noted that to improve the running efficiency of the model, we introduced a multi-head self-attention layer in the inner-layer multi-agent deep deterministic policy gradient model, effectively alleviating the difficulty of training neural networks with a large number of parameters. Multiple attention heads are used, and they share parameters with convolutional layers and fully connected layers. These attention heads divide the observation space according to the correlation intensity, assigning different weights to each vector element in the space. This dynamic weight distribution allows the model to focus on the information most relevant to the decision-making process of the current agent while ignoring the less relevant information from other agents.
[0044] In this embodiment, a corresponding transform domain selected by the Monte Carlo neural virtual self-play model is regarded as a sub-agent. The corresponding anti-interference measure actions are selected, the reward function is observed, the loss function L and the objective function y are updated, and the Adam optimizer is used for optimization.
[0045] In this embodiment, the reward function where w is the number of anti-interference measure actions selected by the sub-agent during a single interaction, and W is the total number of anti-interference measure actions. is the maximum value of the anti-interference decision factor ζ and ρ0 is the evaluation index value when the radio device is not interfered, and ρ J is the evaluation index value when the radio device is interfered, and ρ AJ is the evaluation index value after the radio device applies a certain anti-interference measure, ζ min is the minimum value of the anti-interference decision factor ζ, ζ max is the maximum value of the anti-interference decision factor ζ, ω1 is the first positive hyperparameter, ω2 is the second positive hyperparameter and ω1 + ω2 = 1;
[0046] It should be noted that the reward function r indicates that the sub-agent can confront and select w measures (which can be 1, 2,..., W, not exceeding W) from W measures (the total action library) each time. Each specific measure will correspond to a reward function value. The summation is the sum of the reward function values corresponding to all measures. However, in the confrontation, only when a certain measure is selected, its corresponding reward function has a value; otherwise, it is 0, and even if it is selected, it may be negative. Therefore, the summation will also have a maximum value, and not all can be selected. The purpose of the reward function is to guide the agent to obtain more benefits as much as possible under restricted circumstances (selecting as few anti-interference measures as possible).
[0047] The loss function L = E z,a,r,z' [(Q(z,a) - y) 2, Q(z,a) is the evaluation function centered on the sub-agent. The input of Q(z,a) is the anti-interference measure action a taken by the sub-agent and the environmental information z, and the output of Q(z,a) is the Q value of the sub-agent. E is the expectation, and z' is the environmental update information;
[0048] The objective function is y = r + γQ'(z',a'), where γ is the update coefficient, Q' is the updated evaluation function centered on the sub-agent, and a' is the anti-interference measure action replaced by the sub-agent.
[0049] In this embodiment, the interaction process between the Monte Carlo neural virtual self-play model and the multi-agent deep deterministic policy gradient model: The Monte Carlo neural virtual self-play model determines the transform domain and guides the actions of the multi-agent deep deterministic policy gradient model. The anti-interference measures determined by the multi-agent deep deterministic policy gradient model directly determine the anti-interference effect of the radio device, determine the selection of the next decision-making action, and the decision result will be fed back to the Monte Carlo neural virtual self-play model.
[0050] In this embodiment, by combining Monte Carlo tree search with neural virtual self-play, the performance of large-scale zero-sum imperfect information games is effectively improved; a hierarchical deep reinforcement learning structure is adopted to reduce the action space dimension through a two-layer hierarchical selection and joint optimization process, improving the training efficiency; a reward function is designed specifically to enable the agent to consider both the anti-interference effect and resource constraints when selecting strategies and actions; a multi-head self-attention mechanism is introduced to mitigate the adverse effects of excessive parameter dimensions on the neural network training process and prevent irrelevant agents from interfering with the agent making decisions. It has good policy learning ability in scenarios with large-scale search spaces and can effectively improve the anti-interference decision-making level of radio devices in complex electromagnetic confrontation environments.
[0051] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.
[0052] Each embodiment in this specification is described in a related manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple. For the related parts, reference can be made to the partial description of the method embodiment.
[0053] The above are only the preferred embodiments of the present invention and are not intended to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention are all included in the protection scope of the present invention.
Claims
1. A radio anti-jamming decision-making method based on Monte Carlo neural virtual self-play, characterized in that, Including: Hierarchically optimize the action space of the radio device, where the action space of the radio device is divided into a transform domain subspace and an anti-jamming measure subspace; Use a two-layer radio anti-jamming decision-making model to conduct confrontation between the radio device and the jammer; Among them, the two-layer radio anti-jamming decision-making model includes an outer anti-jamming decision-making layer and an inner anti-jamming decision-making layer; The outer anti-jamming decision-making layer includes a Monte Carlo neural virtual self-play model for selecting the corresponding transform domain in the transform domain subspace; The inner anti-jamming decision-making layer includes a multi-agent deep deterministic policy gradient model for selecting corresponding anti-jamming measures according to the corresponding transform domain; Find the global optimal solution through the interaction between the Monte Carlo neural virtual self-play model and the multi-agent deep deterministic policy gradient model.
2. The radio anti-jamming decision-making method based on Monte Carlo neural virtual self-play according to claim 1, characterized in that: The transform domain subspace includes time domain, frequency domain, space domain, and polarization domain; the anti-jamming measure subspace in the time domain includes pulse compression anti-jamming measures, Doppler processing anti-jamming measures, and waveform agility anti-jamming measures; the anti-jamming measure subspace in the frequency domain includes frequency agility anti-jamming measures, frequency diversity anti-jamming measures, and phase coding anti-jamming measures; the anti-jamming measure subspace in the space domain includes adaptive beamforming anti-jamming measures, sidelobe concealment anti-jamming measures, and sidelobe blanking anti-jamming measures; the anti-jamming measure subspace in the polarization domain includes polarization diversity anti-jamming measures, polarization filtering anti-jamming measures, and variable polarization anti-jamming measures.
3. The radio anti-jamming decision-making method based on Monte Carlo neural virtual self-play according to claim 1, wherein: The jamming techniques in the jammer include noise jamming, deception jamming, and composite jamming. The noise jamming includes blocking jamming, aiming jamming, and frequency sweep jamming; the deception jamming includes range gate pull-off jamming, velocity gate pull-off jamming, range-velocity synchronous pull-off jamming, and intermittent sampling and forwarding jamming; the composite jamming includes frequency sweep and range gate pull-off composite jamming, and aiming and intermittent sampling and forwarding composite jamming.
4. The radio anti-jamming decision-making method based on Monte Carlo neural virtual self-play according to claim 1, characterized in that: The Monte Carlo neural virtual self-play model includes a best response network for Monte Carlo tree search and an average policy network for supervised learning; the best response network and the average policy network are trained through a mixed strategy κ self-play, where κ = (1 - ε)Π + εB, that is, in each action, the agent selects the Monte Carlo tree search result based on the best response network Β with probability ε, or selects the supervised learning result of the average policy network Π with probability 1 - ε.
5. The radio anti-jamming decision-making method based on Monte Carlo neural virtual self-play according to claim 4, wherein: The agent establishes the best response network through reinforcement learning and establishes the average policy network through supervised learning; After each round of adversarial training, store (s, π(s), v) in the replay buffer of reinforcement learning to train the best response network Β; store (s, π(s)) in the replay buffer of supervised learning to train the average policy network Π; update the loss functions L1 of the best response network and L2 of the average policy network, and use the Adam optimizer for optimization. L1(θ B ) = -Σ t (π(s t ) log p t - (v(s t ) - z t ) 2 ), L2(ψ Π ) = -Σ t π(s t ) log p t , where the parameters θ and ψ represent the feature parameters of the best response network and the average policy network respectively, s is the agent state, π(s) is the agent's policy, v is the predicted value of the given state output by the best response network, s t is the agent state at time t, p t is the output of the average policy network, z t is the result of each round of confrontation and z t ∈ [-1, 1].
6. The radio anti-jamming decision-making method based on Monte Carlo neural virtual self-play according to claim 5, wherein: The multi-agent deep deterministic policy gradient model includes an Actor network and a Critic network established based on the Actor-Critic framework. The Actor network and the Critic network update the target network parameters through gradient descent. A multi-head self-attention layer for alleviating the training difficulty of neural networks with a large number of parameters is introduced into the Actor network and the Critic network of the multi-agent deep deterministic policy gradient model.
7. The radio anti-jamming decision-making method based on Monte Carlo neural virtual self-play according to claim 6, wherein: Regard a corresponding transform domain selected by the Monte Carlo neural virtual self-play model as a sub-agent, select the corresponding anti-interference measure action, observe the reward function, update the loss function L and the objective function y, and use the Adam optimizer for optimization.
8. The radio anti-jamming decision-making method based on Monte Carlo neural virtual self-play according to claim 7, wherein: The reward function where w is the number of anti-interference measure actions selected by the sub-agent during a single interaction, W is the total number of anti-interference measure actions is the maximum value of the anti-interference decision factor ζ and ρ0 is the evaluation index value when the radio device is not interfered, ρ J is the evaluation index value when the radio device is interfered, ρ AJ is the evaluation index value after the radio device applies a certain anti-interference measure, ζ min is the minimum value of the anti-interference decision factor ζ, ζ max is the maximum value of the anti-interference decision factor ζ, ω1 is the first positive hyperparameter, ω2 is the second positive hyperparameter and ω1 + ω2 = 1; The loss function L = E z,a,r,z' [(Q(z,a) - y) 2 , where Q(z,a) is the centralized evaluation function of the sub-agent. The input of Q(z,a) is the anti-interference measure action a taken by the sub-agent and the environmental information z, the output of Q(z,a) is the Q-value of the sub-agent, E is the expectation, and z' is the environmental update information; The objective function y = r + γQ'(z', a'), where γ is the update coefficient, Q' is the updated evaluation function centralized by the sub-agent, and a' is the anti-interference measure action replaced by the sub-agent.
9. The radio anti-jamming decision-making method based on Monte Carlo neural virtual self-play according to claim 1, characterized in that: The interaction process between the Monte Carlo neural virtual self-play model and the multi-agent deep deterministic policy gradient model: The Monte Carlo neural virtual self-play model determines the transform domain and guides the actions of the multi-agent deep deterministic policy gradient model. The anti-interference measures determined by the multi-agent deep deterministic policy gradient model directly determine the anti-interference effect of the radio device, determine the selection of the next decision-making action, and the decision result will be fed back to the Monte Carlo neural virtual self-play model.
Citation Information
Patent Citations
Incomplete information game method and system based on reinforcement learning and electronic equipment
CN112926744A
Adversarial task-oriented man-machine symbiosis reinforcement learning method and device, computing equipment and storage medium
CN113688977A
Multi-agent collaborative decision-making method based on deep reinforcement learning under limited communication resources
CN116456480A
Radar intelligent anti-interference decision-making method based on multi-criterion multi-cost function
CN117148286A
Vehicle action decision-making method and device based on Monte Carlo tree, electronic equipment and medium
CN118211145A