A frequency hopping interference resource allocation method based on learning from imperfect examples
By modeling the frequency hopping interference resource allocation problem as a Markov decision process and utilizing the TRPO algorithm and discriminator network optimization strategy, the dependence on fine reward functions in deep reinforcement learning is resolved, and efficient interference resource allocation is achieved in a sparse reward environment.
Patent Information
- Application Number
- CN202410715023.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-04
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2044-06-04
AI Technical Summary
Existing deep reinforcement learning technology requires manual design of sophisticated reward functions in interference resource allocation, which consumes a lot of resources and time.
The frequency hopping interference resource allocation problem is modeled as a Markov decision process. By randomly initializing the policy network and performing multiple iterations, the TRPO algorithm is used to optimize the policy improvement phase. In the policy adversarial imitation phase, the discriminator network is trained with example data to optimize the intermediate allocation strategy and reduce the dependence on the refined reward function.
It realizes forward optimization of strategies in a sparse reward environment, saves resource consumption, and improves the efficiency and effect of interference resource allocation.
Smart Images

Figure CN118740198B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of wireless communication technology, and more particularly to a frequency hopping interference resource allocation method based on learning from imperfect examples. Background Art
[0002] In the field of wireless communications, wireless sensor networks are widely used in military communications. With the development of spread-spectrum communication technology, frequency-hopping spread spectrum (FHSS) communication has become an important means of improving the anti-interference capabilities of wireless sensor networks. Interferers often use partial-band noise jamming (PBNJ) to reduce the efficiency of FHSS systems. For example, by rationally setting multiple non-overlapping interference bands, the utilization of interference resources can be improved. This can expand the interference frequency range and increase the possibility of covering different user channels. Interference resource allocation for FHSS communication is essentially a combinatorial optimization problem.
[0003] In related technologies, with the rapid development of artificial intelligence, deep learning-based methods have emerged to solve combinatorial optimization problems. For example, deep reinforcement learning technology can be applied to the intelligent optimization of interference power allocation or interference link selection.
[0004] Regarding the above technical solution, the inventors found that there are at least the following technical problems: Although deep reinforcement learning technology has been widely used to solve combinatorial optimization problems, it relies heavily on artificially designed sophisticated reward functions, that is, reward features need to be designed specifically for different goals of different tasks. This requires professional domain knowledge as support, and a large number of repeated experiments are required to find a suitable reward function, which is extremely resource-consuming. Summary of the Invention
[0005] The purpose of the embodiments of the present disclosure is to provide a frequency hopping interference resource allocation method based on learning from imperfect examples, which can guide the strategy to be forward optimized in a sparse reward environment without the need for manually designing a sophisticated reward function, thereby saving resource consumption.
[0006] According to an embodiment of the present disclosure, a method for allocating frequency hopping interference resources based on learning from imperfect examples is provided, including:
[0007] Construct the frequency hopping jamming resource allocation problem based on the communication confrontation scenario;
[0008] Modeling the frequency hopping interference resource allocation problem as a Markov decision process;
[0009] Randomly initialize the policy network parameters and the discriminator network;
[0010] The initialized policy network is iterated for multiple times, and in each iteration process, a policy improvement stage and a policy confrontation imitation stage are constructed based on a double confidence domain;
[0011] In the policy improvement stage, an initial allocation policy in the current iteration is optimized based on a TRPO algorithm to obtain an intermediate allocation policy;
[0012] In the policy confrontation imitation stage, the discriminator network is trained by using example data and interaction data of the initial allocation policy in the current iteration to optimize the intermediate allocation policy to obtain a final allocation policy.
[0013] The technical solution provided by the present disclosure can include the following beneficial effects:
[0014] In the embodiments of the present disclosure, by using the frequency hopping interference resource allocation method learned from imperfect examples, on the one hand, the interference resource allocation problem involved in the partial band noise interference of the frequency hopping communication is modeled as a Markov decision process, so that the strategy neural network constructs a dynamic transfer process in the combinatorial optimization in the form of the interference scheme decided by each interference node in time sequence; on the other hand, the initialized policy network is iterated for multiple times, and in each policy iteration process, a policy improvement stage and a policy confrontation imitation stage are constructed based on a double confidence domain, wherein in the policy improvement stage, the initial allocation policy in the current iteration is optimized by using the TRPO algorithm to ensure the monotone robust optimization of the algorithm, to obtain the intermediate allocation policy optimized in the policy improvement stage, and further, in the policy confrontation imitation stage, the discriminator network is trained by using the example data and the interaction data of the initial allocation policy in the current iteration to optimize the intermediate allocation policy, so that the policy imitates the trajectory of the example data in the confidence domain, and guides the policy to be optimized in the positive direction in the sparse reward environment, without the need of designing a fine reward function, thereby saving resource consumption.
[0015] It should be understood that the foregoing general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF DRAWINGS
[0016] The accompanying drawings incorporated in and forming a part of the specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the disclosure. It is apparent that the accompanying drawings in the following description are only some embodiments of the present disclosure, and other drawings can be obtained by those of ordinary skill in the art without creative labor based on these drawings.
[0017] Figure 1 A step schematic diagram of the frequency hopping interference resource allocation method learned from imperfect examples in the exemplary embodiments of the present disclosure is shown.
[0018] Figure 2 A schematic diagram of a ground-to-air interference scenario in an exemplary embodiment of the present disclosure is shown;
[0019] Figure 3 A schematic diagram showing the distribution of channels of different frequency hopping frequency sets on a frequency spectrum in an exemplary embodiment of the present disclosure;
[0020] Figure 4 The k+1th policy iteration process in the exemplary embodiment of the present disclosure is shown;
[0021] Figure 5 A schematic diagram illustrating a discriminator network in an exemplary embodiment of the present disclosure is shown;
[0022] Figure 6 FIG. 1 shows a diagram of an interference resource allocation framework in an exemplary embodiment of the present disclosure;
[0023] Figure 7 A schematic diagram showing a comparison of cumulative interference rewards obtained by different methods in a simulation experiment of an exemplary embodiment of the present disclosure;
[0024] Figure 8 A schematic diagram showing a comparison of total interference bandwidth consumed by different methods in a simulation experiment of an exemplary embodiment of the present disclosure;
[0025] Figure 9 The following figure shows the interference success rate of the method proposed in this application on different frequency sets in the simulation experiment of the exemplary embodiment of the present disclosure;
[0026] Figure 10 The following figure shows the interference success rate of the GAIL method for different frequency sets in the simulation experiment of the exemplary embodiment of the present disclosure;
[0027] Figure 11 A schematic diagram showing the comparison of interference frequency bands constructed after convergence of different methods in a simulation experiment of an exemplary embodiment of the present disclosure is shown. DETAILED DESCRIPTION
[0028] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be embodied in many forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be thorough and complete and will fully convey the concepts of the example embodiments to those skilled in the art. The described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.
[0029] In addition, the accompanying drawings are merely schematic illustrations of the present disclosure and are not necessarily drawn to scale. Identical reference numerals in the figures denote identical or similar parts, and thus repetitive descriptions thereof will be omitted. Some of the block diagrams shown in the accompanying drawings are functional entities that do not necessarily correspond to physically or logically separate entities. These functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.
[0030] This example embodiment first provides a frequency hopping interference resource allocation method based on learning from imperfect examples. Figure 1 As shown in , the frequency hopping interference resource allocation method based on meta-deep reinforcement learning may include the following steps:
[0031] Step S101: construct a frequency hopping interference resource allocation problem according to a communication countermeasure scenario;
[0032] Step S102: Modeling the frequency hopping interference resource allocation problem as a Markov decision process;
[0033] Step S103: Randomly initialize the policy network parameters and the discriminator network;
[0034] Step S104: performing multiple iterations on the initialized policy network, and in each iteration, constructing a policy improvement phase and a policy adversarial imitation phase based on the double reset trust region;
[0035] In the strategy improvement phase, the initial allocation strategy in the current iteration is optimized based on the TRPO algorithm to obtain the intermediate allocation strategy;
[0036] In the strategy adversarial imitation stage, the discriminator network is trained using the interaction data of the example data and the initial allocation strategy in the current iteration to optimize the intermediate allocation strategy and obtain the final allocation strategy of the current iteration.
[0037] By the above frequency hopping interference resource allocation method based on learning from imperfect examples, on the one hand, the interference resource allocation problem involved in the partial band noise interference against frequency hopping communication is modeled as a Markov decision process, so as to let the strategy neural network construct the dynamic transition process in combinatorial optimization in the form of time sequence decision of the interference scheme for each interference node; on the other hand, the initialized strategy network is iterated for multiple times, and in each strategy iteration process, a strategy improvement stage and a strategy counteracting imitation stage are constructed based on double confidence regions, wherein the TRPO algorithm is used in the strategy improvement stage to optimize the initial allocation strategy in the current iteration, so as to ensure the monotone robust optimization of the algorithm, and an intermediate allocation strategy optimized in the strategy improvement stage is obtained, and further, the discriminator network is trained by using the example data and the interaction data of the initial allocation strategy in the current iteration to optimize the intermediate allocation strategy, so as to let the strategy imitate the trajectory of the example data in the confidence region, guide the strategy to optimize in the positive direction in the sparse reward environment, and save resource consumption without designing a fine reward function artificially.
[0038] In the following, the above method in the present example embodiment will be described in more detail with reference to the accompanying drawings. Figures 1 to 11 The above method in the present example embodiment will be described in more detail with reference to the accompanying drawings.
[0039] It should be noted that in the process of iterating the initialized strategy network for multiple times, the final allocation strategy optimized in the last iteration is taken as the initial allocation strategy in the next iteration, so as to realize the iteration of the initialized strategy network for multiple times.
[0040] Further, as shown in FIG. 1, the ground-to-air interference scenario in the present application is that multiple ground interference nodes interfere with multiple ground-to-air frequency hopping communication networks, the ground transmitting nodes and the air receiving nodes communicate by using the FHSS technology, and N frequency hopping communication networks are constructed by using N frequency hopping frequency sets. Figure 2 The attack on the frequency hopping communication network is to attack the frequency hopping frequency set from the energy domain and the frequency domain.
[0041] In one embodiment, the frequency hopping frequency set is represented as wherein n is an index wherein the nth frequency hopping frequency set has p n frequency hopping channels, which can be represented as wherein is the center frequency of the ith frequency hopping channel of the nth frequency hopping frequency set, and the bandwidth of each frequency hopping channel is b u , and the total communication power is uniformly distributed on all frequency hopping channels, which can be represented as:
[0042]
[0043] Wherein, U(f) represents the power spectral density function of the communication signal.
[0044] Optionally, the jammer deploys M ground jammer nodes at the same location, using Represents the index of the interfering node. Each interfering node uses the PBNJ interference pattern to interfere with the N frequency hopping communication networks. Figure 3 The figure shows the distribution diagram of channels of different frequency hopping frequency sets on the spectrum. Through electronic warfare intelligence support and intelligence analysis, the jammer can grasp the channel distribution of the frequency hopping frequency sets used by each frequency hopping communication network in advance.
[0045] Optionally, each interfering node has a rated interference power By selecting the center interference frequency and interference bandwidth Construct a continuous interference frequency band on the spectrum And the interference power of the interfering node is evenly distributed in the interference frequency band, which can be expressed as:
[0046]
[0047] Wherein, J(f) represents the power spectral density function of the interference signal.
[0048] Furthermore, the interference frequency band of the mth interference node is calculated to be The Jamming-plus-Niose-to-Signal Ratio (JNSR) generated when the frequency hopping channel is:
[0049]
[0050] Where N(f) represents the power spectral density of Gaussian white noise, Represents the indicator function, when the center frequency is The interference frequency band can completely cover the center frequency of When the frequency hopping channel is G, the function value is 1, otherwise it is 0; u Represents the channel gain of the communication signal, G j represents the channel gain of the interference signal, F b represents the filtering loss coefficient, ρ represents the polarization loss coefficient; d u Indicates the propagation distance of the communication signal, d j represents the propagation distance of the interference signal, It represents the path transmission loss of the communication signal in free space. Represents the path transmission loss of the interference signal in free space.
[0051] Furthermore, we can get the center frequency of M interference nodes: The combined interference plus noise signal ratio of the frequency hopping channel is if Then the frequency hopping channel is considered to be successfully interfered, otherwise It means that it has not been successfully interfered with.
[0052] In this embodiment, the interference coverage coefficient is introduced Its meaning is that for any frequency hopping frequency set, when there is at least When a certain percentage of the frequency hopping channels are interfered, the frequency hopping frequency set can be considered to be completely interfered. The value of is between 0.3 and 0.4. Therefore, the formula for successfully judging the interference of the nth frequency hopping frequency set can be obtained as follows:
[0053]
[0054] Among them, sgn(·) represents the indicator function. When the independent variable x>0, the function value sgn(x)=1, and when the independent variable x≤0, the function value Indicates that the center frequency of M interference nodes is The joint interference plus noise signal ratio of the frequency hopping channel, p n Indicates the total number of channels in the nth frequency hopping frequency set, is the interference bandwidth, b u is the bandwidth of the hopping channel, M is the number of interfering nodes, Represents the interference coverage factor.
[0055] It should be noted that in order to fully utilize existing interference resources to successfully interfere with as many frequency hopping frequency sets as possible, the interferer needs to carefully select the central interference frequency and interference bandwidth for each interference node, specifically considering the following factors: first, it is necessary to ensure full coverage in the frequency domain and complete suppression in the energy domain for as many frequency hopping channels as possible; at the same time, it is necessary to pay attention to not interfering with all the frequency hopping channels in the frequency hopping set, because as long as the proportion of the number of interfered frequency hopping channels exceeds, it is sufficient to interfere with the corresponding frequency hopping frequency set; finally, it is necessary to coordinate the interference frequency bands constructed by each interference node to avoid excessive overlap of the interference frequency bands of different interference nodes, which would cause waste of interference resources.
[0056] In this embodiment, the interference resource allocation problem in the communication countermeasure scenario can be formulated as follows:
[0057]
[0058] Where sgn(·) represents the indicator function, Indicates that the center frequency of M interference nodes is The joint interference plus noise signal ratio of the frequency hopping channel, p n represents the total number of channels in the nth frequency hopping frequency set, is the interference bandwidth, b u is the bandwidth of the hopping channel, M is the number of interference nodes, represents the interference coverage coefficient.
[0059] In the above formula, which means maximizing the number of frequency hopping sets successfully interfered; (C1) means that the interference power is uniformly distributed on the interference frequency band; (C2) means that the communication power is uniformly distributed on all frequency hopping channels; (C3) restricts the number of frequency hopping frequency sets and interference nodes; (C4) restricts the selection range of the center interference frequency; (C5) restricts the selection range of the interference bandwidth.
[0060] It should be noted that the problem represented in the above formula is a static combinatorial optimization problem, so the size of its solution space is It can be seen that the size of the solution space increases exponentially with the increase of the number of interference nodes M. If a strategy neural network is used to give all interference node schemes at the same time, the number of neurons in the output layer of the strategy network will also reach This deformed neural network structure will bring great trouble to the training of the strategy network. Considering that the interference scheme is composed of the sub-schemes of each interference node, the present application adopts an incremental way to construct the solution to give the interference scheme, that is, the strategy network outputs the sub-schemes of each interference node in turn. This technical solution has two advantages: on the one hand, the number of neurons in the output layer of the network is not affected by the number of interference nodes; on the other hand, a dynamic state transition process can be constructed for the static optimization problem, and a prerequisite condition is established for the modeling of Markov decision process.
[0061] In this embodiment, in order to reasonably and optimally allocate the interference resources of multiple interference devices and realize simultaneous interference on multiple frequency hopping frequency sets, the frequency hopping interference resource allocation problem is modeled as a Markov decision process (MDP).
[0062] Further, step S102 can include the following sub-steps:
[0063] Step S1021: defining the state space, action space, state transition function and reward function of the Markov decision process;
[0064] Step S1022: constructing a discrete Markov decision process according to the state space, action space, state transition function and reward function.
[0065] In one embodiment, the Markov decision process defined in step S1021 can be Indicates that, is the state space, is the action space, is the network, r is the reward function, and the specific physical meanings are as follows:
[0066] In the interference resource allocation problem of this embodiment, the state space S can be expressed as:
[0067] s t =[e1(t),e2(t),…,e n (t),…,e N (t)]∈S (10)
[0068] Among them, e n (t) represents the transient interference effect of the currently constructed interference scheme on the nth frequency hopping frequency set, and
[0069] The interfering party's action a t is the interference action of the interference node that makes a decision at time step t, then the action space A can be expressed as:
[0070]
[0071] Where t∈{0,1,2,…,M-1}, is the central interference frequency of the t-th interference node, is the interference bandwidth of the t-th interfering node;
[0072] The state transition function P can be expressed as:
[0073] P(s t+1 |s t ,a t )→[0,1] (12)
[0074] Among them, P(s t+1 |s t ,a t ) indicates that the interfering end performs action a t After that, according to the state transfer function, the state s of the environmental information t Transition to the next state s t+1 ;
[0075] The reward function is used to guide the agent towards the optimization of the objective function. In this embodiment, a binary reward function is used directly, with the achievement of the task goal as the criterion for obtaining the reward. That is, if the final goal is achieved, the reward is 1, otherwise the reward is 0. The expression of the reward function r is:
[0076]
[0077] Among them, e n (t) represents the transient interference effect of the currently constructed interference scheme on the nth frequency hopping frequency set, represents the interference coverage coefficient, p n represents the total number of channels in the nth frequency hopping frequency set, and N represents the number of frequency hopping frequency sets.
[0078] It should be noted that in order to compare with the artificially designed fine reward function used in existing deep reinforcement learning technology to solve combinatorial optimization problems, an existing fine reward function is given as an example, such as the following formula:
[0079]
[0080] Among them, the reward function provided in this application represents a sparse reward environment, and the refined reward function given in the above example represents a dense reward environment. It can be seen that this refined reward function needs to be supported by professional domain knowledge, and the characteristic parameters in the reward function need to be determined through a large number of verification experiments, which undoubtedly requires a large amount of manpower and computing resources.
[0081] For example, reinforcement learning (RL) is a branch of machine learning that has a general computational framework. Reinforcement learning can control an agent to use a strategy π to interact with the environment, adjust its own strategy through the feedback r given by the environment, and then continue to interact with the environment, repeating the cycle, and finally obtaining an optimized strategy π * Reinforcement learning usually models the policy optimization process as a Markov decision process, which can be defined by a tuple MDP=<S,A,P,r,γ,H> , S is the state space, A is the action space, r is the reward function, P is the state transition function, γ is the discount factor, and H is the decision sequence. Suppose at time step t, the agent observes the state s of the environment t , and from the strategy π(·|s t ) to sample action a t Act on the environment, the state of the environment is affected by the action and becomes s t+1 ~P(·∣s t ,a t ) and generates a reward r t , the agent and the environment repeat the above process until the environment reaches the final state and stops, and generates a trajectory τ=(s0,a0,r0,s1,a1,r1,…). The objective function of reinforcement learning is to find a strategy π *Maximizes the cumulative discounted reward of the trajectory That is π * =argmax π J r (π).
[0082] For example, the state value function is used to measure the state s t The value of can be defined as Indicates that from state s t The expected cumulative reward that can be obtained by starting to follow the policy. The action value function measures the t Next take action a t The value of is defined as In state s t Next take action a t The expected cumulative reward that can be obtained. The advantage function is used to measure the t Take action a t The advantage gained compared to other actions is defined as A r (a t , s t )=Q r (s t , a t )-V r (s t ). For policy π, its discounted state access distribution can be expressed as The access distribution of discounted state-action to policy π can be expressed as ρ π (s, a) = ρ π (s)π(a|s), also known as the occupancy metric of policy π, where the occupancy metric and the policy are in one-to-one correspondence.
[0083] For example, Deep Reinforcement Learning (DRL) is an extension of reinforcement learning. In deep reinforcement learning, neural networks are used to represent the policy π and the state value function V r (s t ) and the action-value function Q r (s t , a t ), which enables deep reinforcement learning algorithms to solve higher-dimensional and more complex problems.
[0084] It should be noted that the method proposed in this application is based on the TRPO algorithm and adopts an actor-critic structure, using an actor neural network to represent the policy and a critic neural network to represent the value function. The TRPO algorithm can find an appropriate update step size for the policy gradient, so that each iteration of the policy is monotonically increasing towards the objective function.
[0085] In an embodiment, before step S103, the following steps can also be included:
[0086] Input the number of interference nodes M, the interference bandwidth set Maximum interference power Iteration number E max , step number step max , confidence domain δ of the strategy countermeasure imitation stage k , confidence domain δ of the strategy improvement stage
[0087] Input the spectrum distribution of the channels in all frequency hopping sets
[0088] Input example data and experience pool
[0089] Referring to Figures 4 to 6 , in each strategy iteration process, the embodiment designs a strategy improvement stage and a strategy countermeasure imitation stage. Taking the k+1th strategy iteration as an example, the initial allocation strategy is denoted as π k , the intermediate allocation strategy is denoted as π k+1 / 2 , the final allocation strategy is denoted as π k+1 , and the confidence domain of the strategy improvement stage is denoted as δ.
[0090] Further, in the strategy improvement stage of the k+1th strategy iteration, the initial allocation strategy π k in the current iteration is optimized based on the TRPO algorithm, including the following steps:
[0091] The initial allocation strategy π k is interacted with the environment to generate an interaction trajectory and an external reward r;
[0092] A target function is constructed based on maximizing the cumulative discounted reward of the interaction trajectory;
[0093] The target function is solved under the KKT condition to obtain the strategy parameters of the intermediate allocation strategy π k+1 / 2 .
[0094] Through the above steps, the intermediate allocation strategy π k that maximizes the target function can be found within the confidence domain δ of the initial allocation strategy π k+1 / 2 .
[0095] Further, the objective function of the policy improvement phase is the same as the TRPO algorithm, and the calculation formula of the objective function to be optimized in the policy improvement phase is:
[0096]
[0097] where D KL represents the Kullback-Leibler (KL) divergence of two distributions, π k represents the initial allocation policy in the k+1th policy iteration, π k+1 / 2 represents the intermediate allocation policy in the k+1th policy iteration, represents the discount state visit distribution of the initial allocation policy π k represents the advantage function, specifically the advantage function of the initial allocation policy π k in state s t takes action a t compared to the advantage that other actions can obtain.
[0098] Optionally, the policy parameters of the intermediate allocation policy obtained by solving the objective function using the KKT condition can be represented as follows:
[0099]
[0100] where (·) T represents the transpose operator, is the Hessian matrix of the average KL distance between policies, is the policy gradient;
[0101] using the final allocation policy π k obtained by the kth policy iteration, a set of trajectories and the policy gradient can be estimated using the set of trajectories, and can be represented as:
[0102]
[0103] where, represents the parameters after completing the kth policy iteration, represents the advantage function, specifically the advantage function of the policy π k in state s t takes action a t compared to the advantage that other actions can obtain.
[0104] However, the inverse of the Hessian matrix in the objective function has a large amount of calculation, so when solving the objective function, the can be regarded as a whole, and only the Then it can be transformed into solving the following equation
[0105]
[0106] in, It can be obtained by conjugate gradient algorithm, which belongs to the existing technology and will not be described here. Substitute the solution formula (15) of the objective function and use backtracking linear search to find the initial allocation strategy π k Update and get the following updated intermediate allocation strategy π k+1 / 2 Parameters:
[0107]
[0108] Where z∈{0, 1, 2, …, Z} is the minimum value that satisfies the KL divergence constraint, Z is the maximum number of backtracking steps, and σ is the backtracking coefficient.
[0109] For example, the Actor-Critic algorithm is a reinforcement learning method that combines policy gradient and value function. It works together through an Actor network and a Critic network to solve reinforcement learning problems in continuous action space and high-dimensional state space.
[0110] In one embodiment, the update of the value function network in the policy improvement phase is the same as the update of the critic network in the general actor-critic algorithm. The value function network is optimized by minimizing the mean square error of the temporal difference error in the samples, which can be achieved by the following formula:
[0111]
[0112] in, represents the parameters of the value function after completing the policy improvement step, Represents the parameters of the value function after completing the kth policy iteration, V r (s t ) indicates that from state s t The expected cumulative reward that can be obtained by starting to follow the strategy, V r (s t+1 ) indicates that from state s t+1 The expected cumulative reward that can be obtained by starting to follow the strategy, γ represents the discount factor;
[0113] Finally, the stochastic gradient descent method is used to complete Updates.
[0114] In one embodiment, in the strategy adversarial imitation phase of the k+1th strategy iteration, using example data and the initial allocation policy π k in the current iteration k+1 / 2 , comprising the following steps:
[0115] sampling data from the example policy π and the initial allocation policy π k in the policy improvement stage, respectively;
[0116] training the discriminator network by stochastic gradient descent using the sampled data;
[0117] using the output of the discriminator network as an intrinsic reward function to optimize the intermediate allocation policy π k+1 / 2 .
[0118] Through the above steps, the example policy π d can play a role in guiding and restricting, so that the intermediate allocation policy π k+1 / 2 optimized after the policy improvement stage can continue to be optimized in a sparse reward environment, or in other words, the intermediate allocation policy π k+1 / 2 can be forced to imitate the actions of the example policy π d in order to effectively explore in an unknown environment.
[0119] In one embodiment, when training the discriminator network, a binary cross-entropy loss is used to calculate the loss function of the discriminator network, and the discriminator network is trained by combining the sampled data with the loss function of the discriminator network, so as to realize the training of the discriminator network.
[0120] Further, the loss function of the discriminator network can be represented as:
[0121]
[0122] where D φ (s t , a t ) represents the discriminator network, represents the discounted state visit distribution of the final allocation policy π k+1 , represents the discounted state visit distribution of the example policy π d , and Loss(D φ ) represents the loss function of the discriminator network.
[0123] In one embodiment, when the output of the discriminator network is used as an intrinsic reward function to optimize the intermediate allocation policy π k+1 / 2 in the policy confrontation imitation stage, the optimization target can be represented as:
[0124]
[0125] Among them, π k+1 / 2 represents the intermediate allocation strategy in the k+1th policy iteration, π k+1 represents the final allocation strategy in the k+1th policy iteration, represents the intermediate allocation strategy π k+1 / 2 The distribution of discount status visits, Denotes the final allocation strategy π k+1 The distribution of discount status visits, represents the intermediate allocation strategy π k+1 / 2 In state s t Take action a t Compared with the advantages that can be obtained by other actions, D KL (·||·) represents the Kullback-Leibler (KL) divergence of the two distributions, δ k Represents the confidence region of the strategy against the imitation phase.
[0126] Optionally, the confidence region δ in the optimization objective of the strategy adversarial imitation phase k is set to decrease geometrically as the number of policy iterations increases. This setting allows the example policy π to be adaptively reduced as the optimization proceeds. d The teaching effect of strategy.
[0127] It should be noted that this application obtains the optimization goal of the above strategy in the anti-imitation stage through the following process:
[0128] First, the example policy π d It can be a non-perfect strategy, that is, J r (π d )≤J r (π * ), where the imperfect strategy consists of the average person's answers to a certain type of problem, and the strategy π * is the optimal strategy that maximizes the cumulative discounted reward of the trajectory. In this embodiment, the example strategy π d It is not necessary to have an advantage over all strategies. Example strategy π d It only needs to be better than the policy at the very beginning. This assumption is reasonable considering that the initial policy is an untrained policy or even a randomly initialized policy.
[0129] Therefore, in the strategy adversarial imitation stage, we set the example strategy π d The sampled action ratio is π from the initial allocation policy k The actions sampled in have advantage values, and the advantage values satisfy:
[0130]
[0131] where ω denotes the advantage value, π k denotes the initial allocation policy in the k+1th policy iteration, π d denotes the example policy, denotes the example policy π d in state s t takes action a the advantage over other actions, denotes the initial allocation policy π k in state s t takes action a t the advantage over other actions.
[0132] Then, the objective function of the policy adversarial imitation phase is set as follows:
[0133]
[0134] where D KL (·||·) denotes the Kullback-Leibler (KL) divergence of two distributions, denotes the intermediate allocation policy π k+1 / 2 , and denotes the final allocation policy π k+1 , and δ k denotes the confidence domain of the policy adversarial imitation phase;
[0135] In order to enable the example policy π d to be effectively imitated by the initial allocation policy π k so as to facilitate policy exploration, the confidence domain δ k is given a larger initial value in the embodiment, and a change mode is designed in which the confidence domain δ k decreases geometrically with the increase of the iteration number k. Through the above setting, the example policy π d can be adaptively reduced in the role of policy teaching as the optimization proceeds.
[0136] It is difficult to directly minimize , because the samples used to optimize the policy π k+1 come from the policy π k+1 to be optimized, which is difficult to achieve. Therefore, the formula (20) is approximately solved by referring to the method of constructing a surrogate function in the TRPO algorithm. First, an internal reward function r I is defined, and the internal reward function r I can be expressed as r I = log[π k+1 (s t , a t) / π d (s t , a t )], using the internal reward function r I Replacing the external reward r in the objective function and advantage function, we can get the internal reward function r I Variables and
[0137] Next, assume that the policy π k+1 / 2 is obtained through the strategy improvement stage, then for any Strategy π k+1 , the following inequalities all hold:
[0138]
[0139] Among them, D KL (·||·) represents the Kullback-Leibler (KL) divergence of the two distributions, π k+1 / 2 represents the intermediate allocation strategy in the k+1th policy iteration, π k+1 represents the final allocation strategy in the k+1th policy iteration, represents the intermediate allocation strategy π k+1 / 2 The distribution of discount status visits, Denotes the final allocation strategy π k+1 The discounted state visit distribution, δ k represents the confidence region of the strategy adversarial imitation phase, represents the advantage function, γ is the discount factor, π d It is important to note that the proof of this assumption will be given later.
[0140] So we can use the right half of inequality (21) to minimize instead of minimizing In inequality (21), and strategy π k+1 Only relevant That is to say, we only need to minimize this term, so we can transform the solution of formula (20) into the solution of the following formula:
[0141]
[0142] Among them, D KL (·||·) represents the Kullback-Leibler (KL) divergence of the two distributions, π k+1 / 2 represents the intermediate allocation strategy in the k+1th policy iteration, π k+1 represents the final allocation strategy in the k+1th policy iteration, represents the intermediate allocation strategy π k+1 / 2The discounted state visit distribution, δ k represents the confidence region of the strategy adversarial imitation phase, represents the advantage function.
[0143] Considering that in practice we can usually only obtain the example strategy π d Generated sample data But it is not possible to obtain the example strategy π in a detailed form d , which is important for the internal reward function r I =log[π k+1 (s t , a t ) / π d (s t , a t )] is very difficult to calculate. This embodiment uses sample data and strategy π k The generated trajectory is used to construct the internal reward. Thanks to the occupancy measure ρ π (s, a) and strategy π(s, a) are in one-to-one correspondence, so we can estimate Instead of estimating r I =log[π k+1 (s t , a t ) / π d (s t , a t )], and it is estimated You can use the sample data To achieve this, see Figure 5 As shown in , it can be estimated by the discriminator network in the Generative adversarial networks GAN
[0144] The above is estimated by the discriminator network in the Generative adversarial networks GAN It should be noted that the strategy π k+1 Think of it as the generator in GAN, and the discriminator D φ (s t , a t ) is the neural network we want to train. Specifically, we will use the occupancy metric The sampled state-action pairs (s t , a t ) is assigned a label of "1" and will be measured based on the occupancy The sampled state-action pairs (s t , a t) is given the label "0". The discriminator takes a state-action pair (s t , a t ) as input and outputs a real number in the range [0, 1], which represents the probability that the discriminator believes that the (s t , a t ) is from the policy π k+1 rather than the example policy π d . In this embodiment, binary cross-entropy loss is used to calculate this, so that the loss function Loss(D φ (s t , a t )) of the discriminator D φ can be obtained. Based on the loss function Loss(D φ ), the discriminator can be trained by Stochastic Gradient Descent (SGD).
[0145] It should also be noted that for the generator, i.e. the policy π k+1 , the goal of the policy π k+1 is to make its (s t , a t ) be mistaken by the discriminator as being from the example policy π d , so for the input the generator, i.e. the policy π k+1 , wants the corresponding output of D φ (s t , a t ) to be as small as possible, so the output of the discriminator D φ (s t , a t ) can be used as an intrinsic reward to optimize the policy π k+1 , so the intrinsic reward is set as r D = -log[D φ (s t , a t )]. This intrinsic reward r D augments the environment reward with example data π d , and when the environment feedback is sparse or exploration is insufficient, this intrinsic reward r D can force the policy π k+1 to produce trajectories similar to the example policy π d , in other words, to minimize the divergence between the policy π k+1 and the policy π d , so the method provided in this embodiment can more effectively explore the environment.
[0146] Finally, as shown in Figure 6 , the intrinsic reward r DSubstitute r in equation (22) with I , and get the optimization objective of the policy confrontation imitation stage, that is, equation (4).
[0147] Based on the fact that equation (4) and equation (1) have similar forms, the structure of the solution of equation (1) can be directly used, that is, replace r in equation (16), equation (17) and equation (18) with r D , replace δ with δ k , and change the gradient ascent to gradient descent, and then get the following results:
[0148]
[0149]
[0150]
[0151] wherein, represents the parameter after completing the policy improvement stage, represents the parameter after completing the k+1th policy iteration, represents the advantage function, z∈{0,1,2,…,Z} is the minimum value satisfying the KL divergence constraint, Z is the maximum backtracking step number, and σ is the backtracking coefficient.
[0152] The parameter update of the value function network in the policy confrontation imitation step and the policy improvement stage is the same, which can be realized by the following equation:
[0153]
[0154] wherein, represents the parameter of the value function after completing the k+1th policy iteration, represents the parameter of the value function after completing the policy improvement stage, represents the expected cumulative reward that can be obtained by following the policy from state s t , represents the expected cumulative reward that can be obtained by following the policy from state s t+1 , and γ represents the discount factor.
[0155] Finally, the update of is completed by the stochastic gradient descent method.
[0156] It should be noted that, regarding the assumption that the policy π k+1 / 2 is obtained through the policy improvement stage, for any policy π satisfying k+1 , the following inequalities are all established:
[0157]
[0158] where D KL (·||·) denotes the Kullback-Leibler (KL) divergence of two distributions, π k+1 / 2 denotes the intermediate allocation policy in the k+1th policy iteration, π k+1 denotes the final allocation policy in the k+1th policy iteration, denotes the discounted state visit distribution of the intermediate allocation policy π k+1 / 2 denotes the discounted state visit distribution of the final allocation policy π k+1 k denotes the confidence region of the policy adversarial imitation phase, denotes the advantage function, γ is the discount factor, π d denotes the example policy. The proof process is given as follows:
[0159] First, the following two lemmas are given:
[0160] Lemma 1: For any two policies π and π', there is
[0161]
[0162] Proof:
[0163]
[0164] And combined with the average KL divergence is less than or equal to the maximum value of KL divergence
[0165] The desired result can be obtained.
[0166] Lemma 2: For any policy π, there is
[0167]
[0168] Proof:
[0169]
[0170] Proof of Theorem 1:
[0171] Combined with the results of Lemma 2 and the inequality given in Lemma 1, then for any policy π and π', the following can be obtained:
[0172]
[0173] Replacing π k+1 / 2 with π' and π k+1 with π, the following can be obtained:
[0174]
[0175] Then, for any π in k+1 , both
[0176]
[0177] Using this inequality in the previous inequality, we can get the final result and complete the proof.
[0178] The specific algorithm of the frequency hopping interference resource allocation method based on learning from imperfect examples proposed in this application is shown in Table 1 below:
[0179] Table 1: Trust region interference strategy optimization algorithm based on imperfect example data
[0180]
[0181] It can be seen that the Imperfect Demonstration-Assisted Trust Region Policy Optimization (IDA-TRPO) algorithm, which is assisted by imperfect example data, constructs a policy improvement phase and a policy adversarial imitation phase based on a double-reset trust region in each policy iteration to achieve policy update optimization, so that the optimization of the interference resource allocation strategy can be completed under the general binary reward function setting.
[0182] It should be noted that in the trust region interference strategy optimization algorithm based on imperfect example data, when the number of iterations does not exceed 10, if the cumulative interference reward obtained in the current iteration exceeds the average cumulative interference reward of the previous iteration, the following is executed: reduce δ according to the attenuation rule k ; When the number of iterations exceeds 10, if the cumulative interference reward obtained in the current iteration exceeds the average cumulative interference reward of the previous 10 iterations, then execute: reduce δ according to the decay rule k .
[0183] In the embodiment of the present application, in order to further verify the effect of the frequency hopping interference resource allocation method based on learning from imperfect examples proposed in the present application, the following simulation experiment was conducted:
[0184] The interference scenarios set in this simulation experiment are as follows:
[0185] In this embodiment, the communicating party has N = 5 frequency-hopping communication networks, and the number of frequency-hopping channels in the frequency-hopping frequency set used by the communicating party is 32, 32, 64, 64, and 128, respectively. The interfering party has M = 12 interfering nodes, each of which uses a PBNJ interference pattern. Due to the non-cooperative nature of the adversarial parties, the interfering party needs to obtain intelligence on the communicating party based on electronic warfare intelligence support methods to estimate the interference effect. To cover the calculation deviation caused by the estimation, the interfering party usually needs to make the interference resources redundant. Therefore, in the simulation experiment, the transmission distance of the interference signal is set to be greater than the transmission distance of the communication signal. In this way, the loss suffered by the interference signal during the propagation process will be greater than the loss of the communication signal.
[0186] The simulation parameters set in this simulation experiment are shown in Table 2 below:
[0187] Table 2: Simulation parameter settings
[0188]
[0189]
[0190] In order to verify the excellent performance of the method proposed in this application, the following is an illustration through the comparison of various algorithms. The various algorithms are as follows:
[0191] Expert: We trained the state-of-the-art LSTM-A3C model in a dense reward setting and brought it to convergence, serving as the optimal baseline.
[0192] Example: The state-of-the-art LSTM-A3C model was trained in the dense reward setting, but was far from convergence, resulting in only mediocre performance for the example policy.
[0193] IDA-TRPO: The method proposed in this application is trained under a sparse reward setting and provides example data generated by example policies as an aid.
[0194] GAIL: The GAIL method disclosed in the prior art is trained under a sparse reward setting, and example data generated by example policies are provided as assistance.
[0195] LSTM-A3C: The LSTM-A3C model is trained in a sparse reward setting.
[0196] TRPO: Train the existing TRPO model in a sparse reward setting.
[0197] Continuous blocking interference: The interference frequency bands determined by all interference nodes are continuous and connected, forming a large continuous interference frequency band. The center frequency of this frequency band switches in a sweeping manner, and the bandwidth is the sum of the maximum bandwidths selectable by all nodes.
[0198] The evaluation indicators for comparing the performance of various algorithms are as follows:
[0199] Evaluation metric 1: Interference success rate: In a certain number of interference tests, count the number of times the interference resource allocation scheme given by the model in these tests can successfully interfere with a specified frequency set, and calculate the proportion of experiments with successful interference.
[0200] Evaluation metric 2: Cumulative interference reward: In an interference test, the model gives sub-interference resource allocation plans for all interference nodes in a time-series manner. First, the sum of the rewards obtained by these sub-interference resource allocation plans is calculated, and then the sum of the rewards in all interference tests is averaged.
[0201] Evaluation metric 3: Interference bandwidth: In an interference test, the model presents sub-interference resource allocation schemes for all interfering nodes in a time-series manner. The total interference bandwidth consumed by these sub-interference resource allocation schemes is first calculated, and then the total interference bandwidth in all interference tests is averaged.
[0202] The model details and hyperparameter settings of this simulation experiment are as follows:
[0203] The strategy, value function, and discriminator in IDA-TRPO all use a three-layer multilayer perceptron (MLP), with 128 neurons in the hidden layer, the activation function is Tanh, the learning rate of the value function is 0.01, the learning rate of the discriminator is 0.0003, and the neural network is optimized using the Adam optimizer. The trust region δ in the strategy improvement phase is 0.0005, and the trust region δ in the strategy adversarial imitation phase is 0.0005. k The initial value of δ0=0.0002, and the adaptive attenuation rule is: δ =5000 iterationsδ k No attenuation occurs, at W δ After iterations, when the reward obtained by the strategy of the current iteration round is higher than the average reward of the previous 10 times, a geometric decay δ is performed k ←αδ k , α = 0.95. The number of conjugate gradient algorithm iterations was 10, the maximum number of linear searches Z was 10, the discount factor γ was 0.98, the backtracking coefficient σ was 0.5, and the generalized advantage estimation (GAE) coefficient was 0.95. To ensure fairness in the comparative experiments, all compared algorithms used the same network structure and optimizer.
[0204] The simulation platform settings of this simulation experiment are as follows:
[0205] The simulation of the interference scenario and the algorithm procedure are built with Python 3.7 and OpenAI's Gym 0.23.1, and the gym uses PyTorch 1.11.0 to build and train the neural network.
[0206] Analysis of the results of the simulation experiment:
[0207] In order to better compare the performance differences of various algorithms, the learning-based algorithm is iteratively trained under the sparse interference reward setting, and then evaluated under the dense interference reward setting. In addition, after every 400 training iterations, 1 evaluation is performed, and each evaluation includes 100 interference tests, and the average interference success rate and the average cumulative interference reward of the 100 interference tests are counted as the evaluation results. The non-perfect example data used to assist the algorithm training in the present application has an interference reward of 79 under the dense reward setting, while the interference reward of the expert data under the dense reward setting is 240.
[0208] Reference Figure 7 as shown in, Figure 7 The evaluation results of different methods in the iterative training process are shown. From Figure 7 It can be seen very intuitively that with the continuous advancement of iterations, under the assistance of the Demo data, the cumulative interference reward of the method proposed in the present application shows an upward trend and can approximate the expert level. The GAIL method which also uses Demo data can only approximate the level of Demo data, because the GAIL method relies on the occupation metric to pull the self policy and the Demo data through imitation learning, so it cannot break through the performance upper bound of the Demo data. The deep reinforcement learning-based method, such as the LSTM-A3C method and the TRPO method, cannot obtain useful feedback information during iterative training, and the policy network training continues to deteriorate, resulting in a continuously decreasing cumulative interference reward. It can be seen that the traditional deep reinforcement learning model relies too much on the setting of the reward function. In addition, it is also noted that the reward curve of the continuous jamming interference is always above the reward curve of the LSTM-A3C method and the TRPO method, but it is at a relatively low level, because in the continuous jamming interference, the interference frequency bands of all interference nodes are connected, and the total interference frequency band formed finally covers a lot of spectrum without user channels, which inevitably leads to the waste of interference resources.
[0209] Figure 8 The comparison diagram of the total interference bandwidth consumed by different methods is shown. From Figure 8It can be seen that the total interference bandwidth used by the expert is 16.5MHz, and the total interference bandwidth used by the method proposed in the present application and the GAIL method is close to the expert level, but the GAIL method shows a large variance due to insufficient training. In addition, the interference bandwidth used by the LSTM-A3C method and the TRPO method is the smallest, which is also the direct reason why the LSTM-A3C method and the TRPO method show the worst interference performance. From the above simulation results, it can also be seen that under the premise of lacking a fine reward function, only the method proposed in the present application and the GAIL method show effectiveness, so in the following experiments, the method proposed in the present application and the GAIL method are mainly compared.
[0210] Figure 9 The interference success rate of the method proposed in the present application on different frequency sets is shown, Figure 10 The interference success rate of the GAIL method on different frequency sets is shown, Figure 9 and Figure 10 It can be directly seen from the above that the interference success rate curves of the two methods on different frequency sets have similar trends with the cumulative interference reward curves of the corresponding methods in Figure 7 And Figure 9 and Figure 10 Both reflect that for the method proposed in the present application and the GAIL method, the interference on the frequency hopping set 5 is the most difficult, because the frequency hopping set 5 has the relatively most frequency hopping channels and the widest spectrum span, so the interference on the frequency hopping set 5 needs more careful planning of the allocation of interference resources.
[0211] Figure 11 The comparison diagram of the interference frequency bands constructed by different methods after convergence is shown. From Figure 11 It can be seen that the interference frequency bands constructed by the expert method, the method proposed in the present application and the GAIL method in the spectrum all span a large spectrum interval, and compared with the interference frequency bands constructed by the expert method, the interference frequency bands constructed by the method proposed in the present application in the spectrum are more concentrated in the part from 350MHz to 400MHz. This is because a large number of channels belonging to the fifth frequency hopping set are distributed in 350MHz to 400MHz, and it is necessary to focus the interference power on interfering some channels of this frequency hopping set to make the number of interfered channels in the fifth frequency hopping set reach the minimum threshold.
[0212] This application uses deep reinforcement learning technology to solve the interference resource allocation problem involved in countering multiple frequency hopping communication networks, and focuses on how to bypass the tedious step of designing a sophisticated reward function to optimize the interference strategy concisely and intuitively. To achieve this goal, an interference resource allocation method based on learning from imperfect example data is proposed. This method divides the interference strategy optimization into two stages based on the idea of trust region optimization, namely the strategy improvement stage and the strategy confrontation imitation stage. The TRPO optimization method is introduced in the strategy improvement stage, and the information contained in the example data is used to guide the further optimization of the strategy in the strategy confrontation imitation stage, thereby improving the exploration efficiency of the strategy in a sparse environment. At the same time, the simulation experimental results confirm that the method proposed in this application can achieve the optimal effect of training under a simple and rough binary interference reward function setting, and greatly surpasses the performance of the example data used for auxiliary training, thereby improving the practicality of the interference resource allocation method based on deep reinforcement learning. It also shows that the generalization of the interference strategy to newly emerging interference tasks can be improved without the need to design a sophisticated reward function, thereby achieving rapid optimization of new tasks.
[0213] It should be noted that although several units of the system for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiment of the present disclosure, the features and functions of two or more units described above can be concretized in one unit. Conversely, the features and functions of a unit described above can be further divided into multiple units for concretization. Some or all of the units can be selected according to actual needs to achieve the purpose of the disclosed solution. Those of ordinary skill in the art can understand and implement it without paying creative work.
[0214] Those skilled in the art will readily appreciate other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, with the true scope and spirit of the present disclosure being indicated by the appended claims.
Claims
1. A frequency hopping interference resource allocation method based on learning from imperfect examples, characterized in that include: Construct the frequency hopping jamming resource allocation problem based on the communication confrontation scenario; Modeling the frequency hopping interference resource allocation problem as a Markov decision process; Randomly initialize the policy network parameters and the discriminator network; The initialized policy network is iterated multiple times, and in each iteration, the policy improvement phase and the policy adversarial imitation phase are constructed based on the double reset trust region; In the strategy improvement phase, the initial allocation strategy in the current iteration is optimized based on the TRPO algorithm to obtain an intermediate allocation strategy; In the strategy adversarial imitation phase, the discriminator network is trained using the example data and the interaction data of the initial allocation strategy in the current iteration to optimize the intermediate allocation strategy to obtain a final allocation strategy; The optimization of the initial allocation strategy in the current iteration based on the TRPO algorithm includes the following steps: Interacting with the environment through initial allocation strategies generates interaction trajectories and external rewards; Construct an objective function based on maximizing the cumulative discounted reward of the interaction trajectory; The objective function is solved under the KKT condition to obtain the parameters of the intermediate allocation strategy.
2. The frequency hopping interference resource allocation method based on learning from imperfect examples according to claim 1, characterized in that The calculation formula of the objective function is: in, represents the Kullback-Leibler (KL) divergence of two distributions, Indicates the The initial allocation strategy in the policy iteration, Indicates the The intermediate allocation strategy in the policy iteration, Represents the initial allocation strategy The distribution of discount status visits, represents the advantage function.
3. The frequency hopping interference resource allocation method based on learning from imperfect examples according to claim 1, characterized in that The method of training the discriminator network using the example data and the interaction data of the initial allocation strategy in the current iteration to optimize the intermediate allocation strategy comprises the following steps: Sampling data from the example strategy and the initial allocation strategy in the strategy improvement phase respectively; Using the sampled data, the discriminator network is trained by stochastic gradient descent; The output of the discriminator network is used as the intrinsic reward function to optimize the intermediate allocation strategy.
4. The method for allocating frequency hopping interference resources based on learning from imperfect examples according to claim 3, wherein: When training the discriminator network, a binary cross entropy loss is used to calculate the loss function of the discriminator network, and the sampled data is combined with the loss function of the discriminator network to train the discriminator network.
5. The method for allocating frequency hopping interference resources based on learning from imperfect examples according to claim 4, characterized in that The loss function of the discriminator network can be expressed as: in, represents the discriminator network, Represents the final allocation strategy The distribution of discount status visits, Represents the final allocation strategy The distribution of discount status visits, represents the loss function of the discriminator network.
6. The method for allocating frequency hopping interference resources based on learning from imperfect examples according to claim 3, wherein: In the strategy adversarial imitation phase, it is assumed that the action sampled from the example strategy has an advantage value over the action sampled from the initial allocation strategy, and the advantage value satisfies: in, represents the advantage value, Indicates the The initial allocation strategy in the policy iteration, represents an example policy, Represents an example policy In state Take action Compared with the advantages that can be obtained by other actions, Represents the initial allocation strategy In state Take action Advantages compared to other actions.
7. The method for allocating frequency hopping interference resources based on learning from imperfect examples according to claim 3, wherein: In the strategy adversarial imitation stage, the output of the discriminator network is used as the intrinsic reward function to optimize the intermediate allocation strategy. The optimization objective can be expressed as: in, Indicates the The intermediate allocation strategy in the policy iteration, Indicates the The final allocation strategy in the policy iteration, Represents the intermediate allocation strategy The distribution of discount status visits, Represents the final allocation strategy The distribution of discount status visits, Represents the intermediate allocation strategy In state Take action Compared with the advantages that can be obtained by other actions, represents the Kullback-Leibler (KL) divergence of two distributions, Represents the confidence region of the strategy against the imitation phase.
8. The method for allocating frequency hopping interference resources based on learning from imperfect examples according to claim 7, characterized in that The confidence region in the optimization objective of the strategy adversarial imitation phase is set to decrease geometrically as the number of strategy iterations increases.
9. The method for allocating frequency hopping interference resources based on learning from imperfect examples according to any one of claims 1 to 8, characterized in that: Modeling the frequency hopping interference resource allocation problem as a Markov decision process includes: defining a state space, an action space, a state transition function, and a reward function for the Markov decision process; A discrete Markov decision process is constructed according to the state space, action space, state transfer function and reward function.
Citation Information
Patent Citations
Method for generating radar intelligent cognitive anti-interference strategy
CN112904290A
Anti-interference zero sum Markov game model and maximum and minimum depth Q learning method
CN116866048A