Intelligent Decision-making Method for Three-party Devices in Physical-layer Secure Wireless Transmission Based on Dynamic Coalition Game

By adopting dynamic alliance game and deep reinforcement learning methods in the physical layer secure wireless transmission environment, the problems of dynamic confrontation and alliance between three-party devices are solved, the optimal strategy and long-term average utility of three-party devices are maximized, and the performance and stability of wireless transmission are improved.

CN116405942BActive Publication Date: 2025-06-10NANJING UNIV OF AERONAUTICS & ASTRONAUTICS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310203365.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-06
Publication Date
2025-06-10
Estimated Expiration
2043-03-06

AI Technical Summary

Technical Problem

The existing physical layer security technology fails to fully consider the dynamic confrontation and alliance issues between three-party devices (legitimate users, eavesdropping devices and jammers) in wireless transmission environments, resulting in increased policy uncertainty and performance optimization difficulties.

Method used

Using a method based on dynamic alliance game, a multi-stage sequential game and dynamic alliance game model is constructed, and the strategic interaction and dynamic alliance behavior of the three-party devices are modeled respectively, and an alliance formation algorithm based on alliance switching criteria is designed and an intelligent decision-making algorithm based on deep reinforcement learning is realized to achieve the optimal alliance selection and intelligent decision-making of the three-party devices.

Benefits of technology

By dynamically analyzing the alliance relationship and strategy evolution of the three-party equipment, the long-term average utility is maximized under system uncertainty, and the performance and stability of secure wireless transmission of the physical layer are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116405942B_ABST
    Figure CN116405942B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for intelligent decision-making of three-party devices in physical-layer secure wireless transmission based on dynamic coalition game. For the dynamic cooperation and confrontation relationships among the eavesdropping device, legitimate user, and jammer in a wireless transmission environment considering physical-layer security, multi-stage sequential game and dynamic coalition game are used for modeling, aiming to maximize the respective utilities of the three parties and help these three-party devices make intelligent decisions. In a wireless transmission system considering physical-layer security, the present invention uses a coalition formation algorithm based on dynamic coalition game to complete the selection of coalition partners for the three-party devices, and trains agents representing each party's device through a deep reinforcement learning algorithm to complete the optimal decisions of the three-party devices, including the base station selection, transmission power allocation, and unit incentive amount formulation of the legitimate user, the single-device activation selection and unit incentive amount formulation of the eavesdropping device, and the interference power allocation of the jammer.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of physical layer security, and particularly relates to the intelligent decision-making of confrontation and alliance among three-party network devices in wireless transmission considering physical layer security. More particularly, the present invention relates to an intelligent decision-making method for three-party devices in physical layer security wireless transmission based on dynamic coalition game. Background Art

[0002] In recent years, Physical Layer Security (PLS), as a promising wireless security technology, has been rapidly developed in the field of 5G and next-generation communications. PLS is widely regarded as an effective method to protect secure transmission in Internet of Things applications with high security requirements such as autonomous driving, remote surgery, and intelligent transportation. Different from traditional cryptography-based methods, PLS has lower computational complexity, resource consumption, and transmission delay, and is more suitable for time-sensitive and power-constrained application scenarios.

[0003] Although PLS has received extensive attention in different aspects, most of the existing research inventions have not fully explored the selfishness of the three parties in PLS, namely legitimate users, eavesdropping devices, and jammers. Specifically, in practice, legitimate users, eavesdropping devices, and jammers may exhibit selfishness for the maximization of their own interests, but their strategies are not always conflicting, and sometimes they are mutually beneficial. On the one hand, legitimate users and jammers may form an alliance to confront eavesdropping devices. Legitimate users can provide rewards (such as monetary rewards) to jammers in exchange for the latter's help in increasing the interference ability against eavesdropping devices, thereby protecting the confidential messages transmitted in the open wireless environment. On the other hand, eavesdropping devices and jammers may form an alliance against legitimate users. In such an alliance, eavesdropping devices can also motivate jammers to interfere with legitimate users, forcing them to increase the data transmission power, thus making legitimate users vulnerable to eavesdropping. Obviously, this complex relationship (i.e., alliance formation) may not be predefined, so the impact on PLS needs to be carefully modeled and analyzed, which is very important but extremely challenging for the following reasons:

[0004] a. From the perspective of the respective interests of the three-party devices in PLS, in addition to the possible alliances, legitimate users can independently decide the target base station for their uplink transmission and allocate data transmission power to improve the transmission rate. At the same time, eavesdropping devices at different geographical locations can choose to be active or dormant at different times to reduce energy consumption. In addition, jammers can better allocate interference power on different links to obtain higher rewards from legitimate users or eavesdropping devices. This requires a multi-stage sequential game with multi-dimensional strategies, which includes a dynamic coalition game as a sub-game to model the decision-making of the three-party devices for alliance selection.

[0005] b. Due to the uncertainties of wireless systems, such as time-varying channel conditions, the strategies of the three parties in the above PLS may change dynamically. Long-term performance optimization requires the study of dynamic games. In particular, the potential coalition game also becomes dynamic, which means that any two of the three devices may temporarily form a coalition and adjust dynamically, that is, merge or split over time. However, to the best of the publicly available information, this key issue has not been solved in previous inventions. Summary of the Invention

[0006] Objective of the Invention: Aiming at the problem that the existing physical layer security technology in the above wireless transmission environment does not fully consider the selfishness and dynamic coalition formation of the three devices, the present invention provides an intelligent decision-making method for three devices in physical layer security wireless transmission based on dynamic coalition game.

[0007] Technical Solution: An intelligent decision-making method for three devices in physical layer security wireless transmission based on dynamic coalition game. This method faces the possible dynamic confrontation and coalition formation behaviors among three parties: legitimate users, eavesdropping devices, and jammers in an open wireless communication environment. It constructs the utility functions of the three devices respectively using physical quantities such as the secrecy transmission rate, eavesdropping rate, and energy consumption of each device required by physical layer security. It uses multi-stage sequential game and dynamic coalition game to model the strategic interaction and dynamic coalition formation behaviors of the three devices respectively. With the goal of maximizing the long-term average utility of each of the three network devices in the open wireless communication environment, it designs a coalition formation algorithm based on coalition switching criteria and an intelligent decision-making algorithm based on deep reinforcement learning to achieve the coalition selection and intelligent decision-making of the three devices;

[0008] Furthermore, the method includes establishing a network model considering physical layer security scenarios in an open wireless communication environment. The network devices therein include legitimate users, eavesdropping devices, jammers, and base stations. The legitimate users transmit secret data to the base station in the uplink, while being eavesdropped by the eavesdropping devices. The jammer selects one of the legitimate users or the eavesdropping devices to form a coalition, that is, jamming the eavesdropping device to help the legitimate user improve the secrecy transmission rate or jamming the base station to improve the interception rate of the eavesdropping device. At the same time, the legitimate user or the eavesdropping device will give the jammer a reward (i.e., incentive amount) to attract the jammer to form a coalition with it. In this wireless transmission environment considering physical layer security, each legitimate user occupies an orthogonal channel with a frequency bandwidth of W for uplink transmission, and its power allocation adopts l-level discrete allocation, expressed as At the same time, the jammer also adopts l-level discrete allocation, expressed as To characterize the time-varying uncertainty, the overall running time of the system is divided into R time slots.

[0009] Furthermore, the establishment of the network model considering physical layer security scenarios in the open wireless communication environment by the method includes the following calculation and processing procedures:

[0010] (1) In each time slice, calculate the relevant physical quantities of the three-party devices, including the upload rate of the legitimate user and the secure transmission rate the eavesdropping rate of the eavesdropping device The calculation method of the upload rate is as follows:

[0011]

[0012]

[0013] Among them, represents the additive Gaussian white noise (AWGN) at base station m, and g nm (t) and g jm (t) respectively represent the instantaneous channel gains of the links from the legitimate user n and the jammer j to base station m; the eavesdropping rate The calculation method is as follows:

[0014]

[0015]

[0016] Among them, represents the AWGN at the eavesdropping device k, and g nk (t) and g ik (t) respectively represent the instantaneous channel gains of the links from the legitimate user n and the jammer j to the eavesdropping device k; the secure transmission rate The calculation method is as follows:

[0017]

[0018] Among them, [x] + = max(x, 0).

[0019] (2) Based on the respective physical quantities of each party's device, construct the utility functions of the three-party devices in each time slice and include the gains and losses of each party during the system operation;

[0020] The utility function of the eavesdropping device in time slice t is expressed as:

[0021]

[0022] Among them, x {EJ} (t) = 1 or 0 indicates whether the eavesdropping device and the jammer form an alliance, and c kDenote the activation cost of a single eavesdropping device within a time slice, which is the performance gain of the eavesdropping device, expressed as:

[0023]

[0024] which is the performance gain of the eavesdropping device without the help of a jammer, expressed as:

[0025]

[0026] The utility function of the legitimate user in time slice t is expressed as:

[0027]

[0028] where x {LJ} (t)=1 or 0 indicates whether the eavesdropping device and the jammer form an alliance, and ξ n represents the unit power consumption cost of the legitimate user, which is the performance gain of the legitimate user, expressed as:

[0029]

[0030] which is the performance gain of the legitimate user without the help of a jammer, expressed as:

[0031]

[0032] The utility function of the jammer in time slice t is expressed as:

[0033]

[0034] where η j represents the unit power consumption cost of the legitimate user, and c conf represents the potential configuration cost caused by the additional connections established by the jammer to notify the alliance change if the jammer chooses to change allies within two consecutive time slices, which is the amount of incentive paid by the legitimate user or the eavesdropping device to the jammer in time slice t, expressed as:

[0035]

[0036] (3) Respectively establish the strategy sets of the three-party devices, generate the long-term average utility maximization optimization problems for each of the three-party devices. For the eavesdropping device, its strategy set is expressed as and its optimization problem is expressed as:

[0037]

[0038] In the formula, represents the activation selection of the eavesdropping device in each time slice, and μ E (t) represents the unit excitation amount of the eavesdropping device in each time slice, represents the upper limit of the unit excitation amount.

[0039] For the legitimate user, its strategy set is expressed as: Its optimization problem is expressed as:

[0040]

[0041]

[0042] In the formula, represents the minimum transmission rate, represents the target base station selection of the legitimate user in each time slice, and μ L (t) represents the unit excitation amount of the legitimate user in each time slice. For the jammer, its strategy set is expressed as Its optimization problem is expressed as:

[0043]

[0044] (4) Construct a multi-stage sequential game to model the strategic interaction of the three-party devices. The expression of the multi-stage sequential game is as follows:

[0045]

[0046] Among them, respectively represent the eavesdropping device, the legitimate user, and the jammer participating in the game, represents the strategies of the three parties, represents the utility functions of the three parties. Each time slice contains three stages. First, the eavesdropping device makes decisions according to the optimization objective and μ E (t). Second, the legitimate user makes decisions according to the optimization objective and μ L (t). Finally, the jammer makes decisions according to the optimization objective The three stages will repeat in each time slice. At the beginning of each time period, the eavesdropping device and the legitimate user can observe the decisions of the jammer in the previous time slice to achieve long-term strategic interaction. For the dynamic alliance of the three-party devices in each time slice, a dynamic coalition game is used to model it, and its expression is as follows:

[0047]

[0048] Among them represents the eavesdropping device, legitimate user, and jammer participating in the game represents all possible coalitions that can be formed by the three-party devices in the dynamic coalition game is a sub-game of, used to transform the problem of solving the optimal coalition selection x {EJ} (t) and x {LJ} (t) into the problem of solving the equilibrium solution for finding the equilibrium solution

[0049] (5) Design a coalition formation algorithm based on the coalition switching criterion to solve the dynamic coalition game in each time slice to achieve the optimal coalition selection of the three-party devices in each time slice (i.e., x {EJ} (t) and x {LJ} (t)), and at the same time generate a stable coalition partition This coalition formation algorithm runs in a distributed manner, that is, each party independently calculates its own coalition selection within the same time slice. Essentially, it is to solve the equilibrium within each time slice i.e., the stable coalition partition This algorithm is implemented based on the following coalition switching criterion

[0050] Criterion 1 If and only if and

[0051]

[0052] Criterion 2 If and only if

[0053] where C a and C b represent two coalitions, and the binary relation symbol represents the coalition preference of a certain party i at time slice t, and the binary relation symbol represents the coalition transfer of a certain party i in time slice t, that is, from the coalition on the left side of the symbol to the coalition on the right side of the symbol

[0054] (6) Design an intelligent decision-making algorithm based on deep reinforcement learning to solve the global equilibrium solution of the multi-stage sequential game in the entire system running time 0 ≤ t ≤ T, and realize the decision variables of the three-party devices other than the coalition selection (i.e., and μ L(t))'s optimal decision. The algorithm for training the agent of the three-party device decision is based on the Proximal Policy Optimization (PPO) and the Actor-Critic (AC) framework. The state space of the reinforcement learning process comprehensively considers the network topology, instantaneous channel gain (including g nm (t), g nk (t), g jm (t) and g jk (t), signal transmission power (including and ), and coalition state (represented by x {EJ} (t) and x {LJ} (t)), and normalizes the environmental state value through the adjacency matrix NT(t). In addition, this intelligent decision-making algorithm based on deep reinforcement learning combines distributed training and centralized training. For different decisions of the three-party device, different agents are used to train the best strategy.

[0055] Beneficial effects: Compared with the prior art, the remarkable features and substantial progress of the present invention include the following three points:

[0056] First, the present invention establishes a hierarchical game model integrating dynamic trilateral coalition formation game to solve the strategic interaction modeling problem among legitimate users, eavesdropping devices, and jammers in PLS under system uncertainty. And in the utility modeling of the three-party device, all possible decisions, benefits, and costs of the three-party device in resource management and coalition selection are fully considered;

[0057] Second, considering the selfishness of the three parties in PLS, the present invention proposes a distributed coalition selection and coalition formation method based on coalition switching criteria to obtain the optimal coalition selection of each device. This method adopts the distributed operation of the three-party device and has high computing efficiency;

[0058] Third, aiming at maximizing the long-term utility of a given game, the present invention proposes an intelligent decision-making method for three-party devices based on deep reinforcement learning. This method can generate the optimal strategic decisions (i.e., equilibria) of all parties in PLS in multiple dynamically evolving time slices, can be applied to a dynamic wireless network system with dynamically changing channel states, and the agent obtained through reinforcement learning has high robustness in the decision-making process. Description of the Drawings

[0059] Figure 1 is a schematic diagram of the system structure and device interaction of the method described in the present invention;

[0060] Figure 2 is a multi-stage sequential game flow chart in the present invention;

[0061] Figure 3 It is a schematic diagram of the framework of the intelligent decision-making method based on reinforcement learning in the present invention;

[0062] Figure 4 It is a comparison chart of the training of the intelligent decision-making method based on reinforcement learning and the existing method in terms of the cumulative utility of the three-party devices. Specific implementation manners

[0063] In order to elaborate in detail the technical solution disclosed by the present invention, the present invention will be further described below in conjunction with the accompanying drawings and specific implementation manners.

[0064] First of all, the key problem solved by the method of the present invention is how to maximize the long-term average utility of each of the three-party devices, namely legitimate users, eavesdropping devices and jammers, with dynamic confrontation and alliance relationships in a wireless communication network considering physical layer security, fully considering the selfishness of the three-party devices and the uncertainty of the channel state, and forming an optimal strategy for the three-party devices in dynamic alliance selection and resource management decision-making.

[0065] The main idea of the present invention is to first propose a multi-stage sequential game framework including dynamic trilateral alliances to characterize the strategic interactions among all parties (legitimate users, eavesdropping devices and jammers) in PLS. In the case of system uncertainty (i.e., time-varying channel gain), a long-term optimization problem is established to maximize the utility of each party respectively. Then a multi-stage sequential game model is constructed to model the decision-making order and correlation of these three-party devices, where each party acts as a player in the game to maximize its expected revenue while minimizing its expected cost. In order to analyze the dynamic alliance relationship among the three-party devices, the present invention uses dynamic alliance game to model the alliance of the three-party devices in each time slice and gives the stability conditions that need to be satisfied for the dynamic change of the alliance in different time slots. Then a distributed alliance selection and alliance formation method for the three-party devices to form a stable alliance partition in each time slice is proposed. Finally, considering the dynamic nature of the strategy evolution of the three-party devices (especially dynamic alliance formation), the method described in the present invention involves a method based on deep reinforcement learning to solve the equilibrium solution of the multi-stage sequential game, that is, to generate the optimal strategy of the three-party devices.

[0066] Specifically, an intelligent decision-making method for three-party devices in physical layer security wireless transmission based on dynamic alliance game can be implemented according to the following steps:

[0067] Step1: Construct a wireless transmission network system model considering physical layer security.

[0068] First, construct a system model. As Figure 1 shown, the present invention considers a wireless uplink communication system, which consists of a set of legitimate users, denoted as Aiming to transmit secret data to a group of base stations, denoted as Each legitimate user occupies an orthogonal channel for its uplink transmission, and the set of uplink channels is also denoted as There is a set of eavesdropping devices, denoted as which may be active or dormant at different locations, and multiple jammers, denoted as It interferes with all links in the system.

[0069] To characterize the time-varying uncertainty, the overall operating time of the system is divided into T time slots, where each time slot \(t\in\{0,1,\ldots,T - 1\}\). Due to the presence of the jammer, from any legitimate user \(n\) in to any base station \(m\) of the uplink transmission link in the time slot \(t\) of the signal-to-noise ratio (Signal to Noise Ratio, SINR) is:

[0070]

[0071] where, represents the additive Gaussian white noise (Additive Gaussian White Noise, AWGN) at base station \(m\), \(g nm (t)\) and \(g jm (t)\) represent the instantaneous channel gains of the links from legitimate user \(n\) and jammer \(j\) to base station \(m\) in time slot \(t\), respectively, and are the \(l\)-level discrete power allocations adopted by legitimate user \(n\) and jammer \(j\), respectively, denoted as Similarly, in time slot \(t\), the SINR from legitimate user \(n\) to eavesdropping device \(k\) is:

[0072]

[0073] where, represents the AWGN at eavesdropping device \(k\), \(g nk (t)\) and \(g jk (t)\) represent the instantaneous channel gains of the links from legitimate user \(n\) and jammer \(j\) to eavesdropping device \(k\) in time slot \(t\), respectively.

[0074] The uplink transmission rate from legitimate user \(n\) to base station \(m\) in time slot \(t\) can be expressed as:

[0075]

[0076] Correspondingly, in time slot \(t\), the eavesdropping rate of eavesdropping device \(k\) on channel can be expressed as:

[0077]

[0078] According to the definition of the secrecy transmission rate in PLS, the secrecy transmission rate refers to the data transmission rate that a legitimate user can safely transmit to its target base station, which is the difference between the transmission rate of the legitimate user and the highest eavesdropping rate eavesdropped on its uplink channel. Therefore, in time slot t, the legitimate user 's secrecy transmission rate is:

[0079]

[0080] where, [x] + = max(x, 0).

[0081] Step2: Construct the long-term average utility optimization problem for the three-party devices.

[0082] Considering that performance gains (such as secrecy transmission rate and eavesdropping rate) and potential costs (such as power consumption and reward payment) both play important roles in the utility of the three parties, the construction of the utility function for the three-party devices comprehensively considers these factors. For the eavesdropping device, in order to improve its successful eavesdropping rate (i.e., the difference between the eavesdropping rate and the legitimate user's uplink transmission rate) and reduce its own eavesdropping cost, in each time slot t, it is necessary to determine:

[0083] (1) The activation status of each eavesdropping device at different positions, denoted as or 0, indicating whether the eavesdropping device k is in the active state or the sleep state respectively;

[0084] (2) The unit incentive amount to attract the cooperation of the jammer, denoted as where is the maximum unit incentive amount of the eavesdropping device.

[0085] In time slot t, the actual reward obtained by the jammer from the eavesdropping device is the product of μ E (t) and the performance gain of the eavesdropping device with the help of the jammer. This performance gain can be expressed as the successful interception rate of the eavesdropping device with the help of the jammer, that is:

[0086]

[0087] and the successful interception rate without the help of the jammer, that is:

[0088]

[0089] The difference between them. Therefore, the utility function of the eavesdropping device in time slot t is expressed as:

[0090]

[0091] where x {EJ} (t) = 1 or 0 indicates whether the eavesdropping device and the jammer form an alliance, and c k represents the activation cost of a single eavesdropping device within a time slot. Denote the set of eavesdropping device strategies as Then its optimization problem is:

[0092]

[0093] For legitimate users, in order to improve their long-term secure transmission performance, in each time slot t, they need to determine:

[0094] (1) Uplink transmission power allocation

[0095] (2) Target base station selection where or 0 indicates whether legitimate user n selects base station m as the target receiving base station;

[0096] (3) Unit incentive used to attract the help of the jammer, where is the maximum unit incentive of the legitimate user.

[0097] Similar to the eavesdropping device, the actual benefit obtained by the jammer from the legitimate user in time slot t is the product of μ L (t) and the performance benefit of the legitimate user when receiving the help of the jammer, and this benefit is the secure transmission rate of the legitimate user with the help of the jammer, that is,

[0098]

[0099] and the secure transmission rate without the help of the jammer, that is,

[0100]

[0101] The difference. Therefore, the utility function of the legitimate user in time slot t is expressed as:

[0102]

[0103] where x {LJ} (t) = 1 or 0 indicates whether the eavesdropping device and the jammer form an alliance, and ξ n represents the power consumption cost per unit transmission power of the legitimate user. Denote the set of legitimate user strategies as Then its optimization problem is:

[0104]

[0105] In the formula, Represents the minimum transmission rate of a single legitimate user.

[0106] For the jammer, given μ E (t) and μ L (t), in each time slot t, it decides (1) the interference power allocation (2) the coalition selection x {EJ} (t) and x {LJ} (t). The reward of the jammer comes from the incentives obtained from the eavesdropping devices or legitimate users within each time slot. The utility function of the jammer in time slot t is:

[0107]

[0108] where η j represents the unit power consumption cost of the jammer, and c conf represents the configuration cost incurred by the jammer for replacing allies. Denote the strategy set of the jammer as Its optimization problem is expressed as

[0109]

[0110] Step3: Construct a multi-stage sequential game to model the strategic interactions of the three-party devices, and a dynamic coalition game to transform the problem of solving the coalition selection of the three-party devices into solving a stable coalition partition.

[0111] The expression of the multi-stage sequential game used to model the strategic interactions of the three-party devices is as follows:

[0112]

[0113] where, respectively represent the eavesdropping device, legitimate user, and jammer participating in the game, represents the strategies of the three parties, represents the utility functions of the three parties. As Figure 2 shown, each time slot contains three stages. First, the eavesdropping device makes decisions and μ and μ E (t) according to the optimization objective, second, the legitimate user makes decisions and μ and μ L (t) according to the optimization objective, and finally, the jammer makes decisions and The three stages repeat in each time slice. At the beginning of each time period, the eavesdropping device and the legitimate user can observe the decisions of the jammer in the previous time slice, enabling long-term strategic interaction. For the dynamic alliance of the three-party devices in each time slice, a dynamic coalition game is used for modeling, and its expression is as follows:

[0114]

[0115] where represents the eavesdropping device, legitimate user, and jammer participating in the game, represents all possible coalitions that can be generated among the three-party devices in the dynamic coalition game. is a sub-game of, used to transform the problem of solving the optimal coalition choices x {EJ} (t) and x {LJ} (t) into finding the equilibrium solution for Step4: Define the coalition preferences and coalition switching criteria of each party's device, and use the distributed coalition selection and coalition formation method to obtain the stable coalition partition and optimal coalition choice of the three-party devices within each time slice.

[0116] First, define the equilibrium solution of the dynamic coalition game

[0117] in each time slice, that is, the stable coalition partition: in each time slice t, if no player can improve their utility by unilaterally switching coalitions (i.e., leaving the original coalition and joining another coalition), then the coalition partition

[0118] is stable, that is, it satisfies the condition

[0119]

[0120] Each party's device has different preferences for joining different coalitions. Specifically, in time slice t, player i∈G is more willing to join a possible coalition rather than another coalition The condition can be expressed as:

[0121] if and only if where the symbol represents the preference order of player i for coalitions in time slice t, and represent the utilities of player i after joining coalition C a and C b respectively. Based on the coalition preferences, the coalition switching criteria are defined as follows:

[0122] Criterion 1: if and only if and

[0123]

[0124] Criterion 2: if and only if

[0125] where the binary relation symbol represents the coalition transition of a certain player i in the time slice t, that is, from the coalition on the left side of the symbol to the coalition on the right side of the symbol.

[0126] Based on the coalition preference and coalition switching criterion, a Distributed Coalition Selection and Coalition Formation (DCSCF) method is adopted in the present invention to obtain a stable coalition partition in each time slot and make an optimal coalition selection for the three parties. Specifically, given the coalition partition of the previous time slot, that is In each coalition in each player i first calculates its own utility, and then decides whether to leave the current coalition and join another coalition existing in according to the coalition switching criterion, or stay in the current coalition C a . This process is repeated until the coalition partition remains unchanged, and the final coalition partition is obtained according to The final coalition selection x {LJ} (t) and x {EJ} (t). This method makes coalition selection for each player in a distributed manner, that is, each player dynamically and independently selects its optimal coalition to join according to its own preference order and coalition switching rules.

[0127] Step5: Use an intelligent decision-making algorithm based on deep reinforcement learning to solve the multi-stage sequential game The global equilibrium solution during the entire system operation time to form the optimal strategies of the three devices respectively.

[0128] First, define the equilibrium solution of the multi-stage sequential game during the overall system operation cycle:

[0129] Use to represent the strategies of all parties in the multi-stage sequential game , that is

[0130]

[0131] and use to represent the alliance selection strategy of the three-party devices in . Then the strategy is called 's equilibrium solution if and only if for any player i ∈ G, the inequality

[0132]

[0133] is satisfied, where and represent the best strategies of the other players except player i. Obviously, when such an equilibrium is reached, the long-term utility of each party can be maximized, and no party will unilaterally deviate from this equilibrium.

[0134] Since the observed quantities of the decisions of the devices on each side in PLS, that is, the decisions of the legitimate users, eavesdropping devices, and jammers in each time slot only depend on the decisions in the previous time slot and the resulting system state (for example, the alliance state and channel state in the previous time slot), this means that the state transition satisfies the Markov property. We can use three separate Markov Decision Processes (MDP) to describe the strategy generation problems of the legitimate users, eavesdropping devices, and jammers. For each device, its corresponding MDP is expressed as The detailed explanation is as follows:

[0135] (1) State space For each device i ∈ G in time slot t, its environmental state is where is the current coalition partition, represents the channel gains of all possible links, represents the actions of other devices. Use to represent the state space of player i. The present invention uses the adjacency matrix NT(t) to normalize the state space, such that The adjacency matrix NT(t) is defined as:

[0136]

[0137] (2) Action space For each device i ∈ G in time slot t, its action is its own strategic decision, that is

[0138] (3) Conditional transition probability Ξ i : The probability that player i makes an action from state The probability of transferring to the state s' ∈ is denoted as

[0139] (4) Reward set For the three-party device in time slice t and their real-time reward values are respectively denoted as:

[0140]

[0141] where ψ is the unit rate penalty for violating the minimum uplink rate. Denote the reward set for i ∈ G.

[0142] As Figure 3 shown, the present invention adopts a deep reinforcement learning algorithm based on Proximal Policy Optimization (PPO) and Actor-Critic (AC) framework to solve these three MDPs, that is, to solve the game equilibrium solution, the process is as follows:

[0143] 1) For each party device i ∈ G, its AC framework includes a critic network with network parameters φ, used to estimate the state value of i where the true state value γ t is the discount factor, and an actor network with network parameters θ to approximate the best policy of i Meanwhile, there is an experience replay pool for storing states, actions and rewards during the training process;

[0144] 2) The decisions of each party device are handed over to multiple agents to complete. Specifically, use agents and to generate the decision variables and μ E (t) of the eavesdropping device respectively; use N agents to generate the decisions P^T_n(t) of the legitimate users and use agent to generate the decision μ L (t) of the legitimate users; use agent to generate the decision of the jammer. The set of all agents is denoted as And each agent has its own AC framework;

[0145] 3) In each training step, all agents interact within the time slice t ∈ [1, 2,..., T], following the game The decision-making order of (eavesdropping device - legitimate user - jammer) defined in []. After generating the actions of each agent through the actor network in each time slice t, NT(t) is updated to the state observation value of the next agent. Then, the three parties are required to iteratively apply the proposed DCSCF method to form a stable coalition partition and calculate their rewards Then, the experience replay pool stores the previous state action subsequent state and the reward of each agent e

[0146] 4) The network parameters of each agent must be updated at a certain frequency. When the experience replay pools of all agents reach the rated capacity, the actor networks and critic networks of all agents are updated. For each agent e in [], this update process includes i) calculating the dequitable reward of e, that is, the discounted reward where γ t is the discount factor; ii) calculating the advantage function of e iii) calculating the loss function of the actor network of e where is the probability that the actor network of e selects action at state and θ' is the original parameter of the actor network of e; and iv) calculating the loss function of the critic network of e Then, the parameters θ and φ can be updated by the stochastic gradient descent method (such as AdamOptimizer) to minimize their corresponding loss functions.

[0147] Figure 3 illustrates an overview of the structure of the deep reinforcement learning method proposed by the present invention, including the interactions of all agents (i.e., the actions taken by each agent, the observations from the environment and other agents, and the decision-making sequence between them), the detailed AC framework inside each agent, and the update process of the environment and network parameters (i.e., θ of the actor network and φ of the critic network). Following the game The action generation process of each agent and the update process of the two types of networks are repeated in each training step to achieve long-term performance guarantee.

[0148] In the performance comparison experiment, this embodiment considers an uplink communication system with a range of 1000m×1000m, randomly scattered with N = 20 legitimate users, M = 5 base stations, K = 5 eavesdropping devices, and J = 2 jammers. The l-level power allocation of legitimate users and jammers is set in the range of [0, 20] dbm, l = 10. The frequency bandwidth W is set to W = 1 MHz. The remaining parameters are set as and In addition, the uncertainty of the channel gain is set to three states, namely the "good" state of the "normal" state of the "bad" state of

[0149] Figure 4 The superiority of the proposed three - party intelligent decision - making method for physical - layer secure wireless network transmission based on dynamic coalition game of the present invention is verified. For the sake of comparison, the EV’sFriendly algorithm in which the jammer always chooses to assist the eavesdropping device for eavesdropping and the LU’sFriendly algorithm in which the jammer always chooses to assist the legitimate user to improve the secrecy transmission rate in the existing invention are used as benchmarks for comparative experiments. As can be seen from Figure 4 (a), the proposed method is superior to the LU’sFriendly algorithm in terms of the cumulative utility of the eavesdropping device. In Figure 4 (b), the proposed method is superior to the EV’sFriendly algorithm in terms of the cumulative utility of the legitimate user. In Figure 4 (c), the proposed method is superior to the LU’sFriendly algorithm and the EV’sFriendly algorithm in terms of the jammer's utility. This is because the proposed method allows the jammer to dynamically form coalitions with the legitimate user or the eavesdropping device over time to obtain more rewards, which enables the legitimate user or the eavesdropping device to also exchange for more help from the jammer in terms of secure transmission or eavesdropping, rather than the jammer and the legitimate user or the eavesdropping device only maintaining a fixed relationship in the LU’sFriendly algorithm and the EV’sFriendly algorithm.

Claims

1. A method for intelligent decision-making of three-party devices in physical-layer secure wireless transmission based on dynamic coalition game, Characterized in that: Based on dynamic coalition game, facing the possible dynamic confrontation and alliance behaviors among legitimate users, eavesdropping devices and jammers in an open wireless communication environment, physical quantities including the secrecy transmission rate, eavesdropping rate and energy consumption of each device required by physical-layer security are used to construct the utility functions of the three-party devices respectively. Multi-stage sequential game and dynamic coalition game are used to model the strategic interaction and dynamic alliance behaviors of the three-party devices respectively. Aiming at maximizing the long-term average utility of each of the three-party network devices in the open wireless communication environment, an alliance formation algorithm based on alliance switching criteria and an intelligent decision-making algorithm based on deep reinforcement learning are designed respectively to realize the alliance selection and intelligent decision-making of the three-party devices; The expression of the multi-stage sequential game is as follows: Among them, respectively represent the eavesdropping device, legitimate user, and jammer participating in the game, represent the strategies of the three parties, represent the utility functions of the three parties; The expression of the dynamic coalition game is as follows: Among them denote the eavesdropping device, legitimate user, and jammer participating in the game, represent all possible coalitions that can occur among the three-party devices in the dynamic coalition game.

2. The method for intelligent decision-making of three-party devices in physical-layer secure wireless transmission based on dynamic coalition game according to claim 1, Characterized in that: The method for the network model considering the physical-layer security scenario in the open wireless communication environment includes the following calculation and processing processes: (1) In each time slice, calculate the relevant physical quantities of the three-party device, including the upload rate of the legitimate user and the secure transmission rate the eavesdropping rate of the eavesdropping device (2)Construct the utility functions of the three-party devices in each time slice based on the physical quantities of their respective devices and include the gains and losses of all parties during the operation of the system; (3) Establish the policy sets of the three parties' devices respectively, and generate the optimization problems for maximizing the long-term average utility of each of the three parties' devices. For the eavesdropping device, its policy set is expressed as Its optimization problem is expressed as: In the formula, represents the activation selection of the eavesdropping device in each time slice, and μ E (t) represents the unit excitation amount of the eavesdropping device in each time slice, represents the upper limit of the unit excitation amount. For legitimate users, its strategy set is expressed as Its optimization problem is expressed as: wherein, represents the minimum transmission rate, represents the target base station selection of legitimate users in each time slot, and μ L (t) represents the unit excitation amount of legitimate users in each time slot, represents the power allocation of legitimate users in each time slot. For the jammer, its strategy set is expressed as Its optimization problem is expressed as: s.t.,x {EJ} (t)∈{0,1}, x {LJ} (t) ∈ {0, 1}, x {EJ} (t) + x {LJ} (t) = 1, where \(x\) {EJ} (t) indicates whether the eavesdropping device allies with the jammer, and \(x\) {LJ} (t) indicates whether the legitimate user allies with the jammer, represents the power allocation of the jammer; (4) Construct a multi-stage sequential game to model the strategic interaction of the three-party devices; (5) Design a coalition formation algorithm based on coalition switching criteria to solve the dynamic coalition game in each time slice for the equilibrium solution to achieve the optimal coalition selection of the three-party devices in each time slice, that is, x {EJ} (t) and x {LJ} (t), and at the same time generate a stable coalition partition (6) Design an intelligent decision-making algorithm based on deep reinforcement learning to solve multi-stage sequential games for the global equilibrium solution during the entire system operation time 0 ≤ t ≤ T, and achieve the optimal decision-making of the decision variables of the three-party devices other than the alliance selection. The decision variables include μ E (t), and μ L (t).

3. The method for intelligent decision-making of three-party devices in physical-layer secure wireless transmission based on dynamic coalition game according to claim 2, Characterized in that: In the network model, each legitimate user occupies an orthogonal channel for uplink transmission, and its power allocation adopts l-level discrete allocation, expressed as Meanwhile, the jammer also adopts l-level discrete allocation, expressed as To characterize the time-varying uncertainty, the overall running time of the system is divided into T time slots, and the frequency bandwidth of each orthogonal uplink channel is W.

4. The method for intelligent decision-making of three-party devices in physical-layer secure wireless transmission based on dynamic coalition game according to claim 2, Characterized in that: In step (1), for the upload rate the calculation method is as follows: Among them, represents the additive white Gaussian noise at base station m, and g nm (t) and g jm (t) respectively represent the instantaneous channel gains of the links from legitimate user n and jammer j to base station m; for the calculation method of the eavesdropping rate is as follows: Among them, denotes the AWGN at the eavesdropping device k, and g nk (t) and g jk (t) respectively denote the instantaneous channel gains of the links from the legitimate user n and the jammer j to the eavesdropping device k; for the calculation method of the secrecy transmission rate is as follows: where [x] + = max(x, 0).

5. The method for intelligent decision-making of three-party devices in physical-layer secure wireless transmission based on dynamic coalition game according to claim 2, Characterized in that: The utility function of the eavesdropping device in time slice t in step (2) is expressed as: where c k represents the activation cost of a single eavesdropping device within a time slice, is the performance gain of the eavesdropping device, expressed as: The performance gain of the eavesdropping device without the assistance of a jammer, expressed as: The utility function of a legitimate user in time slice t It is expressed as: Among them, ξ n represents the unit power consumption cost of a legitimate user, is the performance gain of a legitimate user, expressed as: The performance gain of the legitimate user without the help of a jammer is expressed as: The utility function of the jammer in time slice t It is expressed as: Among them, η j represents the unit power consumption cost of legitimate users, and c conf represents the potential configuration cost caused by the additional connections established by the jammer to notify the alliance change if the jammer chooses to change allies in two consecutive time slots. is the amount of incentive paid by legitimate users or eavesdropping devices to the jammer in time slot t, expressed as:

6. The method for intelligent decision-making of three-party devices in physical-layer secure wireless transmission based on dynamic coalition game according to claim 2, Characterized in that: The multi-stage sequential game in step (4) Each time slot contains three stages. First, the eavesdropping device makes a decision according to the optimization objective decision and μ E (t). Second, the legitimate user makes a decision according to the optimization objective decision and μ L (t). Finally, the jammer makes a decision according to the optimization objective decision The three stages will repeat in each time slot. At the beginning of each time period, the eavesdropping device and the legitimate user can observe the jammer's decision in the previous time slot, enabling long-term strategic interaction and dynamic coalition game is a multi-stage sequential game which is a sub-game used to transform the problem of solving the optimal coalition choices x {EJ} (t) and x {LJ} (t) into the problem of solving the equilibrium solution.

7. The method for intelligent decision-making of three-party devices in physical-layer secure wireless transmission based on dynamic coalition game according to claim 2, Characterized in that: In step (5), the coalition formation algorithm for solving the optimal coalition selection of the three parties in each time slice runs in a distributed manner. In the same time slice, each party independently calculates its coalition selection, which essentially solves the equilibrium within each time slice, that is, the stable coalition partition This algorithm is implemented based on the following coalition switching criterion: Criterion 1: if and only if and Criterion 2: if and only if Among them, C a and C b represent two alliances, and the binary relational symbol represents the alliance preference of a certain party i at time slice t, and the binary relational symbol represents that in time slice t, the alliance transfer of a certain party i, that is, from the alliance on the left side of the symbol to the alliance on the right side of the symbol.

8. The method for intelligent decision-making of three-party devices in physical-layer secure wireless transmission based on dynamic coalition game according to claim 2, Characterized in that: In step (6), the algorithm for training the intelligent agent for the decision-making of the proxy three-party device is based on the proximal policy optimization algorithm and the actor-critic framework; The state space of the reinforcement learning process comprehensively considers the network topology, instantaneous channel gain, signal transmission power and alliance state, and normalizes the environmental state value through the adjacency matrix NT(t). In addition, this intelligent decision-making algorithm based on deep reinforcement learning combines distributed training and centralized training. For different decisions of the three-party devices, different intelligent agents are used to train the best strategies; The instantaneous channel gain described above includes g nm (t), g nk (t), g jm (t) and g jk (t); The signal transmission power described includes and The said coalition state is represented by x {EJ} (t) and x {LJ} (t).