An anti-jamming channel selection method for satellite-ground networks based on the Stackelberg game
Through the anti-interference channel selection method of satellite-ground network based on Steinberg game, the frequency band congestion and malicious interference problems in low-orbit satellite communication are solved, channel selection is optimized, and spectrum utilization efficiency and user satisfaction are improved.
Patent Information
- Application Number
- CN202510265668.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-03-07
AI Technical Summary
There are problems in low-orbit satellite communications with frequency band congestion, user interference and external malicious interference, resulting in shortage of spectrum resources and limited information service capabilities.
The anti-interference channel selection method of the satellite-ground network based on Steinberg game is adopted. By establishing system scenarios and models, optimizing problems are constructed, and equilibrium solutions are obtained using a hierarchical anti-interference learning algorithm, and channel selection is optimized to avoid malicious interference and improve spectrum utilization efficiency.
Effectively avoid external malicious interference, improve spectrum utilization efficiency, enhance the information service capabilities of the satellite-ground network, optimize channel selection to reduce resource waste and improve user satisfaction.
Smart Images

Figure CN119788168B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of satellite communication technology, and particularly relates to an anti-jamming channel selection method for satellite-ground network based on Stackelberg game. Background Art
[0002] Low Earth Orbit (LEO) satellites have advantages such as low latency, high data transmission rate, and good coverage ability, and have received extensive attention from all walks of life in recent years. Their deployment speed and scale have been continuously increasing, and they have become an important part of the next-generation mobile communication. However, in the process of satellite-ground communication, two problems will be faced: on the one hand, with the development of ground information technology, the frequency band is crowded, and multiple users accessing a channel interfere with each other, and the conflict between spectrum resource shortage and the increase in the number of users is becoming increasingly obvious; on the other hand, due to the exposure of satellites and wireless open channels, satellite communication signals are easily intercepted and interfered by the enemy, which brings a huge challenge to the security protection of satellite communication systems. Therefore, researching anti-jamming channel selection technology for satellite-ground networks in a complex electromagnetic environment can effectively avoid external malicious interference and improve spectrum utilization efficiency, which is of great significance for enhancing the information service ability of satellite-ground networks. Summary of the Invention
[0003] In view of the above problems, the purpose of the present invention is to provide an anti-jamming channel selection method for satellite-ground network based on Stackelberg game. Under the current background of wireless open channels and spectrum resource shortage, this method can effectively avoid external malicious interference and improve spectrum utilization efficiency, which is of great significance for enhancing the information service ability of satellite-ground networks.
[0004] The technical solution adopted by the present invention is: an anti-jamming channel selection method for satellite-ground network based on Stackelberg game, including:
[0005] Step S01: Establish a system scenario to simulate the communication between users and satellites, including the interference received by users when users select channels to communicate with satellites;
[0006] Step S02: Construct a system model according to the system scenario, obtain the throughput of users based on the transmit antenna gain of satellites, the receive antenna gain of users, and the interference received by users, and conduct a satisfaction evaluation on the throughput to obtain a satisfaction result;
[0007] Step S03: Optimize the channels selected by users according to the satisfaction result, and construct an optimization problem;
[0008] Step S04: Construct a Stackelberg game model according to the optimization problem, and use a hierarchical anti-interference learning algorithm to obtain the equilibrium solution of the Stackelberg game model, thus completing the anti-interference channel selection for the satellite-ground network based on the Stackelberg game.
[0009] Preferably, the interference received by the user in step S01 includes: co-frequency mutual interference between different users within a beam, co-frequency mutual interference between beams, and malicious interference.
[0010] Preferably, the throughput of the user is:
[0011] ;
[0012] ;
[0013] where is the channel bandwidth, is the signal power received by the user, is the co-channel interference from other users in the same beam of the LEO satellite in the downlink, is the co-channel interference from other beams of the LEO satellite in the downlink, is the malicious interference from the jammer, is the link 's channel noise, is the transmit power of the antenna beam, is the maximum gain of the satellite transmit antenna, is the maximum gain of the user receive antenna, is the wavelength of the communication transmission, is the link 's distance.
[0014] Preferably, the transmit antenna gain of the satellite is:
[0015] ,
[0016] ,
[0017] The receive antenna gain of the user is:
[0018] , ;
[0019] where is the off-axis angle of the satellite antenna, is the antenna half-power beam width, is the level at which the gain at the intersection of the antenna main beam and the near-field sidelobe mask is lower than the peak gain, is the far-field sidelobe level, is the off-axis angle of the user antenna, is the maximum gain of the satellite launch antenna, is the maximum gain of the user receiving antenna, is the wavelength of the communication transmission, is the circular equivalent diameter of the antenna.
[0020] Preferably, ;
[0021] ;
[0022] ;
[0023] ;
[0024] Among them, 、 is the transmission power of the satellite beam, is the transmission power of the jammer, 、 are the satellite antenna interference gains in the direction of the user , is the user antenna reception gain in the direction of the jammer, 、 are the distances of the interference link, is the maximum gain of the user receiving antenna, is the wavelength of the communication transmission, 、 are for the user k 、 selected channels, is a function for judging whether the jammer and the user select the same channel.
[0025] Preferably, the method for obtaining the satisfaction result by evaluating the satisfaction of the throughput includes: calculating the average estimated value, using the average estimated value to describe the subjective QoE feeling, and obtaining the satisfaction result;
[0026] The calculation formula of the average estimated value MOS is:
[0027] MOS = c 1 + exp [ d ( R − h ) ] ;
[0028] The satisfaction of the user is expressed as:
[0029] ;
[0030] Among them, is the throughput of the user, 、 and are constants, is the satisfaction of the user, is for the user The channel selection strategy is the set of channel selection strategies for other users except user . is the channel selection of the jammer, is the score value when the throughput is .
[0031] Preferably, the Stackelberg game model is expressed as:
[0032] ;
[0033] Wherein, is the set of terrestrial users, is the jammer, and are the strategy sets of the user and the jammer respectively, and are the effect functions of the user and the jammer respectively.
[0034] Preferably, finding the best strategy is included in the Stackelberg game model, which is expressed as maximizing the leader's effect function:
[0035] ;
[0036] ;
[0037] , ;
[0038] Wherein, is the channel selected by user , is the set of user channel selection strategies, is the channel selection strategy of the jammer, is the satisfaction threshold for meeting basic communication, is the satisfaction when the throughput is , is the function for judging whether the jammer and the user select the same channel, is the function for judging whether the attack channel destroys data transmission.
[0039] Preferably, finding the best strategy is included in the Stackelberg game model, which is expressed as maximizing the follower's effect function:
[0040] ;
[0041] ;
[0042] ;
[0043] wherein, is the channel selection strategy of the user, is the set of channel selection strategies of other users except the user, is the channel selection of the jammer, is the satisfaction degree of the user, is the neighbor of the user, is the channel selection sequence of the user adjacent to, is the channel selection sequence of the user adjacent to, is other users except the user, is the distance between the users and, is the protection radius of the user. is the channel selection sequence of the user adjacent to, is the channel selection sequence of the user adjacent to, k is the channel selection sequence of the user adjacent to, is other users except the user, is the distance between the users and, is the protection radius of the user. and k is the distance between the users and, is the protection radius of the user.
[0044] Preferably, the equilibrium solution is:
[0045] ;
[0046] ;
[0047] wherein, is the channel allocation strategy that maximizes the user benefit, is the channel allocation strategy that maximizes the jammer benefit, is the effect function of the user, is the effect function of the jammer.
[0048] Beneficial effects of the above technical solution:
[0049] The present invention solves the problem of anti - interference in channel selection for multiple users in the satellite communication downlink, considering that users are simultaneously affected by malicious interference and co - channel interference. At the same time, in order to avoid resource waste caused by blindly increasing throughput, a quality of experience (QoE) model based on the mean opinion score (MOS) value centered on user needs is introduced to improve system performance. And a Stackelberg game model is established to describe the confrontation relationship between satellite users and jammers. Meanwhile, aiming at the characteristics of interference among users, a local altruistic game model is established to describe the cooperative - competitive relationship among users. The proposed game model is proved to have an equilibrium solution. To obtain the equilibrium solution of the game model, an MSAP hierarchical anti - interference learning algorithm is proposed. The upper - layer interference uses the multi - armed bandit algorithm to update its own actions, and the lower - layer users use the space - adaptive algorithm to obtain the optimal strategy. Finally, the simulation results verify the convergence and effectiveness of the proposed anti - interference scheme. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Figure 1 It is a flowchart of the satellite - to - ground network anti - interference channel selection method based on the Stackelberg game provided by the present invention;
[0051] Figure 2 It is a system scenario diagram provided by an embodiment of the present invention;
[0052] Figure 3 It is a schematic diagram of the strategy update iteration of the leader and the follower provided by an embodiment of the present invention;
[0053] Figure 4 It is a schematic diagram of the random distribution of users provided by an embodiment of the present invention;
[0054] Figure 5 It is a curve graph of the change of network satisfaction rate under different numbers of channels provided by an embodiment of the present invention;
[0055] Figure 6 It is a curve graph of the change of the expected regret value of the jammer under different numbers of channels provided by an embodiment of the present invention;
[0056] Figure 7 It is a comparison graph of the network satisfaction rates of different algorithms provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0057] The following further describes the embodiments of the present application in detail. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than an exhaustive list of all embodiments. It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other.
[0058] The terms "first", "second", etc. (if any) in the description and claims are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or are inherent to these processes, methods, products, or devices.
[0059] It should be understood that the term "and / or" used herein is merely a description of the relationship between related objects, and there can be three relationships. For example, A and / or B can be: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " herein generally indicates an "or" relationship between the related objects before and after.
[0060] Embodiment 1
[0061] As Figure 1 The process of the method for selecting an anti-interference channel in a satellite-ground network based on the Steinberg game of the present invention aims to alleviate the problems of low existing resource utilization rate and interference from multiple aspects. The simulation results show the convergence and effectiveness of the method provided by the present invention, and can make the network have a better satisfaction rate.
[0062] 1. Establishment of system scenario
[0063] The present invention considers the problem of channel selection for multiple users under a single beam of a low-earth-orbit multi-beam satellite in the downlink. Each user selects a channel to communicate with the satellite within a time slot. At the same time, there is a UAV carrying a jammer J within the satellite beam. It is equipped with a wide-beam antenna and operates in a time-division multiple access mode. The jammer selects a channel to transmit interference attack signals in each time slot. Therefore, the user will be interfered by three aspects: co-channel interference between different users within the beam, co-channel interference between beams, and malicious interference, as Figure 2 shown. Both the user and the jammer have the functions of environmental perception adaptation and strategy optimization update. For the communication between the satellite and the user, in order to ensure reliable transmission, it is necessary to optimize channel selection to minimize malicious interference and co-channel interference; for the jammer, it is necessary to dynamically adjust the attacked channel to maximize the interference and damage effect.
[0064] The set of satellite beams is ;
[0065] The set of users is ;
[0066] The set of available channels is ;
[0067] Among them, the optional channel sets of the jammer and the user are the same, is the th beam, is the th user, is the to downlink of, is the to co-channel interference link of, is the angle between the link and , is the angle between the link and the jammer .
[0068] 2. Establishment of the system model
[0069] 1) Propagation model:
[0070] In the present invention, only free space loss is considered, and its expression is:
[0071] (1)
[0072] Among them, is the wavelength of the communication transmission, is the distance of the communication transmission.
[0073] For LEO satellites, according to ITU-R S.1528, the transmitting antenna gain of LEO satellites is:
[0074] (2)
[0075] (3)
[0076] Among them, is the maximum gain of the satellite antenna, is the off-axis angle of the satellite antenna, is the half-power beam width of the antenna, is the level at which the gain of the intersection point of the main beam of the antenna and the near-field sidelobe mask is lower than the peak gain, is the far-field sidelobe level. Preferably, , .
[0077] For users, according to ITU-R S.465, the receiving antenna gain of ground users is:
[0078] (4)
[0079] Among them, is the maximum gain of the user antenna, is the off-axis angle of the user antenna.
[0080] (5)
[0081] Among them, is the circular equivalent diameter of the antenna, is the wavelength of the communication transmission.
[0082] 2) Transmission rate model: As Figure 2 , assuming that the user accesses the channel , and the jammer accesses the channel , then there are three potential interferences:
[0083] i) : Co-channel interference from other users in the same beam of the LEO satellite in the downlink;
[0084] (6)
[0085] ii) : Co-channel interference from other beams of the LEO satellite in the downlink;
[0086] (7)
[0087] iii) : Malicious interference from the jammer;
[0088] (8)
[0089] Among them, , are the transmission powers of the satellite beams, is the transmission power of the jammer, , are the interference gains of the satellite antenna in the direction of the user , is the receiving gain of the user antenna in the direction of the jammer, , are the distances of the interference links.
[0090] is a binary indicator function used to determine whether the two select the same channel at the current moment, and its expression is:
[0091] (9)
[0092] Based on the above propagation model, the transmission rate (i.e., throughput) of the user is obtained as follows: :
[0093] (10)
[0094] (11)
[0095] Wherein, is the channel bandwidth, is the channel noise of the link , is the signal power received by the user, is the transmission power of the antenna beam, is the link distance.
[0096] 3) User satisfaction evaluation index
[0097] In many current studies on wireless network channel resource allocation, the main goal is usually to find a channel allocation strategy that can maximize the throughput of the entire network. However, this kind of optimization cannot directly and accurately reflect the actual satisfaction of users with their services. From the user's perspective, if the increase in throughput is not sufficient to significantly improve the user experience, then this increase is actually meaningless. Therefore, if the actual needs of users can be considered in the resource allocation process, the resource utilization rate can be more effectively improved. The present invention uses an optimization idea oriented to user needs, changing the user utility from Quality of Service (QoS) to the subjective QoE level, that is, the average evaluation score MOS. This design will effectively alleviate user competition and optimize resource utilization, that is, when the satisfaction level of a certain user no longer improves, the resources will be allocated to other users.
[0098] Using the average evaluation score MOS to characterize the user's QoE, the present invention adopts the user satisfaction evaluation method proposed by the International Telecommunication Union (ITU), dividing the subjective QoE feeling into six levels: "poor", "relatively poor", "average", "relatively good", "good" and "very good", and using an ordinal scale method from low to high to represent the user's subjective feeling. The larger the MOS value, the higher the user satisfaction; on the contrary, the smaller the MOS value, the lower the user satisfaction. It can be seen that this is a process where quantitative change leads to qualitative change. What users intuitively feel is the qualitative change. When the quantitative change is not sufficient to cause qualitative change, this kind of improvement has little meaning in the eyes of users. In addition, when the throughput of users can already fully meet their needs, pursuing higher throughput will only cause waste.
[0099] The present invention defines the MOS function as:
[0100] MOS = c 1 + exp [ d ( R − h ) ] (12)
[0101] Among them, is the throughput of the user, , and are constants. The specific value range and meaning of MOS are shown in Table 1.
[0102] Table 1 Specific value range and meaning of MOS
[0103] MOS value QoE Satisfaction degree Score S(R)) MOS < 2.6 Poor Dissatisfied 1 2.6 ≤ MOS < 3.1 Relatively poor Relatively dissatisfied 2 3.1 ≤ MOS < 3.6 Average Average 3 3.6 ≤ MOS < 4 Relatively good Relatively satisfied 4 4 ≤ MOS < 4.3 Good Satisfied 5 MOS≥4.3 Very good Very satisfied 6
[0104] 3. Construction of the optimization problem
[0105] Use the function to represent the satisfaction degree of user ,
[0106] (13)
[0107] Among them, represents the score value when the transmission rate is , given according to Table 1 and Equation (12). The total satisfaction degree of all users can be expressed as:
[0108] (14)
[0109] Based on the above analysis, the goal of the user is to dynamically and adaptively adjust the channel selection by perceiving the environment to maximize the satisfaction degree of all users, that is:
[0110] (15)
[0111] To ensure the reliability of data transmission between the satellite and the user, that is or more is required to meet the basic communication. Denote this threshold as . For the user, if , it indicates that the data transmission is successful; otherwise, the data transmission fails.
[0112] Based on this, the effect value C of the jammer is defined as the number of users whose data transmission is damaged by attacking channel , that is:
[0113] (16)
[0114] Among them, is a binary indicator function, and its expression is: .
[0115] Based on the above analysis, the goal of the jammer is to dynamically and adaptively adjust the attack channel by perceiving the environment to maximize the attack effectiveness, that is:
[0116] (17)
[0117] 4. Channel Selection Game
[0118] First, the optimization problem is modeled using a game model, and then the equilibrium solution of the game is given.
[0119] 4.1 Constructing the Stackelberg Game Model
[0120] The Stackelberg game, also known as a hierarchical game, includes two types of participants with different attributes, namely the leader and the follower, and it can well describe the adversarial relationship. The adversarial relationship between the jammer and the ground user is a hierarchical Stackelberg game, denoted as , where is the set of ground users, is the jammer, and are the strategy sets of the user and the jammer respectively, and are the effect functions of the user and the jammer respectively. Taking the jammer as the leader and the ground user as the follower, the jammer first performs environmental perception and strategy update to maximize the corresponding effect function. The ground user performs perception and adjustment after the jammer completes the strategy update to optimize the corresponding effect function.
[0121] 4.1.1 Leader Sub - game
[0122] For the jammer, the goal is to disrupt the data transmission between the satellite and the user as much as possible. Therefore, its corresponding effect function is defined as:
[0123] (18)
[0124] Therefore, the leader sub - game can be mathematically represented as:
[0125] (19)
[0126] Then the leader sub - game problem can be described as finding the optimal strategy to maximize its effect function, that is:
[0127] (20)
[0128] 4.1.2 Follower Sub - game
[0129] In traditional game models, users are "selfish" and the utility function only considers the benefits obtained by themselves. The consequence of this "selfishness" is that the global optimum cannot be achieved because the best benefit of a user is affected by its co-channel users. Therefore, to achieve the global optimum, a local cooperation model is introduced, and each user considers the benefits of itself and adjacent users affected by the co-channel to seek greater benefits. A user may have potential co-channel interference with other users within the radius . The users with potential interference are regarded as the neighbors of user , denoted as:
[0130] (21)
[0131] where is the set of other users excluding user , is the distance between user and k , and is the protection radius of the user.
[0132] Therefore, for a user, its corresponding effect function is defined as:
[0133] (22)
[0134] where is the channel selection sequence of other users except , is the channel selection sequence of the users adjacent to , and is the channel selection sequence of the users adjacent to k .
[0135] Therefore, the follower sub-game can be mathematized as:
[0136] (23)
[0137] Then the follower sub-game can be described as finding the best strategy to maximize its effect function, that is:
[0138] (24)
[0139] 4.2 Game Equilibrium Analysis
[0140] Definition 1 (Nash Equilibrium NE): Given a non-cooperative game , the strategy set is a Nash equilibrium if and only if when , satisfying the condition:
[0141] (25)
[0142] That is, no participant can increase the value of the effect function through a unilateral strategic change. Then, this situation is called a pure-strategy Nash equilibrium.
[0143] Definition 2 (Potential Game): Given a finite-strategy game , for any game participant and two different strategies , , if this game is an exact potential game, it satisfies the condition:
[0144] (26)
[0145] where the function is the potential function of this game.
[0146] In an exact potential game, the change in the utility function caused by any user unilaterally changing the strategy is the same as the change in the potential function.
[0147] Potential games have two important properties:
[0148] (1) Any exact potential game has at least one pure-strategy Nash equilibrium.
[0149] (2) The global or local optimal solution of the potential function is a Nash equilibrium.
[0150] Theorem 1: When the interference strategy of the jammer is given, the follower sub-game is an exact potential game and there exists at least one pure-strategy Nash equilibrium.
[0151] Proof: First, construct the potential function as follows:
[0152] (27)
[0153] When user unilaterally changes the channel strategy from to , the change in the effect function is:
[0154] u i ( a i ′ , α − i , c J ) − u i ( a i , α − i , c J ) = q i ( a i ′ , α I i , c J ) − q i ( a i , α I i , c J ) + ∑ k ∈ I i [ q k ( a k , α I k ′ , c J ) − q k ( a k , α I k , c J ) ] (28)
[0155] At this time, the change in the potential function is:
[0156] (29)
[0157] Since the user only considers the utility of neighboring users, we have:
[0158] (30)
[0159] Therefore, the change in the potential function is further written as:
[0160] Φ ( a i ′ , α − i , c J ) − Φ ( a i , α − i , c J ) = q i ( a i ′ , α I i , c J ) − q i ( a i , α I i , c J ) + ∑ k ∈ I i [ q k ( a k , α I k ′ , c J ) − q k ( a k , α I k , c J ) ] (31)
[0161] In summary, (32)
[0162] It can be seen that the follower sub - game is an exact potential game, that is, the change in the user's utility function caused by a unilateral change in the user's strategy is equal to the change in the potential function. According to the properties of the exact potential game, there exists a pure - strategy Nash equilibrium, that is, the maximum value of the potential function is a pure - strategy Nash equilibrium of the follower sub - game.
[0163] The constructed Stackelberg game has an equilibrium solution, and the proof process is as follows:
[0164] Definition 3 (Stackelberg equilibrium SE):
[0165] If the strategy profile satisfies: (33)
[0166] Then this strategy profile is the Stackelberg equilibrium solution of the proposed game, that is, no participant can improve the overall utility by unilaterally changing the strategy. is the channel allocation strategy that maximizes the user's benefit, is the channel allocation strategy that maximizes the jammer's benefit.
[0167] Theorem 2: The Stackelberg game constructed by the present invention has an SE solution composed of the jammer's stationary strategy and the user's Nash equilibrium strategy.
[0168] Proof: Given a jamming strategy , the Stackelberg game becomes a local cooperation game. Based on Theorem 1, the follower sub - game is an exact potential game, and there exists at least one pure - strategy NE solution . Similarly, when a user strategy is given, there exists at least one NE solution in the leader sub - game. The stable strategies of the leader game and the follower game constitute an SE, that is:
[0169] (34)
[0170] Since every finite-strategy game has a mixed-strategy equilibrium, due to the limited number of users, the available channel schemes are also very limited, and the interference strategies are also restricted. Therefore constitutes the SE in the stationary sense of this problem.
[0171] 5. Equilibrium Solution of the Stackelberg Game Model
[0172] The equilibrium solution of the Stackelberg game model is obtained by using the hierarchical anti-interference learning algorithm.
[0173] 5.1 Algorithm Description
[0174] To obtain the Stackelberg game equilibrium solution of the constructed hierarchical game, based on the multi-armed bandit (MAB) learning mechanism and the spatial adaptive algorithm (SAP), an MSAP hierarchical anti-interference learning algorithm is proposed. In the proposed algorithm, the jammer and the satellite users update their strategy selections at different time scales. The leader updates the channel selection once per period and the follower updates the channel selection once per time slot . Each period contains time slots , that is . The strategy update iterations of the leader and the follower during the execution of the proposed hierarchical channel selection learning algorithm are as shown in Figure 3 .
[0175] In the lower-level subgame, as a learning algorithm, SAP can converge to the pure-strategy Nash equilibrium that maximizes the potential function with an arbitrarily high probability. To characterize SAP, the game is extended to the mixed-strategy form. Let the mixed strategy of user in the k-th iteration be represented by the probability distribution , where is the set of probability distributions on the action set . In SAP, a user is randomly selected according to the mixed strategy to update its channel selection, while the other users remain unchanged. This process is repeated until the stopping rule is satisfied.
[0176] In the upper-level subgame, the channel state and user information are unknown. To obtain the equilibrium solution, a channel selection algorithm is proposed based on MAB. Each channel is regarded as an arm, and the jammer balances exploration and exploitation by selecting the channel with the highest UCB (Upper Confidence Bound) value through interaction with the environment. The UCB algorithm can effectively approximate the best channel while continuously learning the channel performance.
[0177] In the present invention the cumulative reward within
[0178] (35)
[0179] Based on this, the expected cumulative reward is:
[0180] (36)
[0181] Let represent the channel selected in the previous time instants, which is denoted as:
[0182] (37)
[0183] where represents the indicator function, defined as:
[0184] (38)
[0185] In the MAB problem, the policy does not always perform optimally, and the regret value is an important metric that determines the performance of the MAB algorithm.
[0186] The regret value at time instant
[0187] E [ Β ( t ) ] = E [ r * ( t ) ] − E [ r ( t ) ] = ∑ j ≠ j ∗ ( u J j ∗ − u J j ) E [ χ j ( t ) ] (39)
[0188] ;
[0189] where is the upper bound of , , is the number of times channel is selected in the previous time slots. In this problem, the goal is to maximize the long-term reward, that is, to minimize the regret value function.
[0190] The detailed steps of the policy update iteration of the leader and the follower during the execution of the proposed hierarchical anti-jamming learning algorithm are shown in Algorithm 1. First, the jammer selects the channel to attack in the current period; then, the satellite user learns the environment until its sub-game converges to the Nash equilibrium solution; then, the jammer learns the environment and selects the optimal channel; repeat this process until the maximum number of iterations is reached or the channel selection strategies in adjacent time slots are the same. After multiple iterations, both the leader sub-game and the follower sub-game converge to the Nash equilibrium solution, which is the Stackelberg equilibrium solution of the hierarchical game.
[0191] Algorithm 1: MSAP Hierarchical Anti-Jamming Learning Algorithm
[0192] Step1: Set , , , the user randomly generates a channel access strategy.
[0193] Step2: In the th cycle, the interference selects a channel according to Equation (41) .
[0194] Step3: The learning process of the user is as follows, with the loop k = 0, 1, 2, ….
[0195] ① In the th time slot, all users exchange information with their neighbors;
[0196] ② Randomly select a user ;
[0197] ③ Calculate the rewards of all actions of this user according to the information of neighboring users , . The channel selection strategies of other users remain unchanged, that is ;
[0198] ④ Determine the action selected this time according to the Boltzmann update formula (40). The user randomly selects an action according to ;
[0199] (40)
[0200] where is the learning parameter, and . The effect function is calculated according to Equation (22).
[0201] The loop ends
[0202] Step4: The interference obtains the utility according to the user strategy , where .
[0203] Step5: The interference selects the channel with the highest UCB value and updates the channel selection times and the total channel reward.
[0204] (41)
[0205] where , is the total reward obtained by selecting channel at time , is the number of times channel is selected.
[0206] Step6: Update , execute Step2 until the maximum number of loops is reached.
[0207] 5.2 Convergence Analysis
[0208] Analyze the convergence of the SAP algorithm adopted by the follower subgame. The set of all user channel selection combinations is .
[0209] Theorem 3: After a given interference strategy , when the learning parameter is large enough, the SAP learning algorithm can converge to the global optimal solution of the follower subgame with a probability close to 1 .
[0210] Proof: The channel selection combination of all users in time slots is represented as , where is the time slot user selected channel. Since the algorithm SAP only selects one user to randomly probe a certain channel each time, is a non-periodic and irreducible discrete Markov chain, so has a unique steady-state distribution.
[0211] Denote the channel combinations of users in two consecutive iterations as a and b respectively, then the transition probability between them is P a b = P [ a ( k + 1 ) = b | a ( k ) = a ] .
[0212] The Markov balance equation is:
[0213] (42)
[0214] The distribution that satisfies the Markov balance equation is the unique steady-state distribution. Since in the proposed algorithm, only one user is selected to update its selection, therefore, between any two consecutive iterations, at most one element in the network state can be changed, and it can be verified that satisfies the Markov process balance condition, where is the potential function. Therefore is the unique equilibrium stable distribution.
[0215] is the unique action that maximizes the potential function,
[0216] (43)
[0217] That is , when tends to infinity, the following inequality holds:
[0218] (44)
[0219] Then, based on equations (43) and (44), we can obtain:
[0220] (45)
[0221] The lower-layer SAP algorithm will converge to the maximum point of the potential function of the follower sub-game with a probability approaching 1, that is, the global optimal solution.
[0222] The convergence proof of the MAB algorithm adopted by the leader sub-game is as follows:
[0223] Theorem 4: The proposed regret value function grows at a logarithmic rate, that is E [ Β ( t ) ] ∼ O ( l o g t ) , as time goes by, the selected strategy gets closer and closer to the optimal strategy, and the upper bound of the overall regret value function can be expressed as follows:
[0224] E [ Β ( t ) ] ≤ 8 ∑ j : u J j < u J * l n ( t ) Δ j + ( 1 + π 2 3 ) ( ∑ j = 1 M Δ j ) (46)
[0225] Proof:
[0226] E [ χ j ( t ) ] ≤ 8 l n ( t ) Δ j 2 + 1 + π 2 3 (47)
[0227] Among them, .
[0228] Furthermore, the upper bound of the expected regret value function is obtained as follows:
[0229] E [ Β ( t ) ] = ∑ j ≠ j ∗ ( u J j ∗ − u J j ) E [ χ j ( t ) ] ≤ ∑ j ≠ j ∗ ( u J j ∗ − u J j ) ( 8 l n ( t ) Δ j 2 + 1 + π 2 3 ) = ∑ j ≠ j ∗ Δ j ( 8 l n ( t ) Δ j 2 + 1 + π 2 3 ) = ∑ j : u J j < u J * Δ j ( 8 l n ( t ) Δ j 2 + 1 + π 2 3 ) = 8 ∑ j : u J j < u J * l n ( t ) Δ j + ( 1 + π 2 3 ) ( ∑ j = 1 M Δ j ) (48)
[0230] In summary, we can finally obtain E [ Β ( t ) ] ≤ 8 ∑ j : u J j < u J * l n ( t ) Δ j + ( 1 + π 2 3 ) ( ∑ j = 1 M Δ j ) , and the regret value function converges.
[0231] Simulation results of the anti-interference channel selection method for satellite-ground networks based on the Stackelberg game
[0232] The performance of the hierarchical anti-interference learning algorithm adopted in the present invention is simulated and analyzed. Users are randomly distributed within the beam, and the schematic diagram of the random distribution of users is as Figure 4 shown, and the learning parameter increases with the increase of the number of iterations. The specific parameter settings are shown in Table 2:
[0233] Table 2
[0234]
[0235] The present invention uses the network-wide satisfaction rate to measure the overall QoE of the network, which is defined as follows:
[0236] (49)
[0237] where is the QoE value of user , and N is the number of users in the whole network.
[0238] Figure 5 The curves of the network-wide satisfaction rate when the number of channels is 3, 4, and 5 are given respectively. It can be seen that the network-wide satisfaction rate finally converges to the optimal solution as the number of iterations increases. At the same time, the network-wide satisfaction increases with the increase in the number of available channels, which indicates that sufficient channel resources can make the QoE value of users higher, reduce the probability that users and jammers or users and neighboring users use the same channel, bring a better user experience, and improve the overall network satisfaction.
[0239] Figure 6 The curves of the expected regret value of the jammer when the number of channels is 3, 4, and 5 are given respectively. It can be seen that the expected regret value of the jammer increases with the increase in the number of channels, which indicates that the increase in the number of channels makes the user experience better, which reduces the probability that the jammer successfully jams the user, so the regret value increases, resulting in an increase in the expected regret value.
[0240] Figure 7 A comparison chart of the network satisfaction rate under the algorithm proposed in the present invention, the optimal NE, and the random selection algorithm is given. It can be seen that the performance of the proposed hierarchical learning algorithm is better than that of the random selection algorithm and is close to the optimal NE, verifying the effectiveness of the algorithm proposed in the present invention, and finally a higher network satisfaction rate can be obtained.
[0241] The present invention considers the channel selection optimization problem in the satellite-ground network under the malicious interference scenario, and proposes an anti-jamming channel selection mechanism based on the Stackelberg game. First, the confrontation relationship between the user and the jammer is modeled as a Stackelberg game, where the jammer is the leader and its goal is to maximize the attack and destruction efficiency, and the user is the follower and its goal is to maximize the network-wide satisfaction rate. Secondly, considering the influence of neighboring users, the follower sub-game is expressed as a local altruistic game, and it is proved that it is an exact potential game, and the QoE evaluation index is introduced to enhance the user experience. For the constructed game model, the existence of the game equilibrium solution is analyzed. Finally, to obtain the game equilibrium solution, an MSAP hierarchical anti-jamming learning algorithm is proposed. The simulation results show the convergence and effectiveness of the proposed algorithm, and it has good optimization performance in anti-jamming channel selection.
[0242] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, rather than limitations on the implementation manners of the present invention. For those of ordinary skill in the art, other different forms of changes or variations can be made based on the above description. It is impossible to enumerate all the implementation manners here. Any obvious changes or variations derived from the technical solutions of the present invention still fall within the protection scope of the present invention.
Claims
1. A method for anti-interference channel selection in satellite-ground network based on Stackelberg game, characterized in that Including: Step S01: Establish a system scenario to simulate the communication between a user and a satellite, including the interference suffered by the user when the user selects a channel to communicate with the satellite; Step S02: Construct a system model according to the system scenario, obtain the throughput of the user based on the transmit antenna gain of the satellite, the receive antenna gain of the user, and the interference suffered by the user, and perform a satisfaction evaluation on the throughput to obtain a satisfaction result; Step S03: Optimize the channel selected by the user according to the satisfaction result, and construct an optimization problem; Step S04: Construct a Stackelberg game model according to the optimization problem, and use a hierarchical anti-interference learning algorithm to obtain an equilibrium solution for the Stackelberg game model, completing the anti-interference channel selection for the satellite-ground network based on the Stackelberg game.
2. The anti-interference channel selection method for satellite-ground network based on the Steinberg game according to claim 1, characterized in that The interference suffered by the user in Step S01 includes: co-channel interference between different users within a beam, co-channel interference between beams, and malicious interference.
3. The anti-interference channel selection method for satellite-ground network based on the Steinberg game according to claim 1, characterized in that The throughput of the user is: where W is the channel bandwidth, is the signal power received by the user, I intra is the co-channel interference from other users in the same beam of the LEO satellite in the downlink, I inter is the co-channel interference from other beams of the LEO satellite in the downlink, I J is the malicious interference from the jammer, is the link channel noise, p n is the transmit power of the antenna beam, G smax is the maximum gain of the satellite transmit antenna, G umax is the maximum gain of the user receive antenna, λ is the wavelength of the communication transmission, is the link distance.
4. The method for selecting an anti-interference channel in a satellite-ground network based on the Steinberg game according to claim 3, wherein The transmit antenna gain of the satellite is: The receive antenna gain of the user is: Among them, θ1 is the off-axis angle of the satellite antenna, θ b is the half-power beam width of the antenna, Ls is the level at which the gain of the intersection point of the main beam of the antenna and the near-field sidelobe mask is lower than the peak gain, L F is the far-field sidelobe level, θ2 is the off-axis angle of the user antenna, G smax is the maximum gain of the satellite transmitting antenna, G umax is the maximum gain of the user receiving antenna, λ is the wavelength of the communication transmission, and D is the circular equivalent diameter of the antenna.
5. The anti-interference channel selection method for a satellite-ground network based on the Stackelberg game according to claim 3, characterized in that: where p n and p m are the transmission powers of the satellite beams, p J is the transmission power of the jammer, is the satellite antenna interference gain in the direction of user i, is the user antenna reception gain in the direction of the jammer, d Ji is the distance of the interference link, G umax is the maximum gain of the user receiving antenna, λ is the wavelength of the communication transmission, c J is the channel selection strategy of the jammer, a k is the channel selection strategy of user k, a i is the channel selection strategy of user i, and δ(x,y) is a function for judging whether the jammer and the user select the same channel.
6. The method for anti-interference channel selection in satellite-ground network based on the Steinberg game according to claim 1, characterized in that The method for performing a satisfaction evaluation on the throughput to obtain a satisfaction result includes: calculating an average estimated value, using the average estimated value to describe the subjective QoE feeling, and obtaining a satisfaction result; The calculation formula for the average estimated value MOS is: The satisfaction of the user is expressed as: q i (a i ,α -i ,c J )=S(R i ); Among them, R is the throughput of the user, c, d, and h are constants, and q i is the satisfaction of the user, and a i is the channel selection strategy of user i, and α -i is the set of channel selection strategies of other users except user i, and c J is the channel selection strategy of the jammer, and S(R i ) is the score value when the throughput is R i .
7. The anti-interference channel selection method for satellite-ground network based on the Steinberg game according to claim 1, characterized in that The Stackelberg game model is expressed as: G = {U, J, F U , F J , u i , u J}; Among them, U is the set of ground users, J is the jammer, and F U and F J are the strategy sets of the user and the jammer respectively, and u i and u J are the effect functions of the user and the jammer respectively.
8. The method for selecting an anti-interference channel in a space-ground network based on the Steinberg game according to claim 7, wherein The Steinberg game model includes finding the optimal interference strategy F J * , expressed as maximizing the leader effect function: Among them, a i is the channel selection strategy of user i, α is the set of user channel selection strategies, c J is the channel selection strategy of the jammer, S th is the satisfaction threshold for meeting basic communication, S(R i ) is the satisfaction when the throughput is R i , δ(x, y) is a function for judging whether the jammer and the user select the same channel, and l(x, y) is a function for judging whether the channel selection strategy of the jammer disrupts data transmission.
9. The method for anti-interference channel selection in a satellite-ground network based on the Steinberg game according to claim 7, wherein The Steinberg game model includes finding the optimal user strategy F U * , expressed as maximizing the follower effect function: I i = {k ∈ U \ i : d(i, k) ≤ R th}; where a i is the channel selection strategy of user i, and α -i is the set of channel selection strategies of other users except user i, c J is the channel selection strategy of the jammer, is the satisfaction degree of user i, I i is the neighbor of user i, is the channel selection strategy of the user adjacent to i, is the channel selection strategy of the user adjacent to k, U\i is the other users except user i, d(i,k) is the distance between user i and k, and R th is the protection radius of the user.
10. The method for anti-interference channel selection in satellite-ground network based on Stackelberg game according to claim 1, wherein The equilibrium solution is: (c J * ,α * (c J * )) Among them, α * is the channel allocation strategy that maximizes the user benefit, is the channel allocation strategy that maximizes the jammer benefit, u i (α, c J ) is the effect function of the user, u J (α(c J ), c J ) is the effect function of the jammer.
Citation Information
Patent Citations
Fuzzy learning anti-interference method and system for dealing with incomplete channel information
CN117528538A
Satellite anti-interference method based on hierarchical Steinberg game and matching game
CN117768010A