Multi-jammer cooperative beam decision-making method for large-scale communication network

Through the multi-jammer collaborative beam decision-making method and evaluation model based on transmission time variation, the interference beam strategy in large-scale communication networks is optimized, which solves the problem of difficulty in obtaining the opponent's prior information in wireless interference technology, and achieves effective interference to the enemy network and significantly reduces the communication quality.

CN120201451APending Publication Date: 2025-06-24ARMY ENG UNIV OF PLA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510359274.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

In large-scale communication networks, existing wireless interference technology is difficult to effectively carry out because it is difficult to obtain prior information of the other party's communication network, which makes it difficult to achieve the ideal effect of the convex optimization method.

Method used

A multi-jammer collaborative beam decision-making method is adopted to sense the transmission time of users in the network in real time, and an interference performance evaluation model based on changes in transmission time is designed. Under the framework of multi-arm slot machine, a synergistic interference algorithm with extremely miniaturization of confidence upper bound is proposed to optimize the direction and width of the interference beam to find the optimal interference strategy.

Benefits of technology

Without the opponent's prior information, the communication quality of the enemy network is significantly reduced, effective interference to large-scale communication networks is achieved, and the stability and efficiency of the interference effect are improved through the cooperation of the collaborative jammer.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120201451A_ABST
    Figure CN120201451A_ABST
Patent Text Reader

Abstract

The invention provides a multi-jammer cooperative beam decision-making method for a large-scale communication network. The method comprises the following steps: step 1, determining evaluation utility of jammers at any position; 2, the number of iterations of the jammers is given, and all the jammers sequentially traverse each arm in the beam direction and width strategy set of one-time interference; 3, updating the selection frequency of each strategy; 4, obtaining the statistical average return of the jammer; step 5, updating an interference strategy: after each jammer traverses all selectable arms in sequence, seeking expectation based on the cumulative return of the jammer, and selecting the arm of the next interference period; and step 6, stopping iteration when the maximum number of iterations is reached. According to the invention, the communication quality of the opposite network is reduced to the greatest extent.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of wireless interference technology, and particularly to a multi-jammer collaborative beam decision-making method for large-scale communication networks. Background Art

[0002] In the scenario of wireless communication network-level confrontation, using wireless interference technology to disrupt the data transmission of legitimate nodes in the enemy's communication network, and by reducing the communication quality of the other party's network and suppressing the communication ability of the other party, to ensure our information advantage in the confrontation process has become a common means.

[0003] Existing research on wireless interference technology often focuses on single-to-single interference scenarios, and improves the success rate or energy efficiency of interference by adjusting the interference channel or power strategy. However, in the face of large-scale communication networks, the factors affecting interference performance are becoming increasingly complex. In addition to optimizing the interference channel or power strategy, how to deploy the jammers at the best attack positions, or how to find and attack the key mission nodes or relay nodes in the network, or how to use beamforming technology to achieve a more accurate interference range and better interference effect under limited power, are all important research directions of wireless communication interference technology.

[0004] However, most of the existing work uses convex optimization methods to study network interference attack technology and find the theoretically optimal interference strategy solution. However, the biggest limitation of this type of method is that it requires perfect information about the other party's communication network, such as network topology, node positions, and communication parameters, etc. But in the actual interference process, due to the non-cooperative relationship between the two adversarial parties, it is difficult for the interfering party to obtain prior information about the other party's network, resulting in the difficulty of achieving ideal results for this type of solution method. Summary of the Invention

[0005] This application provides a multi-jammer collaborative beam decision-making method for large-scale communication networks, which can be used to solve the technical problem that it is difficult to obtain prior information of the other party, resulting in difficulty in achieving an ideal state.

[0006] A multi-jammer collaborative beam decision-making method for large-scale communication networks, the method includes:

[0007] Step 1, determine the evaluation utility of the jammer at any position;

[0008] Step 2, given the iteration times of the jammer, all jammers traverse each arm in the set of beam direction and width strategies of interference in turn;

[0009] Step 3: Update the selection times of each strategy;

[0010] Step 4: Obtain the statistical average return of the jammer;

[0011] Step 5: Update the interference strategy: After each jammer has traversed all the optional arms in sequence, calculate the expectation based on the cumulative reward C n (t j ) of jammer n, and select the arm for the next interference period;

[0012] Step 6: Reach the maximum number of iterations and stop the iteration.

[0013] Step 1, Determine the evaluation utility of the jammer at any position.

[0014] Step 11, Determine the transmission duration of the user within a time slot;

[0015] In the cooperative interference scenario for a multi-user communication network, there are M user clusters, and the set of user clusters is Each cluster includes L users, so the set of all users within a cluster is Let m l represent the l-th user in the m-th cluster; in the communication network, each cluster has an independent orthogonal channel, so there is no frequency conflict between different clusters; for user cluster m, the L users in the cluster share a channel in a time-division manner, that is, each time slot is further divided into L sub-time slots for the L users within the cluster; the length of a time slot is h slot , then the length of a sub-time slot is h slot / L; let h m,l represent the transmission duration of user m l within a sub-time slot, and h m,l < h slot / L, then {h m,1 (T), h m,2 (T),..., h m,l (T),..., h m,L (T)} represents the set of transmission durations of all users in cluster m within a certain time slot T, and {h m,1 (T + 1), h m,2 (T + 1),..., h m,l (T + 1),..., h m,L (T + 1)} represents the set of transmission durations of all users in cluster m within time slot T + 1 after being interfered in time slot T; for jammer n, the matrix of transmission durations of all users in the entire network within their respective sub-time slots sensed in a time slot T is:

[0016]

[0017] Each row represents the transmission duration of L users in a cluster within their respective sub - time slots. Each cluster has an independent channel, and M user clusters correspond to M channels in total. Therefore, by real - time sensing, the jammer obtains an M×L matrix, which records the transmission durations of all communication users in the network within one time slot.

[0018] Step 12: Determine the transmission duration of users within one interference cycle.

[0019] Since users may reduce or increase their transmission rates under the conditions of being continuously interfered or not being interfered, the interference effect of the jammer within one time slot may not be sufficient to cause obvious changes to communication users. Therefore, to ensure the accuracy of the evaluation results, increase the time length of the interference decision cycle, that is, consider that the decision cycle of the jammer includes Y time slots and the strategy remains unchanged within one cycle. The jammer judges the influence degree of the interference strategy in the previous cycle on all users in the network by comparing the changes in the transmission durations of all users in two adjacent interference cycles, so as to further evaluate the interference effectiveness. Based on formula (1), the jammer obtains a Y×ML matrix, which records the transmission durations of all communication users in the network within one interference cycle, that is, Y time slots, as follows:

[0020]

[0021] Where T′ = T + Y - 1, indicating the interference cycle (T, T′) corresponding to the time slots from the T - th time slot to the T + Y - 1 - th time slot;

[0022] After the jammer n implements interference within the interference cycle (T, T′), the matrix of the transmission durations of all users in their respective sub - time slots in the entire network sensed in the next interference cycle is:

[0023]

[0024] Where T″ = T + Y, T″′ = T + 2Y - 1, indicating the interference cycle (T″, T″′) corresponding to the time slots from the T + Y - th time slot to the T + 2Y - 1 - th time slot;

[0025] Step 13: Evaluate the interference effect achieved by the interference strategy adopted by the jammer in the previous interference cycle on the entire communication network by analyzing the changes in the transmission durations of all users in two adjacent interference cycles.

[0026] The evaluation criterion measures the interference effect based on the changes in the transmission durations of communication users, that is, the greater the change, the more obvious the interference effect. Under this evaluation criterion, convert the time matrices of size Y×ML in the above two adjacent interference cycles (T, T′) and (T″, T″′) into Y×M×L - dimensional vectors H j,n (T, T′) and Hj,n (T″, T″′), where T and T″ respectively correspond to the starting time slots T and T + Y of two interference periods; it is defined as follows:

[0027]

[0028] At the end of the interference period (T″, T″′), jammer n obtains, through sensing, its transmission duration vectors H j,n (T, T′) and H j,n (T″, T″′) in the previous two adjacent interference periods (T, T′) and (T″, T″′), and performs a correlation analysis on the two vectors;

[0029] The cosine similarity function is used to quantify the degree of correlation between the two vectors. The cosine similarity is defined as follows:

[0030]

[0031] where the lower the similarity between the two vectors, the smaller the cosine similarity ξ(H j,n (T, T′), H j,n (T″, T″′)), and correspondingly, the larger 1 - ξ(H j,n (T, T′), H j,n (T″, T″′)) is, indicating that the degree of change of all network users in the interference period (T″, T″′) after being affected by the interference behavior in the interference period (T, T′) is greater;

[0032] For the adjustment mechanism of automatic rate fallback, a judgment function η is designed to analyze the change in the transmission duration of the l-th communication user in the m-th cluster within the interference period (T, T′) after being interfered with, compared to the transmission duration of the corresponding time slot t u within the next interference period (T″, T″′) in the corresponding time slot t u + Y, where t u ∈{T, T + 1,..., T + Y - 1}. The judgment function η1 is defined as follows:

[0033]

[0034] When there is no change in the user transmission duration of the corresponding time slots within two interference periods, the judgment function η2 is given as follows:

[0035]

[0036] where ψ represents the number of times the user transmission durations in the corresponding time slots within two interference periods are equal, and σ represents a positive integer to reduce the magnitude of this evaluation value compared to the case of changes, and to minimize the impact of misjudgment results.

[0037] Step 14, determine the evaluation utility of the jammer's jamming strategy;

[0038] According to Formula (7) and Formula (8), combining the transmission duration vectors H in two adjacent interference cycles (T, T′) and (T″, T″′) j,n (T, T′) and H j,n (T″, T″′), the evaluation utility of the jammer n's jamming strategy in the cycle (T, T′) is given as follows:

[0039]

[0040] where the evaluation utility is composed of the sum of two parts, corresponding to the situations where the communication behaviors of all network users change and do not change in two adjacent interference cycles respectively; in the first half, 1 - ξ(H j,n (T, T′), H j,n (T″, T″′)) is used to represent the degree of change between the two vectors. The larger the value, the greater the change in the communication behaviors of all network users under the influence of interference, and the better the corresponding interference effect; then it is the cumulative sum of the judgment function η1(h m,l (T), h m,l (T + 1)) of the jammer n for all users in the whole network. By accumulating the positive interference effects (i.e., positive return values) and negative interference effects (i.e., negative return values) evaluated for all users, the overall evaluation result is given in the case where the transmission duration of users changes, promoting the jammer to adjust the jamming strategy to generate greater positive interference effects as much as possible and reduce the negative interference effects (i.e., avoid ineffective jamming strategies); similarly, in the second half, the cosine similarity between the two vectors and the cumulative sum of the judgment function η2(h m,l (T), h m,l (T + 1)) are multiplied to evaluate the interference effect in the case where the transmission duration of users does not change. To sum up, this evaluation method finally realizes the mapping from different user change situations to the interference effect.

[0041] However, due to the differences in the position distributions of all network users relative to different jammers and the existence of the signal perception thresholds of jammers for users, different jammers cannot determine whether they accurately perceive the existence of all users' signals, which in turn leads to different evaluation results for different jammers. Therefore, to avoid the evaluation deviation caused by uneven jammer position distributions and incomplete perception, consider that different jammers can exchange information on evaluation results, and take the maximum value among the evaluation results of all jammers as the overall interference effect on all network users, as follows: r j = max{r1, r2,..., r N}.

[0042] Step 2: Given the number of iterations of the jammer, all jammers sequentially traverse each arm in the set of beam direction and width strategies for jamming.

[0043] Each user consists of a set of transmit-receive pairs, and all users have the same transmission power, which is P. u ; There are N randomly distributed jammers in the system, and the set of jammers is All jammers have the same jamming power, which is P. j ; Compared with omnidirectional jamming, beamforming technology releases jamming signals by aiming a directional antenna at a specific direction, enabling the jamming party to effectively jam the densely populated area of users and distant communication users under the condition of constant power. During the jamming process, the jammer can choose V beam directions and W beam widths, and its beam direction strategy set and beam width strategy set are {φ1, φ2,..., φ V} and {θ1, θ2,..., θ W}; For the nth jammer, its beam direction strategy and beam width strategy are φ n ∈ {φ1, φ2,..., φ V} and θ n ∈ {θ1, θ2,..., θ W}, so the jamming area of the nth jammer is a fan-shaped area centered on φ n with a width of θ n ; In addition, when the width of the jamming beam is 360 degrees, it can be regarded as the omnidirectional jamming mode.

[0044] Any jammer n hopes to continuously adjust the beam direction and width strategies for jamming (φ n , θ n ) to achieve the maximum jamming effect. The problem of beam direction and width selection under finite strategies is modeled as a multi-armed bandit problem, where each set of beam direction and width joint strategies (φ v , θ w ), φ v ∈ {φ1, φ2,..., φ V}, θ w ∈ {θ1, θ2,..., θ W} is regarded as an optional arm k, and the total number of optional arms is K = V × W, and the corresponding set of jamming strategies is Update the jamming strategy: Given the maximum number of iterations T max , the set of all arms is

[0045] All jammers sequentially traverse each arm k in the set .

[0046] Step 3: Update the selection times of each strategy.

[0047] The jammer n selects arm B in the i-th interference cycle n (i), and obtains the evaluation return value r of this arm in the (i + 1)-th interference cycle j (B n (i));

[0048] Define the expected return of the jammer n selecting arm as:

[0049] u k,n = E[r j (k n )], (10)

[0050] where E[·] represents the expectation operation;

[0051] For the jammer n, if the expected returns of each arm k n are all known, then the optimal arm of the jammer n, that is, the optimal interference strategy, is:

[0052]

[0053] However, in the actual interference process, the expected return of each arm is completely unknown. Therefore, each jammer hopes to gradually find the optimal interference strategy during the interference, exploration, and learning process of the communication network. Let D k,n (t j ) represent the number of times that arm k n is selected by the jammer n in the past t j interference cycles, as follows:

[0054]

[0055] where λ(B n (i), k n ) is used to determine which arm the jammer selects in the i-th interference cycle, as follows:

[0056]

[0057] Step 4: Obtain the statistical average return of the jammer.

[0058] The cumulative return obtained by the jammer n in the past t j interference cycles is:

[0059]

[0060] Taking the expectation of the cumulative return C n (t j ) gives:

[0061]

[0062] For any jammer n, the statistical average reward of selecting the k-th arm within the past t j interference cycles is as follows:

[0063]

[0064] Step 5: Update the interference strategy. After each jammer has traversed all the available arms in turn, based on formula (15), that is, the cumulative reward C n (t j ), take the expectation and select the arm for the next interference cycle.

[0065] However, in the process of selecting the interference strategy, not only should we use the historical experience reward information to select the arms with higher average reward values more often, but also continuously explore the unknown arms with potential high gains to achieve a compromise between exploration and exploitation. Common MAB algorithms include UCB1, UCB2, and Greedy algorithms based on the upper confidence bound. However, in some cases, such algorithms will also maintain a relatively high selection frequency for sub-optimal arms, resulting in missing the selection of the optimal arm many times in the long-term decision-making process and leading to a relatively high upper bound of the interference regret value.

[0066] To address this problem, in order to maintain the exploration of unknown arms while reducing the selection frequency of non-optimal arms and increasing the selection frequency of the optimal arm in the decision-making process, drawing on the Minimax optimality algorithm, further reducing the upper bound of the interference regret value, the interference strategy update rule based on the improved UCB is given as follows:

[0067]

[0068] where log + (x) = log max{1, x}.

[0069] Step 6: Reach the maximum number of iterations and stop the iteration.

[0070] The present invention studies the multi-jammer cooperative interference method for large-scale unknown communication networks. First, it designs an interference effectiveness evaluation model based on the change of transmission duration to evaluate the interference reward in real time. Secondly, multiple jammers are randomly deployed, and beamforming technology is used to control the direction and width of the interference beam. Under the multi-armed bandit framework, a cooperative interference algorithm of minimizing the upper confidence bound is proposed. By exploring and learning different arms, the optimal interference strategy is found to minimize the communication quality of the enemy network to the greatest extent. Description of the Drawings

[0071] Figure 1 It is a diagram of the implementation steps of the multi-jammer collaborative beamforming strategy optimization for large-scale communication networks provided by the embodiments of this application;

[0072] Figure 2 It is a diagram of the multi-jammer collaborative beamforming interference scenario provided by the embodiments of this application;

[0073] Figure 3 It is a schematic diagram of the distribution scenario of user clusters and multi-jammers provided by the embodiments of this application;

[0074] Figure 4 It is a schematic diagram of the regret value of the jammer provided by the embodiments of this application;

[0075] Figure 5 It is a schematic diagram of the sum of the signal-to-interference-plus-noise ratios (SINRs) of all users within some user clusters provided by the embodiments of this application;

[0076] Figure 6 It is a schematic diagram of the sum of the SINRs of all users in the whole network under different numbers of user clusters provided by the embodiments of this application. Detailed implementation manners

[0077] To make the objectives, technical solutions, and advantages of this application clearer, the following will further describe the embodiments of this application in detail with reference to the accompanying drawings.

[0078] First, the embodiments of this application will be introduced below with reference to the accompanying drawings.

[0079] The present invention discloses a method for optimizing the multi-jammer collaborative beamforming strategy for large-scale communication networks. Considering that communication users can adaptively adjust the transmission rate in an interference environment, an interference effectiveness evaluation model based on the change of transmission duration is designed to evaluate the interference return in real time, providing an information basis for subsequent interference decisions. In addition, without prior information about the other communication network, the limited strategy optimization problem is modeled as a multi-armed bandit problem. Regarding the joint strategy of the direction and width of the interference beam as one arm, a collaborative interference algorithm of minimizing the maximum value of the upper confidence bound is proposed. By exploring and learning different arms, the optimal interference beam strategy is found. Finally, the proposed interference algorithm can significantly reduce the communication quality of all users in the network.

[0080] Figure 1 It is the implementation steps of the multi-jammer collaborative beamforming strategy optimization for large-scale communication networks proposed by the present invention. First, the maximum number of iterations of the jammer is given. Then, all jammers sequentially traverse each arm in the set, and then count the number of times each strategy is selected. According to the number of selections, the unified return value of all jammers is obtained. Then, based on the interference strategy update process, the arm for the next interference time slot is selected. If the number of iterations has not reached the maximum value, continue to update the arm for the next interference cycle according to the above method, otherwise stop.

[0081] Figure 2 This is a collaborative interference scenario for a multi-user communication network considered in the present invention. As Figure 2 shown, there are M user clusters, and the set of user clusters is Each cluster includes L users, so the set of all users within a cluster is Let m l represent the l-th user in the m-th cluster. Each user consists of a set of transmit-receive pairs, and all users have the same transmission power. Each cluster is randomly distributed and enjoys an independent orthogonal channel, so there is no frequency conflict between different clusters, and different users within a cluster share this channel in a time-division manner. There are also N interferers randomly distributed in the system, and the set of interferers is

[0082] Figure 3 This is a schematic diagram of the distribution scenario of user clusters and multiple interferers. There are 8 user clusters and 2 interferers randomly distributed in the system. As Figure 3 shown, where the hexagons represent the centers of the user clusters and the triangles represent the interferers, and all user clusters and interferers are randomly distributed within a 400m×400m area. For any user cluster, 10 user transmit-receive pairs are evenly distributed on a circle with the center of the user cluster as the center and a radius of 20m.

[0083] Figure 4 Compares the regret value functions of interferers under different algorithms. The regret value represents the interference performance loss caused by missing the selection of the optimal arm during the long-term decision-making process, and can reflect the quality of the interference strategy selection throughout the process. Since the evaluation reward values of each interferer are considered the same, the regret value functions of different interferers are also the same. It can be seen from the figure that the regret value function under the random interference algorithm (where the interferer randomly selects each strategy) shows a linear growth, while under the traditional UCB algorithm and the proposed interference algorithm, the regret value of the interferer grows logarithmically and the growth rate gradually decreases, indicating that the interferer can increasingly select the optimal arm to reduce the reward loss, but at the same time will still maintain the exploration of non-optimal arms, so the regret value will still continue to grow slowly. In addition, compared with the traditional UCB algorithm, the regret value of the proposed algorithm is lower, because the proposed algorithm can reduce the number of times the interferer explores sub-optimal arms and other arms, and improve the cumulative interference reward during the long-term decision-making process.

[0084] Figure 5The sum of the signal-to-interference-plus-noise ratios (SINRs) of all users within some user clusters under different algorithms is compared. Each curve in the figure is obtained by statistical averaging every 200 time slots. Considering that the number of users, transmission power, and relative distribution positions within each user cluster are the same, the severity of the interference impact on different user clusters can be judged by analyzing the degree of decrease in the sum of the SINRs of different user clusters. Compared with the random interference algorithm, the sum of the SINRs of all users in different clusters has a significant decrease, indicating that the jammer achieves a good interference effect on the entire communication network. Among them, the sum of the SINRs of user cluster 3 decreases the most severely, and the sum of the SINRs of user cluster 6 decreases relatively less because this cluster is located in two relatively remote positions in the northwest and northeast corners and is less affected by interference.

[0085] Figure 6 The sum of the SINRs of all users within all clusters under different numbers of user clusters is given. Each curve in the figure is obtained by statistical averaging every 200 time slots. To ensure the fairness of the comparison as much as possible, two additional user clusters are randomly generated on the basis of the original 8 clusters. It can be seen that as the number of user clusters increases, the proposed interference algorithm can always significantly reduce the sum of the SINRs of all users in the whole network, indicating that the proposed method can still achieve an ideal interference effect when facing a large-scale unknown communication network.

[0086] In view of the problems that it is difficult for the jammer to obtain prior information such as the number of nodes, location distribution, and transmission power of the communication party in the non-cooperative confrontation scenario and a single jammer is difficult to achieve an ideal interference effect, this application designs an interference effectiveness evaluation model based on the change of transmission duration and a multi-jammer cooperative interference method combined with beamforming. By optimizing the direction and width of their respective interference beams, the optimal interference beam strategy is found to maximize the long-term expected return of the jammer and achieve effective interference on the unknown communication network.

[0087] The above-described embodiments of this application do not constitute a limitation on the protection scope of this application.

Claims

1. A multi-jammer coordinated beam decision method for large-scale communication networks, characterized in that: The method comprises: Step 1, determine the evaluation utility of the jammer at any location; Step 2, given the number of jammer iterations, all jammers traverse each arm in the jamming beam direction and width strategy set in turn; Step 3: Update the selection times of each strategy; Step 4: Get the statistical average return of the jammer; Step 5: Update the jamming strategy: After each jammer has traversed all optional arms in turn, based on the cumulative reward C of jammer n n (t j ) Find the expectation and select the arm of the next interference cycle; Step 6: When the maximum number of iterations is reached, stop iterating.

2. The method according to claim 1, characterized in that Step 1, determine the estimated utility of the jammer at any location, including: Step 11, determining the transmission duration of the user in a time slot; There are M user clusters in the cooperative interference scenario for multi-user communication networks, and the set of user clusters is Each cluster includes L users, so the set of all users in a cluster is Let m l represents the lth user in the mth cluster; in a communication network, each cluster has an independent orthogonal channel, so the frequencies used by different clusters do not conflict with each other; for user cluster m, the L users in the cluster share a channel in a time-division manner, that is, each time slot is divided into L sub-time slots for use by the L users in the cluster; the length of a time slot is h slot , then the length of a sub-slot is h slot / L; let h m,l Indicates user m l The transmission duration in a sub-slot, and h m,l <h slot / L, then {h m,1 (T),h m,2 (T),...,h m,l (T),...,h m,L (T)} represents the set of transmission durations of all users in cluster m in a certain time slot T, {h m,1 (T+1),h m,2 (T+1),...,h m,l (T+1),...,h m,L (T+1)} represents the set of transmission durations of all users in cluster m in time slot T+1 after being interfered in time slot T; for the interference machine n in a time slot T, the transmission duration matrix of all users in the entire network in their respective sub-time slots is: Each row represents the transmission duration of L users in a cluster in their respective sub-time slots. Each cluster has an independent channel, and M user clusters correspond to M channels in total. Therefore, the jammer obtains an M×L matrix through real-time perception, which records the transmission duration of all communication users in the network in a time slot. Step 12, determining the transmission duration of the user in an interference period; Consider that the decision cycle of the jammer includes Y time slots and the strategy remains unchanged within one cycle; the jammer determines the impact of the previous cycle's jamming strategy on all network users by comparing the changes in the transmission duration of all users in two adjacent jamming cycles; based on formula (1), the jammer obtains a Y×ML matrix, which records the transmission duration of all communication users in the network within one jamming cycle, i.e., Y time slots, as shown below: Where T′=T+Y-1, which represents the interference period (T, T′) corresponding to the Tth time slot to the T+Y-1th time slot; When the jammer n implements interference in the interference period (T, T′), the transmission duration matrix of all users in the entire network in their respective sub-time slots is perceived in the next interference period as follows: Wherein T″=T+Y,T″′=T+2Y-1, indicating the interference period (T,T"′) corresponding to the T+Yth time slot to the T+2Y-1th time slot; Step 13, by analyzing the changes in the transmission duration of all users in two adjacent interference periods, evaluate the interference effect of the interference strategy adopted by the jammer in the previous interference period on the entire communication network; The evaluation criterion is to measure the interference effect based on the change of the transmission time of the communication user, that is, the greater the change, the more obvious the interference effect; under the evaluation criterion, the time matrices of size Y×ML in the above two adjacent interference periods (T, T′) and (T″, T″′) are converted into vectors H of dimension Y×M×L j,n (T,T′) and H j,n (T",T"′), where T and T" correspond to the starting time slots T and T+Y of the two interference periods respectively; the definitions are as follows: At the end of the interference period (T", T″′), the jammer n obtains its transmission duration vector H in the first two adjacent interference periods (T, T′) and (T″, T″′) by sensing. j,n (T,T′) and H j,n (T″,T″′), and perform correlation analysis on the two vectors; The cosine similarity function is used to quantify the degree of correlation between two vectors. The cosine similarity is defined as follows: The lower the similarity between two vectors, the cosine similarity ξ(H j,n (T,T′),H j,n (T″,T″′)) is smaller, and accordingly 1-ζ(H j,n (T,T′),H j,n The larger the (T″,T″′)), the greater the degree of change of users in the entire network in the interference period (T,T′) after being affected by the interference behavior in the interference period (T″,T″′); According to the automatic rate fallback adjustment mechanism, a judgment function η is designed to analyze the lth communication user in the mth cluster after being interfered within the interference period (T, T′). u The transmission duration is compared to the corresponding time slot t in the next interference period (T″, T″′). u + Changes in transmission time within Y, where t u ∈{T,T+1,...,T+Y-1}, the judgment function η1 is defined as follows: When the user transmission duration of the corresponding time slot does not change in two interference periods, the judgment function η2 is given as follows: Where ψ represents the number of times that the user transmits with equal duration in the corresponding time slots during two interference periods, and σ represents a positive integer; Step 14, determining the evaluation utility of the jammer jamming strategy; According to formula (7) and formula (8), combining the transmission duration vector H in two adjacent interference periods (T, T′) and (T″, T″′) j,n (T,T′) and H j,n (T″,T″′), gives the evaluation utility of the jammer n’s jamming strategy in period (T,T′), as follows: The evaluation utility consists of two parts, corresponding to the situation where the communication behaviors of all network users change and remain unchanged in two adjacent interference periods. In the first half, 1-ξ(H j,n (T,T′),H j,n (T″,T″′)) is used to indicate the degree of change between the two vectors. The larger the value, the greater the change in the communication behavior of all network users under the influence of interference, and the better the corresponding interference effect; is the judgment function η1(h m,l (T),h m,l (T+1)) cumulative sum, by accumulating the positive interference effects (positive reward values) and negative interference effects (negative reward values) evaluated for all users, the overall evaluation result is given when the user transmission time changes, which promotes the jammer to adjust the jamming strategy to produce a greater positive interference effect as much as possible and reduce the negative interference effect, that is, to avoid invalid jamming strategies; in the second half, the cosine similarity between the two vectors and the judgment function η2(h m,l (T),h m,l (T+1)) is multiplied by the cumulative sum to evaluate the interference effect when the user transmission duration does not change; To avoid evaluation bias caused by uneven distribution of jammer locations and incomplete perception, consider the information that different jammers can interact with each other in evaluation results, and take the maximum value of all jammer evaluation results as the overall interference effect on all network users, as shown below: j =max{r1,r2,...,r N }.

3. The method according to claim 2, characterized in that Step 2, given the number of jammer iterations, all jammers traverse each arm in the jamming beam direction and width strategy set in turn; include: Each user consists of a set of transmit-receive pairs, and the transmission power of all users is the same, which is P u ; There are N jammers randomly distributed in the system, and the set of jammers is The interference power of all jammers is the same, which is P j ; In the process of implementing interference, the jammer has V beam directions and W beam widths to choose from, and its beam direction strategy set and beam width strategy set are {φ1,φ2,...,φ V } and {θ1,θ2,...,θ W }; For the nth jammer, its beam direction strategy and beam width strategy are φ n ∈{φ1,φ2,...,φ V } and θ n ∈{θ1,θ2,...,θ W }, so the interference area of ​​the nth jammer is is φ n is the center and the width is θ n sector-shaped area; in addition, when the width of the interference beam is 360 degrees, it can be regarded as an omnidirectional interference mode. Any jammer n hopes to continuously adjust the jamming beam direction and width strategy (φ n ,θ n ), the beam direction and width selection problem under finite strategies is modeled as a multi-armed bandit problem, where each set of beam direction and width joint strategy (φ v ,θ w ),φ v ∈{φ1,φ2,...,φ V },θ w ∈{θ1,θ2,...,θ W } is regarded as an optional arm k, the total number of optional arms is K = V × W, and the corresponding set of interference strategies is Update the interference strategy: given the maximum number of iterations T max , the set of all arms is All jammers traverse the set in turn For each arm k in .

4. The method according to claim 3, characterized in that Step 3: Update the selection times of each strategy, including: Jammer n selects arm B in the i-th jamming cycle n (i) After that, the evaluation reward value r of the arm is obtained in the i+1th interference cycle j (B n (i)); define jammer n select arm The expected return is: Where E[·] represents the expected operation; For jammer n, if each arm k n The expected returns of are all known, then the optimal arm of jammer n, that is, the optimal jamming strategy, is: Let D k,n (t j ) represents arm k n In the past j The number of times a jammer n is selected by a jammer is as follows: Where λ(B n (i),k n ) is used to determine which arm the jammer selects in the i-th jamming cycle, as shown below:

5. The method according to claim 4, characterized in that Step 4: Get the statistical mean return of the jammer. Jammer n in the past t j The cumulative reward obtained from the interference cycle is: Cumulative reward C to jammer n n (t j ) to find the expected value: For any jammer n, in the past t j The statistical average reward of selecting the kth arm within the interference period is:

6. The method according to claim 5, characterized in that Step 5: Update the interference strategy, including: By referring to the Minimax optimality algorithm, the upper bound of the interference regret value is lowered, and the interference strategy update rule based on the improved UCB is given as follows: where log + (x)=logmax{1,x}.