Intelligent game network anti-interference method and device, equipment and medium
Patent Information
- Application Number
- CN202310788337.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-29
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2043-06-29
AI Technical Summary
[0002]无人机通信系统辅助的空中计算网络将移动边缘计算与无人机使能的网络集成在一起,可以为物联网提供连接,辅助物联网设备收集处理数据,但是无人机通信系统辅助的空中计算网络容易受到干扰源的攻击,因此需要高效且有效的抗干扰技术保障无人机在空中计算网络受到干扰源攻击时,可正常与其他无人机进行通信;
[0037]本申请提供的智能博弈网络抗干扰方法、装置、设备及介质,首先构建无人机通信系统的网络拓扑模型,基于空中计算网络计算在所述网络拓扑模型中每个无人机与其他无人机之间进行通信的通信参数,得到多个无人机之间的通信模型;所述网络拓扑模型反映无人机通信系统中,多个无人机之间的实时相对位置与通信交互;然后构建Dec-POMDP模型、构建第一干扰源种群并构建MARL抗干扰算法,将所述MARL抗干扰算法应用于所述Dec-POMDP模型,并将所述第一干扰源种群与所述通信模型输入到所述Dec-POMDP模型中进行对抗训练,并基于所述MARL抗干扰算法得到不同时刻下对抗训练的结果,基于不同时刻下对抗训练的结果对所述Dec-POMDP模型进行迭代训练直到完成训练;所述对抗训练的结果包括每个无人机与其他无人机之间通信的频率与功率;最后将完成训练的Dec-POMDP模型应用于所述无人机通信系统中,并基于信息时序平滑机制优化所述Dec-POMDP模型的输出结果,得到恶意干扰存在下的无人机通信系统中每个无人机与其他无人机之间通信的目标频率与功率。本方案通过构建联合频率选择和功率控制的Dec-POMDP模型,并将其应用于实际的无人机通信系统,并优化Dec-POMDP模型的输出结果,可控制无人机通信系统中的每个无人机在干扰源的攻击下,选择最优的通信频率与通信功率,提高了保障无人机正常通信的能力与效率。
Smart Images

Figure CN116915310B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of intelligent communication, and in particular to an anti-interference method, apparatus, device and medium for intelligent game networks. Background Technology
[0002] Drone communication system-assisted aerial computing networks integrate mobile edge computing with drone-enabled networks, providing connectivity for the Internet of Things (IoT) and assisting IoT devices in collecting and processing data. However, drone communication system-assisted aerial computing networks are vulnerable to interference attacks, so efficient and effective anti-interference technologies are needed to ensure that drones can communicate normally with other drones when the aerial computing network is attacked by interference sources.
[0003] Current traditional anti-jamming methods are based on frequency hopping spread spectrum technology using sequences, employing fixed or preset modes. However, they lack intelligent decision-making capabilities. Alternatively, when using anti-jamming models with intelligent decision-making capabilities, random strategies are employed to resist interference attacks before training is complete, requiring the model to learn from scratch. This approach takes a long time to converge in practical deployments, resulting in poor anti-jamming performance in latency-sensitive environments. Therefore, current anti-jamming methods are inefficient and impractical in ensuring normal communication for drones when facing interference attacks. Summary of the Invention
[0004] This application provides a method, apparatus, device, and medium for anti-interference in intelligent gaming networks, which can improve the ability and efficiency of ensuring normal communication for unmanned aerial vehicles (UAVs).
[0005] Firstly, this application provides an anti-interference method, the method comprising:
[0006] A network topology model of an unmanned aerial vehicle (UAV) communication system is constructed. Based on an aerial computing network, communication parameters for communication between each UAV and other UAVs in the network topology model are calculated to obtain a communication model between multiple UAVs. The network topology model reflects the real-time relative positions and communication interactions between multiple UAVs in the UAV communication system.
[0007] A Dec-POMDP model is constructed, a first interference source population is constructed, and a MARL anti-jamming algorithm is constructed. The MARL anti-jamming algorithm is applied to the Dec-POMDP model, and the first interference source population and the communication model are input into the Dec-POMDP model for adversarial training. The adversarial training results at different time points are obtained based on the MARL anti-jamming algorithm. The Dec-POMDP model is iteratively trained based on the adversarial training results at different time points until training is completed. The adversarial training results include the frequency and power of communication between each UAV and other UAVs.
[0008] The trained Dec-POMDP model is applied to the UAV communication system, and the output of the Dec-POMDP model is optimized based on the information time-series smoothing mechanism to obtain the target frequency and power of communication between each UAV and other UAVs in the UAV communication system under malicious interference.
[0009] In one example, constructing the Dec-POMDP model includes:
[0010] Define the state space, observation space, action space, and reward function, and construct the Dec-POMDP model based on the defined state space, observation space, action space, and reward function;
[0011] The construction of the first interference source population includes:
[0012] Multiple interference source instances are initialized, and each initialized interference source instance is randomly assigned an identification code. The multiple initialized interference source instances with identification codes are combined into an initial interference source population, and the initial interference source population is updated based on the first algorithm to obtain the first interference source population.
[0013] The construction of the MARL anti-interference algorithm includes:
[0014] Design an observation encoder, a graph convolutional layer, and a Q-network, and fuse the observation encoder, graph convolutional layer, and Q-network into the MARL anti-interference algorithm;
[0015] The observation encoder is used to encode the observation data of the local spectrum environment of the state space observed by each UAV in the observation space in the UAV communication system into a low-dimensional feature vector, and send all the low-dimensional feature vectors to the convolutional layer. The convolutional layer obtains the weighted sum of each UAV and the interference source population based on the obtained low-dimensional feature vectors, and transmits the weighted sum of each UAV and the interference source population to the Q network. The Q network can obtain the frequency and power of communication between each UAV and other UAVs based on the weighted sum of each UAV and the interference source population.
[0016] In one example, updating the initial interference source population based on the first algorithm to obtain the first interference source population includes:
[0017] A first initial interference source instance is randomly selected from the initial interference source population, and the first initial interference source instance and the communication model are input into the Dec-POMDP model, so that the first initial interference source instance performs adversarial training with the communication model based on the first identification code of the first initial interference source instance, and based on the reward function in the Dec-POMDP model, a first evolved interference source instance with a second identification code is obtained from the first initial interference source instance.
[0018] The difference between the second identification code and the identification codes of all initial interference source instances in the initial interference source population is compared with a preset first threshold.
[0019] If the difference between the second identification code and the identification code of any initial interference source instance is less than or equal to the first threshold, then the first evolved interference source instance is ignored.
[0020] If the difference between the second identification code and the identification codes of all initial interference source instances is greater than the first threshold, then the first evolved interference source instance is added to the initial interference source population until all initial interference source instances in the initial interference source population are input into the Dec-POMDP model to obtain the first interference source population.
[0021] In one example, the iterative training of the Dec-POMDP model based on the results of adversarial training at different time points until training is complete includes:
[0022] Add a loss function to the MARL anti-interference algorithm and set a threshold for the number of iterations during iterative training;
[0023] The results of adversarial training at different times are used as training samples in sequence to iteratively train the Dec-POMDP model. The value of the loss function output corresponding to each iteration is recorded. Iterative training stops when the value of the loss function output is less than a preset value or when the number of iterations reaches the iteration threshold.
[0024] In one example, optimizing the output of the Dec-POMDP model based on the information time-series smoothing mechanism includes:
[0025] Each UAV in the UAV communication system is assigned an individual Q-function, and the individual Q-functions assigned to each UAV are summed to obtain the global Q-function of the UAV communication system, so that each UAV can generate shared information based on the global Q-function;
[0026] A timing smoothing mechanism is created, which consists of a message encoder, a combination block, a send message buffer, and a receive message buffer.
[0027] The UAV communication system is configured based on the aforementioned timing smoothing mechanism. The configuration includes: when information interaction occurs between UAVs, the UAV that performs information transmission generates a message to be transmitted based on the message encoder and places the message to be transmitted in the transmission message buffer; and when accessing an available channel, it sends the message to be transmitted to an adjacent UAV. The UAV that performs message reception stores the received messages in the reception message buffer and marks each message in the reception message buffer with a valid bit. When messages with the same valid bit no longer change, all information with that valid bit is sent to the combination block, so that the combination block generates a corresponding joint Q value based on all information with that valid bit.
[0028] The joint Q-value is applied to the Dec-POMDP model so that the Dec-POMDP model, in conjunction with the joint Q-value, optimizes the target frequency and power of the communication between the UAVs output by the Dec-POMDP model.
[0029] On the other hand, this application provides an anti-interference device, the device comprising:
[0030] The construction module is used to construct a network topology model of the UAV communication system. Based on the airborne computing network, it calculates the communication parameters between each UAV and other UAVs in the network topology model to obtain a communication model between multiple UAVs. The network topology model reflects the real-time relative positions and communication interactions between multiple UAVs in the UAV communication system.
[0031] The construction module is also used to construct a Dec-POMDP model, construct a first interference source population, and construct a MARL anti-interference algorithm. The MARL anti-interference algorithm is applied to the Dec-POMDP model, and the first interference source population and the communication model are input into the Dec-POMDP model for adversarial training. Based on the MARL anti-interference algorithm, the adversarial training results at different times are obtained. The Dec-POMDP model is iteratively trained based on the adversarial training results at different times until training is complete. The adversarial training results include the frequency and power of communication between each UAV and other UAVs.
[0032] The processing module is used to apply the trained Dec-POMDP model to the UAV communication system and optimize the output of the Dec-POMDP model based on the information time-series smoothing mechanism to obtain the target frequency and power of communication between each UAV and other UAVs in the UAV communication system under malicious interference.
[0033] In another aspect, this application provides an electronic device, including: a processor, and a memory communicatively connected to the processor;
[0034] The memory stores computer-executed instructions;
[0035] The processor executes computer execution instructions stored in the memory to implement the method described above.
[0036] In another aspect, this application provides a computer-readable storage medium storing computer-executable instructions that, when executed by a processor, are used to implement the method as described in the preceding claim.
[0037] The intelligent game network anti-interference method, apparatus, device, and medium provided in this application first construct a network topology model of an unmanned aerial vehicle (UAV) communication system. Based on an aerial computing network, communication parameters for communication between each UAV and other UAVs in the network topology model are calculated to obtain a communication model between multiple UAVs. The network topology model reflects the real-time relative positions and communication interactions between multiple UAVs in the UAV communication system. Then, a Dec-POMDP model is constructed, a first interference source population is constructed, and a MARL anti-interference algorithm is constructed. The MARL anti-interference algorithm is applied to the Dec-POMDP model, and the first interference source population and the communication model are input into... Adversarial training is performed on the Dec-POMDP model, and the results of the adversarial training at different time points are obtained based on the MARL anti-jamming algorithm. The Dec-POMDP model is then iteratively trained based on these results until training is complete. The results of the adversarial training include the frequency and power of communication between each UAV and other UAVs. Finally, the trained Dec-POMDP model is applied to the UAV communication system, and the output of the Dec-POMDP model is optimized based on an information time-series smoothing mechanism to obtain the target frequency and power for communication between each UAV and other UAVs in the presence of malicious interference. This scheme, by constructing a Dec-POMDP model with joint frequency selection and power control, applying it to a practical UAV communication system, and optimizing the output of the Dec-POMDP model, can control each UAV in the UAV communication system to select the optimal communication frequency and power under the attack of interference sources, improving the ability and efficiency of ensuring normal UAV communication. Attached Figure Description
[0038] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0039] Figure 1This is a schematic diagram illustrating an application scenario for this application.
[0040] Figure 2 A flowchart illustrating an anti-interference method provided in Embodiment 1 of this application;
[0041] Figure 3 A flowchart illustrating another anti-interference method provided in Embodiment 1 of this application;
[0042] Figure 4 This is a schematic diagram of the spectrum waterfall.
[0043] Figure 5 This is a schematic diagram of the MARL anti-interference algorithm;
[0044] Figure 6 A flowchart illustrating another anti-interference method provided in Embodiment 1 of this application;
[0045] Figure 7 This is a schematic diagram of the population update of the interference source.
[0046] Figure 8 A flowchart illustrating another anti-interference method provided in Embodiment 1 of this application;
[0047] Figure 9 This is a schematic diagram of a parallel Q-network;
[0048] Figure 10 A flowchart illustrating another anti-interference method provided in Embodiment 1 of this application;
[0049] Figure 11 This is a schematic diagram of the message smoothing mechanism;
[0050] Figure 12 An anti-interference device provided in Embodiment 2 of this application;
[0051] Figure 13 This is a schematic diagram of the structure of an electronic device provided in Embodiment 3 of this application.
[0052] The accompanying drawings illustrate specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to particular embodiments. Detailed Implementation
[0053] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.
[0054] The specific application scenario for this application is in the field of geophysics. Figure 1 This is a schematic diagram of an application scenario for this application. In an aerial computing network assisted by a drone communication system, each drone (Unmanned Aerial Vehicle, or UAV) is equipped with an intelligent agent. Interference sources around the drone will send interference attacks to the drone. The intelligent agent on the drone will select an appropriate communication frequency and communication power based on the attack signals sent by the interference sources.
[0055] Current traditional anti-jamming methods are based on frequency hopping spread spectrum technology, using fixed or preset modes to control communication on UAVs. These methods lack intelligent decision-making capabilities regarding interference signals emitted by the source. Alternatively, they directly apply anti-jamming models with intelligent decision-making capabilities to the UAV communication system before they are fully trained, employing random strategies to resist interference attacks and requiring the model to learn from scratch. This approach requires a long convergence time in practical deployment, resulting in poor anti-jamming performance in latency-sensitive environments. Therefore, current anti-jamming methods are inefficient and impractical in ensuring normal UAV communication when facing interference attacks.
[0056] However, the method provided in this application constructs and trains a Dec-POMDP model with joint frequency selection and power control, applies it to a real UAV communication system, and optimizes the output of the Dec-POMDP model through an information timing smoothing mechanism. This allows each UAV in the UAV communication system to select the optimal communication frequency and communication power under the attack of interference sources, thereby improving the ability and efficiency of ensuring normal communication of UAVs.
[0057] It should be noted that the brief descriptions of terms in this application are only for the convenience of understanding the embodiments described below, and are not intended to limit the embodiments of this application. Unless otherwise stated, these terms should be understood in their ordinary and common meaning.
[0058] The technical solutions of this application will be described in detail below with reference to specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. In the description of this application, unless otherwise expressly specified and limited, the terms should be broadly understood within the art. The embodiments of this application will now be described with reference to the accompanying drawings.
[0059] Example 1
[0060] Figure 2 This is a flowchart illustrating an anti-interference method provided in Embodiment 1 of this application, as shown below. Figure 2 As shown, the method includes:
[0061] Step 201: Construct a network topology model of the UAV communication system. Based on the aerial computing network, calculate the communication parameters between each UAV and other UAVs in the network topology model to obtain a communication model between multiple UAVs. The network topology model reflects the real-time relative positions and communication interactions between multiple UAVs in the UAV communication system.
[0062] Step 202: Construct a Dec-POMDP model, construct a first interference source population and a MARL anti-interference algorithm, apply the MARL anti-interference algorithm to the Dec-POMDP model, and input the first interference source population and the communication model into the Dec-POMDP model for adversarial training. Obtain the adversarial training results at different times based on the MARL anti-interference algorithm, and iteratively train the Dec-POMDP model based on the adversarial training results at different times until training is complete. The adversarial training results include the frequency and power of communication between each UAV and other UAVs.
[0063] Step 203: Apply the trained Dec-POMDP model to the UAV communication system, and optimize the output of the Dec-POMDP model based on the information time-series smoothing mechanism to obtain the target frequency and power of communication between each UAV and other UAVs in the UAV communication system under malicious interference.
[0064] The execution subject of this embodiment is an anti-interference device, which can be implemented by a computer program, such as application software; or it can be implemented as a medium storing relevant computer programs, such as a USB flash drive or cloud drive; or it can be implemented by a physical device that integrates or installs relevant computer programs, such as a chip.
[0065] Based on scenario examples and the relative positions and communication interactions between UAVs in an unmanned aerial vehicle (UAV) communication system, a network topology model is established. To consider the mobility of UAVs, the network topology is abstracted as a dynamic graph that changes over time. In this graph, vertices represent interference-resistant UAVs. The communication parameters for communication between UAVs, as well as the interference signals experienced by each UAV, can be calculated using an aerial computing network. Specifically, we assume that each UAV needs to send information to other UAVs, which can be the coordination information calculated by the unloading task from the mobile phone on the IoT device or by the intelligent agent. Time is divided into equal and fixed intervals. A specific time interval t is selected, and the frequency and power calculated by the intelligent agent on the UAV in the previous time interval based on spectrum observation are defined as f. i t and Define the frequency of the drone as The transmission power of the drone is discretized into P+1 available levels, represented as P = {0, ..., P}. The minimum power is P. min The maximum power is P max Therefore, the actual transmission power of the drone can be calculated using the following first formula:
[0066]
[0067] For each UAV UAVi, the potential attack interference can be expressed by the following second formula:
[0068]
[0069] in, This represents the frequency selected by UAVI. This represents the frequency selection configuration for all UAVs except UAV. The policy configuration for interference sources is defined by c = [c1, c2, ..., c...]. j ] indicates that c j The attack frequency of interference source j, the mutual interference between drones, and the impact of interference attacks are determined by... and This indicates that α is the interference attack factor, which reflects the severity of the attack's impact on the communication link.
[0070] In the network topology model, each UAV (UAVi) has a set of homogeneous neighbors Bi. When any two UAVs, UAVi and UAVj, belong to Bi, and UAVi and UAVj are within a communicable range, and a communication link is established between them on the same interference-free channel, then UAVi and UAVj can communicate. Potential interference with UAVi can be defined using the following third formula:
[0071]
[0072] Where δ(x) is the indicator function, Specify A subset of the set, where the radio states of any UAVi and UAVj are in the state of transmitting information in the current time slot t, g j This is the channel power gain of UAV j. Similarly, the interference attack signal received by UAV j can be expressed by the following fourth formula:
[0073]
[0074] In summary, the signal-to-noise ratio received by the drone can be represented by the following fifth formula:
[0075]
[0076] β th If the minimum SINR threshold for successful transmission is set, the transmission rate is calculated using the following sixth formula:
[0077]
[0078] Where W is the channel bandwidth, and according to the fifth formula, the communication success rate can be defined based on the following seventh formula:
[0079]
[0080] Where M represents the number of data packets transmitted within a time slot, it is worth noting that parameter M reflects the network load. To focus on the performance of the proposed anti-interference solution and ignore the impact of heavy load, we set M to a small value in the table. To evaluate power efficiency, we use the average transmit power P... ave This can be expressed as the eighth formula:
[0081]
[0082] The average hopping frequency k can be defined by the following ninth formula:
[0083]
[0084] It is worth noting that drones lack prior knowledge of interference sources, including interference patterns and source locations. Therefore, each drone must perform anti-jamming communication in unknown and dynamic environments. Thus, on the one hand, since interference sources employ different attack channels and fixed or intelligent sensing modes, each drone should predict interference strategies and bypass potentially attacked channels. On the other hand, it also needs to explore the trade-off between high SINR and high transmission power, as well as the resulting severe mutual interference and energy waste.
[0085] Based on the various parameters obtained from the above formulas for communication between UAVs, including: the actual transmission power of the UAV, possible attack interference, potential interference, received interference attack signals, signal-to-noise ratio, transmission rate, communication success rate, average transmission power and average hopping frequency, a communication model for the UAV communication system is established.
[0086] Multiple high-quality interference sources with diversity are designed through an evolutionary and selection process. These multiple high-quality interference sources are constructed into a first interference source population, which is then applied to the communication model to interfere with the drones within the model. Furthermore, a graph convolution-based Multi-Agent Reinforcement Learning (MARL) anti-interference algorithm can be constructed to calculate the frequency and power selection for communication between drones when facing interference from the first interference source population. A Dec-POMDP model is constructed, and the first interference source population and the communication model are input into the Dec-POMDP model. The MARL anti-interference algorithm is then applied to the Dec-POMDP model, allowing the frequency and power selections obtained by the drones when facing interference from the first interference source population to serve as training data for training the Dec-POMDP model. This allows the trained Dec-POMDP model to control the drones to select appropriate transmission frequencies and powers based on the infection information from the interference sources.
[0087] The trained Dec-POMDP model is deployed on the agent of a drone in a real drone communication system, and the Dec-POMDP model is optimized for the drone communication system used in the actual application. This allows the Dec-POMDP model to provide more accurate transmission frequency and transmission power when the drone communication system is facing interference from interference sources.
[0088] This example provides an intelligent game-theoretic network anti-interference method, apparatus, device, and medium. First, a network topology model of an unmanned aerial vehicle (UAV) communication system is constructed. Based on an aerial computing network, communication parameters for communication between each UAV and other UAVs in the network topology model are calculated, resulting in a communication model among multiple UAVs. Then, the MARL anti-interference algorithm is applied to the Dec-POMDP model, and the first interference source population and the communication model are input into the Dec-POMDP model for adversarial training. The results of adversarial training at different times are obtained, and the Dec-POMDP model is iteratively trained based on these results until training is complete. Finally, the trained Dec-POMDP model is applied to the UAV communication system, and the output of the Dec-POMDP model is optimized based on an information time-series smoothing mechanism to obtain the target frequency and power for communication between each UAV and other UAVs in the UAV communication system under malicious interference. This solution constructs a Dec-POMDP model for joint frequency selection and power control, applies it to a practical UAV communication system, and optimizes the output of the Dec-POMDP model. This allows each UAV in the UAV communication system to select the optimal communication frequency and power under the attack of interference sources, thereby improving the ability and efficiency of ensuring normal communication of UAVs.
[0089] Optional, Figure 3 A flowchart illustrating another anti-interference method provided in Embodiment 1 of this application is shown below. Figure 3 As shown, in step 202, constructing the Dec-POMDP model includes:
[0090] Step 301: Define the state space, observation space, action space, and reward function, and construct the Dec-POMDP model based on the defined state space, observation space, action space, and reward function;
[0091] In step 202, constructing the first interference source population includes:
[0092] Step 302: Initialize multiple interference source instances and randomly assign an identification code to each initial interference source instance. Combine the multiple initial interference source instances with identification codes into an initial interference source population, and update the initial interference source population based on the first algorithm to obtain the first interference source population.
[0093] In step 202, constructing the MARL anti-interference algorithm includes:
[0094] Step 303: Design the observation encoder, graph convolutional layer and Q network, and fuse the observation encoder, graph convolutional layer and Q network into the MARL anti-interference algorithm;
[0095] The observation encoder is used to encode the observation data of the local spectrum environment of the state space observed by each UAV in the observation space in the UAV communication system into a low-dimensional feature vector, and send all the low-dimensional feature vectors to the convolutional layer. The convolutional layer obtains the weighted sum of each UAV and the interference source population based on the obtained low-dimensional feature vectors, and transmits the weighted sum of each UAV and the interference source population to the Q network. The Q network can obtain the frequency and power of communication between each UAV and other UAVs based on the weighted sum of each UAV and the interference source population.
[0096] With a scenario example, in an interference-resistant environment, the representation of the state space is one of the key factors affecting model performance and convergence. The state space can be characterized as a spectral waterfall, which represents continuous observations of the spectral environment over a continuous time period and within a fixed frequency range. Figure 4 This is a schematic diagram of a spectrum waterfall. To achieve comprehensive awareness of communication frequency bands, UAVs continuously acquire and record signal strength using channel energy detection technology. Each detected value is stored, ultimately forming a spectrum waterfall diagram. To prevent the information transmitted between UAVs from becoming too large, continuous time is discretized into time slots, and the frequencies in the spectrum waterfall are sampled into multiple channels, so that frequencies correspond to channels, resulting in a discretized spectrum waterfall. This discretized spectrum waterfall can be represented as:
[0097]
[0098] in This represents the power level on channel l at time t.
[0099] In the observation space, each UAV's own observations can only provide a partial understanding of the current environmental state, due to its limited sensing range and position relative to interference sources. Specifically, each UAV obtains its own discretized spectral waterfall by utilizing historical observations, which can be represented as:
[0100]
[0101] Their format is the same as the environment state, but the specific values are different.
[0102] In the pre-built Dec-POMDP model, actions and states are closely related. Since actions in a single domain may not be sufficient to effectively resist malicious interference sources, we adopt a multi-domain joint anti-interference action strategy. In practice, due to the power limitations of UAVs, we also consider transmission power in the model. Specifically, the joint action for each time slot is represented as follows: Among them, f i t∈{1,2,…,L} represents the channel selection action, that is, the frequency selection for each UAV. This represents the power adjustment actions of the drone within time slot t, i.e., the power selection corresponding to each drone. The space between the frequency and power selections made by the drone is defined as the action space. A reward function can be used to reflect the quality of the joint actions performed by the drone, which can be expressed by the following tenth formula:
[0103]
[0104] Here, ∈ and ρ are two system parameters.
[0105] The Dec-POMDP model is constructed based on the defined state space, observation space, action space, and reward function.
[0106] For example, when constructing the first interference source population, multiple interference sources can be initialized, and each interference source can be assigned a corresponding identification code. The identification code can characterize the interference strategy of the interference source, and different interference strategies will generate different trajectories during the interference process. The multiple initialized interference sources are combined into an initialized interference source population. The first interference source population can be obtained by setting a distance function and updating the initialized interference source population based on a first algorithm and the distance function.
[0107] For example, in constructing the MARL anti-jamming algorithm, the algorithm aims to take cooperative anti-jamming actions to maximize network throughput while minimizing hop count and energy consumption in the presence of malicious interference attacks. Therefore, information interaction can be limited to neighboring agents within a one-hop range. In the UAV communication system, each agent on the UAV possesses its own Q-function. During the training phase of the Dec-POMDP model, each individual's Q-function Q... i The i∈[1,…,N] form the global Q-function, which promotes cooperation among agents. During the execution phase of the Dec-POMDP model, control message interaction between agents can be introduced to encourage collaboration among agents and improve overall anti-interference performance based on their joint experience.
[0108] The entire MARL anti-interference algorithm can be composed of an observation encoder, convolutional layers, and a Q-network. Figure 5 This is a schematic diagram of the MARL anti-interference algorithm. Figure 5 It can be seen that the observation encoder is based on the local observation of the spectral environment by the UAV in the observation space. As input, it outputs a low-dimensional feature vector. Transforming local observations of the spectral environment into low-dimensional feature vectors can shorten the size of transmitted information and improve transmission efficiency. We can use a multi-layer infectedr as the observation encoder, and the output of the observation encoder is:
[0109]
[0110] in and Let represent the weights and biases of the linear network, respectively, and d be the dimension of the observation encoder.
[0111] After obtaining all low-dimensional feature vectors output by the observation encoder, these vectors are merged into the graph convolutional layer. A multi-head dot product attention mechanism can be used as the convolution kernel to extract the spectral observation relationships between agents and suppress input data redundancy. By stacking more convolutional layers, we can extract higher-order relationships and promote cooperative anti-interference behavior. That is, through a single convolutional layer, agent i can directly obtain the encoded spectral observations of agents within a one-hop range. By stacking two convolutional layers, agent i can obtain the outputs of the upper convolutional layers of other agents within a one-hop range, which contain spectral information from agents within a two-hop range. However, regardless of the number of convolutional layers used, agent i only needs to communicate with its one-hop neighbors. Taking agents i and j as an example, the importance of the information exchanged between the two agents can be calculated using the following eleventh formula:
[0112]
[0113] Where d Ki It is K i The dimensions. Each agent can perform the following calculations using the twelfth formula to obtain the query Γ, key K, and value V:
[0114] Γ i =W Γ h i ,K i =W K h i V i =W V h i
[0115] Where W is a learnable weight matrix.
[0116] Subsequently, for each attention head, we perform a weighted combination of the information from the adversarial agents, ultimately obtaining a weighted sum. The weighted sum for each agent can be expressed as:
[0117]
[0118] Here, σ can be a ReLU function. Weighted sums can broaden the agent's field of vision by incorporating information from neighboring agents in the network environment. Specifically, since drones and interference sources may be in different relative positions, each drone may have different spectral observations. To capture neighboring features in other subspaces and improve training stability, single-head attention can be extended to a multi-head structure, and the output of the convolutional layer can be recalculated as follows:
[0119]
[0120] Where H represents the number of attention heads, h′ i This can be expressed as a weighted sum of the features of neighboring agents, with h′ corresponding to the drone agent. i The larger the value, the more important the information that the drone sends to other drones. Each agent connects the outputs of all attention heads and then feeds them back into the Q-network.
[0121] The Q-network can determine the importance of the information emitted by each UAV based on the weighted sum of each UAV and the interference source population, and give the joint action selection of each UAV, including the frequency and power of communication with other UAVs.
[0122] This example demonstrates how to obtain the communication frequency and power selected by each UAV by constructing a first interference source population, using the MARL anti-jamming algorithm, and building a Dec-POMDP model.
[0123] Optional, Figure 6 A flowchart illustrating another anti-interference method provided in Embodiment 1 of this application is shown below. Figure 6 As shown, in step 302, updating the initial interference source population based on the first algorithm to obtain the first interference source population includes:
[0124] Step 601: Randomly select a first initial interference source instance from the initial interference source population, and input the first initial interference source instance and the communication model into the Dec-POMDP model, so that the first initial interference source instance performs adversarial training with the communication model based on the first identification code of the first initial interference source instance, and obtain a first evolved interference source instance with a second identification code based on the reward function in the Dec-POMDP model;
[0125] Step 602: Compare the difference between the second identification code and the identification codes of all initial interference source instances in the initial interference source population with a preset first threshold.
[0126] Step 603: If the difference between the second identification code and the identification code of any initial interference source instance is less than or equal to the first threshold, then ignore the first evolving interference source instance.
[0127] Step 604: If the difference between the second identification code and the identification codes of all initial interference source instances is greater than the first threshold, then the first evolved interference source instance is added to the initial interference source population until all initial interference source instances in the initial interference source population are input into the Dec-POMDP model to obtain the first interference source population.
[0128] In a scenario example, from the constructed initial interference source population, an initial interference source instance is arbitrarily selected as the first initial interference source instance. This first initial interference source instance and the communication model are then input into the Dec-POMDP model. During the interference process of the first initial interference source instance on the communication model, the communication model, based on the interference mode of the first identification code, utilizes different anti-interference strategies to generate distinctly different trajectories. The generated trajectories are input into a trajectory encoder, and the identification code of the first evolved interference source instance, derived from the first initial interference source instance, is obtained using the following thirteenth formula. The first initial interference source instance is designated as j, and the first evolved interference source instance as j'.
[0129]
[0130] Where f η The identifier of the first initial interference source instance j, represented by the η-parameterized coding network, can be equivalently represented as z. j The identification code of the first evolutionary interference source instance j' can be equivalently represented as z. j′ Define a distance function D(π) j ,π j′ The degree of difference between interference sources j and j' is represented by ), and the distance between the first initial interference source instance and the first evolved interference source instance can be defined as:
[0131] D(π j ,π j′ )=‖z j -z j′ ||
[0132] A first threshold is set for the distance function. The distance between the identifier of the first evolved interference source instance and the identifiers of all initial interference source instances in the initial interference source population is calculated according to the defined distance function. If the distance between the first initial interference source instance and the identifiers of all initial interference source instances in the initial interference source population is greater than the first threshold, the first evolved interference source instance is added to the initial interference source population. If the distance between the first initial interference source instance and the identifiers of all initial interference source instances in the initial interference source population is not greater than the first threshold (i.e., at least one initial interference source instance has a distance between its identifier and the first evolved interference source instance that is less than or equal to the first threshold), the first evolved interference source instance is ignored. According to the method described above, the evolved interference source instance of each interference source instance in the initial interference source population is obtained, and the distance between each interference source instance and the identifier code of each initial interference source instance is compared in turn. Evolved interference source instances whose distances to the identifier codes of each initial interference source instance are all greater than the first threshold are added to the initial interference source population, and otherwise the evolved interference source instances are ignored, so as to complete the update of the initial interference source population and obtain the first interference source population.
[0133] Optionally, when designing a trajectory encoder network, the most important capability is to identify specific features within a trajectory. The encoder network architecture can be designed as a Transformer model, where the attention mechanism can extract features from the input. A corresponding decoder network is proposed. Figure 7 This is a schematic diagram of the population update of the interference source, such as... Figure 7 As shown, the identification code z in a trajectory faced by interference source j. j and observed values As input, it outputs the predicted observations for the next moment of this trajectory. Finally, the decoder network parameters ζ and encoder network parameters v can be updated by minimizing the following loss function, in conjunction with the following fourteenth formula:
[0134]
[0135]
[0136] Among them, f ζ This represents the decoder network. Following this line of thought, we define the performance of the interference source j as follows:
[0137]
[0138] Where R k Let $\frac{k}{k}$ represent the average reward of our proposed system when using the $k$-th trajectory.
[0139] In this example, by updating the initial interference source population, a first interference source population with diversity is obtained.
[0140] Optional, Figure 8 A flowchart illustrating another anti-interference method provided in Embodiment 1 of this application is shown below. Figure 8 As shown, in step 202, the iterative training of the Dec-POMDP model based on the results of adversarial training at different times until training is completed includes:
[0141] Step 801: Add a loss function to the MARL anti-interference algorithm and set a threshold for the number of iterations during iterative training;
[0142] Step 802: Use the results of the adversarial training at different times as training samples to iteratively train the Dec-POMDP model, and record the value of the loss function output corresponding to each iteration. Stop iterative training when the value of the loss function output is less than a preset value, or when the number of iterations reaches the iteration threshold.
[0143] Based on the scenario example, a deep parallel Q-network with a memory playback mechanism can be designed in the aforementioned parallel Q-network. Figure 9 This is a schematic diagram of a parallel Q-network, as shown below. Figure 9 As shown, during the training of the Dec-POMDP model, tuples can be... Stored in buffer D, where o′ is the next observation. This is the adjacency matrix describing the connection relationships of agent i. The time step t can be omitted. When the buffer size reaches S, the anti-interference algorithm will use uniformly sampled mini-batch data from D to initiate iterative training of the Dec-POMDP model by the Q network. The network parameters θ in the Q network can be updated to guide the UAV's joint actions on frequency and power. Specifically, the target value can first be calculated using the following formula (fifteenth equation):
[0144]
[0145] The target Q-network is composed of parameter θ - "Determined" represents the estimated value of the current parameter θ, where γ is the discount factor. The set of observations within the receiving range of agent i is represented by the adjacency matrix. This matrix indicates the combined observations within the neighborhood. A loss function can be introduced, and the goal of iterative training of the Dec-POMDP model is to minimize this loss function. Specifically, the loss function can be expressed as:
[0146]
[0147] Where 's' refers to the size of the mini-batch. The parameter θ can be updated using gradient descent using the following formula (equation sixteen):
[0148]
[0149] Furthermore, the target network parameters θ - It is done through the following seventeenth formula, per N r Update once per step:
[0150] θ - =βθ+(1-β)θ -
[0151] Set a minimum threshold for the output value of the loss function and a threshold for the number of iterations. When the output of the loss function is less than the set minimum threshold, or when the number of iterations for training the Dec-POMDP model is greater than the threshold, stop the iterative training of the Dec-POMDP model.
[0152] This application introduces a loss function to monitor the iterative training of the Dec-POMDP model, making the output of the trained Dec-POMDP model more accurate.
[0153] Optional, Figure 10 A flowchart illustrating another anti-interference method provided in Embodiment 1 of this application is shown below. Figure 10 As shown, in step 203, optimizing the output of the Dec-POMDP model based on the information time-series smoothing mechanism includes:
[0154] Step 1001: Assign an individual Q function to each UAV in the UAV communication system, and add the individual Q functions assigned to each UAV to obtain the global Q function of the UAV communication system, so that each UAV can generate shared information based on the global Q function;
[0155] Step 1002: Create a timing smoothing mechanism, which consists of a message encoder, a combination block, a send message buffer, and a receive message buffer;
[0156] Step 1003: Configure the UAV communication system based on the timing smoothing mechanism. The configuration includes: when information interaction occurs between UAVs, the UAV that performs information transmission generates a message to be transmitted based on the message encoder and places the message to be transmitted in the transmission message buffer; and when accessing an available channel, it sends the message to be transmitted to an adjacent UAV. The UAV that performs message reception stores the received messages in the reception message buffer and marks each message in the reception message buffer with a valid bit. When messages with the same valid bit no longer change, all information with that valid bit is sent to the combination block so that the combination block generates a corresponding joint Q value based on all information with that valid bit.
[0157] Step 1004: Apply the joint Q value to the Dec-POMDP model so that the Dec-POMDP model, in conjunction with the joint Q value, optimizes the target frequency and power of the communication between the UAVs output by the Dec-POMDP model.
[0158] With a scenario example, when applying the Dec-POMDP model to a real-world UAV communication system, an information timing smoothing mechanism can be used to optimize the output of the Dec-POMDP model. Specifically, an individual Q-function can be assigned to each agent on the UAV, and the various Q-functions can be summed to obtain a global Q-function, which is expressed as:
[0159]
[0160] Q i Generated by a parallel Q-network, therefore, the content of the shared information can be represented as x. i =(h i ,h i ′,h i ",Q i ).
[0161] Figure 11 This is a schematic diagram of the message smoothing mechanism. Figure 11 It is understood that the message smoothing mechanism consists of a message encoder, a combining block, a sending buffer, and a received buffer. When information exchange occurs between drones, the agent of the drone executing the message sending is based on the x input from the message encoder. i Generate message m i =f msg (x iThe message is then sent to the send message buffer and broadcast to neighboring drone agents when accessing an available channel. The receive message buffer of the receiving drone stores historically received messages and marks each message with a valid bit, indicating whether the received message has expired. Specifically, when agent i receives a message from its neighboring drone agent in time slot t, it updates its receive buffer and selects a message that conforms to {m}. i The message m of |i∈N,val(i=1} i continuous t s If a time slot is not updated, its valid bit is set to 0, indicating that the message has expired. These selected messages m... i Along with the individual's local Q-value (Q loc Together, they are passed to the combining block to produce the joint Q-value (Q). tot To guide the agent in outputting actions Where, m i [Q] and Q loc The dimensions are the same. The combinatorial block kernel only needs to perform element-wise addition, and the joint Q value can be obtained using the following formula (Equation 18):
[0162]
[0163] Where m i [Q] represents the Q-value portion of the message.
[0164] Introducing the obtained joint Q value into the Dec-POMDP model can better optimize the selection of frequency and power for the UAV agent.
[0165] The intelligent game-theoretic network anti-interference method provided in this embodiment first constructs a network topology model of an unmanned aerial vehicle (UAV) communication system. Based on an aerial computing network, it calculates the communication parameters between each UAV and other UAVs in the network topology model, obtaining a communication model among multiple UAVs. Then, it applies the MARL anti-interference algorithm to the Dec-POMDP model, inputting a first interference source population and the communication model into the Dec-POMDP model for adversarial training. The training results at different times are obtained, and the Dec-POMDP model is iteratively trained until training is complete. Finally, the trained Dec-POMDP model is applied to the UAV communication system, and its output is optimized based on an information time-series smoothing mechanism to obtain the target frequency and power for communication between each UAV and other UAVs in the presence of malicious interference. This scheme, by constructing a Dec-POMDP model with joint frequency selection and power control, applying it to a practical UAV communication system, and optimizing the output of the Dec-POMDP model, can control each UAV in the UAV communication system to select the optimal communication frequency and power under interference attacks, improving the ability and efficiency of ensuring normal UAV communication.
[0166] Example 2
[0167] Figure 12 An anti-interference device provided in Embodiment 2 of this application, such as Figure 12 As shown, the device includes:
[0168] Module 121 is used to construct a network topology model of the UAV communication system. Based on the airborne computing network, it calculates the communication parameters between each UAV and other UAVs in the network topology model to obtain a communication model between multiple UAVs. The network topology model reflects the real-time relative positions and communication interactions between multiple UAVs in the UAV communication system.
[0169] The construction module 121 is also used to construct a Dec-POMDP model, construct a first interference source population, and construct a MARL anti-interference algorithm. The MARL anti-interference algorithm is applied to the Dec-POMDP model, and the first interference source population and the communication model are input into the Dec-POMDP model for adversarial training. Based on the MARL anti-interference algorithm, the adversarial training results at different times are obtained. The Dec-POMDP model is iteratively trained based on the adversarial training results at different times until training is complete. The adversarial training results include the frequency and power of communication between each UAV and other UAVs.
[0170] The processing module 122 is used to apply the trained Dec-POMDP model to the UAV communication system and optimize the output of the Dec-POMDP model based on the information time-series smoothing mechanism to obtain the target frequency and power of communication between each UAV and other UAVs in the UAV communication system under malicious interference.
[0171] Based on scenario examples, Module 121 establishes a network topology model for UAVs based on the relative positions and communication interactions between UAVs in the UAV communication system. To consider the mobility of UAVs, the network topology is abstracted as a dynamic graph that changes over time, with vertices representing interference-resistant UAVs. The communication parameters for communication between UAVs, as well as the interference signals received by each UAV, can be calculated through an aerial computing network. Specifically, assuming that each UAV needs to send information to other UAVs, this can be coordinated by offloading tasks from mobile phones on IoT devices or by intelligent agents calculating completed coordination information. Time can be divided into fixed intervals of equal length. A specific time slot t is selected, and the intelligent agent on the UAV in the previous time slot is used to calculate the frequency and power selected by the intelligent agent on the UAV based on spectrum observation. Based on the various parameters for communication between UAVs obtained from the above formulas, including: the actual transmission power of the UAV, possible attack interference, potential interference, received interference attack signals, signal-to-noise ratio, transmission rate, communication success rate, average transmission power, and average hop frequency, Module 121 establishes a communication model for the UAV communication system. Multiple high-quality interference sources with diversity are designed through an evolutionary and selection process. These multiple high-quality interference sources are constructed into a first interference source population, which is then applied to the communication model to interfere with the drones within the model. Furthermore, a graph convolution-based MARL anti-interference algorithm can be constructed to calculate the frequency and power selection for communication between drones when facing interference from the first interference source population. A Dec-POMDP model is constructed, and the first interference source population and the communication model are input into the Dec-POMDP model. The MARL anti-interference algorithm is then applied to the Dec-POMDP model, allowing the frequency and power selections obtained by the drones when facing interference from the first interference source population to serve as training data for training the Dec-POMDP model. This allows the trained Dec-POMDP model to control the drones to select appropriate transmission frequencies and powers based on the infection information from the interference sources.
[0172] The processing module 122 deploys the trained Dec-POMDP model on the intelligent agent of the drone in the actual drone communication system, and optimizes the Dec-POMDP model for the drone communication system used in actual applications, so that the Dec-POMDP model can provide more accurate transmission frequency and transmission power when the drone communication system faces interference from interference sources.
[0173] Optionally, module 121 is used to define the state space, observation space, action space and reward function, and to construct the Dec-POMDP model based on the defined state space, observation space, action space and reward function;
[0174] The construction module 121 is further used to initialize multiple interference source instances, randomly assign an identification code to each initialized interference source instance, combine multiple initialized interference source instances with identification codes into an initial interference source population, and update the initial interference source population based on the first algorithm to obtain the first interference source population.
[0175] Module 121 is further used to design an observation encoder, a graph convolutional layer and a Q network, and to integrate the observation encoder, the graph convolutional layer and the Q network into the MARL anti-interference algorithm;
[0176] The observation encoder is used to encode the observation data of the local spectrum environment of the state space observed by each UAV in the observation space in the UAV communication system into a low-dimensional feature vector, and send all the low-dimensional feature vectors to the convolutional layer. The convolutional layer obtains the weighted sum of each UAV and the interference source population based on the obtained low-dimensional feature vectors, and transmits the weighted sum of each UAV and the interference source population to the Q network. The Q network can obtain the frequency and power of communication between each UAV and other UAVs based on the weighted sum of each UAV and the interference source population.
[0177] In the pre-built Dec-POMDP model, actions and states are closely related, illustrated by a scenario example. Since actions in a single domain may not be sufficient to effectively resist malicious interference sources, we adopt a multi-domain joint anti-interference action strategy. For instance, when constructing the first interference source population, multiple interference sources can be initialized, and each source can be assigned a corresponding identification code. This identification code characterizes the interference strategy of the interference source, and different interference strategies will generate different trajectories during the interference process.
[0178] For example, the entire MARL anti-interference algorithm can be composed of an observation encoder, a convolutional layer, and a Q-network. The observation encoder takes the local observations of the spectrum environment observed by the UAV in the observation space as input and outputs a low-dimensional feature vector. Transforming the local observations of the spectrum environment into a low-dimensional feature vector can shorten the size of the transmitted information and make the transmission more efficient.
[0179] After obtaining all low-dimensional feature vectors output by the observation encoder, these vectors are merged into the graph convolutional layer, using a multi-head dot product attention mechanism as the convolution kernel. By stacking more convolutional layers, higher-order relationships can be extracted, promoting cooperative anti-interference behavior. Subsequently, for each attention head, the information of the anti-interference agent is weighted and combined to obtain a weighted sum. This weighted sum can expand the agent's field of vision by incorporating information from neighboring agents in the network environment. The Q-network can determine the importance of the information emitted by each drone based on the weighted sum of each drone and the interference source population, and provide joint actions and choices for each drone, including the frequency and power of communication with other drones. Each agent concatenates the outputs of all attention heads and feeds them back into the Q-network. The Q-network can determine the importance of the information emitted by each drone based on the weighted sum of each drone and the interference source population, and provide joint actions and choices for each drone, including the frequency and power of communication with other drones.
[0180] Optionally, the construction module 121 is further configured to arbitrarily select a first initial interference source instance from the initial interference source population, and input the first initial interference source instance and the communication model into the Dec-POMDP model, so that the first initial interference source instance performs adversarial training with the communication model based on the first identification code of the first initial interference source instance, and obtains a first evolved interference source instance with a second identification code generated by the first initial interference source instance based on the reward function in the Dec-POMDP model;
[0181] The construction module 121 is further configured to compare the difference between the second identification code and the identification codes of all initial interference source instances in the initial interference source population with a preset first threshold.
[0182] The construction module 121 is further configured to ignore the first evolving interference source instance if the difference between the second identification code and the identification code of any initial interference source instance is less than or equal to the first threshold.
[0183] The construction module 121 is further configured to add the first evolved interference source instance to the initial interference source population if the difference between the second identification code and the identification codes of all initial interference source instances is greater than the first threshold, until all initial interference source instances in the initial interference source population are input into the Dec-POMDP model to obtain the first interference source population.
[0184] Based on a scenario example, module 121 arbitrarily selects one initial interference source instance from the constructed initial interference source population as the first initial interference source instance, and inputs the first initial interference source instance and the communication model into the Dec-POMDP model. During the interference process of the first initial interference source instance on the communication model, the communication model will use different anti-interference strategies based on the interference mode of the first identification code, thereby generating distinctly different trajectories, and inputting the generated trajectories into a trajectory encoder. The construction module 121 sets a first threshold for the distance function. According to the defined distance function, it calculates the distance between the identification code of the first evolved interference source instance and the identification codes of all initial interference source instances in the initial interference source population. If the distance between the first initial interference source instance and the identification codes of all initial interference source instances in the initial interference source population is greater than the first threshold, then the first evolved interference source instance is added to the initial interference source population. If the distance between the first initial interference source instance and the identification codes of all initial interference source instances in the initial interference source population is not greater than the first threshold, that is, at least one initial interference source instance has a distance between its identification code and the first evolved interference source instance that is less than or equal to the first threshold, then the first evolved interference source instance is ignored. Following the method described above, evolved interference source instances are obtained for each interference source instance in the initial interference source population. The distance between each interference source instance and the identifier of each initial interference source instance is compared sequentially. Evolved interference source instances whose distances to each initial interference source instance identifier are greater than the first threshold are added to the initial interference source population; otherwise, they are ignored. This process updates the initial interference source population to obtain the first interference source population. In this example, by updating the initial interference source population, a first interference source population with diversity is obtained.
[0185] Optionally, module 121 is further used to add a loss function to the MARL anti-interference algorithm and set a threshold for the number of iterations of iterative training.
[0186] The construction module 121 is further used to take the results of the adversarial training at different times as training samples in sequence, to iteratively train the Dec-POMDP model, and to record the value of the loss function output corresponding to each iteration training. When the value of the loss function output is less than a preset value, or when the number of iterations reaches the iteration number threshold, the iterative training is stopped.
[0187] In the parallel Q-network, the construction module 121 can introduce a loss function. The goal of iterative training of the Dec-POMDP model is to minimize the loss function. A minimum threshold for the output value of the loss function and a threshold for the number of iterations are set. When the output of the loss function is less than the set minimum threshold, or when the number of iterations of the Dec-POMDP model exceeds the threshold, the iterative training of the Dec-POMDP model is stopped.
[0188] Optionally, the processing module 122 is specifically used to assign an individual Q function to each UAV in the UAV communication system, and to add the individual Q functions assigned to each UAV to obtain the global Q function of the UAV communication system, so that each UAV generates shared information based on the global Q function;
[0189] The processing module 122 is further used to create a timing smoothing mechanism, which consists of a message encoder, a combination block, a sending message buffer, and a receiving message buffer.
[0190] The processing module 122 is further configured to configure the UAV communication system based on the timing smoothing mechanism. The configuration includes: when information interaction occurs between UAVs, the UAV performing information transmission generates a message to be transmitted based on the message encoder and places the message to be transmitted in the transmission message buffer, and sends the message to be transmitted to neighboring UAVs when accessing an available channel; the UAV performing message reception stores the received messages in the reception message buffer and marks each message in the reception message buffer with a valid bit. When messages with the same valid bit no longer change, all information with that valid bit is sent to the combination block, so that the combination block generates a corresponding joint Q value based on all information with that valid bit.
[0191] The processing module 122 is further configured to apply the joint Q value to the Dec-POMDP model, so that the Dec-POMDP model can optimize the target frequency and power of the communication between the UAVs output by the Dec-POMDP model in combination with the joint Q value.
[0192] In a scenario example, when applying the Dec-POMDP model to a real-world UAV communication system, an information timing smoothing mechanism can be used to optimize the model's output. When UAVs interact, the agent of the UAV sending the message generates a message based on the message encoder and sends it to the message sending buffer. Upon accessing an available channel, it broadcasts the message to the agents of neighboring UAVs. The message receiving buffer of the agent of the UAV receiving the message stores historically received messages and marks each message with a valid bit, indicating whether the received message has expired. Introducing the obtained joint Q-value into the Dec-POMDP model can better optimize the selection of frequency and power for the UAV agents.
[0193] The intelligent game-theoretic network anti-interference device provided in this embodiment first constructs a network topology model of the UAV communication system. Based on the aerial computing network, it calculates the communication parameters between each UAV and other UAVs in the network topology model to obtain a communication model between multiple UAVs. Then, it applies the MARL anti-interference algorithm to the Dec-POMDP model and inputs the first interference source population and the communication model into the Dec-POMDP model for adversarial training to obtain the adversarial training results at different times. Based on the adversarial training results at different times, it iteratively trains the Dec-POMDP model until training is completed. Finally, the processing module applies the trained Dec-POMDP model to the UAV communication system and optimizes the output of the Dec-POMDP model based on the information time-series smoothing mechanism to obtain the target frequency and power of communication between each UAV and other UAVs in the UAV communication system under malicious interference. This solution constructs a Dec-POMDP model for joint frequency selection and power control, applies it to a practical UAV communication system, and optimizes the output of the Dec-POMDP model. This allows each UAV in the UAV communication system to select the optimal communication frequency and power under the attack of interference sources, thereby improving the ability and efficiency of ensuring normal communication of UAVs.
[0194] Example 3
[0195] Figure 13 This is a schematic diagram of the structure of an electronic device provided in Embodiment 3 of this application, as shown below. Figure 13 As shown, the electronic device includes:
[0196] The server includes a processor 291 and a memory 292; it may also include a communication interface 293 and a bus 294. The processor 291, memory 292, and communication interface 293 can communicate with each other via the bus 294. The communication interface 293 can be used for information transmission. The processor 291 can call logical instructions stored in the memory 292 to execute the method described in Embodiment 1.
[0197] Furthermore, the logic instructions in the aforementioned memory 292 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium.
[0198] The memory 292, as a computer-readable storage medium, can be used to store software programs and computer-executable programs, such as program instructions / modules corresponding to the methods in the embodiments of this application. The processor 291 executes functional applications and data processing by running the software programs, instructions, and modules stored in the memory 292, that is, implementing the method of Embodiment 1 described above.
[0199] The memory 292 may include a program storage area and a data storage area. The program storage area may store the operating system and application programs required for at least one function; the data storage area may store data created based on the use of the terminal device. Furthermore, the memory 292 may include high-speed random access memory and may also include non-volatile memory.
[0200] This application provides a non-transitory computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the methods described in the foregoing embodiments.
[0201] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the invention filed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not claimed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this application are indicated by the following claims.
[0202] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.
Claims
1. An anti-interference method, characterized in that, The method includes: A network topology model of an unmanned aerial vehicle (UAV) communication system is constructed. Based on an aerial computing network, communication parameters for communication between each UAV and other UAVs in the network topology model are calculated to obtain a communication model between multiple UAVs. The network topology model reflects the real-time relative positions and communication interactions between multiple UAVs in the UAV communication system. A Dec-POMDP model is constructed, a first interference source population is constructed, and a MARL anti-jamming algorithm is constructed. The MARL anti-jamming algorithm is applied to the Dec-POMDP model, and the first interference source population and the communication model are input into the Dec-POMDP model for adversarial training. The adversarial training results at different time points are obtained based on the MARL anti-jamming algorithm. The Dec-POMDP model is iteratively trained based on the adversarial training results at different time points until training is completed. The adversarial training results include the frequency and power of communication between each UAV and other UAVs. The construction process of the MARL anti-jamming algorithm includes: An observation encoder, a graph convolutional layer, and a Q-network are designed and integrated into the MARL anti-jamming algorithm. The observation encoder encodes the observation data of each UAV in the observation space of the local spectral environment of the state space observed in the UAV communication system into low-dimensional feature vectors, and sends all low-dimensional feature vectors to the convolutional layer. The convolutional layer calculates a weighted sum of each UAV and the interference source population based on the acquired low-dimensional feature vectors, and transmits this weighted sum to the Q-network. The Q-network can then calculate the frequency and power of communication between each UAV and other UAVs based on the weighted sum of each UAV and the interference source population. A timing smoothing mechanism is created, which consists of a message encoder, a combination block, a send message buffer, and a receive message buffer. The UAV communication system is configured based on the aforementioned timing smoothing mechanism. The configuration includes: when information interaction occurs between UAVs, the UAV that performs information transmission generates a message to be transmitted based on the message encoder and places the message to be transmitted in the transmission message buffer; and when accessing an available channel, it sends the message to be transmitted to an adjacent UAV. The UAV that performs message reception stores the received messages in the reception message buffer and marks each message in the reception message buffer with a valid bit. When messages with the same valid bit no longer change, all information with that valid bit is sent to the combination block, so that the combination block generates a corresponding joint Q value based on all information with that valid bit. The trained Dec-POMDP model is applied to the UAV communication system, and the output of the Dec-POMDP model is optimized based on the time-series smoothing mechanism to obtain the target frequency and power for communication between each UAV and other UAVs in the UAV communication system under malicious interference.
2. The method according to claim 1, characterized in that, The construction of the Dec-POMDP model includes: Define the state space, observation space, action space, and reward function, and construct the Dec-POMDP model based on the defined state space, observation space, action space, and reward function; The construction of the first interference source population includes: Multiple interference source instances are initialized, and each initialized interference source instance is randomly assigned a corresponding identification code. The multiple initialized interference source instances with identification codes are combined into an initial interference source population, and the initial interference source population is updated based on a first algorithm to obtain the first interference source population.
3. The method according to claim 2, characterized in that, The step of updating the initial interference source population based on the first algorithm to obtain the first interference source population includes: A first initial interference source instance is randomly selected from the initial interference source population, and the first initial interference source instance and the communication model are input into the Dec-POMDP model, so that the first initial interference source instance performs adversarial training with the communication model based on the first identification code of the first initial interference source instance, and based on the reward function in the Dec-POMDP model, a first evolved interference source instance with a second identification code is obtained from the first initial interference source instance. The difference between the second identification code and the identification codes of all initial interference source instances in the initial interference source population is compared with a preset first threshold. If the difference between the second identification code and the identification code of any initial interference source instance is less than or equal to the first threshold, then the first evolved interference source instance is ignored. If the difference between the second identification code and the identification codes of all initial interference source instances is greater than the first threshold, then the first evolved interference source instance is added to the initial interference source population until all initial interference source instances in the initial interference source population are input into the Dec-POMDP model to obtain the first interference source population.
4. The method according to claim 1, characterized in that, The iterative training of the Dec-POMDP model based on the results of adversarial training at different time points until training is complete includes: Add a loss function to the MARL anti-interference algorithm and set a threshold for the number of iterations during iterative training; The results of adversarial training at different times are used as training samples in sequence to iteratively train the Dec-POMDP model. The value of the loss function output corresponding to each iteration is recorded. Iterative training stops when the value of the loss function output is less than a preset value or when the number of iterations reaches the iteration threshold.
5. The method according to claim 1, characterized in that, The optimization of the output of the Dec-POMDP model based on the time-series smoothing mechanism includes: Each UAV in the UAV communication system is assigned an individual Q-function, and the individual Q-functions assigned to each UAV are summed to obtain the global Q-function of the UAV communication system, so that each UAV can generate shared information based on the global Q-function; The joint Q-value is applied to the Dec-POMDP model so that the Dec-POMDP model, in conjunction with the joint Q-value, optimizes the target frequency and power of the communication between the UAVs output by the Dec-POMDP model.
6. An anti-interference device, characterized in that, The device includes: The construction module is used to construct a network topology model of the UAV communication system. Based on the airborne computing network, it calculates the communication parameters between each UAV and other UAVs in the network topology model to obtain a communication model between multiple UAVs. The network topology model reflects the real-time relative positions and communication interactions between multiple UAVs in the UAV communication system. The construction module is also used to construct a Dec-POMDP model, construct a first interference source population, and construct a MARL anti-jamming algorithm. The MARL anti-jamming algorithm is applied to the Dec-POMDP model, and the first interference source population and the communication model are input into the Dec-POMDP model for adversarial training. Based on the MARL anti-jamming algorithm, the adversarial training results at different times are obtained. The Dec-POMDP model is iteratively trained based on the adversarial training results at different times until training is complete. The adversarial training results include the frequency and power of communication between each UAV and other UAVs. The construction process of the MARL anti-jamming algorithm includes: The construction module is also used to design an observation encoder, a graph convolutional layer, and a Q-network, and to integrate the observation encoder, graph convolutional layer, and Q-network into the MARL anti-interference algorithm. Specifically, the observation encoder is used to encode the observation data of each UAV in the observation space of the local spectrum environment of the state space observed in the UAV communication system into low-dimensional feature vectors, and then sends all the low-dimensional feature vectors to the convolutional layer. The convolutional layer obtains a weighted sum of each UAV and the interference source population based on the obtained low-dimensional feature vectors, and transmits the weighted sum of each UAV and the interference source population to the Q-network. The Q-network can obtain the frequency and power of communication between each UAV and other UAVs based on the weighted sum of each UAV and the interference source population. The building module is also used to create a timing smoothing mechanism, which consists of a message encoder, a combination block, a send message buffer, and a receive message buffer. The construction module is also used to configure the UAV communication system based on the timing smoothing mechanism. The configuration includes: when information interaction occurs between UAVs, the UAV that performs information transmission generates a message to be transmitted based on the message encoder and places the message to be transmitted in the transmission message buffer, and sends the message to be transmitted to neighboring UAVs when accessing an available channel; the UAV that performs message reception stores the received messages in the reception message buffer and marks each message in the reception message buffer with a valid bit. When messages with the same valid bit no longer change, all information with that valid bit is sent to the combination block, so that the combination block generates a corresponding joint Q value based on all information with that valid bit. The processing module is used to apply the trained Dec-POMDP model to the UAV communication system and optimize the output of the Dec-POMDP model based on the time-series smoothing mechanism to obtain the target frequency and power of communication between each UAV and other UAVs in the UAV communication system under malicious interference.
7. The apparatus according to claim 6, characterized in that, The construction module is specifically used to define the state space, observation space, action space, and reward function, and to construct the Dec-POMDP model based on the defined state space, observation space, action space, and reward function. The construction module is further configured to initialize multiple interference source instances, randomly assign an identification code to each initialized interference source instance, combine multiple initialized interference source instances with identification codes into an initial interference source population, and update the initial interference source population based on a first algorithm to obtain the first interference source population.
8. The apparatus according to claim 7, characterized in that, The construction module is further configured to arbitrarily select a first initial interference source instance from the initial interference source population, and input the first initial interference source instance and the communication model into the Dec-POMDP model, so that the first initial interference source instance performs adversarial training with the communication model based on the first identification code of the first initial interference source instance, and obtains a first evolved interference source instance with a second identification code generated by the first initial interference source instance based on the reward function in the Dec-POMDP model; The construction module is further configured to compare the difference between the second identification code and the identification codes of all initial interference source instances in the initial interference source population with a preset first threshold. The construction module is further configured to ignore the first evolving interference source instance if the difference between the second identification code and the identification code of any initial interference source instance is less than or equal to the first threshold. The construction module is further configured to add the first evolved interference source instance to the initial interference source population if the difference between the second identification code and the identification codes of all initial interference source instances is greater than the first threshold, until all initial interference source instances in the initial interference source population are input into the Dec-POMDP model to obtain the first interference source population.
9. An electronic device, characterized in that, include: A processor, and a memory communicatively connected to the processor; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory to implement the method as described in any one of claims 1-5.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the method as described in any one of claims 1-5.