Anti-interference unmanned aerial vehicle scheduling method, device, equipment, storage medium and program product
By employing multi-agent deep reinforcement learning and a three-layer air-ground collaborative architecture, intelligent scheduling and spectrum resource management of UAVs in the air-ground integrated network are realized. This solves the problems of interference and limited computing power of UAVs in complex electromagnetic environments, improves the system's anti-interference capability and spectrum utilization efficiency, and meets the high-efficiency coverage requirements of the Internet of Things.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-17
- Publication Date
- 2026-03-10
AI Technical Summary
In air-to-ground integrated networks, intelligent scheduling and communication resource management of multiple UAVs face challenges such as difficulty in ensuring link stability and interference in complex electromagnetic environments, which affect communication performance and spectrum utilization. Furthermore, the limited computing power of UAVs makes it difficult to achieve global optimization.
By employing a multi-agent deep reinforcement learning framework combined with a UAV wireless communication service model, the system dynamically allocates resources through real-time perception of the spectrum environment, optimizes UAV sub-band allocation and secondary user access strategies, and performs intelligent scheduling using a three-layer air-ground collaborative architecture, thereby achieving dynamic management of spectrum resources and anti-interference capabilities.
It improves the system's anti-interference capability and spectrum utilization efficiency in complex electromagnetic environments, enhances network coverage performance and operational stability, meets the balanced service needs of massive IoT terminals, reduces service gaps for edge users, and improves the system's recovery speed in dynamic scenarios.
Smart Images

Figure CN121645532A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of wireless communication, and in particular to an anti-interference unmanned aerial vehicle scheduling method and device, computer equipment, computer readable storage medium and computer program product. BACKGROUND
[0002] In recent years, information and communication technology continues to evolve at a high speed, and mobile communication networks have made significant breakthroughs in capacity, speed and coverage capabilities. The fifth generation mobile communication technology (5G) has been commercially deployed worldwide, providing a solid foundation for diversified services such as ultra-high speed, low latency and large connectivity. On this basis, the research and exploration of the sixth generation mobile communication technology (6G) has been launched worldwide, aiming to meet the needs of future smart society for global coverage, extreme experience and intelligent services. The academic and industrial communities around the world generally believe that 6G networks will integrate terrestrial wireless networks, satellite networks and near-space unmanned aerial vehicle networks to build an integrated space-ground terrestrial all-scenario communication architecture, achieving truly "signal everywhere, network seamless interconnection".
[0003] Under this background, near-space unmanned aerial vehicle networks are considered as an important bridge connecting the ground and the air due to their flexibility, rapid deployment and dynamic coverage adjustment. A large number of domestic and foreign researches show that the integration of near-space unmanned aerial vehicle networks and ground base station networks can form an air-ground integrated network, which can effectively address the coverage deficiency and capacity bottleneck problems of ground networks in remote areas, disaster emergency, hot spot areas and other scenarios. At the same time, the air-ground integrated network can also realize dynamic networking and flexible resource scheduling according to business needs, with strong adaptability and scalability.
[0004] In the air-ground integrated network, unmanned aerial vehicles equipped with cognitive radio technology can actively sense the surrounding wireless environment, dynamically identify available spectrum resources, and adaptively adjust communication parameters in a changing electromagnetic environment, thereby improving spectrum utilization and anti-interference capability. Such unmanned aerial vehicles not only can quickly switch to available channels when the communication link is interfered, but also can coordinate with other unmanned aerial vehicles and ground network nodes to realize spectrum sharing and intelligent scheduling, playing an important role in improving system capacity and ensuring communication quality.
[0005] However, the current multi-UAV air-ground integrated network still faces many challenges in practical applications. On the one hand, the mobility of UAVs in the airspace and the dynamic changes of network topology make it difficult to guarantee link stability; on the other hand, in a complex electromagnetic environment, factors such as malicious interference, same frequency competition and multiple user access at the same time can significantly reduce communication performance. In addition, the intelligent scheduling of multiple UAVs and the management of communication resources need to consider the quality of service, fairness, energy efficiency and interference suppression at the same time, which puts higher requirements on algorithm design, real-time computing capability and cross-layer cooperation mechanism. Therefore, how to design an adaptive anti-interference, multi-UAV intelligent scheduling and communication resource management method for air-ground integrated network has become a key technical problem that needs to be solved. SUMMARY
[0006] Therefore, it is necessary to provide an anti-interference UAV scheduling method and device, computer equipment, computer readable storage medium and computer program product to realize real-time perception of the radio environment and dynamic allocation of spectrum resources, improve the anti-interference ability and spectrum utilization efficiency of the system in a complex electromagnetic environment, thereby improving the overall coverage performance and operation stability of the network while ensuring the quality of service.
[0007] In a first aspect, the present application provides an anti-interference UAV scheduling method, comprising:
[0008] obtaining local environment information and user state information collected by each cognitive UAV;
[0009] Based on the local environment information and the user state information, the channel gain of each cognitive UAV to each user, the channel gain of each interfering UAV to each user, and the channel gain of each base station to each user in the tth time slot are analyzed by using a pre-constructed UAV wireless communication service model, and the received data rate of each user in each sub-band in the tth time slot is determined.
[0010] Based on the received data rate of each user in each sub-band in the tth time slot, a pre-constructed UAV scheduling and communication resource management model is used to solve and determine the sub-band allocation strategy of each cognitive UAV and the secondary user access strategy; the UAV scheduling and communication resource management model is used to maximize the effective cumulative transmission rate and service fairness under the preset constraint condition;
[0011] The action value is determined according to the observation value by using a multi-agent deep reinforcement learning framework; the observation value includes the position information of each cognitive UAV, the channel gain of each cognitive UAV to each user, and the cumulative service rate of the secondary user, and the action value includes the flight steering angle of each cognitive UAV.
[0012] Scheduling is performed on each cognitive UAV based on sub-frequency band allocation strategy, secondary user access strategy, and flight turning angle.
[0013] In one embodiment, the UAV scheduling and communication resource management model includes a cognitive UAV spectrum resource allocation problem model and a secondary user access problem model; based on the received data rate of each user in each sub-frequency band within the t-th time slot, the pre-built UAV scheduling and communication resource management model is used to solve the problem, determining the sub-frequency band allocation strategy and secondary user access strategy for each cognitive UAV, including:
[0014] Based on a greedy strategy, the model of the spectrum resource allocation problem for cognitive UAVs is solved according to the received data rate of each user in each sub-band within the t-th time slot, and the sub-band allocation strategy for each cognitive UAV is determined.
[0015] Based on the sub-band allocation strategy of each cognitive UAV, determine the channel interference signal-to-noise ratio matrix;
[0016] Based on the branch and bound algorithm, the secondary user access problem model is solved according to the channel interference signal-to-noise ratio matrix to determine the secondary user access strategy.
[0017] In one embodiment, the cognitive unmanned aerial vehicle (UAV) spectrum resource allocation problem model is expressed as follows:
[0018] ;
[0019] in:
[0020] ;
[0021] ;
[0022] The main user set is represented as The set of secondary users is represented as The cognitive drone set is represented as The spectrum resource set is represented as ;
[0023] This represents the sub-band allocation strategy for each cognitive UAV within the t-th time slot. This indicates the time slot t-1. The first cognitive drone to the first The cumulative number of services provided to each user Indicates the first A cognitive drone in the t-th time slot Upwards from the first sub-frequency band Channel interference signal-to-noise ratio for each user's transmitted data;
[0024] The secondary user access problem model is represented as follows:
[0025] ;
[0026] in:
[0027] ;
[0028] ;
[0029] ;
[0030] ;
[0031] This represents the access strategy for secondary users within the t-th time slot. This represents the channel interference signal-to-noise ratio matrix. Indicates that in determining In the case of the first A cognitive drone in the t-th time slot Upwards from the first sub-frequency band The transmission rate of each user Indicates the penalty coefficient. Indicates a penalty item. Indicates the first Minimum total transmission rate requirement for each primary user Used to describe normalization priority The weight representing the access priority of secondary users. Indicates the first The cumulative service rate of each user from the first time slot to the t-th time slot. These are the preset adjustment parameters. Indicates the first The primary user in the t-th time slot... The received data rate of each sub-band.
[0032] In one embodiment, after scheduling each cognitive UAV based on a sub-band allocation strategy, a secondary user access strategy, and a flight turn angle, the method further includes:
[0033] The reward value is determined based on the data reception rate, service fairness ratio, sub-band allocation strategy, and secondary user access strategy of each user in each sub-band within the t-th time slot.
[0034] Observations, action values, and reward values are stored as experience data;
[0035] The parameters of the policy network and evaluation network of the deep reinforcement learning framework are updated periodically based on stored empirical data.
[0036] In one embodiment, the reward value is determined by the following equation:
[0037] ;
[0038] wherein, denotes the frequency band allocated by the i-th cognitive UAV to the j-th user in the t-th time slot, and is denoted as:
[0039] ;
[0040] denotes the transmission rate of the i-th cognitive UAV to the j-th secondary user in the t-th time slot on the i-th sub-band, under the condition that denotes the service fairness ratio, denotes the punishment coefficient, denotes the punishment term, denotes the sub-band allocation strategy value corresponding to the i-th cognitive UAV, denotes the secondary user access strategy value corresponding to the j-th secondary user.
[0041] In one embodiment, the UAV wireless communication service model is denoted as:
[0042] ;
[0043] ;
[0044] denotes the signal-to-noise ratio of the i-th cognitive UAV transmitting data to the j-th secondary user in the t-th time slot through the i-th sub-band; is a binary variable, used to represent whether the i-th cognitive UAV transmits data to the j-th secondary user in the t-th time slot using the i-th sub-band; denotes the white noise interference on the channel transmission, , , and respectively denote the wireless transmission power of the cognitive UAV, the interfering UAV, and the base station; channel gain from the base station to the kth secondary user in the tth time slot; channel gain from the kth interfering drone to the jth secondary user in the tth time slot; channel gain from the base station to the kth secondary user in the tth time slot; channel gain from the kth interfering drone to the jth secondary user in the tth time slot; channel gain from the base station to the kth secondary user in the tth time slot; channel gain from the base station to the kth secondary user in the tth time slot; channel gain from the base station to the kth secondary user in the tth time slot; is a binary variable indicating whether the kth interfering drone interferes with all users in the jth sub-band in the tth time slot; is a binary variable indicating whether the kth interfering drone interferes with all users in the jth sub-band in the tth time slot;
[0045] signal-to-noise ratio of the kth primary user in the jth sub-band in the tth time slot, channel gain from the base station to the kth primary user in the tth time slot; channel gain from the base station to the kth primary user in the tth time slot; channel gain from the base station to the kth primary user in the tth time slot; channel gain from the base station to the kth primary user in the tth time slot; channel gain from the kth cognitive drone to the jth primary user in the tth time slot; channel gain from the kth cognitive drone to the jth primary user in the tth time slot; channel gain from the kth cognitive drone to the jth primary user in the tth time slot; received data rate of the kth primary user in the jth sub-band in the tth time slot is determined by the following equation: received data rate of the kth primary user in the jth sub-band in the tth time slot is determined by the following equation:
[0046] received data rate of the kth secondary user in the jth sub-band in the tth time slot is determined by the following equation:
[0047] ;
[0048] received data rate of the kth secondary user in the jth sub-band in the tth time slot is determined by the following equation:
[0049] ;
[0050] wherein, received data rate of the kth primary user in the jth sub-band in the tth time slot, received data rate of the kth secondary user in the jth sub-band in the tth time slot.
[0051] In a second aspect, the application further provides an anti-interference unmanned aerial vehicle scheduling device, comprising:
[0052] An acquisition module is configured to acquire local environment information and user state information collected by each cognitive unmanned aerial vehicle;
[0053] A communication processing module is configured to analyze channel gains of each cognitive unmanned aerial vehicle to each user, channel gains of each interfering unmanned aerial vehicle to each user, and channel gains of each base station to each user in the tth time slot based on the local environment information and the user state information by using a pre-constructed unmanned aerial vehicle wireless communication service model, and determine a receiving data rate of each user in each sub-band in the tth time slot.
[0054] A solving module is configured to determine a sub-band allocation strategy of each cognitive unmanned aerial vehicle and a secondary user access strategy by using a pre-constructed unmanned aerial vehicle scheduling and communication resource management model based on the receiving data rate of each user in each sub-band in the tth time slot, wherein the unmanned aerial vehicle scheduling and communication resource management model is used to maximize effective cumulative transmission rate and service fairness under a preset constraint condition.
[0055] An optimization module is configured to determine an action value according to an observation value by using a multi-agent deep reinforcement learning framework, wherein the observation value comprises position information of each cognitive unmanned aerial vehicle, channel gains of each cognitive unmanned aerial vehicle to each user, and a cumulative service rate of a secondary user, and the action value comprises a flight steering angle of each cognitive unmanned aerial vehicle.
[0056] A scheduling module is configured to schedule each cognitive unmanned aerial vehicle based on the sub-band allocation strategy, the secondary user access strategy, and the flight steering angle.
[0057] In a third aspect, the application further provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements steps of the method according to the first aspect when executing the computer program.
[0058] In a fourth aspect, the application further provides a computer readable storage medium having a computer program stored thereon, wherein the computer program is executed by a processor to implement steps of the method according to the first aspect.
[0059] In a fifth aspect, the application further provides a computer program product comprising a computer program, wherein the computer program is executed by a processor to implement steps of the method according to the first aspect.
[0060] The aforementioned anti-interference drone scheduling method, apparatus, computer equipment, computer-readable storage medium, and computer program product acquire local environmental information and user status information collected by each cognitive drone; based on the local environmental information and user status information, and using a pre-constructed drone wireless communication service model, analyze the channel gain from each cognitive drone to each user, the channel gain from each interfering drone to each user, and the channel gain from each base station to each user within the t-th time slot, and determine the received data rate of each user in each sub-frequency band within the t-th time slot; based on the received data rate of each user in each sub-frequency band within the t-th time slot, utilize the pre-constructed drone scheduling... The algorithm solves the degree and communication resource management model to determine the sub-frequency band allocation strategy and secondary user access strategy for each cognitive UAV. The UAV scheduling and communication resource management model is used to maximize the effective cumulative transmission rate and service fairness under preset constraints. Through a multi-agent deep reinforcement learning framework, action values are determined based on observations. The observations include the position information of each cognitive UAV, the channel gain from each cognitive UAV to each user, and the cumulative service rate of the secondary user. The action values include the flight turning angle of each cognitive UAV. Based on the sub-frequency band allocation strategy, secondary user access strategy, and flight turning angle, the cognitive UAVs are scheduled. Through the above methods, based on the UAV wireless communication service model, UAV scheduling, and communication resource management model, real-time perception of the radio environment and dynamic allocation of spectrum resources are achieved. An optimization objective of maximizing the effective cumulative transmission rate and service fairness is introduced, reducing service gaps for edge users and significantly improving service fairness for secondary users, thus meeting the balanced service needs of massive IoT terminals. UAV scheduling is based on sub-band allocation strategies, secondary user access strategies, and flight turning angles, improving the system's anti-interference capability and spectrum utilization efficiency in complex electromagnetic environments. This ensures service quality while enhancing overall network coverage performance and operational stability. A multi-agent collaborative reinforcement learning framework effectively improves the trajectory optimization efficiency of UAVs in dynamic scenarios, accelerates the system's recovery speed during network topology changes, and meets the requirements of full coverage and low latency. Attached Figure Description
[0061] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments of this application or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0062] Figure 1 This is a schematic diagram of the layered architecture of an integrated air-ground network in one embodiment;
[0063] Figure 2 This is an application environment diagram of an anti-interference drone scheduling method in one embodiment;
[0064] Figure 3 This is a schematic diagram of the interaction flow of an anti-interference drone scheduling method in one embodiment;
[0065] Figure 4 This is a flowchart illustrating an anti-interference drone scheduling method in one embodiment;
[0066] Figure 5 This is a comparative diagram of cumulative reward values in one embodiment;
[0067] Figure 6 This is a schematic diagram illustrating how the cumulative reward for different drone flight strategies changes with each round time slot in one embodiment;
[0068] Figure 7 This is a schematic diagram illustrating how the cumulative reward for different computation strategies changes with each round slot in one embodiment;
[0069] Figure 8 This is a diagram illustrating the effect of the next user fairness coefficient changing with time slots in one embodiment.
[0070] Figure 9 This is a structural block diagram of an anti-interference drone scheduling device in one embodiment;
[0071] Figure 10 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation
[0072] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0073] It should be noted that the terms "first," "second," etc., used in this application can be used to describe various elements, but these elements are not limited by these terms. These terms are only used to distinguish the first element from the second element. The terms "comprising" and "having," and any variations thereof, used in this application, are intended to cover non-exclusive inclusion. The term "multiple" used in this application refers to two or more. The term "and / or" used in this application refers to one of the embodiments, or any combination of multiple embodiments.
[0074] To facilitate understanding of the embodiments provided in this application, some terms that appear in the detailed description are introduced below:
[0075] Cognitive radio (CR) technology: Intelligent radio technology with spectrum sensing, monitoring and management capabilities. It can sense spectrum occupancy status, interference level and channel quality in real time, providing core technical support for dynamic spectrum sharing.
[0076] Cognitive Unmanned Aerial Vehicle (C-UAV): An unmanned aerial vehicle platform integrating cognitive radio technology, possessing real-time spectrum perception, wireless data transmission, dynamic resource scheduling, and three-dimensional maneuverability. It can autonomously optimize communication links in complex wireless environments, providing flexible and reliable wireless transmission services to ground users.
[0077] Jamming Unmanned Aerial Vehicle (J-UAV): A drone device that intentionally interferes with wireless communication systems. By emitting jamming signals in specific frequency bands, it reduces or destroys the quality of communication links and is a typical source of interference in emergency and military communications.
[0078] Air-Ground Integrated Network (AGIN): A network architecture that integrates aerial platforms (such as drones and balloons) with ground communication infrastructure. Its goal is to achieve air-ground collaboration, improve coverage and communication capabilities, and it is widely used in emergency communications, smart transportation, and the Internet of Things (IoT) scenarios.
[0079] Orthogonal Frequency Division Multiplexing (OFDM) is a spectrum utilization technique that divides a broadband spectrum into multiple orthogonal sub-spectrums for parallel data transmission. The sub-spectrums maintain orthogonality, avoiding inter-spectrum interference, improving spectrum utilization and resistance to multipath fading, and is widely used in wireless communication spectrum transmission.
[0080] Dynamic spectrum resource allocation: Based on a real-time spectrum awareness-based intelligent resource allocation paradigm, wireless communication service users are divided into primary users (PUs) and secondary users (SUs). The cognitive UAV first senses the channel quality of the spectrum and then feeds it back to the edge server. The server calculates the optimal user access scheme and sends the corresponding control signals to the cognitive UAV. The UAV then uses the designated spectrum resources to provide wireless communication services to the secondary users according to the received control.
[0081] Multi-Agent Deep Reinforcement Learning (MA-DRL): An intelligent decision-making framework that integrates multi-agent systems and deep reinforcement learning. Through collaborative exploration and experience sharing among multiple agents (such as multiple drones), it achieves global optimization goals in dynamic environments and possesses distributed decision-making and adaptive learning capabilities.
[0082] Multi-Agent Deep Deterministic Policy Gradient (MADDPG) is a reinforcement learning algorithm for multi-agent systems. It is an extension of the single-agent DDPG method, employing a "centralized training, distributed execution" framework. During training, each agent's policy is guided by its own observations and global information (such as the actions of other agents) to improve stability and cooperation; during execution, agents make independent decisions based solely on local observations. MADDPG is suitable for complex scenarios such as multi-agent cooperation and competition, and is commonly used in tasks such as UAV scheduling, game theory, and resource allocation.
[0083] Branch and Bound: An exact algorithm for solving combinatorial optimization or integer programming problems. Its basic idea is to progressively divide the solution space of the original problem into several subspaces (branches), and by calculating the upper and lower bounds of the possible optimal values achievable in each subspace (bounding), subspaces that cannot contain the global optimum are eliminated in advance (pruning), thereby reducing the search size and ultimately obtaining the global optimum.
[0084] Understandably, the integrated air-ground information network mainly includes terrestrial communication networks (such as terrestrial vehicle networks, maritime communication networks, and small-scale terrestrial internet) and near-ground unmanned aerial vehicle (UAV) networks. Near-ground UAVs, with their line-of-sight communication and near-ground wireless transmission capabilities, can provide high-quality wireless transmission services to ground mobile device nodes, such as data forwarding and relay communication. The equipment nodes of the terrestrial communication network meet their own communication needs by receiving services from ground base stations or near-ground UAVs.
[0085] With the development of 6G networks, the number of IoT devices is surging, placing higher demands on wireless communication coverage and connection quality. To achieve broad coverage of massive IoT terminals, drones, as a flexible and mobile aerial platform, can assist ground base stations in providing wireless communication services to users, playing a crucial role, especially in environments with insufficient ground network coverage or complex environments.
[0086] However, drones are limited by energy constraints and computing power. After collecting local data about the communication environment, they struggle to fully grasp the global information of the entire network and cannot formulate optimal resource allocation and service strategies in real time through their own calculations. This limitation restricts the intelligent scheduling and efficient resource management capabilities of drones.
[0087] To effectively address the limitations of single-point computing power and incomplete information in drones, such as Figure 2 As shown, this application introduces a three-layer architecture design combining the airborne layer, the ground layer, and the network control and management layer. The anti-interference UAV scheduling method provided in the embodiments of this application can be applied to, for example... Figure 2 In the application environment shown, this architecture not only fully leverages the advantages of UAVs' flexible perception and maneuverability, but also achieves intelligent scheduling and dynamic management of communication resources for UAVs through the coordinated efforts of the airborne layer, network control, and management layer. The specific structure and detailed functions of each layer are described below:
[0088] The airborne layer (multi-UAV near-ground network layer): This layer consists of multiple UAVs equipped with cognitive radio devices. These UAVs possess real-time perception and interference detection capabilities of the surrounding spectrum environment, dynamically detecting and collecting multi-dimensional information such as signal strength and interference levels on different spectrum channels. Due to the limited computing and processing capabilities of the UAVs themselves, the collected environmental data and user service quality information are fed back to the network control and management layer in real time, and then await control commands. Upon receiving control commands, each UAV dynamically adjusts its communication parameters (such as frequency band selection, transmit power, and access device set) and flight position according to the control layer instructions, achieving intelligent interference avoidance, improved link stability, and service fairness. This closed-loop collaboration of perception and control ensures the airborne layer's efficient adaptive capabilities in complex electromagnetic environments.
[0089] The ground layer (terrestrial wireless communication network layer) consists of multiple ground base stations and diverse user terminals. User terminals, as service receivers, can directly access ground base stations or obtain communication services via UAV relays, meeting connectivity needs in different environments. Ground base stations, as service providers, are responsible for receiving, processing, and forwarding user signals, while collecting user status information, traffic demands, and service quality indicators, feeding this data back to the network control and management layer. Based on scheduling instructions from the control layer, base stations optimize user access strategies and resource allocation schemes to achieve accurate responses to user needs and ensure fair service. The ground layer plays a crucial role in the overall architecture in user access and information collection, providing fundamental data support for intelligent scheduling by the network control layer.
[0090] Control and Management Layer (Intelligent Scheduling and Resource Optimization Center): The network control and management layer is typically deployed on edge servers within the terrestrial wireless communication network, undertaking the functions of global situational awareness, comprehensive analysis, and intelligent scheduling decision-making for the entire system. This layer receives real-time spectrum sensing and interference environment data from UAVs in the airspace layer, as well as user status, traffic demand, and quality of service information from ground-level base stations, constructing a dynamic network model covering the entire network. Based on multi-objective optimization algorithms, the control layer formulates comprehensive scheduling strategies encompassing UAV scheduling, communication parameter adjustment, and resource allocation, balancing maximizing the overall system transmission rate with ensuring fairness in user services. Through continuous closed-loop feedback and online optimization, the network control layer can dynamically respond to environmental changes and fluctuations in user demand, improving the system's robustness and overall performance, ensuring stable communication and efficient resource utilization in complex electromagnetic environments.
[0091] This three-layer architecture, through multi-level collaboration and intelligent resource management, significantly enhances the anti-interference capability and transmission efficiency of the air-ground integrated network in complex electromagnetic environments, laying a solid foundation for achieving fair and high-quality communication for large-scale IoT devices, and is of great significance in promoting the widespread deployment and continuous evolution of future 6G networks.
[0092] It should be noted that, in Figure 2 Under the three-layer collaborative architecture shown, the interaction process of the anti-interference UAV scheduling method follows a closed-loop logic of "perception—fusion—optimization—execution". Specifically, the ground base station and the UAV in the air respectively undertake the functions of user information collection and spectrum perception, while the network control and management layer integrates and optimizes multi-source information for decision-making. Finally, the UAV and the base station execute scheduling and allocation commands. Through this closed-loop mechanism, the system can achieve efficient adaptive scheduling and resource management under complex electromagnetic environments and dynamic demand conditions. (Refer to...) Figure 3 The specific interaction process includes:
[0093] ① Data acquisition phase: The cognitive drone swarm uses its onboard cognitive radio equipment to detect the surrounding spectrum environment in real time and collect data such as signal strength, interference level and user service quality in each frequency band;
[0094] ②Perception data upload stage: The UAV uploads the collected local environmental information and user status information to the edge server in real time, providing basic data support for the global decision-making of the control layer;
[0095] ③ Information integration and strategy optimization stage: After receiving the environmental information uploaded by the UAV, the edge server integrates information such as the communication needs of ground network nodes, the usage of communication resources, and the quality of channel communication. Through optimization algorithms and intelligent algorithms, it calculates resource allocation and UAV scheduling schemes locally.
[0096] ④ Control signal transmission phase: The edge server converts the execution plan into control commands and sends them to each UAV in the air layer;
[0097] ⑤ Service Phase: The UAV swarm moves to the execution position based on the received control quality and uses the designated communication resources to serve the ground terminal equipment.
[0098] Through a three-layer integrated air-ground architecture and closed-loop mechanism, dynamic collaboration between UAVs and ground networks is achieved. The air layer is responsible for real-time spectrum perception and interference detection, the ground layer is responsible for user access and status acquisition, and the control and management layer completes global optimization decisions based on edge computing, solving the problems of limited computing power and lack of global information for individual UAVs.
[0099] In one exemplary embodiment, such as Figure 4 As shown, an anti-interference drone scheduling method is provided, which can be applied to... Figure 1 Taking edge servers as an example, this will be explained as follows:
[0100] Step 402: Obtain local environmental information and user status information collected by each cognitive UAV.
[0101] Among them, local environmental information refers to the data of the surrounding environment perceived by each cognitive drone from its own perspective and within the range of its sensors, while user status information refers to the real-time status information of users accessing the network perceived by each cognitive drone.
[0102] Step 404: Using a pre-built UAV wireless communication service model, based on local environmental information and user state information, analyze the channel gain from each cognitive UAV to each user, the channel gain from each interfering UAV to each user, and the channel gain from each base station to each user in the t-th time slot, and determine the received data rate of each user in each sub-frequency band in the t-th time slot.
[0103] Among them, a UAV wireless communication service model is pre-built, and the currently acquired local environmental information and user status information are substituted into the model to analyze the channel gain caused by different distances between objects (UAVs, base stations or users). Then, the received data rate of each user in each sub-frequency band in the t-th time slot is determined by the preset rules.
[0104] Step 406: Based on the received data rate of each user in each sub-frequency band within the t-th time slot, the pre-built UAV scheduling and communication resource management model is used to solve the problem and determine the sub-frequency band allocation strategy and secondary user access strategy for each cognitive UAV. The UAV scheduling and communication resource management model is used to maximize the effective cumulative transmission rate and service fairness under preset constraints.
[0105] The UAV scheduling and communication resource management model is a mixed-integer optimization model with multi-objective constraints. Its objective is to maximize the effective cumulative transmission rate and service fairness. By substituting data into the solution, it can determine the sub-frequency band allocation strategy and the secondary user access strategy for each cognitive UAV. The sub-frequency band allocation strategy specifies which spectrum resources the cognitive UAV chooses to occupy, and the secondary user access strategy specifies which users the cognitive UAV provides wireless transmission services to.
[0106] Step 408: Using a multi-agent deep reinforcement learning framework, determine action values based on observations. Observations include the location information of each cognitive UAV, the channel gain from each cognitive UAV to each user, and the cumulative service rate of the secondary user. Action values include the flight turning angle of each cognitive UAV.
[0107] The cognitive drone flight strategy is defined as follows: within each time slot, the cognitive drone flies at a constant speed, travels a distance of d, and maintains a fixed altitude. Based on this, this embodiment utilizes Multi-Agent Deep Reinforcement Learning (MADDPG) to incorporate the data estimated in the aforementioned steps into the observations, and uses a policy network to determine the action values of each cognitive drone.
[0108] Step 410: Schedule each cognitive UAV based on the sub-frequency band allocation strategy, the secondary user access strategy, and the flight turning angle.
[0109] The sub-frequency band allocation strategy, secondary user access strategy, and flight turning angle of each cognitive UAV are sent to the corresponding cognitive UAV. This allows the cognitive UAV to move according to the flight turning angle, select the spectrum resources to occupy according to the sub-frequency band allocation strategy, and select the users to access according to the secondary user access strategy. In this way, directional wireless transmission services are provided to designated service users on designated frequency bands. This can effectively avoid the problem of spectrum interference and serve as many users as possible. While improving service quality, it can also ensure the fairness of secondary users receiving cognitive UAV services.
[0110] In the aforementioned anti-interference UAV scheduling method, local environmental information and user status information collected by each cognitive UAV are acquired. Using a pre-constructed UAV wireless communication service model, based on the local environmental information and user status information, the channel gain from each cognitive UAV to each user, the channel gain from each interfering UAV to each user, and the channel gain from each base station to each user within the t-th time slot are analyzed to determine the received data rate of each user in each sub-frequency band within the t-th time slot. Based on the received data rate of each user in each sub-frequency band within the t-th time slot, a pre-constructed UAV scheduling and communication resource management model is used to calculate... The solution involves determining the sub-frequency band allocation strategy and secondary user access strategy for each cognitive UAV; a UAV scheduling and communication resource management model is used to maximize the effective cumulative transmission rate and service fairness under preset constraints; action values are determined based on observations using a multi-agent deep reinforcement learning framework; observations include the location information of each cognitive UAV, the channel gain from each cognitive UAV to each user, and the cumulative service rate of the secondary user; action values include the flight turning angle of each cognitive UAV; and scheduling of each cognitive UAV is performed based on the sub-frequency band allocation strategy, secondary user access strategy, and flight turning angle. Through the above methods, based on the UAV wireless communication service model, UAV scheduling, and communication resource management model, real-time perception of the radio environment and dynamic allocation of spectrum resources are achieved. An optimization objective of maximizing the effective cumulative transmission rate and service fairness is introduced, reducing service gaps for edge users and significantly improving service fairness for secondary users, thus meeting the balanced service needs of massive IoT terminals. UAV scheduling is based on sub-band allocation strategies, secondary user access strategies, and flight turning angles, improving the system's anti-interference capability and spectrum utilization efficiency in complex electromagnetic environments. This ensures service quality while enhancing overall network coverage performance and operational stability. A multi-agent collaborative reinforcement learning framework effectively improves the trajectory optimization efficiency of UAVs in dynamic scenarios, accelerates the system's recovery speed during network topology changes, and meets the requirements of full coverage and low latency.
[0111] In one exemplary embodiment, the drone wireless communication service model is represented as follows:
[0112] ;
[0113] ;
[0114] Indicates the first A cognitive drone passes through the first time slot t. Sub-band to the first Signal-to-noise ratio of data transmitted by each user; A binary variable used to represent the first... Does the cognitive drone use the [missing information] in the t-th time slot? Sub-band to the first Each user transmits data; This represents white noise interference in channel transmission. , and These represent the wireless transmission power of the cognitive drone, the jamming drone, and the base station, respectively. Indicates the t-th time slot. The first cognitive drone to the first Channel gain for each user; Indicates the t-th time slot. The jamming drone to the first Channel gain for each user; Indicates the first The base station within the first time slot to the first Channel gain for each user; A binary variable used to represent the first... Does the interfering drone affect the first time slot in the t-th time slot? This causes interference to all users within the sub-band;
[0115] Indicates the t-th time slot. The first main user in the Signal-to-noise ratio of each sub-band This indicates white noise interference in channel transmission; Indicates the first The base station within the first time slot to the first Channel gain for each primary user; Indicates the t-th time slot. The first cognitive drone service In the case of a single user, up to the first Channel gain for each primary user; Indicates the t-th time slot. The jamming drone to the first Channel gain for each primary user;
[0116] The following formula is used to determine the time slot within the t-th time slot. The first main user in the The received data rate of each sub-band:
[0117] ;
[0118] The following formula is used to determine the time slot within the t-th time slot. The user in the first The received data rate of each sub-band:
[0119] ;
[0120] in, Indicates the t-th time slot. The first main user in the The received data rate of each sub-band Indicates the t-th time slot. The user in the first The received data rate of each sub-band.
[0121] Among them, the pre-constructed UAV wireless communication service model can effectively capture the essential characteristics of the system, providing solid theoretical support for the subsequent optimization methods and scheduling strategies.
[0122] Understandably, considering real-world scenarios where wireless communication networks are paralyzed due to interference from drones and temporary base station failures, this application's embodiments model a realistic and complex wireless communication network environment. Specifically, the wireless communication employs orthogonal frequency division multiplexing (OFDM) coding, meaning the total spectrum resources can be divided into multiple non-overlapping sub-spectrum resources, with communication on different sub-spectrum resources not interfering with each other. The terrestrial wireless communication network consists of primary users (PU), secondary users (SU), and base stations (BS), while the near-ground drone network comprises multiple cognitive drones (C-UAVs). Furthermore, there are interfering drones (J-UAVs) that continuously transmit interference noise on certain sub-spectrum resources. Specifically, the set of primary users is represented as... The set of secondary users is represented as The cognitive drone set is represented as The set of interfering drones is represented as The spectrum resource set is represented as Due to the limited service capacity of base stations, they will only provide wireless communication transmission services to primary users, while secondary users will have their wireless transmission services provided by auxiliary drones.
[0123] This application uses a three-dimensional Cartesian coordinate system to represent the locations of the user, the cognitive drone, the jamming drone, and the base station. The data transmission service process is divided into... Each time slot (or time slot). The fixed location of the base station is defined as... Furthermore, the t-th time slot can acquire the th... The primary user and the first The location of each user is denoted as follows: and To better characterize the user's location characteristics, this embodiment of the application assumes that the user's location remains unchanged during the t-th time slot.
[0124] This embodiment of the application assumes that each cognitive drone flies at a constant speed within each time slot, and the flight distance is... Each jamming drone, however, remains in a fixed position. The coordinates of the cognitive drone in the t-th time slot are: , No. The fixed position of an interfering drone is defined as ,in To understand the fixed flight altitude of drones, To disrupt the fixed flight altitude of the drone. Accordingly, the first The coordinates of a cognitive UAV in the t-th time slot can be represented as:
[0125] ;
[0126] ;
[0127] in, Indicates the first The two-dimensional flight direction vector angle of the first cognitive UAV, i.e., the flight turning angle. Based on the above definition, the distance from the base station to the first... The distance to the primary user, the base station to the first Distance of each user, the first The first cognitive drone to the first The distance between the primary users, and the first The first cognitive drone to the first The distance between each user can be expressed by the following formulas (1) to (4):
[0128] (1);
[0129] (2);
[0130] (3);
[0131] (4);
[0132] The primary communication channel between cognitive drones and ground users is the line-of-sight (LoS) link. Cognitive drones are equipped with directional radios and use top-wire wireless transmission to serve secondary users. This represents the channel power gain (i.e., channel gain) at a reference distance of 1 meter. Considering that cognitive UAV services for secondary users may cause sidelobe interference to primary users, this embodiment will... The first cognitive drone service In the case of the first user The channel power gain for a primary user at a reference distance of 1 meter is defined as follows: It can be expressed by the following formula (5):
[0133] (5);
[0134] in, This represents the channel gain of the primary user when the reference distance is 1 meter within the sidelobe range of the secondary user served by the drone. This represents the channel gain of the primary user when the reference distance is 1 meter outside the sidelobe range of the secondary user served by the drone. In real-world scenarios, the former is usually hundreds of times larger than the latter. Indicates from the first The drone to the first The orientation and angle of each main user. Indicates from the first The drone to the first The orientation and angle of each user This indicates the main lobe beam angle.
[0135] Within the t-th time slot, the first... The first cognitive drone service In the case of a single user, up to the first The channel gain for each primary user can be expressed as... Similarly, in the t-th time slot, the th... The first cognitive drone to the first The channel gain for each user can be expressed as... Unlike the channel model between drones and ground users, base stations use omnidirectional antennas to transmit wireless signals, and the channel model between base stations and ground users uses distance-dependent path loss (path loss exponent). Characterized by ) and small-scale Rayleigh fading. Therefore, the first The base station within the first time slot to the first The channel gain for each primary user can be expressed as... Similarly, the first The base station within the first time slot to the first The channel gain for each user can be expressed as... ,in and It is an exponentially distributed random variable with a mean of 1, used to describe Rayleigh fading.
[0136] To better describe how cognitive drones transmit data to secondary users using spectrum resources within each time slot, 0-1 variables (i.e., binary variables) are introduced. , used to indicate the first Does the cognitive drone use the [missing information] in the t-th time slot? Sub-band to the first Each user transmits data. Similarly, 0-1 variables... Indicates the first Does the interfering drone affect the first time slot in the t-th time slot? This causes interference to all users within each sub-band. Therefore, the first... A cognitive drone passes through the first time slot t. Sub-band to the first The signal-to-noise ratio (SINR) of each user's transmitted data can be expressed as:
[0137] (6);
[0138] in Indicates the t-th time slot. The jamming drone to the first The channel gain of each user (its definition is...) similar), This represents white noise interference in channel transmission. , and Let represent the wireless transmission power of the cognitive drone, the jamming drone, and the base station, respectively. Similarly, consider the impact of cognitive drone communication and jamming drone interference on the signal-to-noise ratio of the main user's received data transmission service. Introduce... Indicates the t-th time slot. The first main user in the The signal-to-noise ratio of each sub-band. In this case, the power gain of the base station providing transmission services to the primary user is considered the effective gain, while the effects of white noise, cognitive drone communication, and interference from interfering drones are considered the interference gain. Specifically, It can be represented as:
[0139] (7);
[0140] in, This represents white noise interference in the channel transmission. Referring to Shannon's formula, the white noise interference within the t-th time slot... The first main user in the The received data rate of each sub-band, the data rate of the t-th time slot The user in the first The received data rates of each sub-band can be expressed as follows: and .
[0141] In an exemplary embodiment, the UAV scheduling and communication resource management model includes a cognitive UAV spectrum resource allocation problem model and a secondary user access problem model; step 406 includes: based on a greedy strategy, solving the cognitive UAV spectrum resource allocation problem model according to the received data rate of each user in each sub-frequency band within the t-th time slot to determine the sub-frequency band allocation strategy for each cognitive UAV; determining the channel interference signal-to-noise ratio matrix according to the sub-frequency band allocation strategy for each cognitive UAV; and based on a branch and bound algorithm, solving the secondary user access problem model according to the channel interference signal-to-noise ratio matrix to determine the secondary user access strategy.
[0142] In an exemplary embodiment, the cognitive unmanned aerial vehicle (UAV) spectrum resource allocation problem model is represented as follows:
[0143] ;
[0144] in:
[0145] ;
[0146] ;
[0147] The main user set is represented as The set of secondary users is represented as The cognitive drone set is represented as The spectrum resource set is represented as ;
[0148] This represents the sub-band allocation strategy for each cognitive UAV within the t-th time slot. This indicates the time slot t-1. The first cognitive drone to the first The cumulative number of services per user Indicates the first A cognitive drone in the t-th time slot Upwards from the first sub-frequency band Channel interference signal-to-noise ratio for each user's transmitted data;
[0149] The secondary user access problem model is represented as follows:
[0150] ;
[0151] in:
[0152] ;
[0153] ;
[0154] ;
[0155] ;
[0156] This represents the access strategy for secondary users within the t-th time slot. This represents the channel interference signal-to-noise ratio matrix. Indicates that in determining In the case of the first A cognitive drone in the t-th time slot Upwards from the first sub-frequency band The transmission rate of each user Indicates the penalty coefficient. Indicates a penalty item. Indicates the first Minimum total transmission rate requirement for each primary user Used to describe normalization priority The weight representing the access priority of secondary users. Indicates the first The cumulative service rate of each user from the first time slot to the t-th time slot. These are the preset adjustment parameters. Indicates the first The primary user in the t-th time slot... The received data rate of each sub-band.
[0157] In this embodiment, a multi-objective optimization model is established with the goal of maximizing cumulative transmission rate and ensuring service fairness. By decoupling the original optimization problem into three sub-problems—UAV communication resource allocation, secondary user access, and intelligent UAV scheduling—the difficulty of solving the problem is effectively reduced, and system efficiency and user fairness are balanced. The model incorporates practical constraints such as primary user interference protection, UAV antenna quantity limitations, and co-channel interference avoidance to ensure its rationality in real-world scenarios.
[0158] In this embodiment, to ensure efficient use of spectrum resources and protect users from excessive interference in the drone network, a drone scheduling and communication resource management model is pre-constructed. Considering the real-world need to maximize transmission quality and fairness of user services, the above model is modeled as a multi-objective problem, including the following two optimization objectives:
[0159] Effective cumulative transmission rate: The effective cumulative transmission rate can be measured by the total rate at which the UAV transmits data to secondary users across multiple sub-bands, and is expressed as: ;
[0160] Service fairness: According to the definition of the Jain Fairness Index, the service fairness ratio is expressed as:
[0161] ;
[0162] in Record number The cumulative service rate of a user from the first time slot to the t-th time slot is expressed as: .
[0163] Understandably, the problem of data transmission and trajectory planning assisted by cognitive UAVs in spectrum-sharing networks under interference conditions is modeled as a multi-objective optimization problem, aiming to maximize the effective cumulative transmission rate and the service fairness ratio.
[0164] In solving practical problems, these two optimization objectives often conflict with each other. To facilitate unified modeling and solving, these two objectives are integrated into a single maximization objective, namely:
[0165] ;
[0166] The optimization problem P1 is defined as follows:
[0167] (8);
[0168] Considering the interference caused by cognitive drone services to primary users by secondary users, the number of directional antennas equipped by cognitive drones, and the need to avoid co-band interference, constraints are defined. – The explanation is as follows:
[0169] Considering that secondary users of cognitive drone services may interfere with primary users, in order to protect the transmission quality of primary user services from potential interference, the following measures are introduced: ,in Indicates the first Minimum total transmission rate requirement for each primary user.
[0170] From the perspective of service fairness, each secondary user is limited to accessing a maximum of one sub-band, which helps cognitive drones provide services to more secondary users.
[0171] Because cognitive drones use directional wireless communication to provide transmission services, each spectrum access session occupies one directional antenna. To reflect the device load constraints in a real-world scenario, it is assumed that each cognitive drone is equipped with... There are only one directional antenna, therefore limiting its simultaneous use to no more than [number missing]. Data is transmitted through each sub-band.
[0172] To prevent resource inefficiency caused by cognitive drones remaining idle and not serving any secondary users, we introduce... To ensure that each cognitive drone serves at least one secondary user.
[0173] Considering that multiple cognitive drones transmitting data on the same frequency band can cause co-channel interference, we introduce... To limit the number of cognitive drones serving a single sub-user to a maximum of one per given sub-band.
[0174] :constraint It is a binary variable.
[0175] Constraints on the flight direction angle of cognitive drones Within the range.
[0176] Understandably, P1 is a mixed-integer optimization problem with combinatorial constraints. To better handle its combinatorial constraints, ρ is decomposed into μ and μ. The original three-dimensional spectrum allocation constraints are reformulated into two-stage constraints, including cognitive UAV and sub-band pairing, and secondary user spectrum access. The transformed problem is denoted as P2 and expressed by the following formula:
[0177] (9);
[0178] In its implementation, P2 is decomposed into three sub-problems: multi-UAV spectrum resource allocation, secondary user access, and multi-UAV scheduling. For each sub-problem, a spectrum allocation algorithm based on a greedy strategy, an optimal spectrum access algorithm based on a branch-and-bound algorithm, and a multi-UAV intelligent scheduling algorithm based on MADDPG are designed respectively.
[0179] It should be noted that because jerking unmanned aerial vehicles (J-UAVs) continuously output interference signals on fixed sub-bands, the transmission rate and quality provided by cognitive UAVs vary across different sub-bands. Therefore, to maximize the overall transmission rate, the goal is to allocate less-interfered spectrum resources to cognitive UAVs capable of providing high-quality service, thus obtaining a better initial solution. However, solving the spectrum allocation problem is extremely complex. On one hand, the secondary user access scheduling of cognitive UAVs is unknown, making their transmission capabilities uncertain; on the other hand, the spectrum allocation problem is a constrained combinatorial optimization problem with an exceptionally large solution space. Based on this, the cognitive UAV spectrum allocation problem is formally defined as P2-1, expressed by the following formula:
[0180] (10);
[0181] in, It is used to evaluate the first A cognitive drone in the t-th time slot The utility value of transmission quality on each sub-band. Considering that historical sub-user access data can reflect the service preferences of cognitive UAVs, a weighted summation method based on the number of service visits is used to calculate the utility value, expressed by the following formula:
[0182] (11);
[0183] in:
[0184] ;
[0185] This indicates the time slot t-1. The first cognitive drone to the first The cumulative number of services served per user. Specifically, Indicates the first A cognitive drone in the t-th time slot Upwards from the first sub-frequency band The channel interference signal-to-noise ratio of each user's transmitted data, Under these conditions, its definition is the same as the aforementioned. The same applies. Considering that the distance between the cognitive drone and the secondary user is relatively large in the early time slots, and that historical secondary user service information has little impact on the multi-cognitive drone trajectory strategy, a time slot threshold can be introduced. .when When the utility matrix is calculated, it is estimated by a simple weighted average of the current channel measurements; when At that time, a weighted estimation method that incorporates historical service frequency is used to determine the utility matrix.
[0186] The optimization objective of P2-1 is to maximize the sum of utility values under combined constraints. Therefore, Algorithm 1 (utility matrix generation algorithm) and Algorithm 2 (UAV spectrum resource allocation algorithm) are proposed to solve P2-1 and obtain the matrix. The implementation of Algorithm 1 is shown below:
[0187] enter:
[0188] Output:
[0189] 1:if
[0190] 2:
[0191] 3:end if
[0192] 4: Defined according to formula (11)
[0193] 5:
[0194] 6:for
[0195] 7:
[0196] 8: for
[0197] 9:
[0198] 10: end for
[0199] 11: if
[0200] 12:
[0201] 13: end if
[0202] 14:end for
[0203] The implementation of Algorithm 2 is shown below:
[0204] enter:
[0205] Output:
[0206] 1:
[0207] 2: According to sort in descending order Use Record
[0208] 3:
[0209] 4:for
[0210] 5: if then
[0211] 6:
[0212] 7:
[0213] 8:
[0214] 9: end if
[0215] 10: if then
[0216] 11: break
[0217] 12: end if
[0218] 13:end for
[0219] 14:for
[0220] 15: if then
[0221] 16:
[0222] 17: end if
[0223] 18:end for
[0224] 19: Sort by utility value in ascending order And adopt save
[0225] 20:
[0226] 21:for
[0227] 22: if then
[0228] twenty three:
[0229] 24: end if
[0230] 25: if then
[0231] 26: break
[0232] 27: end if
[0233] 28:end for
[0234] 29: Using the Hungarian algorithm in sets and Solving for the optimal matching and using Record
[0235] 30:for
[0236] 31:
[0237] 32:endfor
[0238] Understandably, in Algorithm 1, the input includes... Specifically, This represents the matching relationship between the sub-band and the secondary user. Due to the integration of cognitive radio technology, in the initial stage of the t-th time slot, the cognitive UAV is able to perceive and monitor three-dimensional channel interference signal-to-noise ratio information through a polling-based sensing method. Output It is the result of sub-band allocation Definite The channel interference signal-to-noise ratio matrix of size. At the beginning of Algorithm 1, if... First, initialize Subsequently, the utility value matrix is calculated according to formula (11). Then, Algorithm 2 is used to obtain the sub-band allocation strategy. Finally, based on a given detailed sub-band allocation strategy... The channel interference signal-to-noise ratio (SNR) for each user in different access sub-bands can be estimated based on the SNR information monitored by the cognitive UAV. Based on a greedy strategy, Algorithm 2 is proposed to solve P2-1. The input of Algorithm 2 is... The output is the pairing results. Considering that in real-world scenarios, the number of available sub-band resources typically exceeds the number of cognitive drones, the following conditions are set... Following the greedy principle, first... All elements are sorted in descending order, and then the constraints are satisfied. and Under the premise of [unclear], sub-frequency bands are sequentially allocated to the corresponding cognitive UAVs. Obviously, this initial allocation may result in some cognitive UAVs not being allocated a sub-frequency band. To address this issue and satisfy the constraints... A two-stage allocation process is employed. Specifically, the initially allocated pairs are determined based on their... The utility values are sorted in ascending order, and then those assigned sub-bands that are not uniquely assigned to their corresponding cognitive drones are removed sequentially until the number of reselected sub-bands equals the number of unassigned cognitive drones. Finally, the Hungarian algorithm is used to find the optimal match between the reselected sub-bands and the unassigned cognitive drones.
[0239] After solving P2-1, the two-dimensional channel interference signal-to-noise ratio matrix is obtained. Sub-band allocation strategy ,in Indicates in The channel interference signal-to-noise ratio matrix is given below. To better achieve the goal of maximizing the total transmission rate and service fairness, the sub-problem P2-2 of secondary user access is expressed by the following formula:
[0240] (12);
[0241] in, Indicates that in determining In the case of the first A cognitive drone in the t-th time slot Upwards from the first sub-frequency band The transmission rate of each user. The variables in P2-2 essentially correspond to the transmission rate of each user. This is a matching scheme under certain conditions. However, due to constraints... And Jain fairness coefficient P2-2 is an NP-hard problem, and there is no polynomial-time solution.
[0242] To find the approximate optimal solution of P2-2 more efficiently, P2-2 is transformed into P2-2' as defined below:
[0243] (13);
[0244] In P2-2', in order to explore a larger solution space, we first use a penalty term. hard constraints This is transformed into a soft constraint in the optimization objective, where It's the penalty coefficient. Then, considering... The calculation not only depends on It also depends on ,design This describes the normalization priority under the goal of maximizing service fairness. The definition is as follows:
[0245] (14);
[0246] in:
[0247] ;
[0248] The weight representing the access priority of secondary users is crucial. In scenarios where only maximizing the total transmission rate is considered, some secondary users with poor transmission quality are rarely served, resulting in excessively low cumulative transmission rates and thus reducing the service fairness coefficient. By introducing Users with lower cumulative transmission rates are assigned higher priority weights, thereby increasing their probability of being served. This effectively balances the goals of maximizing total transmission rate and maximizing service fairness.
[0249] Essentially, feasible solutions to P2-2' satisfy the matching relation. This is without considering integer constraints. In this case, the problem can be solved in polynomial time using the simplex method or linear programming. Furthermore, the binary integer constraints... It is well-suited for branch-based solution processes. Based on the above observations, the branch and bound algorithm solves optimization problems by systematically branching and pruning the search space, making it suitable for solving the problem P2-2'.
[0250] The implementation of Algorithm 3 (Branch and Bound Algorithm) is shown below:
[0251] enter:
[0252] Output:
[0253] 1: Randomly select an integer solution that satisfies the matching constraints and set it as the initial optimal solution. and its corresponding function value Set as the initial optimal value
[0254] 2: Generate initial nodes ,in and Initialize an empty queue. .
[0255] 3: Random selection and ,Will according to The values are divided into and Two branches. Settings. ,set up ,in , .Will and Push the two branches into the queue middle.
[0256] 4:while
[0257] 5: From the queue pop-up ,get and With fixed parameters In this case, through the Relax integer constraints on medium parameters and find the branch. The upper realm Let the solution corresponding to this upper bound be .
[0258] 6: if It is an integer solution.
[0259] 7: if
[0260] 8: Bounding operations
[0261] 9: end if
[0262] 10: else
[0263] 11: if
[0264] 12: Randomly from Select And will Branches are and Two sub-branches. (Settings) ,set up ,in , .Will and Push the two branches into the queue middle. Branch operations
[0265] 13: else
[0266] 14: Prune the branches Pruning operations
[0267] 15: end if
[0268] 16: end if
[0269] 17: end while
[0270] 18: Return
[0271] It should be noted that the branch and bound algorithm mainly consists of three parts: branching, bounding, and pruning. In the initial stage, a feasible solution is obtained by generating random matchings, and then the optimal solution and optimal value are set... and To better describe branches, nodes are designed. To store branch information. Specifically, Used to identify the current node's position in the binary tree. Store a fixed set of parameters. This stores indices for variable parameters. After the initial node is generated, a variable is randomly selected, and two sub-branches are created based on its values of 0 and 1. It's important to note that due to matching constraints, when a variable is assigned the value 1, additional constraints must be applied to ensure that all other elements in the corresponding row and column are set to 0. The two sub-branches are then pushed into a queue. In the main loop, a branch is popped from the queue. The upper bound is calculated by relaxing the integer variables. If the obtained solution is an integer solution, then perform a bounding operation: if If the solution is not an integer, then update the optimal solution and the optimal value. and Compare. If If it's smaller, there's no need for further branching and pruning. Otherwise, for Further branch, and push its sub-branches into the queue.
[0272] In this embodiment, the combination of dynamic spectrum allocation and greedy strategy optimizes the utilization efficiency of spectrum resources; multi-UAV collaborative scheduling expands the network coverage, effectively alleviates capacity pressure in dense user scenarios, and improves the overall transmission capacity of the system.
[0273] In an exemplary embodiment, after step 410, the method further includes: determining a reward value based on the received data rate, service fairness ratio, sub-band allocation strategy, and secondary user access strategy of each user in each sub-band within the t-th time slot; storing the observation value, action value, and reward value as empirical data; and periodically updating the parameters of the policy network and evaluation network of the deep reinforcement learning framework based on the stored empirical data.
[0274] In one exemplary embodiment, the reward value is determined using the following formula:
[0275] ;
[0276] in, Indicates the first The set of frequency bands and user pairs allocated to a cognitive drone in the t-th time slot is represented as:
[0277] ;
[0278] Indicates that in determining In the case of the first A cognitive drone in the t-th time slot Upwards from the first sub-frequency band The transmission rate of each user Indicates the service fairness ratio, Indicates the penalty coefficient. Indicates a penalty item. Indicates the first The sub-band allocation strategy value corresponding to each cognitive drone. Indicates the first The user access policy value corresponding to each user.
[0279] It is understandable that the multi-awareness drone intelligent scheduling problem can be defined as P2-3, expressed by the following formula:
[0280] (15);
[0281] For problems P2-3, a multi-agent reinforcement learning framework, MADDPG (Multi-Agent Deep Deterministic Policy Gradient), is constructed. Based on this, each cognitive UAV is modeled as an independent agent, and the near-ground communication network is considered as the interaction environment for deep reinforcement learning. A Markov decision process is used to describe the framework of this method, which consists of three core components:
[0282] Observation First, cognitive drones In the time slot Two-dimensional coordinates Add to agent In the observations. Furthermore, to better explore, the channel gain of all secondary users served by the cognitive UAV was analyzed. Incorporate this into the observation. To better explore the impact of drone scheduling on fairness objectives, the observation will also be conducted from the initial time slot to the next time slot. Cumulative service rate for secondary users Add it to the observation.
[0283] action : Drone The flight direction is defined as its flight direction in the time slot. Flight maneuvers, namely .
[0284] reward function Considering that the number of frequency bands allocated to each cognitive drone may vary each time, the cognitive drone... In the time slot The average rate of service to secondary users is included as part of the reward. Furthermore, constraints are taken into consideration. The average penalty is taken into account when calculating the reward. The frequency band and user pair set allocated to a cognitive drone in the t-th time slot are defined as follows:
[0285] ;
[0286] The definition of a reward is:
[0287] ;
[0288] In addition, state It contains the observation information of each agent, representing the global state observed from the environment. Similarly, the global action can be represented as... .
[0289] In this framework, each agent has a deterministic policy network (Actor). Observing from its own local perspective Output Action During training, a centralized evaluation network (Critic) is configured for each agent. Its input includes global state Actions of all intelligent agents Experience is gained through interaction with the environment. This data is then stored in the experience replay pool. During the training phase, a batch of data with a size of [missing value] is randomly sampled from the experience replay pool. Based on experience, first use a target Actor network. Generate the target action for the next moment. and form Then, the target Critic network is used to calculate the Temporal Difference (TD):
[0290] (16);
[0291] Then, the mean squared loss of the Critic set is minimized to... Perform gradient descent updates:
[0292] (17);
[0293] The Actor network employs a deterministic policy gradient improvement mechanism. Its goal is to maximize the current evaluation of the ontology's actions by the Critic, i.e., to update the policy along the direction that improves the evaluation of the set of Critics. The gradient is expressed by the following formula:
[0294] (18);
[0295] Finally, to ensure training stability, a soft update mechanism is used to update the parameters of the target network, namely: and .
[0296] In one alternative implementation, after determining the sub-band allocation strategy and the secondary user access strategy, the server first sends an estimated channel interference signal-to-noise ratio (SNR) observation vector to the cognitive UAV. Then, based on this observation vector, a reinforcement learning network derives the flight turn angle. Finally, the server sends control signals for the flight turn angle, sub-band allocation strategy, and secondary user access strategy to the cognitive UAV. The cognitive UAV updates its position based on the flight turn angle. Then, in the new UAV coordinates, it provides wireless transmission services using the control signals based on the sub-band allocation strategy and the secondary user access strategy. Subsequently, rewards are calculated based on metrics such as rate estimates. The actions, observation vectors, and reward values are stored as empirical data, and the parameters of the policy network and evaluation network are periodically updated using gradient descent.
[0297] Based on the above, the joint solution process for UAV intelligent scheduling and communication resource management is shown in Algorithm 4 below:
[0298] 1: Initialization: Randomly initialize the parameters of the policy network and target policy network in DDPG. and Randomly initialize the parameters of the evaluation network and the target evaluation network in DDPG. and The experience replay area of each agent is initialized to be empty.
[0299] 2: for do
[0300] 3: Initialize the reinforcement learning environment and obtain the initial state.
[0301] 4: for do
[0302] 5: Drone swarms acquire information through cognitive radio technology.
[0303] 6: Obtain the sub-band allocation strategy according to Algorithm 1
[0304] 7: Obtain the secondary user access strategy based on Algorithm 3.
[0305] 8: for DDPG agents do
[0306] 9: Based on action networks Select Action
[0307] 10: end for
[0308] 11: Receive immediate rewards after the agent interacts with the environment. and the state of the next time slot
[0309] 12: for DDPG agents do
[0310] 13: Experience Store in the corresponding experience replay area
[0311] 14: if DDPG agent The experience replay area is full.
[0312] 15: Take a size from the experience replay area. Small batch experience
[0313] 16: Update the parameters of the DDPG policy network according to formula (17).
[0314] 17: Update the parameters of the DDPG evaluation network according to formula (18).
[0315] 18: Update the parameters of the DDPG target policy network and target evaluation network using a soft update mechanism: and
[0316] 19: end if
[0317] 20: end
[0318] 21: end for
[0319] 22: end
[0320] It is understood that, in this embodiment of the application, a joint solution framework is designed based on the three sub-problems, combining a greedy strategy, a branch and bound method, and multi-agent deep reinforcement learning (MADDPG):
[0321] A drone utility evaluation mechanism was designed, which uses a greedy strategy to select high-efficiency drone-spectrum resource combinations and satisfies the service constraints of drones through secondary allocation.
[0322] The branch and bound method is used to solve the secondary user access scheme, and the balance between transmission rate and service fairness is optimized through branching and pruning.
[0323] A multi-agent learning framework based on MADDPG is constructed, which uses the flight direction of the UAV as the action and combines observation and reward functions to achieve dynamic trajectory optimization, thereby improving adaptability to complex environments.
[0324] In an exemplary embodiment, taking the collaborative scheduling and communication restoration of multiple drones in a disaster emergency communication scenario as an example, after major natural disasters such as earthquakes, floods, and typhoons, ground communication base stations often fail due to physical damage or power outages, causing user terminals in the disaster-stricken area to be unable to access the communication network normally, severely affecting emergency rescue and information transmission. To address this problem, the anti-interference drone scheduling method provided in this application embodiment can quickly construct a temporary emergency communication network. Specifically, after a disaster occurs, multiple cognitive drones are urgently deployed to the disaster-stricken area, using onboard cognitive radio equipment to perceive available spectrum resources in real time and feeding the perception results back to an edge server. After integrating the access needs and spectrum status information of ground users, the edge server generates a drone scheduling and communication resource allocation scheme through an optimization algorithm, and then distributes it to each drone for execution. The drone swarm adjusts its flight position and service range according to control commands, providing relay or direct communication services to ground-affected users, ensuring wide coverage and high reliability of emergency communication. This embodiment demonstrates that this method can quickly restore basic communication connections in extreme environments where ground networks temporarily fail, significantly improving communication support capabilities in disaster scenarios.
[0325] In an exemplary embodiment, taking anti-interference multi-drone communication service in a smart transportation scenario as an example, in the vehicle-to-everything (V2X) environment of smart cities or highways, a large number of vehicle terminals need to maintain real-time communication with base stations or drones to support autonomous driving, traffic control, and multimedia services. However, in such high-density access scenarios, some unauthorized drones may enter the airspace and interfere with specific frequency bands. To address this problem, the method provided in this application embodiment uses a cognitive drone to monitor the spectrum status in real time, identify the interfered sub-frequency bands, and report the perception results to an edge server. The edge server solves the communication resource allocation and drone scheduling scheme, prioritizing the allocation of less interfered sub-frequency bands to drones with stronger service capabilities and directing drones closer to terminals with poor service quality. This embodiment demonstrates that this method can ensure the stability and reliability of V2X communication in complex electromagnetic environments with interfering drones, providing strong support for the safe operation of smart transportation scenarios.
[0326] To verify the feasibility of the method described in this application in a real-world scenario, refer to... Figures 5 to 8This study simulates an integrated air-ground network environment and conducts detailed simulation experiments within this environment. In this environment, the x and y coordinate ranges are set to [0,100] and [0,100] respectively. Four cognitive drones are used, with a fixed flight altitude of 70m. Furthermore, the number of divisible spectrum resources is set to 20, and the base station coordinates in the scenario are (50,50). Twenty primary users are randomly generated within a circle with a radius of 30m centered on the base station, followed by 50 secondary users randomly generated within the coordinate range. Three jamming drones are also randomly generated within the scenario, randomly occupying five frequency bands to transmit jamming signals. Before each training round, the initial positions of the cognitive drones are set to (0,0), (0,100), (100,0), and (100,100), respectively, and the fixed flight distance of the drones in each round is 10m.
[0327] In the above experimental environment, we first compare the total communication gains under different reinforcement learning frameworks, where each round has a fixed time slot of 20. For example... Figure 5 As shown, comparing the multi-agent reinforcement learning framework IDPG with the MADDPG framework used in this application embodiment, it was found that the cumulative reward value after convergence of MADDPG is about 11.54% higher than that of IDPG. This is mainly because the value network part of MADDPG introduces a global splicing mechanism of local information, which enables the network to learn information about the overall environment better; while IDPG, due to the independent information between agents, can only explore through local information, which limits the exploration effect.
[0328] Subsequently, MADDPG was compared with common drone flight strategies, selecting two comparative flight strategies: circling flight and random flight. Figure 6 This demonstrates how the cumulative rewards for different drone flight strategies change with each round's time slot. For example... Figure 6 As shown, the MADDPG method outperforms the other two algorithms in all time slots, followed by the orbital flight strategy, while the random flight strategy performs the worst. This is mainly because the orbital flight strategy only considers the fairness of user choice but fails to take into account factors such as user distribution and distance, resulting in a lower overall transmission rate; while the random factor fails to utilize environmental information, thus performing the worst. Furthermore, when there are more than 10 time slots, the cumulative reward increases linearly with the number of time slots. This is because when there are more than 10 time slots, the fairness coefficient tends to stabilize, and as the number of time slots increases, the transmission rate provided by the drone to the user increases, thus the cumulative reward grows linearly.
[0329] In addition, the communication service performance of different user access methods under the MADDPG flight strategy is compared. The specific user access methods are explained as follows:
[0330] Greedy strategy: After acquiring communication resources, the UAV only provides services to the user with the highest signal-to-noise ratio; Greedy strategy (introducing a fairness coefficient): When a UAV user accesses the network, it comprehensively considers the user's historical service rate and signal-to-noise ratio, but does not consider interference to the PU user; Branch and bound method: The optimization objective is defined as maximizing the transmission rate, and the interference to the PU is set as a soft constraint. The branch and bound algorithm is used to satisfy the constraints of the combined conditions; Branch and bound method (introducing a fairness coefficient): The method proposed in this application introduces a fairness coefficient on the basis of the optimization objective of the branch and bound method, comprehensively considering the optimal transmission rate and fairness; Random strategy: The UAV randomly accesses users using a scheme that satisfies the combined constraints.
[0331] Comparison chart as follows Figure 7 and Figure 8 As shown, they illustrate the effects of different calculation strategies on cumulative rewards and secondary user fairness coefficients as a function of time slots. (Refer to...) Figure 7 Under different time slots, the cumulative reward from largest to smallest is as follows: Branch and bound (introducing a fairness coefficient) > Greedy strategy (introducing a fairness coefficient) > Branch and bound > Random strategy > Greedy strategy. The reasons for this can be summarized as follows:
[0332] Algorithms that incorporate fairness considerations offer higher fairness in user service, while other algorithms that do not consider fairness, although improving the overall transmission rate to some extent, reduce user service fairness, resulting in lower reward values. Compared to greedy strategies, branch and bound methods consider interference with the primary user, thus having lower penalty terms in the reward and outperforming greedy strategies in terms of reward performance. Compared to greedy strategies, random strategies, while providing a lower transmission rate, maintain a higher fairness coefficient, whereas greedy strategies prioritize serving certain users, resulting in an extremely low fairness coefficient and therefore a very small total reward value.
[0333] Reference Figure 8 , Figure 8 This chart shows the trend of user fairness coefficients changing with each time slot under different access methods. From Figure 8 The results show that when the number of time slots per round is greater than 10, the fairness changes little under different user access methods. The figure also reveals that user access methods incorporating a fairness coefficient outperform those without considering fairness. The branch-and-bound method, because it addresses the issue of interference with the PU (User Unit) and causes changes in the UAV access scheme for secondary users, also maintains high fairness in user services.
[0334] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated in the embodiments of this application, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages in other steps. It is understood that the steps in different embodiments can be freely combined as needed, and all non-contradictory solutions formed by such combinations are within the scope of protection of this application.
[0335] Based on the same inventive concept, this application also provides an anti-interference drone scheduling device for implementing the anti-interference drone scheduling method described above. The solution provided by this device is similar to the implementation described in the above method; therefore, the specific limitations in one or more embodiments of the anti-interference drone scheduling device provided below can be found in the limitations of the anti-interference drone scheduling method described above, and will not be repeated here.
[0336] In one exemplary embodiment, such as Figure 9 As shown, an anti-interference drone scheduling device is provided, comprising:
[0337] The acquisition module 902 is used to acquire local environmental information and user status information collected by each cognitive UAV;
[0338] The communication processing module 904 is used to analyze the channel gain from each cognitive UAV to each user, the channel gain from each interfering UAV to each user, and the channel gain from each base station to each user in the t-th time slot based on local environmental information and user status information through a pre-built UAV wireless communication service model, and to determine the received data rate of each user in each sub-frequency band in the t-th time slot.
[0339] The solution module 906 is used to solve the sub-frequency band allocation strategy and secondary user access strategy for each cognitive UAV based on the received data rate of each user in each sub-frequency band within the t-th time slot using a pre-built UAV scheduling and communication resource management model. The UAV scheduling and communication resource management model is used to maximize the effective cumulative transmission rate and service fairness under preset constraints.
[0340] The optimization module 908 is used to determine action values based on observations through a multi-agent deep reinforcement learning framework. The observations include the position information of each cognitive UAV, the channel gain from each cognitive UAV to each user, and the cumulative service rate of the secondary user. The action values include the flight turning angle of each cognitive UAV.
[0341] The scheduling module 910 is used to schedule each cognitive UAV based on sub-frequency band allocation strategy, secondary user access strategy, and flight turning angle.
[0342] The aforementioned anti-interference drone scheduling device achieves real-time perception of the radio environment and dynamic allocation of spectrum resources based on the drone wireless communication service model, drone scheduling, and communication resource management model. It introduces an optimization objective to maximize the effective cumulative transmission rate and service fairness, reducing service gaps for edge users and significantly improving service fairness for secondary users, thus meeting the balanced service needs of a massive number of IoT terminals. Drone scheduling is based on sub-band allocation strategies, secondary user access strategies, and flight turning angles, improving the system's anti-interference capability and spectrum utilization efficiency in complex electromagnetic environments. This ensures service quality while enhancing overall network coverage performance and operational stability. Through a multi-agent collaborative reinforcement learning framework, it effectively improves the trajectory optimization efficiency of drones in dynamic scenarios, accelerates the system's recovery speed during network topology changes, and meets the requirements of full coverage and low latency.
[0343] In an exemplary embodiment, the UAV scheduling and communication resource management model includes a cognitive UAV spectrum resource allocation problem model and a secondary user access problem model. The solution module 906 is further configured to solve the cognitive UAV spectrum resource allocation problem model based on a greedy strategy, according to the received data rate of each user in each sub-frequency band within the t-th time slot, to determine the sub-frequency band allocation strategy for each cognitive UAV; determine the channel interference signal-to-noise ratio matrix based on the sub-frequency band allocation strategy for each cognitive UAV; and solve the secondary user access problem model based on a branch-and-bound algorithm, according to the channel interference signal-to-noise ratio matrix, to determine the secondary user access strategy.
[0344] In an exemplary embodiment, the anti-interference drone scheduling device further includes an update module, which is used to determine the reward value based on the received data rate, service fairness ratio, sub-frequency band allocation strategy, and secondary user access strategy of each user in each sub-frequency band within the t-th time slot; store the observation value, action value, and reward value as empirical data; and periodically update the parameters of the policy network and evaluation network of the deep reinforcement learning framework based on the stored empirical data.
[0345] The modules in the aforementioned anti-interference drone dispatching device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the corresponding operations of each module.
[0346] In one exemplary embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 10 As shown, this computer device includes a processor, memory, input / output interfaces (I / O), and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The database stores empirical data. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communicating with external terminals via a network connection. When the computer program is executed by the processor, it implements an interference-resistant UAV scheduling method.
[0347] Those skilled in the art will understand that Figure 10 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0348] In one exemplary embodiment, a computer device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above-described method embodiments.
[0349] In one exemplary embodiment, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements the steps in the above-described method embodiments.
[0350] In one exemplary embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above-described method embodiments.
[0351] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.
[0352] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.
[0353] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.
[0354] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. An anti-interference unmanned aerial vehicle scheduling method, characterized in that, The method comprises: acquiring local environment information collected by each cognitive unmanned aerial vehicle and user state information; based on the local environment information and the user state information, analyzing channel gains of each cognitive unmanned aerial vehicle to each user, channel gains of each interfering unmanned aerial vehicle to each user, and channel gains of each base station to each user in a tth time slot, and determining a receiving data rate of each user in each sub-band in the tth time slot; based on the receiving data rate of each user in each sub-band in the tth time slot, solving a pre-constructed unmanned aerial vehicle scheduling and communication resource management model to determine a sub-band allocation strategy of each cognitive unmanned aerial vehicle and a secondary user access strategy; the unmanned aerial vehicle scheduling and communication resource management model is used to maximize effective cumulative transmission rate and service fairness under a preset constraint condition; based on a multi-agent deep reinforcement learning framework, determining an action value according to an observation value; the observation value comprises position information of each cognitive unmanned aerial vehicle, channel gains of each cognitive unmanned aerial vehicle to each user, and a cumulative service rate of a secondary user, and the action value comprises a flight steering angle of each cognitive unmanned aerial vehicle; based on the sub-band allocation strategy, the secondary user access strategy, and the flight steering angle, scheduling each cognitive unmanned aerial vehicle.
2. The method of claim 1, wherein, The unmanned aerial vehicle scheduling and communication resource management model comprises a cognitive unmanned aerial vehicle spectrum resource allocation problem model and a secondary user access problem model; based on the receiving data rate of each user in each sub-band in the tth time slot, solving the pre-constructed unmanned aerial vehicle scheduling and communication resource management model to determine the sub-band allocation strategy of each cognitive unmanned aerial vehicle and the secondary user access strategy, comprising: based on a greedy strategy, solving the cognitive unmanned aerial vehicle spectrum resource allocation problem model according to the receiving data rate of each user in each sub-band in the tth time slot to determine the sub-band allocation strategy of each cognitive unmanned aerial vehicle; determining a channel interference signal-to-noise ratio matrix according to the sub-band allocation strategy of each cognitive unmanned aerial vehicle; based on a branch and bound algorithm, solving the secondary user access problem model according to the channel interference signal-to-noise ratio matrix to determine a secondary user access strategy.
3. The method of claim 2, wherein, The cognitive unmanned aerial vehicle spectrum resource allocation problem model is represented as: ; wherein: ; ; The primary user set is denoted as , the secondary user set is denoted as , the cognitive UAV set is denoted as , and the spectrum resource set is denoted as ; a sub-band allocation strategy of each cognitive unmanned aerial vehicle in the tth time slot, a cumulative service number of the tth cognitive unmanned aerial vehicle to the nth secondary user in the t-1th time slot, a cumulative service number of the tth cognitive unmanned aerial vehicle to the nth secondary user in the t-1th time slot, a cumulative service number of the tth cognitive unmanned aerial vehicle to the nth secondary user in the t-1th time slot, a cumulative service number of the tth cognitive unmanned aerial vehicle to the nth secondary user in the t-1th time slot, a channel interference signal-to-noise ratio of the tth cognitive unmanned aerial vehicle transmitting data to the nth secondary user on the tth sub-band in the tth time slot, a channel interference signal-to-noise ratio of the tth cognitive unmanned aerial vehicle transmitting data to the nth secondary user on the tth sub-band in the tth time slot, a channel interference signal-to-noise ratio of the tth cognitive unmanned aerial vehicle transmitting data to the nth secondary user on the tth sub-band in the tth time slot. The secondary user access problem model is represented as: ; wherein: ; ; ; ; This represents the access strategy for secondary users within the t-th time slot. This represents the channel interference signal-to-noise ratio matrix. Indicates that in determining In the case of the first A cognitive drone in the t-th time slot Upwards from the first sub-frequency band The transmission rate of each user Indicates the penalty coefficient. Indicates a penalty item. Indicates the first Minimum total transmission rate requirement for each primary user Used to describe normalization priority The weight representing the access priority of secondary users. Indicates the first The cumulative service rate of each user from the first time slot to the t-th time slot. These are the preset adjustment parameters. Indicates the first The primary user in the t-th time slot... The received data rate of each sub-band.
4. The method of claim 1, wherein, After scheduling each cognitive unmanned aerial vehicle based on the sub-band allocation strategy, the secondary user access strategy, and the flight steering angle, the method further comprises: determining a reward value according to the receiving data rate of each user in each sub-band in the tth time slot, a service fairness ratio, the sub-band allocation strategy, and the secondary user access strategy; storing the observation value, the action value, and the reward value as experience data; periodically updating parameters of a policy network and an evaluation network of the deep reinforcement learning framework based on the stored experience data.
5. The method of claim 4, wherein, The reward value is determined by the following formula: ; wherein, denotes the frequency band allocated to the cognitive UAV at the t-th time slot and the set of users is denoted as: ; represents the transmission rate of the tth cognitive user to the tth secondary user on the tth sub-band at the tth time slot, represents the transmission rate of the tth cognitive user to the tth secondary user on the tth sub-band at the tth time slot, represents the transmission rate of the tth cognitive user to the tth secondary user on the tth sub-band at the tth time slot, represents the transmission rate of the tth cognitive user to the tth secondary user on the tth sub-band at the tth time slot, represents the transmission rate of the tth cognitive user to the tth secondary user on the tth sub-band at the tth time slot, represents the service fairness ratio, represents the penalty coefficient, represents the penalty term, represents the sub-band allocation strategy value corresponding to the tth cognitive user, represents the sub-band allocation strategy value corresponding to the tth cognitive user, represents the secondary user access strategy value corresponding to the tth secondary user, represents the secondary user access strategy value corresponding to the tth secondary user.
6. The method according to any one of claims 1 to 5, characterized in that, The unmanned aerial vehicle wireless communication service model is represented as: ; ; Indicates the first A cognitive drone passes through the first time slot t. Sub-band to the first Signal-to-noise ratio of data transmitted by each user; A binary variable used to represent the first... Does the cognitive drone use the [missing information] in the t-th time slot? Sub-frequency band to the first Each user transmits data; This represents white noise interference in channel transmission. , and These represent the wireless transmission power of the cognitive drone, the jamming drone, and the base station, respectively. Indicates the t-th time slot. The first cognitive drone to the first Channel gain for each user; Indicates the t-th time slot. The jamming drone to the first Channel gain for each user; Indicates the first The base station within the first time slot to the first Channel gain for each user; A binary variable used to represent the first... Does the interfering drone affect the first time slot in the t-th time slot? This causes interference to all users within the sub-band; Indicates the t-th time slot. The first main user in the Signal-to-noise ratio of each sub-band This indicates white noise interference in channel transmission; Indicates the first The base station within the first time slot to the first Channel gain for each primary user; Indicates the t-th time slot. The first cognitive drone service In the case of a single user, up to the first Channel gain for each primary user; Indicates the t-th time slot. The jamming drone to the first Channel gain for each primary user; The received data rate of the tth time slot of the th primary user in the th sub-band is determined by the following equation: ; The received data rate of the tth time slot for the nth sub-band of the mth secondary user is determined by the following equation: ; wherein, denotes the received data rate of the t-th time slot for the i-th primary user in the j-th sub-band, denotes the received data rate of the t-th time slot for the i-th primary user in the j-th sub-band, denotes the received data rate of the t-th time slot for the i-th primary user in the j-th sub-band, denotes the received data rate of the t-th time slot for the i-th primary user in the j-th sub-band. denotes the received data rate of the t-th time slot for the i-th primary user in the j-th sub-band. denotes the received data rate of the t-th time slot for the i-th 7. An anti-interference unmanned aerial vehicle scheduling device, characterized in that, The device comprises: An acquisition module is configured to acquire local environment information and user state information collected by each cognitive UAV; A communication processing module is configured to analyze channel gains from each cognitive UAV to each user, channel gains from each interfering UAV to each user, and channel gains from each base station to each user in the tth time slot based on the local environment information and the user state information through a pre-constructed UAV wireless communication service model, and determine a receiving data rate of each user in each sub-band in the tth time slot; A solving module is configured to determine a sub-band allocation strategy of each cognitive UAV and a secondary user access strategy based on the receiving data rate of each user in each sub-band in the tth time slot and by using a pre-constructed UAV scheduling and communication resource management model for solving, the UAV scheduling and communication resource management model being configured to maximize effective cumulative transmission rate and service fairness under a preset constraint condition; An optimization module is configured to determine an action value according to an observation value through multi-agent based on a deep reinforcement learning framework, the observation value including position information of each cognitive UAV, channel gains from each cognitive UAV to each user, and cumulative service rate of a secondary user, and the action value including a flight steering angle of each cognitive UAV; A scheduling module is configured to schedule each cognitive UAV based on the sub-band allocation strategy, the secondary user access strategy, and the flight steering angle.
8. A computer device comprising a memory and a processor, the memory storing a computer program, characterized in that, The processor executes the computer program to implement the steps of the method in any one of claims 1 to 6.
9. A computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the method in any one of claims 1 to 6.
10. A computer program product comprising a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the method in any one of claims 1 to 6.