Frequency spectrum sharing method and system for assisting mobile edge calculation of fixed unmanned aerial vehicle group
By constructing an intelligent collaborative decision-making model based on a multi-agent deep deterministic policy gradient algorithm, the static high concurrency problem of spectrum resource allocation in a fixed UAV formation is solved, efficient sharing and multi-objective optimization of spectrum resources are achieved, and the system's coverage capability and resource utilization are improved.
Patent Information
- Application Number
- CN202510905789.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-02
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2045-07-02
AI Technical Summary
Existing technologies are unable to effectively handle the problem of static, highly concurrent spectrum resource allocation in fixed drone formations, resulting in increased co-frequency interference. Centralized optimization algorithms have high computational complexity and are unable to respond to dynamic environmental changes in real time, resulting in insufficient fairness in resource allocation.
A spectrum sharing method using fixed drone swarms to assist mobile edge computing is adopted. By constructing an intelligent collaborative decision-making model based on a multi-agent deep deterministic policy gradient algorithm, the joint optimization objectives of coverage efficiency, spectrum efficiency, energy efficiency and task delay are defined. A centralized training and distributed execution mechanism are adopted to achieve efficient sharing of spectrum resources and multi-objective joint optimization.
It significantly improves the efficiency of spectrum resource utilization and the overall performance of the system, taking into account high coverage, low-interference communication and low-latency task offloading, and is particularly suitable for high-concurrency scenarios such as dense urban communications and emergency services.
Smart Images

Figure CN120751393A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of wireless communication technologies, and in particular to a spectrum sharing method and system for fixed drone swarm-assisted mobile edge computing. Background Art
[0002] With the rapid development of drone-based edge computing technology, fixed drone fleets, with their stable topology and controllable coverage, have become a key deployment model for edge service networks in dense urban areas, emergency services, and other scenarios. However, the dense deployment of airborne nodes leads to intense competition for spectrum resources, making traditional dynamic spectrum access (DSA) methods difficult to adapt to the high-concurrency service requirements of static fleets.
[0003] Zhou, Hu, and others explored three drone-enabled MEC architectures. These architectures significantly improved computing performance and reduced execution latency by integrating drones into MEC networks. However, these architectures have low cost-benefit, are difficult to implement, and decision theory has great difficulties in determining thresholds.
[0004] Chen Nan and others conducted in-depth research on the energy efficiency optimization problem in MEC systems, using compressed sensing algorithms and one-dimensional search. Compared with conventional methods, the computational complexity of this method is reduced, but its feature selection is more difficult and engineering practice is difficult.
[0005] Zhang Cheng and others used a multi-objective optimization framework to analyze the optimal trade-off between energy and time, providing flexible strategies for different scenarios. For non-convex constraints, they constructed a local convex approximation using a first-order Taylor expansion, gradually approaching the optimal solution. However, offline training relies on historical data, and dynamic environmental changes (such as sudden channel changes) can affect online execution. Furthermore, the algorithm's high complexity makes real-time application difficult in large-scale device scenarios.
[0006] It can be seen from the prior art that there are the following deficiencies: (1) Dynamic spectrum access method: This method relies on a dynamic adjustment mechanism and cannot effectively handle the static and highly concurrent spectrum resource allocation problem in a fixed formation, resulting in increased co-channel interference. (2) Centralized optimization algorithm: The computational complexity is high and it is difficult to respond to dynamic changes in the environment in real time; (3) Resource allocation fairness: Existing methods do not fully consider the collaborative optimization of multiple objectives such as user coverage efficiency (CE), spectrum efficiency (SE), and energy efficiency (EE), resulting in insufficient service fairness. Summary of the Invention
[0007] The present invention provides a spectrum sharing method and system for fixed drone swarm-assisted mobile edge computing, which can achieve efficient sharing of spectrum resources and multi-objective joint optimization under complex electromagnetic environments.
[0008] In a first aspect, the present invention provides a spectrum sharing method for mobile edge computing assisted by a fixed drone swarm, comprising: S1. Build a mobile edge computing network assisted by a fixed drone formation, define coverage efficiency CE, spectrum efficiency SE, energy efficiency EE, and offload delay The fixed formation consists of several drone clusters. A drone cluster consists of a cluster head drone and several cluster member drones. Each drone in the formation is an independent intelligent entity. The cluster head is responsible for performing offloading tasks and calculating the drone's internal deep learning and resource optimization tasks. All drones can be used as aerial base stations to provide resources and network coverage services. S2. Build an agent collaborative decision-making model based on a multi-agent deep deterministic policy gradient algorithm, defining the state space, action space, and reward function. The state space consists of three parts: the state of the drone formation, the state of the user device, and the dynamics of the environment. The action space covers three types of decisions: formation layout adjustment, resource allocation, and task scheduling. The reward function consists of four parts: coverage efficiency reward, resource efficiency reward, fairness reward, and constraint penalty. S3. Use centralized training and distributed execution mechanisms to learn and train the spectrum sharing strategy of the drone swarm to obtain the optimal spectrum sharing solution; the solution includes: drone location, allocation coefficient, power control coefficient and access strategy.
[0009] According to the spectrum sharing method for fixed drone swarm assisted mobile edge computing provided by the present invention, the mobile edge computing network includes: a macro base station equipped with a MEC server, an Internet of Things network composed of multiple base stations equipped with computing servers; a set of edge Internet of Things devices distributed on the ground, denoted as There are M drones divided into C clusters distributed in the air. Each communication terminal adopts a single antenna working mode. There are UAVs, the UAV cluster head is denoted as .
[0010] According to the spectrum sharing method for fixed drone swarm assisted mobile edge computing provided by the present invention, the coverage efficiency CE, spectrum efficiency SE, energy efficiency EE, fairness in step S1 and uninstall delays The joint optimization objective is expressed as: ; ; ; ; ; in, represents a fixed formation set, is the maximum coverage radius of a single machine, is the edge device location, is the coordinate of the UAV’s position on the ground, is the indicator function, Indicates drone and edge devices The access relationship between Indicates drone The frequency allocation scheme in Indicates when =1, UAV Use frequency n and edge devices i Signal-to-noise ratio during communication, represents the transmission power of UAV m, Indicates drone bandwidth, Indicates the transmission delay, Indicates processing delay, Indicates the latency of local computation.
[0011] According to the spectrum sharing method for fixed drone swarm-assisted mobile edge computing provided by the present invention, in step S1, the joint optimization objective constraint is expressed as follows:
[0012]
[0013]
[0014]
[0015]
[0016]
[0017]
[0018]
[0019] in, Indicates drone The maximum transmit power, Indicates drone Maximum bandwidth, Indicates the minimum distance of the drone, Indicates the maximum frequency of the CPU, Indicates the maximum computing power of the drone.
[0020] According to the spectrum sharing method for fixed drone swarm-assisted mobile edge computing provided by the present invention, the state space in step S2 is composed of three parts: drone formation state, user device state, and environment dynamics, which are specifically defined as:
[0021] The action space in step S2 covers three types of decisions: formation layout adjustment, resource allocation, and task scheduling, and is defined as:
[0022] in, Representation device The CPU frequency allocated in time slot n; The reward function consists of four parts: coverage efficiency reward, resource efficiency reward, fairness reward, and constraint penalty. The specific definitions are as follows:
[0023] in, ,satisfy , reflecting the priority of each goal, For penalty items.
[0024] According to the spectrum sharing method for fixed drone swarm assisted mobile edge computing provided by the present invention, in step S3, a centralized training and distributed execution mechanism is adopted to learn and train the drone swarm spectrum sharing strategy, specifically including: In the offline phase, the agent generates formation distribution through global information Sharing strategies with resources , ensuring that coverage CE and balance goals are met; In the online phase, each agent dynamically adjusts the spectrum allocation coefficient based on local observations , power control coefficient , access strategy To respond to users' real-time task offloading needs; Among them, global information includes: user device distribution, channel status and energy constraints; local observations include: number of connected devices, remaining energy and channel interference.
[0025] According to the spectrum sharing method for fixed drone swarm-assisted mobile edge computing provided by the present invention, the penalty term is designed as follows: Insufficient Coverage Penalty:
[0026] Collision Avoidance Violation Penalties:
[0027] Delay Exceeding Limit Penalty:
[0028] in, Indicates the maximum allowed delay.
[0029] In a second aspect, the present invention further provides a spectrum sharing system for mobile edge computing assisted by a fixed drone swarm, comprising: Network building module, used to build a mobile edge computing network assisted by fixed drone formation, defining coverage efficiency CE, spectrum efficiency SE, energy efficiency EE, and offloading delay The fixed formation consists of several drone clusters. A drone cluster consists of a cluster head drone and several cluster member drones. Each drone in the formation is an independent intelligent entity. The cluster head is responsible for performing offloading tasks and calculating the drone's internal deep learning and resource optimization tasks. All drones can be used as aerial base stations to provide resources and network coverage services. The decision model construction module is used to build an agent collaborative decision model based on a multi-agent deep deterministic policy gradient algorithm, defining the state space, action space, and reward function. The state space consists of three parts: the state of the drone formation, the state of the user device, and the dynamics of the environment; the action space covers three types of decisions: formation layout adjustment, resource allocation, and task scheduling; and the reward function consists of four parts: coverage efficiency reward, resource efficiency reward, fairness reward, and constraint penalty. The strategy execution module is used to learn and train the spectrum sharing strategy of the drone swarm using a centralized training and distributed execution mechanism to obtain the optimal spectrum sharing solution; the solution includes: drone location, allocation coefficient, power control coefficient and access strategy.
[0030] In a third aspect, the present invention provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the steps of the spectrum sharing method for fixed drone swarm-assisted mobile edge computing as described above are implemented.
[0031] In a fourth aspect, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the spectrum sharing method for fixed drone swarm-assisted mobile edge computing as described in any of the above.
[0032] The spectrum sharing method and system for fixed drone swarm-assisted mobile edge computing provided by the present invention significantly improve the efficiency of spectrum resource utilization and the overall performance of the system through multi-objective joint optimization and intelligent collaborative decision-making mechanism. The present invention innovatively constructs a fixed drone formation network, defines the joint optimization objectives of coverage efficiency, spectrum efficiency, energy efficiency and task delay, and realizes dynamic collaborative decision-making of spectrum sharing strategies based on a multi-agent deep reinforcement learning framework. It adopts a centralized training and distributed execution mechanism, uses global information to generate formation layout and resource allocation plans in the offline phase, and adjusts spectrum allocation, power control and access strategies in real time through local observation in the online phase, effectively taking into account high coverage, low-interference communication and low-latency task offloading. Its constraint-driven design (such as coverage guarantee, anti-collision mechanism, and delay control) ensures system stability, and is particularly suitable for high-concurrency scenarios such as dense urban communications and emergency services. It is significantly superior to traditional dynamic spectrum sharing methods in terms of coverage capability, resource utilization and algorithm convergence efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction is given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0034] Figure 1 Schematic diagram of the flow of the spectrum sharing method for fixed drone swarm-assisted mobile edge computing provided by the present invention; Figure 2 is a schematic diagram of a mobile edge computing network provided by the present invention; Figure 3 This is a flow chart of the multi-agent deep deterministic policy gradient algorithm in the present invention; Figure 4 It is a sub-goal trade-off diagram of the present invention under different weight combinations; Figure 5 This is a comparison chart of the cumulative rewards and the number of training rounds in the present invention; Figure 6 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION
[0035] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0036] It should be noted that, in the description of the embodiments of the present invention, the terms "include", "comprising" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprising a ..." do not exclude the presence of other identical elements in the process, method, article or device comprising the elements. Unless otherwise expressly specified and defined, the terms "installed", "connected" and "connected" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or an indirect connection through an intermediate medium, or it can be a communication between the internal parts of two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to the specific circumstances.
[0037] The terms "first," "second," and the like in this application are used to distinguish similar objects, and are not used to describe a specific order or precedence. It should be understood that the terms used in this manner are interchangeable, where appropriate, so that the embodiments of this application can be implemented in an order other than that illustrated or described herein. Furthermore, the objects distinguished by "first," "second," and the like generally refer to a class of objects and do not limit the number of objects; for example, the first object can be one or more. Furthermore, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the objects connected are in an "or" relationship.
[0038] The following combination Figures 1-6 The present invention describes a spectrum sharing method and system for fixed drone swarm-assisted mobile edge computing.
[0039] Figure 1 This is a flow chart of the spectrum sharing method for fixed drone swarm assisted mobile edge computing provided by the present invention, as shown in FIG. Figure 1 As shown, including but not limited to the following steps: S1: Build a mobile edge computing network assisted by a fixed drone formation, define coverage efficiency CE, spectrum efficiency SE, energy efficiency EE, and offload delay joint optimization goal.
[0040] 1.1) Figure 2 is a schematic diagram of the mobile edge computing network provided by the present invention, such as Figure 2As shown in Figure 1, the mobile edge computing network consists of a macro base station equipped with a MEC server and many small base stations to form a large-scale IoT network. The small base station in each cell has its own computing server. Let L represent a set of edge IoT devices (sensors, mobile users, etc.), denoted as In this multi-drone assisted network communication facility, there are Cluster Each communication terminal uses a single antenna working mode and the frequency reuse factor is set to 1 to maximize the utilization of wireless network resources. is divided into B orthogonal subcarrier channels, expressed as . Each cluster covers a network with drones, of which The UAV cluster head is represented as Generally speaking, UCH has better computing performance. Within the same cluster, there is U2U transmission between the UAV cluster node and the UCH head, where , is a collection of drones that perform U2U transmissions, used to execute offloaded edge device tasks and compute deep learning and resource optimization tasks within the drones. All drones can be used as aerial base stations to provide resources and network coverage services to compute edge device workloads and allocate transmission power. If the ground base station is overloaded or fails, the quality of the communication channel will degrade and may fall below the minimum threshold, resulting in a failure to communicate normally. At this point, edge devices must connect to drones to offload their computationally intensive tasks and request resources.
[0041] 1.2) Coverage The drone formation is fixed in position and cannot adjust its coverage range through dynamic flight. It is necessary to optimize the formation topology to improve the regional coverage. The service coverage radius is limited by the static layout and must meet the minimum coverage radius requirements:
[0042] in, is the maximum coverage radius of a single machine, is the edge device location. Define coverage efficiency CE: .in, is an indicator function that measures user coverage.
[0043] 1.3) Spectral efficiency Spectral efficiency SE refers to the number of information bits transmitted per second in a unit bandwidth communication channel in the network, and its unit is In the given network model, the spectrum efficiency SE can be expressed as ; in, 、 and Represent the set of drones, edge devices, and available frequencies, respectively. Indicates drone and edge devices The access relationship between them is: ; in, Indicates drone The frequency allocation scheme in , namely: ; also, Indicates when When the drone Use frequency n and edge devices Signal-to-Interference-Plus-Noise Ratio (SINR) during communication. Respectively represent drones Use frequency n and edge devices The transmission power of the device during communication and the square of the channel gain, Represents drone m and edge device The variance of the additive white Gaussian noise during communication is Denoted as: ; in, When the drone Adopting frequency n and edge devices The interval interference signal during communication is: ; Channel gain Taking into account path loss and shadow fading, it can be expressed as: ; in, is the carrier frequency (Hz), c is the speed of light (m / s), is the path loss index, which is related to the environment.
[0044] 1.4) Energy Efficiency Energy efficiency EE refers to the number of information bits that can be transmitted per unit of energy consumed in the network, and its unit is In the given network model, the energy efficiency EE can be expressed as: ; in, Represents the drone m using frequency n and edge device The bandwidth used for communication.
[0045] 1.5) Fairness The variance of information throughput is used to measure fairness. Therefore, the smaller the variance of user throughput, the better the fairness. In the given network model, fairness is defined as Expressed as: ; Wherein, D{} represents a variance operation. It should be noted that n in the superscript represents frequency n, and n in the bracket () represents time slot n.
[0046] 1.6) Task Delay Assuming edge devices There is a computationally intensive task R that needs to be executed. Let R = {1, 2, 3, ..., R} represent the set of tasks generated by a single edge device. The task profile of each edge device is defined as: .in Indicates the amount of data that needs to be transferred to the UCN when deciding to offload. represents the total number of CPU cycles required to complete this task, and 𝑇max represents the maximum tolerable delay of the task. The task arrival process is modeled as a Poisson process, and the number of tasks arriving in time slot n is The local computing task is represented as , the task offloaded to the UAV network is represented as The task queue update formula is: ; in, , [x′]+=max(x′,0).
[0047] When drones assist ground edge devices in air-to-ground (A2G) communication infrastructure, the achievable data transmission rate can be expressed as: ; The latency of a task being computed locally is expressed as: ; in, It is a device In the time slot Assigned CPU frequency.
[0048] If the task needs to be offloaded, the transmission delay and processing delay Respectively expressed as: ; ; in, It is the computing power of the drone.
[0049] Total task processing delay It can be expressed as: ; SE, EE, and fairness are determined by the distance between nodes, frequency allocation scheme, and SINR. Based on the characteristics of fixed formation, the multi-objective joint optimization problem is described as follows:
[0050] in, represents a fixed formation set, is the multi-objective weight coefficient, satisfying .
[0051] S2. Build an agent collaborative decision-making model based on a multi-agent deep deterministic policy gradient algorithm, defining the state space, action space, and reward function. The state space consists of three parts: the state of the drone formation, the state of the user device, and the dynamics of the environment; the action space covers three types of decisions: formation layout adjustment, resource allocation, and task scheduling; and the reward function consists of four parts: coverage efficiency reward, resource efficiency reward, fairness reward, and constraint penalty.
[0052] Figure 3 This is the flow chart of the multi-agent deep deterministic policy gradient algorithm in the present invention, such as Figure 3 As shown in the figure, it mainly includes six steps: initialization, interaction, experience storage, sampling and updating, soft updating of target network, and repetition. The basic idea of the process implementation is: (1) Initialization: Initialize the Actor network and Critic network, as well as the corresponding target network for each agent; (2) Interaction: The agent interacts with the environment. At each time t, each agent selects an action according to its strategy, and then the environment returns the next state and reward; (3) Experience storage: Store the state, action, reward, and next state in the experience replay pool. (4) Sampling and updating: Sample a batch from the experience replay pool, use the temporal difference error to construct the MSE loss function, and update the Critic and Actor networks of each agent through gradient descent. (5) Soft updating of target network: Update the parameters of the target network in a slow manner relative to the Critic and Actor networks. (6) Repetition: Repeat the above process until training is completed. The ultimate goal of the MADDPG algorithm is to enable each agent to learn the optimal strategy in a multi-agent environment through centralized training and decentralized execution, thereby achieving the best global performance in a collaborative or competitive environment.
[0053] (a) State space The state space consists of three parts: UAV formation state, user equipment state and environment dynamics, which are specifically defined as
[0054] (b) Action space The action space covers three types of decisions: formation layout adjustment, resource allocation and task scheduling, which are defined as
[0055] (c) Reward function The reward function consists of four parts: coverage efficiency reward, resource efficiency reward, fairness reward and constraint penalty. The specific definitions are as follows
[0056] in ,satisfy , reflecting the priority of each goal.
[0057] For penalty items The design is as follows: 1) Penalty for insufficient coverage: ; 2) Penalties for collision avoidance violations: ; 3) Penalty for exceeding the delay limit: .
[0058] in, They represent the penalty coefficients for insufficient coverage, collision avoidance violations, and excessive delays, respectively; Indicates the maximum allowed delay.
[0059] S3: Using a centralized training and distributed execution mechanism to learn and train the spectrum sharing strategy of the drone swarm to obtain the optimal spectrum sharing solution. The solution includes: drone location, allocation coefficient, power control coefficient, and access strategy.
[0060] Aiming at the multi-objective optimization requirements of fixed drone group formation in assisted mobility scenarios. In the offline phase, the agent generates formation distribution based on global information (including user device distribution, channel status and energy constraints). Sharing strategies with resources , ensuring that coverage CE and balance goals are met. In the online phase, each agent dynamically adjusts the spectrum allocation coefficient based on local observations (number of connected devices, remaining energy, and channel interference) , power control coefficient , access strategy and computing resources to respond to users' real-time task offloading needs. It mainly includes six steps: initialization, interaction, experience storage, sampling and updating, soft update of target network and repetition.
[0061] The training of the entire network can be decomposed into a three-stage optimization process, each stage contains a specific parameter update mechanism: (1) Critic network parameter optimization According to the characteristics of UAV swarm collaborative tasks, The construction of the function needs to consider the heterogeneity constraints of the state space. Value Definition ,in Represents a set of policies, is the composite state space, is the multi-agent action vector. Since the MADDPG algorithm is model-independent, the algorithm directly replays the experience buffer Random sampling Group state transfer data, achieve parameters by minimizing the loss function The mathematical expression of the loss function is: ; Where the input contains the cluster joint state and the set of actions of each agent ,Target The values are generated by the target critic network: ; in, is the target critic network, is the target actor network, It is The nth observation in the sample.
[0062] (2) Actor Network Parameter Optimization The deterministic policy gradient ascent method is used to update the parameters, and the gradient calculation expression is: ; in, Indicates action Take-by strategy According to observation The strategy only requires local observation information. You can generate actions without relying on global state.
[0063] (3) Target network parameter optimization To ensure training stability, a soft update mechanism is used instead of direct parameter copying: .
[0064] The effect of the present invention can be further illustrated by simulation: 1. Simulation conditions: A laptop with a memory of no less than 8GB.
[0065] 2. Simulation Content: The simulation was conducted using TensorFlow 2.5.0 on an Ubuntu server. Each UAV's actor and critic networks used a multilayer perceptron architecture. The intermediate layers of the actor and critic networks used ReLU as the nonlinear activation function, and the output layer of the actor network used the Tanh activation function.
[0066] Control the deployment of drone formations and monitor the dynamic changes in cumulative rewards over the training cycle. In simulation, the objective weight coefficient determines the priority of the optimization direction. Different weight combinations will cause the cumulative rewards to converge to different Pareto optimal solutions.
[0067] Figure 4 It is a sub-goal trade-off diagram of the present invention under different weight combinations; Figure 5 This is a comparison chart of the cumulative rewards and the number of training rounds in the present invention.
[0068] In another aspect, the present invention further provides a spectrum sharing system for mobile edge computing assisted by a fixed drone swarm, comprising: Network building module, used to build a mobile edge computing network assisted by fixed drone formation, defining coverage efficiency CE, spectrum efficiency SE, energy efficiency EE, and offloading delay The fixed formation consists of several drone clusters. A drone cluster consists of a cluster head drone and several cluster member drones. Each drone in the formation is an independent intelligent entity. The cluster head is responsible for performing offloading tasks and calculating the drone's internal deep learning and resource optimization tasks. All drones can be used as aerial base stations to provide resources and network coverage services. The decision model construction module is used to build an agent collaborative decision model based on a multi-agent deep deterministic policy gradient algorithm, defining the state space, action space, and reward function. The state space consists of three parts: the state of the drone formation, the state of the user device, and the dynamics of the environment; the action space covers three types of decisions: formation layout adjustment, resource allocation, and task scheduling; and the reward function consists of four parts: coverage efficiency reward, resource efficiency reward, fairness reward, and constraint penalty. The strategy execution module is used to learn and train the spectrum sharing strategy of the drone swarm using a centralized training and distributed execution mechanism to obtain the optimal spectrum sharing solution; the solution includes: drone location, allocation coefficient, power control coefficient and access strategy.
[0069] It should be noted that the spectrum sharing system for fixed drone swarm assisted mobile edge computing provided in an embodiment of the present invention can, during specific operation, execute the spectrum sharing method for fixed drone swarm assisted mobile edge computing described in any of the above embodiments, which will not be elaborated in this embodiment.
[0070] The spectrum sharing method and system for fixed drone swarm-assisted mobile edge computing provided by the present invention significantly improve the efficiency of spectrum resource utilization and the overall performance of the system through multi-objective joint optimization and intelligent collaborative decision-making mechanism. The present invention innovatively constructs a fixed drone formation network, defines the joint optimization objectives of coverage efficiency, spectrum efficiency, energy efficiency and task delay, and realizes dynamic collaborative decision-making of spectrum sharing strategies based on a multi-agent deep reinforcement learning framework. It adopts a centralized training and distributed execution mechanism, uses global information to generate formation layout and resource allocation plans in the offline phase, and adjusts spectrum allocation, power control and access strategies in real time through local observation in the online phase, effectively taking into account high coverage, low-interference communication and low-latency task offloading. Its constraint-driven design (such as coverage guarantee, anti-collision mechanism, and delay control) ensures system stability, and is particularly suitable for high-concurrency scenarios such as dense urban communications and emergency services. It is significantly superior to traditional dynamic spectrum sharing methods in terms of coverage capability, resource utilization and algorithm convergence efficiency.
[0071] Figure 6 Schematic diagram of the structure of the electronic device provided by the present invention, such as Figure 6 As shown, the electronic device may include: a processor 610, a communication interface 620, a memory 630, and a communication bus 640, wherein the processor 610, the communication interface 620, and the memory 630 communicate with each other via the communication bus 640. The processor 610 may call the logic instructions in the memory 630 to execute the spectrum sharing method for fixed drone swarm-assisted mobile edge computing.
[0072] Furthermore, the logic instructions in the aforementioned memory 630 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0073] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the spectrum sharing method for fixed drone swarm-assisted mobile edge computing provided in the above-mentioned embodiments. Through the above description of the embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a necessary general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods described in each embodiment or certain portions of the embodiments.
[0074] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A spectrum sharing method for fixed drone swarm-assisted mobile edge computing, characterized in that: include: S1. Build a mobile edge computing network assisted by a fixed drone formation, define coverage efficiency CE, spectrum efficiency SE, energy efficiency EE, and offload delay The fixed formation consists of several drone clusters. A drone cluster consists of a cluster head drone and several cluster member drones. Each drone in the formation is an independent intelligent entity. The cluster head is responsible for performing offloading tasks and calculating the drone's internal deep learning and resource optimization tasks. All drones can be used as aerial base stations to provide resources and network coverage services. S2. Build an agent collaborative decision-making model based on a multi-agent deep deterministic policy gradient algorithm, defining the state space, action space, and reward function. The state space consists of three parts: the state of the drone formation, the state of the user device, and the dynamics of the environment. The action space covers three types of decisions: formation layout adjustment, resource allocation, and task scheduling. The reward function consists of four parts: coverage efficiency reward, resource efficiency reward, fairness reward, and constraint penalty. S3. Use centralized training and distributed execution mechanisms to learn and train the spectrum sharing strategy of the drone swarm to obtain the optimal spectrum sharing solution; the solution includes: drone location, allocation coefficient, power control coefficient and access strategy.
2. The spectrum sharing method for fixed drone swarm assisted mobile edge computing according to claim 1 is characterized in that: The mobile edge computing network includes: A macro base station equipped with a MEC server, and an IoT network consisting of multiple base stations equipped with computing servers; The set of edge IoT devices distributed on the ground is denoted as ; There are M drones divided into C clusters distributed in the air. Each communication terminal adopts a single antenna working mode. There are UAVs, the UAV cluster head is .
3. The spectrum sharing method for fixed drone swarm assisted mobile edge computing according to claim 2 is characterized in that: In step S1, coverage efficiency CE, spectrum efficiency SE, energy efficiency EE, fairness and uninstall delays The joint optimization objective is expressed as: in, represents a fixed formation set, is the maximum coverage radius of a single machine, is the edge device location, is the coordinate of the UAV’s position on the ground, is the indicator function, Indicates drone and edge devices The access relationship between Indicates drone The frequency allocation scheme in Indicates when =1, UAV Use frequency n and edge devices i Signal-to-noise ratio during communication, represents the transmission power of UAV m, Indicates drone bandwidth, Indicates the transmission delay, Indicates processing delay, represents the delay of local calculation, D{} represents the variance operation, n as a superscript represents the frequency n, and n in the brackets represents the time slot.
4. The spectrum sharing method for fixed drone swarm assisted mobile edge computing according to claim 3 is characterized in that: In step S1, the joint optimization objective constraint is expressed as follows: in, Indicates drone The maximum transmit power, Indicates drone Maximum bandwidth, Indicates the minimum distance of the drone, Indicates the maximum frequency of the CPU, Indicates the maximum computing power of the drone.
5. The spectrum sharing method for fixed drone swarm assisted mobile edge computing according to claim 4 is characterized in that: The state space in step S2 consists of three parts: UAV formation state, user device state, and environment dynamics, which are specifically defined as: The action space in step S2 covers three types of decisions: formation layout adjustment, resource allocation, and task scheduling, and is defined as: in, Representation device The CPU frequency allocated in time slot n; The reward function consists of four parts: coverage efficiency reward, resource efficiency reward, fairness reward, and constraint penalty. The specific definitions are as follows: in, ,satisfy , reflecting the priority of each goal, For penalty items.
6. The spectrum sharing method for fixed drone swarm assisted mobile edge computing according to claim 5, characterized in that: In step S3, a centralized training and distributed execution mechanism is used to learn and train the spectrum sharing strategy of the drone swarm, specifically including: In the offline phase, the agent generates formation distribution through global information Sharing strategies with resources , ensuring that coverage CE and balance goals are met; In the online phase, each agent dynamically adjusts the spectrum allocation coefficient based on local observations , power control coefficient , access strategy To respond to users' real-time task offloading needs; Among them, global information includes: user device distribution, channel status and energy constraints; local observations include: number of connected devices, remaining energy and channel interference.
7. The spectrum sharing method for fixed drone swarm assisted mobile edge computing according to claim 6, characterized in that: The penalty item is designed as follows: Insufficient Coverage Penalty: Collision Avoidance Violation Penalties: Delay Exceeding Limit Penalty: in, Indicates the maximum allowed delay.
8. A spectrum sharing system for fixed drone swarm-assisted mobile edge computing, characterized in that: include: Network building module, used to build a mobile edge computing network assisted by fixed drone formation, defining coverage efficiency CE, spectrum efficiency SE, energy efficiency EE, and offloading delay The fixed formation consists of several drone clusters. A drone cluster consists of a cluster head drone and several cluster member drones. Each drone in the formation is an independent intelligent entity. The cluster head is responsible for performing offloading tasks and calculating the drone's internal deep learning and resource optimization tasks. All drones can be used as aerial base stations to provide resources and network coverage services. The decision model construction module is used to build an agent collaborative decision model based on a multi-agent deep deterministic policy gradient algorithm, defining the state space, action space, and reward function. The state space consists of three parts: the state of the drone formation, the state of the user device, and the dynamics of the environment; the action space covers three types of decisions: formation layout adjustment, resource allocation, and task scheduling; and the reward function consists of four parts: coverage efficiency reward, resource efficiency reward, fairness reward, and constraint penalty. The strategy execution module is used to learn and train the spectrum sharing strategy of the drone swarm using a centralized training and distributed execution mechanism to obtain the optimal spectrum sharing solution; the solution includes: drone location, allocation coefficient, power control coefficient and access strategy.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the spectrum sharing method for fixed drone swarm assisted mobile edge computing as described in any one of claims 1 to 7 are implemented.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the spectrum sharing method for fixed drone swarm assisted mobile edge computing as described in any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Unmanned aerial vehicle group spectrum resource allocation method and system based on cooperative spectrum sensing
CN119485695A
Cognitive unmanned aerial vehicle network spectrum sharing method based on multi-agent reinforcement learning
CN119697646A
Unmanned aerial vehicle group spectrum resource sharing method and device under centralized architecture
CN119743762A
Method and apparatus for assigning frequency resource in non-terrestrial network
US20240022317A1
Task offloading and resource allocation method based on mobile edge computing
WO2024174426A1