A method and device for sharing spectrum resources of drone swarm under centralized architecture
By constructing the intelligent spectrum resource model of the drone swarm under a centralized architecture and using the Markov decision-making process model and other technical means, the problems of rapid convergence and low energy efficiency in the spectrum resource sharing of the drone swarm are solved, and efficient and fair spectrum sharing is achieved.
Patent Information
- Application Number
- CN202510255569.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-05
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-03-05
AI Technical Summary
Under the centralized architecture, the sharing of drone cluster spectrum resources has problems such as rapid convergence and low energy efficiency, and lacks dynamic and fast response capabilities.
By constructing an intelligent spectrum resource model of the drone group, the drone group is divided into multiple clusters, each cluster consisting of cluster head drone and cluster member drone, and cluster head drone communicates directly with the central drone. The Markov decision-making process model and optimization target model are adopted, and resource sharing training is combined with independent two-layer competitive architecture Q network and hybrid exploration strategy, and the spectrum sharing strategy is generated using the weighted aggregate weight of the model contribution rate.
It improves the dynamicity and rapid response capabilities of spectrum sharing in drone groups, improves energy efficiency, and effectively improves spectrum sharing efficiency and fairness.
Smart Images

Figure CN119743762B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of spectrum resource sharing for drone swarms, and more specifically, to a method and device for sharing spectrum resources for drone swarms under a centralized architecture. Background Art
[0002] As drone swarms show great potential in civil and military applications, research on spectrum sharing for drone swarms has become particularly important. The centralized spectrum sharing architecture, developed from a single drone system, has a simple structure and is easy to implement in engineering. The centralized spectrum sharing architecture usually has a central facility, and other nodes are connected to the central facility and directly transmit information to the central facility. The centralized architecture can fully obtain the environmental information perceived by other nodes and find the optimal spectrum sharing solution. At present, the main methods of the centralized spectrum sharing architecture for drone swarms include: frequency division method, time division multiple access (TDMA), frequency division multiple access (FDMA), code division multiple access (CDMA), orthogonal frequency division multiple access (OFDMA), dynamic spectrum access (DSA), spectrum pool, etc.
[0003] Shen Xiangzhi et al. proposed a joint power and spectrum resource allocation method to address the problem of insufficient spectrum resource sharing when drone clusters perform multiple tasks under limited spectrum resources. The drone cluster is divided into multiple groups, which perform tasks separately and are optimized in two dimensions: channel allocation and power allocation. The goal is to maximize the minimum group transmission throughput and ensure the fairness of all group transmission throughputs. A joint power and spectrum resource optimization algorithm is proposed, which decomposes the problem into two sub-problems: channel allocation and transmission power optimization, and uses an improved genetic algorithm and a convex optimization method to solve them respectively. The algorithm improves the utilization of spectrum resources, ensures transmission fairness, and has high algorithm efficiency. However, the algorithm still has a certain degree of complexity and the problem of untimely dynamic allocation of spectrum resources.
[0004] Wu Yizheng and others studied the problem of spectrum resource optimization for multi-UAV information transmission in complex environments. A dynamic information transmission model and spectrum access model were constructed, and a spectrum access method based on multi-user non-coupled queuing was proposed. Through independent decision-making by UAVs, frequency conflicts can be effectively alleviated, and channel utilization and system throughput can be improved. A multi-UAV covert transmission model was constructed, and a joint UAV alliance selection and spectrum resource optimization method was proposed. Through alliance formation game and particle swarm algorithm, the satisfaction of UAV information transmission was maximized under the premise of ensuring covert transmission. A hierarchical game model based on Steinberg game was constructed, and a hierarchical confrontation algorithm was designed. Through UAV alliance formation game and particle swarm algorithm, the satisfaction of UAV transmission was maximized in the presence of intelligent interference, and the confrontation parties reached an equilibrium solution. The algorithm takes into account complex environments, uses game theory methods to improve the efficiency of spectrum resource utilization, considers multi-dimensional optimization, and uses heuristic algorithms to improve the efficiency and convergence speed of the algorithm. However, the algorithm needs to further verify the applicability of the model, which may not be suitable for large-scale UAV networks. The security guarantee mechanism has not been deeply explored, and the cost of UAV mission execution, such as energy consumption and damage, has not been considered.
[0005] Zhao Mingfei and others proposed a method for evaluating the data transmission performance of drone swarms based on weighted processing, aiming to solve the problem that the existing evaluation method focuses too much on the improvement of a single indicator and lacks a comprehensive evaluation of the transmission system. This method divides the indicators into focused indicators, comprehensive indicators and secondary indicators according to the type and characteristics of drone information, and assigns weights to finally obtain the performance scores and comprehensive scores. The algorithm takes into account comprehensiveness, objective evaluation, flexible weights and strong versatility, but the algorithm weight assignment is complex, the evaluation indicators are limited, and the simulation scenario has certain limitations.
[0006] Li Wei and others proposed a spectrum allocation method based on non-cooperative game to address the shortage of spectrum resources and the self-interference and mutual interference problems in drone swarm communication networks. This method adaptively allocates spectrum from three aspects: frequency domain, time domain, and energy domain, aiming to effectively save spectrum resources, avoid interference, and reduce energy consumption. The algorithm saves spectrum resources, avoids interference, reduces energy consumption, and has a good convergence speed, but the algorithm has high computational complexity and high requirements on network topology, and further research is needed on dynamic spectrum allocation strategies. Summary of the invention
[0007] In response to at least one defect or improvement need in the prior art, the present invention provides a method and device for sharing spectrum resources of a swarm of drones under a centralized architecture, which solves the problem of sharing spectrum resources of a swarm of drones under a centralized architecture, has fast convergence and high energy efficiency, and can effectively improve the dynamics and rapid response capabilities of spectrum resource sharing of a swarm of drones.
[0008] To achieve the above-mentioned purpose, according to the first aspect of the present invention, a method for sharing spectrum resources of a drone swarm under a centralized architecture is provided, the method comprising: constructing an intelligent spectrum resource model of a drone swarm, wherein the drone swarm is divided into multiple clusters, each cluster in the drone swarm comprises a cluster head drone and multiple cluster member drones, and the cluster head drone communicates directly with the central drone; constructing a Markov decision process model for describing the drone swarm intelligent spectrum resource sharing process, and constructing an optimization target model, wherein the optimization targets in the optimization target model include the drone swarm transmission rate, cluster fairness rate and inter-cluster fairness rate; constructing an independent double-layer competition architecture Q network input and output model and adopting a hybrid exploration strategy to perform resource sharing training for each cluster, and initializing data training to generate a cluster model; using the weighted aggregation weight based on the model contribution rate of the cluster model to perform drone swarm resource sharing training, and generate a drone swarm spectrum sharing strategy.
[0009] In an exemplary embodiment, the characteristic parameters in the drone swarm intelligent spectrum resource model include the number of channels and the number of available frequency points for each channel. The number of channels includes the number of channels required for communication between the cluster head drones and the central drone, the number of channels required within each cluster, and the number of channels required for communication between cluster heads.
[0010] In an exemplary embodiment, the Markov decision process model is used to describe the intelligent spectrum resource sharing process of the drone swarm, and the optimization target model is constructed, including: determining the state space of the drone swarm spectrum resource sharing plan, determining the action space for changing the current resource sharing plan; determining the state transition probability and introducing a discount factor, rewarding the Q network for interacting with the environment, obtaining the strategy corresponding to the highest reward value, and generating the decayed sum of the rewards; calculating the action value function, calculating the optimal action value function, and solving the learning goal.
[0011] In an exemplary embodiment, the construction of an independent double-layer competitive architecture Q network input-output model and the use of a hybrid exploration strategy to perform resource sharing training for each cluster include: the first layer of the constructed independent double-layer competitive architecture Q network input-output model is a channel selection agent, which is used to convert a state vector containing a frequency allocation table and a transmission power table into a Boolean-valued channel state vector representing a channel working state as an input vector, obtain a conflicting channel after learning, and output a channel selection action; the second layer of the independent double-layer competitive architecture Q network input-output model is a frequency switching agent, which includes an independent environment coding layer and a state analysis layer; wherein the input of the independent environment coding layer is an environment feature vector containing the number of available frequencies of the current channel and the channel noise power of each frequency, and the environment coding layer transmits the compressed environment information to the state analysis layer; the state analysis layer encodes the channel selection action output by the channel selection agent, and combines the coding vector with the current spectrum sharing scheme to form a comprehensive coding vector, the frequency switching agent learns the comprehensive coding vector and the environment feature vector, and finally outputs a frequency change action.
[0012] In an exemplary embodiment, the hybrid exploration strategy is expressed as:
[0013]
[0014] in, It's action The current estimated value of is the total number of attempts so far, It's action The number of times it has been attempted, is the exploration constant used to describe the degree of control exploration.
[0015] In an exemplary embodiment, the using of the weighted aggregation weight based on the model contribution rate of the cluster model to perform drone swarm resource sharing training and generate a drone swarm spectrum sharing strategy includes:
[0016] Determine the weighted aggregation weight , calculated as,
[0017]
[0018] in, is the global model weight, Cluster The weight of Cluster The contribution rate is calculated as:
[0019]
[0020] in, Indicates that the cluster head local model is in the spectrum data The environmental score of the decision result, Representation Cluster spectral data.
[0021] According to the second aspect of the present invention, there is also provided a spectrum resource sharing device for a drone swarm under a centralized architecture, comprising: a first construction unit, used to construct an intelligent spectrum resource model for a drone swarm, wherein the drone swarm is divided into multiple clusters, each cluster in the drone swarm comprises a cluster head drone and multiple cluster member drones, and the cluster head drone communicates directly with the central drone; a second construction unit, used to construct a Markov decision process model for describing the drone swarm intelligent spectrum resource sharing process, and to construct an optimization target model, wherein the optimization target in the optimization target model includes the drone swarm transmission rate, cluster fairness rate and inter-cluster fairness rate; a shared training unit, used to construct an independent double-layer competitive architecture Q network input and output model and adopt a hybrid exploration strategy to perform resource sharing training for each cluster, and initialize data training to generate a cluster model; a generation unit, used to perform drone swarm resource sharing training using a weighted aggregation weight based on the model contribution rate of the cluster model, and to generate a drone swarm spectrum sharing strategy.
[0022] According to a third aspect of the present invention, there is also provided a computer-readable storage medium, in which a computer program is stored, wherein the computer program is configured to execute the above-mentioned method for sharing spectrum resources of a swarm of drones under a centralized architecture when running.
[0023] According to a fourth aspect of the present invention, there is also provided an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the above-mentioned method for sharing spectrum resources of a swarm of drones under a centralized architecture through the computer program.
[0024] In general, the above technical solutions conceived by the present invention can achieve the following beneficial effects compared with the prior art:
[0025] (1) The present invention provides a spectrum resource sharing method for a drone swarm under a centralized architecture, wherein a drone swarm is divided into multiple clusters, each cluster comprising a cluster head drone and cluster member drones; the drone swarm intelligent spectrum resource sharing process is modeled as a Markov decision process, and the drone swarm transmission rate, cluster fairness rate and inter-cluster fairness rate are taken as optimization targets; an independent double-layer competitive architecture Q network is designed, and a hybrid exploration strategy is adopted to perform resource sharing training for each cluster; drone swarm resource sharing training is performed using weighted aggregation weights based on model contribution rate to obtain a drone swarm spectrum sharing strategy, and the drone swarm transmission rate and cluster fairness rate are taken as the optimization targets of the algorithm, thereby effectively improving the drone swarm spectrum sharing efficiency and improving sharing fairness.
[0026] (2) By adopting a spectrum resource sharing method for drone swarms under a centralized architecture provided by the present invention, an independent double-layer competitive architecture Q network (Dueling Deep Q Network, DuelingDQN) is adopted, and a strategy and upper confidence bound (UCB) method are integrated, a hybrid exploration strategy is designed to perform local training, improve the model's adaptability to the environment, reduce the difficulty of model training, and achieve an efficient training process; a weighted aggregation method based on model contribution rate is designed to perform global training, which promotes experience sharing among cluster head agents and achieves efficient aggregation of centralized models. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0028] Figure 1 A schematic diagram of a flow chart of an optional method for sharing spectrum resources of a drone swarm under a centralized architecture provided in an embodiment of the present application;
[0029] Figure 2 A schematic diagram of a flow chart of another optional method for sharing spectrum resources of a drone swarm under a centralized architecture provided in an embodiment of the present application;
[0030] Figure 3 A schematic diagram of an optional DuelingDQN network structure provided in an embodiment of the present application;
[0031] Figure 4 A schematic diagram of an optional comparison of average rewards of various algorithms for environmental changes provided in an embodiment of the present application;
[0032] Figure 5A schematic diagram of comparing the average losses of various algorithms to environmental changes provided in an optional embodiment of the present application;
[0033] Figure 6 A schematic diagram of the structure of a spectrum resource sharing device for a drone swarm in an optional centralized architecture provided in an embodiment of the present application;
[0034] Figure 7 A schematic diagram of the structure of an optional electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0035] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0036] The terms "first", "second", "third", etc. in the specification and claims of this application and the above-mentioned drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally includes steps or units that are not listed, or optionally includes other steps or units inherent to these processes, methods, products or devices.
[0037] According to one aspect of the embodiments of the present application, a method for sharing spectrum resources of a drone swarm under a centralized architecture is provided. Figure 1 The present invention describes a method for sharing spectrum resources of a swarm of drones under a centralized architecture provided in an embodiment of the present application.
[0038] Figure 1 is a flow chart of an optional method for sharing spectrum resources of a drone swarm under a centralized architecture provided in an embodiment of the present application, such as Figure 1 As shown, the process of the method may include the following steps:
[0039] S102, constructing an intelligent spectrum resource model of a drone swarm, wherein the drone swarm is divided into a plurality of clusters, each cluster in the drone swarm comprises a cluster head drone and a plurality of cluster member drones, and the cluster head drone communicates directly with the central drone;
[0040] S104, constructing a Markov decision process model to describe the UAV swarm intelligent spectrum resource sharing process, and constructing an optimization target model, wherein the optimization targets in the optimization target model include the UAV swarm transmission rate, cluster fairness rate, and inter-cluster fairness rate;
[0041] S106, constructing an independent double-layer competition architecture Q network input and output model and adopting a hybrid exploration strategy to perform resource sharing training for each cluster, initializing data training to generate a cluster model;
[0042] S108, performing drone swarm resource sharing training using a weighted aggregation weight based on a model contribution rate of the cluster model to generate a drone swarm spectrum sharing strategy.
[0043] A method for sharing spectrum resources of a swarm of drones under a centralized architecture provided in an embodiment of the present application can be applied to a scenario of sharing spectrum resources of a swarm of drones under a centralized architecture.
[0044] Optionally, combined Figure 1 and Figure 2 As shown in the figure, in order to obtain the spectrum resource sharing scheme of drone swarm under centralized architecture, the drone swarm can be divided into multiple clusters, each cluster contains cluster head drones and cluster member drones; the drone swarm intelligent spectrum resource sharing process is modeled as a Markov decision process, and the drone swarm transmission rate, cluster fairness rate and inter-cluster fairness rate are used as optimization targets; an independent double-layer competition architecture Q network is designed, and a hybrid exploration strategy is used to train the resource sharing of each cluster; the drone swarm resource sharing is trained using weighted aggregation weights based on model contribution rate, and the drone swarm spectrum sharing strategy is obtained. The independent double-layer competition architecture Q network and federated reinforcement learning method are used to solve the spectrum resource sharing problem of drone swarm under centralized architecture, which not only has fast convergence and high energy efficiency, but also can effectively improve the dynamic and rapid response capabilities of drone swarm spectrum resource sharing.
[0045] Through the above steps S102 to S108, an intelligent spectrum resource model of a drone swarm is constructed, wherein the drone swarm is divided into multiple clusters, each cluster in the drone swarm includes a cluster head drone and multiple cluster member drones, and the cluster head drone communicates directly with the central drone; a Markov decision process model is constructed to describe the drone swarm intelligent spectrum resource sharing process, and an optimization target model is constructed, wherein the optimization targets in the optimization target model include the drone swarm transmission rate, cluster fairness rate and inter-cluster fairness rate; an independent double-layer competition architecture Q network input and output model is constructed and a hybrid exploration strategy is adopted to perform resource sharing training for each cluster, and the cluster model is generated by initialization data training; the drone swarm resource sharing training is performed using the weighted aggregation weight of the model contribution rate based on the cluster model, and a drone swarm spectrum sharing strategy is generated, which solves the problem of drone swarm spectrum resource sharing under a centralized architecture, has fast convergence and high energy efficiency, and can effectively improve the dynamics and rapid response capabilities of drone swarm spectrum resource sharing.
[0046] In an exemplary embodiment, the characteristic parameters in the drone swarm intelligent spectrum resource model include the number of channels and the number of available frequency points for each channel. The number of channels includes the number of channels required for communication between the cluster head drones and the central drone, the number of channels required within each cluster, and the number of channels required for communication between cluster heads.
[0047] In this embodiment, the drone swarm intelligent spectrum resources mainly include the number of channels 、Number of available frequencies per channel Two characteristic parameters. They are expressed as follows:
[0048]
[0049] in, is the number of channels required for the cluster head UAV to communicate with the central UAV, Cluster The number of channels required within is the number of channels required for communication between cluster heads. It is expressed as follows:
[0050]
[0051] in, represents the number of cluster head drones, Represents a cluster The number of inner cluster member drones.
[0052] In an exemplary embodiment, the Markov decision process model is constructed to describe the intelligent spectrum resource sharing process of the drone swarm, and the optimization target model is constructed including:
[0053] Determine the state space of the spectrum resource sharing scheme for the drone swarm and determine the action space for changing the current resource sharing scheme;
[0054] Determine the state transition probability and introduce a discount factor to get rewards from the interaction between the Q network and the environment, get the strategy corresponding to the highest reward value, and generate the decay sum of the rewards;
[0055] Calculate the action-value function, calculate the optimal action-value function, and solve the learning objective.
[0056] In this embodiment, optionally, the Markov decision process can be constructed to determine the state space of the spectrum scheme of the drone swarm: , determine the action space for changes to the current spectrum plan , determine the state transition probability , introducing the discount factor , the Q network interacts with the environment to get rewards, and the strategy corresponding to the highest value is obtained , the decaying sum of the generated rewards , calculate the action value function , calculate the optimal action-value function , that is, solving the learning goal. It is expressed as follows:
[0057]
[0058] in, For the The sub-spectrum scheme is defined as a state, For action.
[0059]
[0060] in, is the discount factor, for All rewards of the moment.
[0061]
[0062] in, is the action value function, For status, For strategy, For action, To represent the immediate reward after performing the action, is the expectation under strategy π, is the expected value of the immediate reward obtained by selecting action a in state s, is the discount factor, is the probability of transitioning to state s' after selecting action a in state s.
[0063]
[0064] in, is the optimal action value function, is the action value function.
[0065]
[0066] in, To take action in state s After receiving the instant reward, is the discount factor, is the state transition probability, For the next state Take action The optimal action-value function.
[0067] Furthermore, the transmission rate of the drone group in the above steps is , which is expressed as follows:
[0068]
[0069] in, is the UAV transmission power, is the instantaneous random quantity, is the distance between the communicating drones, is the current channel noise power, Jammers for UAVs The interference power, It is co-channel interference.
[0070] Furthermore, the cluster fairness ratio in the above steps , which is expressed as follows:
[0071]
[0072] in, is the effective transmission rate, Represents a cluster The number of UAVs in the inner cluster indicates that represents the mean value, Represents a collection of cluster node drones.
[0073] Furthermore, the inter-cluster fairness ratio in the above steps , which is expressed as follows:
[0074]
[0075] in, is the sum of the effective transmission rates of UAVs between clusters, expressed as follows:
[0076]
[0077] in, is the effective transmission rate.
[0078] Through this embodiment, the transmission rate of the drone swarm and the cluster fairness rate are used as the optimization goals of the algorithm, which effectively improves the spectrum sharing efficiency of the drone swarm and improves the sharing fairness.
[0079] In an exemplary embodiment, constructing an independent double-layer competitive architecture Q network input-output model and adopting a hybrid exploration strategy to perform resource sharing training for each cluster includes:
[0080] See also Figure 3 The first layer of the constructed independent double-layer competition architecture Q network input-output model is a channel selection agent, which is used to convert the state vector containing the frequency allocation table and the transmission power table into a Boolean value channel state vector representing the channel working state as an input vector, obtain the conflicting channel after learning, and output the channel selection action;
[0081] The second layer of the independent double-layer competition architecture Q network input-output model is a frequency switching agent, which includes an independent environment coding layer and a state analysis layer; wherein the input of the independent environment coding layer is an environment feature vector including the number of available frequencies of the current channel and the channel noise power of each frequency, and the environment coding layer transmits the compressed environment information to the state analysis layer;
[0082] The state analysis layer encodes the channel selection action output by the channel selection agent, and combines the coding vector and the current spectrum sharing scheme to form a comprehensive coding vector. The frequency switching agent learns the comprehensive coding vector and the environmental feature vector, and finally outputs a frequency change action.
[0083] In an exemplary embodiment, the hybrid exploration strategy is expressed as:
[0084]
[0085] in, It's action The current estimate of represents the “utilization” value of the action. is the total number of attempts made so far. It's action The number of times it has been attempted. is an exploration constant that controls the degree of exploration. For example, it can be set to .when Smaller (i.e. action has been tried less times), this value will be larger, encouraging the agent to explore this action. As increases, this value gradually decreases, making the agent more inclined to exploit actions that are known to perform well.
[0086] In an exemplary embodiment, the using of the weighted aggregation weight based on the model contribution rate of the cluster model to perform drone swarm resource sharing training and generate a drone swarm spectrum sharing strategy includes:
[0087] Determine weighted aggregation weights , calculated as,
[0088]
[0089] in, is the global model weight, Cluster The weight of Cluster The contribution rate is calculated as:
[0090]
[0091] in, Indicates that the cluster head local model is in the spectrum data The environmental score of the decision result, Representation Cluster spectral data.
[0092] In an optional embodiment, Python language can be used as a simulation tool to implement spectrum sharing simulation of clustered drone swarms, and the clustered federated distributed spectrum sharing algorithm of drone swarms can be verified with examples, and the recognition model can be further optimized based on the training results.
[0093] Specifically, combined Figure 4 and Figure 5 As shown, the spectrum resource sharing method of drone swarm under the centralized architecture of the present invention is adopted ( Figure 4 and Figure 5 FL+DuDQN, namely Federated Learning + Dueling Deep Q Network, DuelingDQN, federated learning + independent double-layer competitive architecture Q network) and traditional DQN algorithm ( Figure 4 and Figure 5 DQN, namely DeepQ Network, Deep Q Network), DuelingDQN algorithm ( Figure 4 and Figure 5The response ability of the DuDQN (Dueling Deep Q Network, DuelingDQN, independent double-layer competitive architecture Q network) to environmental changes is simulated and compared. It can be seen from the simulation results that due to the change of channel noise, the decision-making efficiency of the pre-trained models of the above three algorithms has dropped significantly. However, after 5000 rounds of training, the decision-making ability of these three algorithms in the new environment has been improved to a certain extent. Among them, the improvement of the DQN algorithm is relatively small, only 80.7%, while DuelingDQN and the spectrum resource sharing method of the drone swarm under the centralized architecture of the present invention show similar improvement levels, which are 198.30% and 191.52% respectively. It is worth noting that the spectrum resource sharing method of the drone swarm under the centralized architecture of the present invention has reached a convergence state after about 2000 rounds of training, and its convergence speed is about 3000 rounds faster than DuelingDQN. The spectrum resource sharing method of drone swarm under the centralized architecture of the present invention is significantly better than DQN and DuelingDQN algorithms in terms of environmental adaptation speed, showing more outstanding environmental adaptability.
[0094] According to another aspect of an embodiment of the present application, a spectrum resource sharing device is also provided for implementing the spectrum resource sharing method of a swarm of drones under the above-mentioned centralized architecture. Figure 6 is a schematic diagram of the structure of a spectrum resource sharing device for a drone swarm under an optional centralized architecture according to an embodiment of the present application, such as Figure 6 As shown, the device may include:
[0095] A first construction unit 602 is used to construct an intelligent spectrum resource model of a drone swarm, wherein the drone swarm is divided into a plurality of clusters, each cluster in the drone swarm comprises a cluster head drone and a plurality of cluster member drones, and the cluster head drone communicates directly with the central drone;
[0096] The second construction unit 604 is used to construct a Markov decision process model for describing the UAV swarm intelligent spectrum resource sharing process, and to construct an optimization target model, wherein the optimization targets in the optimization target model include the UAV swarm transmission rate, cluster fairness rate, and inter-cluster fairness rate;
[0097] A shared training unit 606 is used to construct an independent double-layer competition architecture Q network input and output model and adopt a hybrid exploration strategy to perform resource sharing training for each cluster, and initialize data training to generate a cluster model;
[0098] The generating unit 608 is used to perform drone swarm resource sharing training by using the weighted aggregation weight based on the model contribution rate of the cluster model to generate a drone swarm spectrum sharing strategy.
[0099] It should be noted that the first construction unit 602 in this embodiment can be used to execute the above step S102, the second construction unit 604 in this embodiment can be used to execute the above step S104, the shared training unit 606 in this embodiment can be used to execute the above step S106, and the generation unit 608 in this embodiment can be used to execute the above step S108.
[0100] Through the above modules, an intelligent spectrum resource model of drone swarm is constructed, wherein the drone swarm is divided into multiple clusters, each cluster in the drone swarm contains a cluster head drone and multiple cluster member drones, and the cluster head drone communicates directly with the central drone; a Markov decision process model is constructed to describe the intelligent spectrum resource sharing process of the drone swarm, and an optimization target model is constructed, wherein the optimization targets in the optimization target model include the drone swarm transmission rate, cluster fairness rate and inter-cluster fairness rate; an independent double-layer competitive architecture Q network input and output model is constructed and a hybrid exploration strategy is adopted to carry out resource sharing training for each cluster, and the cluster model is generated by initialization data training; the drone swarm resource sharing training is carried out using the weighted aggregation weight of the model contribution rate based on the cluster model, and the drone swarm spectrum sharing strategy is generated, which solves the spectrum resource sharing problem of drone swarm under the centralized architecture, has fast convergence and high energy efficiency, and can effectively improve the dynamic and rapid response capabilities of drone swarm spectrum resource sharing.
[0101] In an exemplary embodiment, the second construction unit comprises:
[0102] The first determination module is used to determine the state space of the spectrum resource sharing scheme of the drone swarm and determine the action space for changing the current resource sharing scheme;
[0103] The second determination module is used to determine the state transition probability and introduce a discount factor to obtain rewards by interacting with the Q network and the environment, obtain the strategy corresponding to the highest reward value, and generate the decay sum of the rewards;
[0104] Calculate the action-value function, calculate the optimal action-value function, and solve the learning objective.
[0105] In an exemplary embodiment, the shared training unit comprises:
[0106] A construction module, for constructing the first layer of the independent double-layer competition architecture Q network input-output model as a channel selection agent, for converting a state vector containing a frequency allocation table and a transmission power table into a Boolean valued channel state vector representing a channel working state as an input vector, obtaining a conflicting channel after learning, and outputting a channel selection action;
[0107] The switching agent module, the second layer of the independent double-layer competition architecture Q network input and output model is a frequency switching agent, which includes an independent environment coding layer and a state analysis layer; wherein the input of the independent environment coding layer is an environment feature vector including the number of available frequencies of the current channel and the channel noise power of each frequency, and the environment coding layer transmits the compressed environment information to the state analysis layer;
[0108] The encoding module is used for the state analysis layer to encode the channel selection action output by the channel selection agent, and to combine the encoding vector and the current spectrum sharing scheme to form a comprehensive encoding vector. The frequency switching agent learns the comprehensive encoding vector and the environmental feature vector, and finally outputs the frequency change action.
[0109] It should be noted here that the examples and scenarios implemented by the above-mentioned modules and corresponding steps are the same, but are not limited to the contents disclosed in the above-mentioned embodiments. It should be noted that the above-mentioned modules as part of the device can run in a hardware environment and can be implemented by software or hardware, wherein the hardware environment includes a network environment.
[0110] According to another aspect of the embodiment of the present application, a storage medium is also provided. Optionally, in this embodiment, the storage medium can be used to execute the program code of the drone swarm spectrum resource sharing method under the centralized architecture in the embodiment of the present application.
[0111] Optionally, in this embodiment, the storage medium is configured to store program codes for executing the following steps:
[0112] S1, constructing a UAV swarm intelligent spectrum resource model, where the UAV swarm is divided into multiple clusters, each cluster in the UAV swarm contains a cluster head UAV and multiple cluster member UAVs, and the cluster head UAV communicates directly with the center UAV;
[0113] S2, construct a Markov decision process model to describe the UAV swarm intelligent spectrum resource sharing process, and construct an optimization target model, where the optimization targets in the optimization target model include the UAV swarm transmission rate, cluster fairness rate, and inter-cluster fairness rate;
[0114] S3, construct an independent double-layer competition architecture Q network input and output model and adopt a hybrid exploration strategy to carry out resource sharing training for each cluster, and initialize data training to generate a cluster model;
[0115] S4, using the weighted aggregation weight of the model contribution rate based on the cluster model to train the drone swarm resource sharing and generate the drone swarm spectrum sharing strategy.
[0116] Optionally, the specific examples in this embodiment may refer to the examples described in the above embodiments, which will not be described in detail in this embodiment.
[0117] Among them, computer-readable storage media may include, but are not limited to, any type of disk, including floppy disks, optical disks, DVDs, CD-ROMs, microdrives and magneto-optical disks, ROMs, RAMs, EPROMs, EEPROMs, DRAMs, VRAMs, flash memory devices, magnetic or optical cards, nanosystems (including molecular memory ICs), or any type of medium or device suitable for storing instructions and / or data.
[0118] According to another aspect of an embodiment of the present application, an electronic device for implementing the above-mentioned method for sharing spectrum resources of a swarm of drones under a centralized architecture is also provided. The electronic device may be a server, a terminal, or a combination thereof.
[0119] Figure 7 is a schematic diagram of the structure of an optional electronic device according to an embodiment of the present application, such as Figure 7 As shown, it includes a processor 702, a communication interface 704, a memory 706 and a communication bus 708, wherein the processor 702, the communication interface 704, and the memory 706 communicate with each other through the communication bus 708, wherein:
[0120] Memory 706, used to store computer programs;
[0121] The processor 702 is used to implement the following steps when executing the computer program stored in the memory 706:
[0122] S1, constructing a UAV swarm intelligent spectrum resource model, where the UAV swarm is divided into multiple clusters, each cluster in the UAV swarm contains a cluster head UAV and multiple cluster member UAVs, and the cluster head UAV communicates directly with the center UAV;
[0123] S2, construct a Markov decision process model to describe the UAV swarm intelligent spectrum resource sharing process, and construct an optimization target model, where the optimization targets in the optimization target model include the UAV swarm transmission rate, cluster fairness rate, and inter-cluster fairness rate;
[0124] S3, construct an independent double-layer competition architecture Q network input and output model and adopt a hybrid exploration strategy to carry out resource sharing training for each cluster, and initialize data training to generate a cluster model;
[0125] S4, using the weighted aggregation weight of the model contribution rate based on the cluster model to train the drone swarm resource sharing and generate the drone swarm spectrum sharing strategy.
[0126] Optionally, the communication bus may be a PCI (Peripheral Component Interconnect) bus, or an EISA (Extended Industry Standard Architecture) bus, etc. The communication bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 7 The communication interface is used for communication between the electronic device and other devices.
[0127] The memory may include RAM, or may include non-volatile memory, such as at least one disk memory. Optionally, the memory may also be at least one storage device located away from the aforementioned processor.
[0128] As an example, the memory 706 may include, but is not limited to, the first construction unit 602, the second construction unit 604, the shared training unit 606, and the generation unit 608 in the above-mentioned device for sharing spectrum resources of a drone swarm under a centralized architecture. In addition, other module units in the above-mentioned device for sharing spectrum resources of a drone swarm under a centralized architecture may also be included but are not limited to, which will not be repeated in this example.
[0129] The above-mentioned processor can be a general-purpose processor, which can include but not be limited to: CPU (Central Processing Unit), NP (Network Processor), etc.; it can also be DSP (Digital Signal Processing), ASIC (Application Specific Integrated Circuit), FPGA (Field-Programmable Gate Array) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0130] Optionally, the specific examples in this embodiment may refer to the examples described in the above embodiments, and this embodiment will not be described in detail here.
[0131] It should be noted that, for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the present application is not limited by the described order of actions, because according to the present application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present application.
[0132] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0133] In the several embodiments provided in the present application, it should be understood that the disclosed devices can be implemented in other ways. For example, the device embodiments described above are only schematic, such as the division of the units, which is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some service interfaces, and the indirect coupling or communication connection of devices or units can be electrical or other forms.
[0134] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0135] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0136] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable memory. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a memory, including a number of instructions to enable a computer device (which can be a personal computer, server or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned memory includes: U disk, read-only memory (ROM), random access memory (RAM), mobile hard disk, disk or optical disk and other media that can store program codes.
[0137] A person skilled in the art may understand that all or part of the steps in the various methods of the above embodiments may be completed by instructing related hardware through a program, and the program may be stored in a computer-readable memory, which may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.
[0138] The above is only an exemplary embodiment of the present disclosure, and the scope of the present disclosure cannot be limited thereto. That is, any equivalent changes and modifications made according to the teachings of the present disclosure are still within the scope of the present disclosure. After considering the specification and practicing the disclosure here, those skilled in the art will easily think of the implementation scheme of the present disclosure. This application is intended to cover any modification, use or adaptation of the present disclosure, which follows the general principles of the present disclosure and includes common knowledge or customary technical means in the technical field not recorded in the present disclosure. The description and examples are regarded as exemplary only, and the scope and spirit of the present disclosure are defined by the claims.
[0139] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0140] It will be easily understood by those skilled in the art that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A method for sharing spectrum resources of drone swarms under a centralized architecture, characterized in that: include: Constructing an intelligent spectrum resource model of a drone swarm, wherein the drone swarm is divided into multiple clusters, each cluster in the drone swarm comprises a cluster head drone and multiple cluster member drones, and the cluster head drone communicates directly with the central drone; A Markov decision process model is constructed to describe the UAV swarm intelligent spectrum resource sharing process, and an optimization target model is constructed, wherein the optimization targets in the optimization target model include the UAV swarm transmission rate, cluster fairness rate, and inter-cluster fairness rate; Construct an independent double-layer competitive architecture Q network input and output model and adopt a hybrid exploration strategy to train the resource sharing of each cluster, and initialize data training to generate a cluster model; Using the weighted aggregation weight based on the model contribution rate of the cluster model to perform drone swarm resource sharing training, and generate a drone swarm spectrum sharing strategy; The construction of an independent double-layer competition architecture Q network input and output model and the use of a hybrid exploration strategy to perform resource sharing training for each cluster include: The first layer of the constructed independent double-layer competition architecture Q network input-output model is a channel selection agent, which is used to convert the state vector containing the frequency allocation table and the transmission power table into a Boolean valued channel state vector representing the channel working state as an input vector, obtain the conflicting channel after learning, and output the channel selection action; The second layer of the independent double-layer competition architecture Q network input-output model is a frequency switching agent, which includes an independent environment coding layer and a state analysis layer; wherein the input of the independent environment coding layer is an environment feature vector including the number of available frequencies of the current channel and the channel noise power of each frequency, and the environment coding layer transmits the compressed environment information to the state analysis layer; The state analysis layer encodes the channel selection action output by the channel selection agent, and combines the coding vector with the current spectrum sharing scheme to form a comprehensive coding vector. The frequency switching agent learns the comprehensive coding vector and the environmental feature vector, and finally outputs a frequency change action. The hybrid exploration strategy is expressed as: in, It's action The current estimated value of is the total number of attempts so far, It's action The number of times it has been attempted, is the exploration constant used to describe the degree of control exploration.
2. The method for sharing spectrum resources of drone swarms under a centralized architecture as claimed in claim 1, characterized in that: The characteristic parameters in the drone swarm intelligent spectrum resource model include the number of channels and the number of available frequency points for each channel. The number of channels includes the number of channels required for communication between the cluster head drone and the central drone, the number of channels required within each cluster, and the number of channels required for communication between cluster heads.
3. The method for sharing spectrum resources of drone swarms under a centralized architecture as claimed in claim 1, characterized in that: The Markov decision process model is used to describe the intelligent spectrum resource sharing process of the drone swarm, and the optimization target model is constructed including: Determine the state space of the spectrum resource sharing scheme for the drone swarm and determine the action space for changing the current resource sharing scheme; Determine the state transition probability and introduce a discount factor to get rewards from the interaction between the Q network and the environment, get the strategy corresponding to the highest reward value, and generate the decay sum of the rewards; Calculate the action-value function, calculate the optimal action-value function, and solve the learning objective.
4. The method for sharing spectrum resources of drone swarms under a centralized architecture as claimed in claim 1, characterized in that: The method of using the weighted aggregation weight based on the model contribution rate of the cluster model to perform the drone swarm resource sharing training and generate the drone swarm spectrum sharing strategy includes: Determine weighted aggregation weights , calculated as, in, is the global model weight, Cluster The weight of is the total number of attempts so far, Cluster The contribution rate is calculated as: in, Indicates that the cluster head local model is in the spectrum data The environmental score of the decision result, Representation Cluster spectral data.
5. A spectrum resource sharing device for drone swarms under a centralized architecture, executing a spectrum resource sharing method for drone swarms under a centralized architecture as claimed in claim 1, characterized in that: include: A first construction unit is used to construct an intelligent spectrum resource model of a drone swarm, wherein the drone swarm is divided into a plurality of clusters, each cluster in the drone swarm comprises a cluster head drone and a plurality of cluster member drones, and the cluster head drone communicates directly with the central drone; The second construction unit is used to construct a Markov decision process model for describing the UAV swarm intelligent spectrum resource sharing process and construct an optimization target model, wherein the optimization targets in the optimization target model include the UAV swarm transmission rate, cluster fairness rate and inter-cluster fairness rate; The shared training unit is used to construct an independent double-layer competitive architecture Q network input and output model and adopt a hybrid exploration strategy to perform resource sharing training for each cluster, and initialize data training to generate a cluster model; A generating unit is used to perform resource sharing training of a swarm of drones by using a weighted aggregation weight based on a model contribution rate of the cluster model, and to generate a spectrum sharing strategy for the swarm of drones.
6. The device for sharing spectrum resources of drone swarms in a centralized architecture as claimed in claim 5, characterized in that: The second construction unit comprises: The first determination module is used to determine the state space of the spectrum resource sharing scheme of the drone swarm and determine the action space for changing the current resource sharing scheme; The second determination module is used to determine the state transition probability and introduce a discount factor to obtain rewards by interacting with the Q network and the environment, obtain the strategy corresponding to the highest reward value, and generate the decay sum of the rewards; Calculate the action-value function, calculate the optimal action-value function, and solve the learning objective.
7. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored program, wherein the program executes the method according to any one of claims 1 to 4 when executed.
8. An electronic device comprising a memory and a processor, characterized in that: A computer program is stored in the memory, and the processor is configured to execute the method according to any one of claims 1 to 4 through the computer program.
Citation Information
Patent Citations
Unmanned aerial vehicle auxiliary network path optimization method and system based on deep reinforcement learning
CN119094975A
Unmanned aerial vehicle group spectrum resource allocation method and system based on cooperative spectrum sensing
CN119485695A