A method and device for sharing spectrum resources of drone swarm under distributed architecture
By adopting a two-layer decision-making network method under a distributed architecture in the drone cluster, the problems of high computational complexity and complex parameter settings in the spectrum resource sharing of the drone cluster are solved, the spectrum resource utilization and anti-destruction performance are improved, and intelligent dynamic spectrum sharing is realized.
Patent Information
- Application Number
- CN202510265594.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-03-07
AI Technical Summary
In the prior art, the calculation complexity and parameter settings are high in the sharing of spectrum resources of the drone group, and the internal mutual interference problems of the drone group have low spectral resource utilization and insufficient damage resistance.
Using the two-layer decision network method under a distributed architecture, data samples are collected and initialized two-layer decision network is trained through the environment perception layer and the decision generation layer. The drones in the drone swarm judge the changes in the spectrum environment, freeze some network structures, establish reward functions, and retrain local two-layer decision networks. The model weights are aggregated through the voting mechanism, a global two-layer decision-making network is generated and a spectrum resource sharing strategy is formulated.
It reduces the computational complexity and parameter setting complexity, improves the efficiency of the UAV cluster to utilize limited spectrum resources, enhances the anti-destruction performance, and realizes intelligent dynamic spectrum sharing of the UAV cluster.
Smart Images

Figure CN119789096B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of spectrum resource sharing of drone swarms, and more specifically, to a method and device for sharing spectrum resources of drone swarms under a distributed architecture. Background Art
[0002] As the application of drone swarms shows great potential, the research on spectrum sharing of drone swarms has become particularly important. In order to improve the flexibility and anti-destruction ability of the architecture, a distributed spectrum sharing architecture has been proposed. In a distributed architecture, the internal communication of the swarm does not rely on the central facility, and the communication between the swarm and the central facility is connected through a gateway node. Each node makes decisions based on local spectrum data and allocates and uses spectrum through collaboration and communication. The distributed architecture is decentralized and robust. At present, the main methods of the distributed spectrum sharing architecture of drone swarms include competition mechanism, coordination mechanism, virtualization technology, etc.
[0003] In the related technology, a clustering optimization method based on the improved gray wolf algorithm (Generalized Wave Continuity Optimization Algorithm, GWCOA) is proposed to solve the problems of unstable network structure and short network life cycle of drone clusters. Clustering is performed according to the speed similarity and distance similarity of nodes, and the number of nodes in each cluster is guaranteed to be moderate; the improved gray wolf algorithm is used to elect the best cluster head by comprehensively considering the remaining energy of the nodes, the highest node degree, the communication status and the task type; a periodic maintenance mechanism is used to update the cluster head according to the frequency of network topology changes to ensure the stability of the network structure. The algorithm has high clustering balance, stable network structure and long network life cycle, but the calculation amount is relatively large and the parameter setting is complex.
[0004] Aiming at the problem of low spectrum resource utilization of heterogeneous drone swarms, a spectrum allocation method based on the improved whale optimization algorithm is proposed. The main contents include spectrum allocation model construction, spectrum benefit calculation, algorithm improvement, and simulation experiments. The algorithm model is comprehensive, the benefit calculation is accurate, the algorithm performance is improved, and the experimental verification is effective. However, the algorithm assumes that the drone communication environment parameters and communication topology structure will not change. In practical applications, dynamic changes need to be considered, and the algorithm complexity is relatively high.
[0005] Aiming at the problem of spectrum scarcity and perception delay constraint in cognitive UAV networks, a fast spectrum sensing scheme based on sequential probability ratio test (SPRT) is proposed. A cognitive UAV network framework with three-dimensional spatiotemporal spectrum sensing is established, including primary users, UAVs and relay nodes. An intra-frame cooperative spectrum sensing structure is designed, which replaces the node-to-node cooperation in traditional inter-frame cooperation by cooperation between multiple micro-sensing time slots, avoiding the communication overhead of the common control channel. A sequential maximum truncation method is proposed, which achieves fast spectrum sensing while meeting the delay constraint by introducing the perception time coefficient and soft truncation. The algorithm has fast perception, superior performance, flexible deployment, and low communication overhead, avoiding the communication overhead of the common control channel and improving spectrum utilization. However, the selection of the algorithm's soft truncation trade-off threshold needs to be adjusted according to the actual network requirements, and the impact of the number of UAVs on the perception performance is not considered.
[0006] At present, the related clustering algorithm for heterogeneous wireless sensor networks based on genetic algorithms aims to extend the network life cycle and reduce energy consumption. The algorithm mainly improves the cluster head election method and selects the optimal cluster head set with the best number in each round through genetic algorithms. At the same time, in order to solve the energy loss caused by periodic clustering, a cluster head update strategy is established to allow qualified cluster head nodes to continue to be elected. The algorithm extends the network life cycle, reduces energy consumption, and balances energy consumption. However, there are still problems such as high computational complexity and dependence of parameter settings on experience. Summary of the invention
[0007] In response to at least one defect or improvement need in the prior art, the present invention provides a method and device for sharing spectrum resources of a swarm of drones under a distributed architecture, which solves the problems of high computational complexity and complex parameter settings in the spectrum resource sharing of a swarm of drones in the related technology. While solving the mutual interference problem within the swarm of drones, it can improve the utilization efficiency of the swarm of drones for limited spectrum resources, further enhance the anti-destruction performance of the swarm of drones, and realize intelligent dynamic spectrum sharing of the swarm of drones.
[0008] To achieve the above-mentioned purpose, according to the first aspect of the present invention, a method for sharing spectrum resources of a swarm of drones under a distributed architecture is provided, the method comprising: training and generating an initialized two-layer decision network after collecting data samples, wherein the initialized two-layer decision network comprises an environment perception layer and a decision generation layer; using drones in the drone swarm to determine whether the spectrum environment has changed, freezing part of the network structure in the initialized two-layer decision network and establishing a reward function, and retraining to generate a local two-layer decision network corresponding to each cluster of drones; the cluster head drone corresponding to each cluster of drones broadcasts the local two-layer decision network, and adopts a voting mechanism to determine the model weight based on the local model weights of each local two-layer decision network; the drones aggregate and generate a global two-layer decision network based on the model weights, and generate a drone spectrum resource sharing strategy based on the global two-layer decision network.
[0009] In an exemplary embodiment, the input data of the environmental perception layer in the initialized two-layer decision network is spectrum environment information, wherein the spectrum environment information includes the current occupancy status, signal strength, interference level, and noise conditions of each channel, and the output of the environmental perception layer is the currently available channel; the input data of the decision generation layer in the initialized two-layer decision network is the performance parameters of the current channel and the spectrum environment information, and the output of the decision generation layer is a channel transfer decision.
[0010] In an exemplary embodiment, the method of using drones in a drone swarm to determine whether the spectrum environment has changed, freezing part of the network structure in the initialized two-layer decision network and establishing a reward function, and retraining to generate a local two-layer decision network corresponding to each cluster of drones includes: when a drone perceives that the spectrum environment has changed, the cluster head drone freezes all layers except the environment perception layer in the initialized two-layer decision network, establishes a reward function and calculates the cumulative reward, and retrains to generate a local two-layer decision network corresponding to each cluster of drones; when the drone perceives that the spectrum environment has not changed, the cluster head drone freezes all layers except the decision generation layer in the two-layer decision network, establishes a reward function and calculates the cumulative reward, and retrains to generate a local two-layer decision network corresponding to each cluster of drones.
[0011] In an exemplary embodiment, the reward function It is expressed as:
[0012]
[0013] in, is the reward weight, is the external environment function, Break down rewards for goals, For intrinsic rewards, Reinventing for rewards, Rewards for energy efficiency.
[0014] In an exemplary embodiment, the method of using a voting mechanism to determine the model weight based on the local model weight of each local two-layer decision network includes: determining the model weight based on the determined global model weight, the local model weight of each cluster of drones, and the contribution rate corresponding to each cluster of drones; cluster Contribution rate , expressed as:
[0015]
[0016] in, For Model The number of votes received, For Model The number of votes received.
[0017] In an exemplary embodiment, after freezing part of the network structure in the initialized two-layer decision network and establishing a reward function, the method further includes: calculating a cumulative reward, wherein the cumulative reward is expressed as:
[0018]
[0019] in, is the immediate reward after taking action a in state s, is the discount factor, is the state transition probability, is the optimal action-value function.
[0020] According to the second aspect of the present invention, there is also provided a spectrum resource sharing device for a swarm of drones under a distributed architecture, comprising: a collection unit, for collecting data samples and then training and generating an initialized two-layer decision network, wherein the initialized two-layer decision network comprises an environment perception layer and a decision generation layer; a generation unit, for using drones in the drone swarm to determine whether the spectrum environment has changed, freezing part of the network structure in the initialized two-layer decision network and establishing a reward function, and retraining to generate a local two-layer decision network corresponding to each cluster of drones; a determination unit, for the cluster head drones corresponding to each cluster of drones to broadcast the local two-layer decision network, and adopting a voting mechanism to determine the model weight based on the local model weights of each local two-layer decision network; an aggregation unit, for the drones to aggregate and generate a global two-layer decision network based on the model weights, and generate a drone spectrum resource sharing strategy based on the global two-layer decision network.
[0021] According to the third aspect of the present invention, there is also provided a computer-readable storage medium, in which a computer program is stored, wherein the computer program is configured to execute the above-mentioned method for sharing spectrum resources of a swarm of drones under a distributed architecture when running.
[0022] According to a fourth aspect of the present invention, there is also provided an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the above-mentioned method for sharing spectrum resources of a swarm of drones under a distributed architecture through the computer program.
[0023] In general, the above technical solutions conceived by the present invention can achieve the following beneficial effects compared with the prior art:
[0024] (1) The present invention provides a method for sharing spectrum resources of drone swarms under a distributed architecture. After collecting data samples and training to generate an initialized two-layer decision network, the drone is used to determine whether the environment has changed, freeze part of the network structure of the initialized two-layer decision network, establish a reward function, and retrain to generate a two-layer decision network. The drone broadcasts the two-layer decision network to surrounding drones and receives the two-layer decision networks broadcast by surrounding drones. Using a voting mechanism, drones aggregate to generate a new two-layer decision network; after generating a drone spectrum resource sharing strategy, the problem that the central drone cannot upload and aggregate the model in the event of failure is solved. The algorithm proposed in the present invention can still robustly update the model parameters, thereby realizing decentralized global model updates, enhancing the system's anti-destruction ability, and effectively overcoming the limitations of traditional reinforcement learning algorithms in such scenarios.
[0025] (2) By adopting the spectrum resource sharing method of drone swarm under a distributed architecture provided by the present invention, transfer learning technology is introduced in local training, and two targeted training methods, environmental learning and decision learning, are proposed. Transfer learning is performed on the environmental layer and the decision layer respectively, which accelerates the convergence speed of the algorithm after the failure of the central drone; in the global model aggregation process, a voting mechanism is designed. The cluster head drone calculates the aggregation ratio according to the vote statistics information, and aggregates the global model synchronously locally, thereby realizing a decentralized model aggregation method, solving the problem of drone swarm dependence on the central drone, and focusing on the energy efficiency of drone swarm communication, which helps to extend the network life of the drone swarm. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0027] Figure 1 A schematic diagram of a flow chart of an optional method for sharing spectrum resources of a drone swarm under a distributed architecture provided in an embodiment of the present application;
[0028] Figure 2 A schematic diagram of a flow chart of another optional method for sharing spectrum resources of a drone swarm under a distributed architecture provided in an embodiment of the present application;
[0029] Figure 3 A flowchart of an optional transfer learning process provided in an embodiment of the present application;
[0030] Figure 4 A schematic diagram of a flow chart of an optional voting aggregation global model provided in an embodiment of the present application;
[0031] Figure 5 A schematic diagram for comparing the response capabilities of different algorithms to environmental changes provided in an optional embodiment of the present application;
[0032] Figure 6 Another optional schematic diagram for comparing the responsiveness of different algorithms to environmental changes provided in an embodiment of the present application.
[0033] Figure 7 A schematic diagram of the structure of a spectrum resource sharing device for a drone swarm under an optional distributed architecture provided in an embodiment of the present application;
[0034] Figure 8 A schematic diagram of the structure of an optional electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0035] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0036] The terms "first", "second", "third", etc. in the specification and claims of this application and the above-mentioned drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally includes steps or units that are not listed, or optionally includes other steps or units inherent to these processes, methods, products or devices.
[0037] According to one aspect of the embodiments of the present application, a method for sharing spectrum resources of a drone swarm under a distributed architecture is provided. Figure 1 The present invention describes a method for sharing spectrum resources of a swarm of drones under a distributed architecture provided in an embodiment of the present application.
[0038] Figure 1 is a flow chart of an optional method for sharing spectrum resources of a drone swarm under a distributed architecture provided in an embodiment of the present application, such as Figure 1 As shown, the process of the method may include the following steps:
[0039] S102, after collecting data samples, training and generating an initialized two-layer decision network, wherein the initialized two-layer decision network includes an environment perception layer and a decision generation layer;
[0040] S104, using drones in the drone group to determine whether the spectrum environment has changed, freezing and initializing part of the network structure in the two-layer decision network and establishing a reward function, and retraining to generate a local two-layer decision network corresponding to each cluster of drones;
[0041] S106, the cluster head drone corresponding to each cluster drone group broadcasts the local double-layer decision network, and uses a voting mechanism to determine the model weight based on the local model weights of each local double-layer decision network;
[0042] S108, the UAV generates a global two-layer decision network based on model weight aggregation, and generates a UAV spectrum resource sharing strategy based on the global two-layer decision network.
[0043] The embodiment of the present application provides a method for sharing spectrum resources of a drone swarm under a distributed architecture. The method provided in the present application can be used in scenarios where the central drone fails and the model cannot be uploaded and aggregated. The model parameters can still be updated robustly by using the method provided in the present application, thereby realizing decentralized global model updates, enhancing the system's anti-destruction capabilities, and effectively overcoming the limitations of traditional reinforcement learning algorithms in such scenarios.
[0044] Optionally, combined Figure 1 and Figure 2 As shown, the global model spectrum resource sharing method of the drone swarm under the distributed architecture provided in this application can be implemented by the following steps:
[0045] S1, collect data samples, train and generate an initialized two-layer decision network;
[0046] S2, when the drone senses that the environment has changed, freeze and initialize all layers except the environment perception layer in the two-layer decision network, establish a reward function, calculate the cumulative reward, and retrain to generate a two-layer decision network; when the drone senses that the environment has not changed, freeze all layers except the decision generation layer in the two-layer decision network, establish a reward function, calculate the cumulative reward, and retrain to generate a two-layer decision network;
[0047] S3, model aggregation is performed. The drone broadcasts the two-layer decision network to the surrounding drones and receives the two-layer decision network broadcast by the surrounding drones at the same time. Then, a voting mechanism can be used to calculate the model weights. The drones are aggregated to generate a new two-layer decision network, and finally a drone spectrum resource sharing strategy is generated.
[0048] Furthermore, the generated two-layer decision network includes two layers of architecture: the environment perception layer and the decision generation layer. Each layer is a deep neural network. The input data of the deep neural network of the environment perception layer is the spectrum environment information such as the current occupancy of each channel, signal strength, interference level, noise condition, etc., and the output result is the current available channel. The input of the deep neural network of the decision generation layer is the performance parameters and spectrum environment information of the current channel, and the output result is the channel transfer decision.
[0049] Through the above steps S102 to S108, after collecting data samples, an initialized two-layer decision network is trained and generated, wherein the initialized two-layer decision network includes an environment perception layer and a decision generation layer; the drones in the drone group are used to determine whether the spectrum environment has changed, part of the network structure in the initialized two-layer decision network is frozen and a reward function is established, and the local two-layer decision network corresponding to each cluster of drones is retrained and generated; the cluster head drone corresponding to each cluster of drone groups broadcasts the local two-layer decision network, and a voting mechanism is used to determine the model weight based on the local model weights of each local two-layer decision network; the drone generates a global two-layer decision network based on model weight aggregation, and generates a drone spectrum resource sharing strategy based on the global two-layer decision network, which solves the problems of high computational complexity and complex parameter setting in the spectrum resource sharing of drone groups in related technologies, and can improve the utilization efficiency of limited spectrum resources of the drone group while solving the internal interference problem of the drone group, further improve the anti-destruction performance of the drone group, and realize intelligent dynamic spectrum sharing of the drone group.
[0050] In an exemplary embodiment, the input data of the environmental perception layer in the initialized two-layer decision network is spectrum environment information, wherein the spectrum environment information includes the current occupancy status, signal strength, interference level, and noise conditions of each channel, and the output of the environmental perception layer is the currently available channel; the input data of the decision generation layer in the initialized two-layer decision network is the performance parameters of the current channel and the spectrum environment information, and the output of the decision generation layer is a channel transfer decision.
[0051] In an exemplary embodiment, the method of using drones in a drone group to determine whether the spectrum environment has changed, freezing part of the network structure in the initialized two-layer decision network and establishing a reward function, and retraining to generate a local two-layer decision network corresponding to each cluster of drones includes:
[0052] When the drone senses that the spectrum environment has changed, the cluster head drone freezes and initializes all layers except the environment perception layer in the two-layer decision network, establishes a reward function and calculates the cumulative reward, and retrains to generate a local two-layer decision network corresponding to each cluster of drones;
[0053] When the drone senses that the spectrum environment has not changed, the cluster head drone freezes all layers except the decision generation layer in the two-layer decision network, establishes a reward function and calculates the cumulative reward, and retrains and generates a local two-layer decision network corresponding to each cluster of drones.
[0054] In the examples of this application, refer to Figure 3 As shown, after the reward function is established, transfer learning training can be performed. When the drone senses that the spectrum environment has changed, the cluster head drone freezes all layers in the network except the environment perception layer and uses the data in the target domain for learning.
[0055] When the UAV senses that the spectrum environment has not changed, the cluster head UAV will freeze all layers in the network except the decision generation layer and search for the optimal strategy in the target domain.
[0056] Through this embodiment, transfer learning technology is introduced in local training, and two targeted training methods, environmental learning and decision learning, are proposed. Transfer learning is performed on the environmental layer and the decision layer respectively, which accelerates the convergence speed of the algorithm after the failure of the central UAV.
[0057] In an exemplary embodiment, the reward function It is expressed as:
[0058]
[0059] in, is the reward weight, is the external environment function, Break down rewards for goals, For intrinsic rewards, Reinventing for rewards, Rewards for energy efficiency.
[0060] In the embodiment of the present application, after designing the reward function, a transfer learning algorithm can be used for implementation. Furthermore, the external environment function can be expressed as:
[0061]
[0062] in, For the The sub-spectrum scheme is defined as a state, is the sum of the effective transmission rates of UAVs between clusters.
[0063] The target decomposition reward can be expressed as, ,in, For the The number of channels of the sub-spectrum scheme, For the The number of channels in the sub-spectrum scheme.
[0064] The intrinsic reward can be expressed as, ,in, Status Count.
[0065] The reward reshaping can be expressed as,
[0066]
[0067] in, is the sum of the effective transmission rates of UAVs between clusters, For the The sub-spectrum scheme is defined as a state, is the cluster fairness rate.
[0068] The energy efficiency reward can be expressed as, ,in, For energy efficiency.
[0069] Furthermore, the cluster fairness rate It can be expressed as follows:
[0070]
[0071] in, is the effective transmission rate, Represents a cluster The number of UAVs in the inner cluster indicates that represents the mean value, Represents a collection of cluster node drones.
[0072] Cumulative rewards during model training It can be expressed as follows:
[0073]
[0074] in, is the immediate reward obtained after taking action a in state s, is the discount factor, is the state transition probability, is the optimal action-value function.
[0075] In an exemplary embodiment, the method of using a voting mechanism to determine the model weight based on the local model weights of each local two-layer decision network includes:
[0076] Based on the determined global model weight, the local model weight of each cluster of drones, and the contribution rate of each cluster of drones, the model weight is determined; Contribution rate , expressed as:
[0077]
[0078] in, For Model The number of votes received, For Model The number of votes received.
[0079] In the examples of this application, refer to Figure 4 After determining the model weight, model aggregation can be performed. The cluster head drone first uses the P2P mechanism to broadcast the local model. After receiving each local model, the cluster head uses the local environment to evaluate and vote on the model, and then obtains the model weight of the global model algorithm based on the vote weight. It can be expressed as follows:
[0080]
[0081] in, is the global model weight, Cluster The weight of Cluster The contribution rate can be expressed as:
[0082]
[0083] in, For Model The number of votes received, For Model The number of votes obtained can be expressed as:
[0084]
[0085] in, is an indicator function, which returns 1 when the intra-cluster spectrum requirement is met, otherwise 0; For the t The sub-spectrum scheme is defined as a state, Indicates that the current spectrum solution meets the demand status.
[0086] Through this embodiment, a voting mechanism is designed in the global model aggregation process. The cluster head drone calculates the aggregation ratio according to the vote statistics information and performs the global model aggregation locally synchronously, thereby realizing a decentralized model aggregation method and solving the problem of drone group dependence on the central drone.
[0087] See also Figure 5 and Figure 6 In an optional example, the spectrum resource sharing method of drone swarm under the distributed architecture of the present invention (FL+DuDQN in the figure, i.e. Federated Learning + Dueling Deep Q Network, DuelingDQN, federated learning + independent double-layer competitive architecture Q network) and the traditional DQN algorithm (DQN in the figure, i.e. Deep QNetwork, deep Q network) and DuelingDQN algorithm (DuDQN in the figure, i.e. Dueling Deep Q Network, DuelingDQN, independent double-layer competitive architecture Q network) are simulated and compared in terms of their response capabilities to environmental changes.
[0088] For example, Figure 5 When the central drone fails, the drone swarm uses the pre-trained model to make decisions and adjustments, and the decision-making environment rewards of the three algorithm models. After the central drone fails, the drone swarm spectrum resource sharing method under the distributed architecture of the present invention adopts a distributed horizontal training algorithm and transfer learning for training. After every 50 rounds of training, the updated model is verified. However, for the DQN algorithm model and DuelingDQN algorithm model, because the model parameters cannot be updated, the pre-trained model is directly used for verification in the verification phase. In a fixed environment, the rewards obtained after the three algorithms make decisions, this curve reflects the decision-making ability of the model. Since the DQN and DuelingDQN algorithms lose the ability to update the model, the decision-making ability does not change with the rounds.
[0089] The algorithm provided in this application freezes the environment perception layer of the frequency switching network and updates the model parameters of the remaining layers. As the model is trained, the decision-making ability improves, and the model effect is significantly improved at 600 rounds. At 1000 rounds, the decision-making ability of the algorithm in this paper is 11.2% higher than the pre-trained model, 23.8% higher than the DQN model, and 17.6% higher than DuelingDQN.
[0090] See also Figure 6 , which reflects the change in the decision-making ability of the model when the channel noise changes. At 50 rounds, the channel noise list changes from N0 = [0.01, 0.03, 0.1, 0.3, 1, 0.01, 0.03, 0.1, 0.3, 1,0.01,0.03, 0.1, 0.3, 1,0.01, 0.03, 0.1, 0.3, 1,0.01, 0.03, 0.1, 0.3, 1] to N0 = [0.01, 1, 0.1, 0.3, 1,0.01, 0.03, 0.1, 0.3, 1,0.01, 0.03, 0.1, 0.3, 1,0.01, 0.03, 0.1, 0.3, 1], that is, the noise of channel 2 increases to 1W. The decision-making capabilities of the above three algorithms deteriorated after 50 rounds. The algorithm in this application was used to freeze all layers except the environmental perception layer, and only update the parameters of the environmental perception layer. After 500 rounds of training, the adverse effects of environmental deterioration were overcome and the decision-making effect of the model was gradually improved.
[0091] Through this embodiment, the problem of being unable to upload and aggregate the model after the failure of the central drone is avoided, and the proposed algorithm can still robustly update the model parameters, thereby realizing decentralized global model updates, enhancing the system's anti-destruction capabilities, and effectively overcoming the limitations of traditional reinforcement learning algorithms in such scenarios.
[0092] According to another aspect of an embodiment of the present application, a spectrum resource sharing device is also provided for implementing the spectrum resource sharing method of a swarm of drones under the above-mentioned distributed architecture. Figure 7 is a schematic diagram of a structure of a spectrum resource sharing device for a drone swarm under an optional distributed architecture according to an embodiment of the present application, such as Figure 7 As shown, the device may include:
[0093] A collection unit 702 is used to collect data samples and then train and generate an initialized two-layer decision network, wherein the initialized two-layer decision network includes an environment perception layer and a decision generation layer;
[0094] A generating unit 704 is used to use drones in the drone group to determine whether the spectrum environment has changed, freeze part of the network structure in the initialized two-layer decision network and establish a reward function, and retrain and generate a local two-layer decision network corresponding to each cluster of drones;
[0095] A determination unit 706 is used for the cluster head drones corresponding to each cluster drone group to broadcast the local double-layer decision network, and to determine the model weight based on the local model weights of each local double-layer decision network by a voting mechanism;
[0096] The aggregation unit 708 is used for the drone to generate a global two-layer decision network based on the model weight aggregation, and generate a drone spectrum resource sharing strategy based on the global two-layer decision network.
[0097] It should be noted that the collection unit 702 in this embodiment can be used to execute the above step S102, the generation unit 704 in this embodiment can be used to execute the above step S104, the determination unit 706 in this embodiment can be used to execute the above step S106, and the aggregation unit 708 in this embodiment can be used to execute the above step S108.
[0098] Through the above modules, after collecting data samples, an initialized two-layer decision network is trained and generated, wherein the initialized two-layer decision network includes an environment perception layer and a decision generation layer; the drones in the drone swarm are used to judge whether the spectrum environment has changed, freeze part of the network structure in the initialized two-layer decision network and establish a reward function, and retrain to generate a local two-layer decision network corresponding to each cluster of drones; the cluster head drone corresponding to each cluster of drones broadcasts the local two-layer decision network, and adopts a voting mechanism to determine the model weight based on the local model weights of each local two-layer decision network; the drone generates a global two-layer decision network based on model weight aggregation, and generates a drone spectrum resource sharing strategy based on the global two-layer decision network, which solves the problems of high computational complexity and complex parameter setting in the spectrum resource sharing of drone swarms in related technologies, and can improve the utilization efficiency of limited spectrum resources of drone swarms while solving the internal interference problem of drone swarms, further improve the anti-destruction performance of drone swarms, and realize intelligent dynamic spectrum sharing of drone swarms.
[0099] In an exemplary embodiment, the input data of the environmental perception layer in the initialized two-layer decision network is spectrum environment information, wherein the spectrum environment information includes the current occupancy status, signal strength, interference level, and noise conditions of each channel, and the output of the environmental perception layer is the currently available channel; the input data of the decision generation layer in the initialized two-layer decision network is the performance parameters of the current channel and the spectrum environment information, and the output of the decision generation layer is a channel transfer decision.
[0100] In an exemplary embodiment, the generating unit comprises:
[0101] The first freezing module is used for, when the drone senses that the spectrum environment has changed, the cluster head drone freezes and initializes all layers except the environment perception layer in the two-layer decision network, establishes a reward function and calculates the cumulative reward, and retrains and generates a local two-layer decision network corresponding to each cluster of drones;
[0102] The second freezing module is used for freezing all layers except the decision generation layer in the double-layer decision network of the cluster head drone when the drone senses that the spectrum environment has not changed, establishing a reward function and calculating the cumulative reward, and retraining and generating a local double-layer decision network corresponding to each cluster drone.
[0103] In an exemplary embodiment, the reward function is expressed as:
[0104]
[0105] in, is the reward weight, is the external environment function, Break down rewards for goals, For intrinsic rewards, Reinventing for rewards, Rewards for energy efficiency.
[0106] In an exemplary embodiment, the determining unit includes:
[0107] A determination module is used to determine the model weight based on the determined global model weight, the local model weight of each cluster of drones, and the contribution rate corresponding to each cluster of drones; cluster Contribution rate , expressed as:
[0108]
[0109] in, For Model The number of votes received, For Model The number of votes received.
[0110] In an exemplary embodiment, the apparatus further comprises:
[0111] A calculation unit, used to calculate the cumulative reward, the cumulative reward is expressed as,
[0112]
[0113] in, is the immediate reward obtained after taking action a in state s, is the discount factor, is the state transition probability, is the optimal action-value function.
[0114] It should be noted here that the examples and scenarios implemented by the above-mentioned modules and corresponding steps are the same, but are not limited to the contents disclosed in the above-mentioned embodiments. It should be noted that the above-mentioned modules as part of the device can run in a hardware environment and can be implemented by software or hardware, wherein the hardware environment includes a network environment.
[0115] According to another aspect of the embodiments of the present application, a storage medium is also provided. Optionally, in this embodiment, the storage medium can be used to execute the program code of any of the above-mentioned methods for sharing spectrum resources of drone swarms under a distributed architecture in the embodiments of the present application.
[0116] Optionally, in this embodiment, the storage medium is configured to store program codes for executing the following steps:
[0117] S1, after collecting data samples, training and generating an initialized two-layer decision network, wherein the initialized two-layer decision network includes an environment perception layer and a decision generation layer;
[0118] S2, using the drones in the drone group to determine whether the spectrum environment has changed, freezing and initializing part of the network structure in the two-layer decision network and establishing a reward function, and retraining to generate the local two-layer decision network corresponding to each cluster of drones;
[0119] S3, the cluster head drone corresponding to each cluster of drones broadcasts the local two-layer decision network, and uses a voting mechanism to determine the model weight based on the local model weight of each local two-layer decision network;
[0120] S4, the drone generates a global two-layer decision network based on model weight aggregation, and generates a drone spectrum resource sharing strategy based on the global two-layer decision network.
[0121] Optionally, the specific examples in this embodiment may refer to the examples described in the above embodiments, which will not be described in detail in this embodiment.
[0122] Among them, computer-readable storage media may include, but are not limited to, any type of disk, including floppy disks, optical disks, DVDs, CD-ROMs, microdrives and magneto-optical disks, ROMs, RAMs, EPROMs, EEPROMs, DRAMs, VRAMs, flash memory devices, magnetic or optical cards, nanosystems (including molecular memory ICs), or any type of medium or device suitable for storing instructions and / or data.
[0123] According to another aspect of an embodiment of the present application, an electronic device for implementing the above-mentioned method for sharing spectrum resources of a swarm of drones under a distributed architecture is also provided. The electronic device may be a server, a terminal, or a combination thereof.
[0124] Figure 8 is a schematic diagram of the structure of an optional electronic device according to an embodiment of the present application, such as Figure 8 As shown, it includes a processor 802, a communication interface 804, a memory 806 and a communication bus 808, wherein the processor 802, the communication interface 804, and the memory 806 communicate with each other through the communication bus 808, wherein:
[0125] Memory 806, used for storing computer programs;
[0126] The processor 802 is used to execute the computer program stored in the memory 806 to implement the following steps:
[0127] S1, after collecting data samples, training and generating an initialized two-layer decision network, wherein the initialized two-layer decision network includes an environment perception layer and a decision generation layer;
[0128] S2, using the drones in the drone group to determine whether the spectrum environment has changed, freezing and initializing part of the network structure in the two-layer decision network and establishing a reward function, and retraining to generate the local two-layer decision network corresponding to each cluster of drones;
[0129] S3, the cluster head drone corresponding to each cluster of drones broadcasts the local two-layer decision network, and uses a voting mechanism to determine the model weight based on the local model weight of each local two-layer decision network;
[0130] S4, the drone generates a global two-layer decision network based on model weight aggregation, and generates a drone spectrum resource sharing strategy based on the global two-layer decision network.
[0131] Optionally, the communication bus may be a PCI (Peripheral Component Interconnect) bus, or an EISA (Extended Industry Standard Architecture) bus, etc. The communication bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 8 The communication interface is used for communication between the electronic device and other devices.
[0132] The memory may include RAM, or may include non-volatile memory, such as at least one disk memory. Optionally, the memory may also be at least one storage device located away from the aforementioned processor.
[0133] As an example, the memory 806 may include, but is not limited to, the collection unit 702, the generation unit 704, the determination unit 706, and the aggregation unit 708 in the above-mentioned unmanned aerial vehicle swarm spectrum resource sharing device under the distributed architecture. In addition, it may also include, but is not limited to, other module units in the above-mentioned unmanned aerial vehicle swarm spectrum resource sharing device under the distributed architecture, which will not be repeated in this example.
[0134] The above-mentioned processor can be a general-purpose processor, which can include but not be limited to: CPU (Central Processing Unit), NP (Network Processor), etc.; it can also be DSP (Digital Signal Processing), ASIC (Application Specific Integrated Circuit), FPGA (Field-Programmable Gate Array) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0135] Optionally, the specific examples in this embodiment may refer to the examples described in the above embodiments, and this embodiment will not be described in detail here.
[0136] It should be noted that, for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the present application is not limited by the described order of actions, because according to the present application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present application.
[0137] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0138] In several embodiments provided by the present application, it should be understood that the disclosed device can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some service interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical or other form.
[0139] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0140] In addition, each functional unit in various embodiments of the present application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0141] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable memory. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present application. And the aforementioned memory includes: USB flash drives, read-only memory (ROM), random access memory (RAM), mobile hard disks, magnetic disks, or optical disks and other media that can store program codes.
[0142] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing relevant hardware through a program. This program can be stored in a computer-readable memory. The memory can include: flash drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks, etc.
[0143] The above is only an exemplary embodiment of the present disclosure, and the scope of the present disclosure cannot be limited thereto. That is, any equivalent changes and modifications made according to the teachings of the present disclosure are still within the scope of the present disclosure. After considering the specification and practicing the disclosure here, those skilled in the art will easily think of the implementation scheme of the present disclosure. This application is intended to cover any modification, use or adaptation of the present disclosure, which follows the general principles of the present disclosure and includes common knowledge or customary technical means in the technical field not recorded in the present disclosure. The description and examples are regarded as exemplary only, and the scope and spirit of the present disclosure are defined by the claims.
[0144] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0145] It will be easily understood by those skilled in the art that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A method for sharing spectrum resources of drone swarms under a distributed architecture, characterized in that: include: After collecting data samples, training and generating an initialized two-layer decision network, wherein the initialized two-layer decision network includes an environment perception layer and a decision generation layer; Using drones in the drone group to determine whether the spectrum environment has changed, freezing part of the network structure in the initialized two-layer decision network and establishing a reward function, and retraining to generate a local two-layer decision network corresponding to each cluster of drones; The cluster head drone corresponding to each cluster drone group broadcasts the local two-layer decision network, and adopts a voting mechanism to determine the model weight based on the local model weights of each local two-layer decision network; The UAV generates a global two-layer decision network based on the model weight aggregation, and generates a UAV spectrum resource sharing strategy based on the global two-layer decision network; The method of using drones in the drone group to determine whether the spectrum environment has changed, freezing part of the network structure in the initialized two-layer decision network and establishing a reward function, and retraining to generate a local two-layer decision network corresponding to each cluster of drones includes: When the drone senses that the spectrum environment has changed, the cluster head drone freezes and initializes all layers except the environment perception layer in the two-layer decision network, establishes a reward function and calculates the cumulative reward, and retrains to generate a local two-layer decision network corresponding to each cluster of drones; When the drone senses that the spectrum environment has not changed, the cluster head drone freezes all layers except the decision generation layer in the two-layer decision network, establishes a reward function and calculates the cumulative reward, and retrains to generate a local two-layer decision network corresponding to each cluster of drones.
2. The method for sharing spectrum resources of drone swarms under a distributed architecture as claimed in claim 1, characterized in that: The input data of the environmental perception layer in the initialized two-layer decision network is spectrum environment information, wherein the spectrum environment information includes the current occupancy status, signal strength, interference level, and noise conditions of each channel, and the output of the environmental perception layer is the currently available channel; the input data of the decision generation layer in the initialized two-layer decision network is the performance parameters of the current channel and the spectrum environment information, and the output of the decision generation layer is the channel transfer decision.
3. The method for sharing spectrum resources of drone swarms under a distributed architecture as claimed in claim 1, characterized in that: The reward function It is expressed as: in, is the reward weight, is the external environment function, Break down rewards for goals, For intrinsic rewards, Reinventing for rewards, Rewards for energy efficiency.
4. The method for sharing spectrum resources of drone swarms under a distributed architecture as claimed in claim 1, characterized in that: The method of using a voting mechanism to determine the model weight based on the local model weights of each local two-layer decision network includes: Based on the determined global model weight, the local model weight of each cluster of drones, and the contribution rate of each cluster of drones, the model weight is determined; Contribution rate , expressed as: in, For Model The number of votes received, For Model The number of votes received.
5. The method for sharing spectrum resources of drone swarms under a distributed architecture as claimed in claim 1, characterized in that: After freezing part of the network structure in the initialized two-layer decision network and establishing a reward function, the method further includes: Calculate the cumulative reward, which is expressed as, in, is the immediate reward obtained after taking action a in state s, is the discount factor, is the state transition probability, is the optimal action-value function, S is the state space, A Space for action.
6. A spectrum resource sharing device for drone swarms under a distributed architecture, characterized in that: include: A collection unit, used for collecting data samples and then training and generating an initialized two-layer decision network, wherein the initialized two-layer decision network includes an environment perception layer and a decision generation layer; A generating unit, used to use drones in the drone group to determine whether the spectrum environment has changed, freeze part of the network structure in the initialized two-layer decision network and establish a reward function, and retrain and generate a local two-layer decision network corresponding to each cluster of drones; A determination unit, configured for the cluster head drones corresponding to each cluster of drone groups to broadcast the local two-layer decision network, and to determine the model weight based on the local model weights of each local two-layer decision network using a voting mechanism; An aggregation unit, used for the UAV to generate a global two-layer decision network based on the model weight aggregation, and generate a UAV spectrum resource sharing strategy based on the global two-layer decision network; The generating unit comprises: The first freezing module is used for, when the drone senses that the spectrum environment has changed, the cluster head drone freezes and initializes all layers except the environment perception layer in the two-layer decision network, establishes a reward function and calculates the cumulative reward, and retrains and generates a local two-layer decision network corresponding to each cluster of drones; The second freezing module is used for freezing all layers except the decision generation layer in the double-layer decision network of the cluster head drone when the drone senses that the spectrum environment has not changed, establishing a reward function and calculating the cumulative reward, and retraining and generating a local double-layer decision network corresponding to each cluster drone.
7. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored program, wherein the program executes the method according to any one of claims 1 to 5 when executed.
8. An electronic device comprising a memory and a processor, characterized in that: A computer program is stored in the memory, and the processor is configured to execute the method according to any one of claims 1 to 5 through the computer program.
Citation Information
Patent Citations
Non-Gaussian noise assisted unmanned aerial vehicle covert communication method and system
CN119172774A
Unmanned aerial vehicle group spectrum resource allocation method and system based on cooperative spectrum sensing
CN119485695A