A method for scheduling spectrum resources in the absence of interaction information for unmanned aerial vehicles
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-07
- Publication Date
- 2026-08-11
AI Technical Summary
[0007]有鉴于此,本发明的目的在于提出一种面向无人机交互信息缺失条件下的频谱资源调度方法,以解决信息不完备条件下多无人机协同效率低、频谱分配不合理及通信时延难以保障的问题
Smart Images

Figure CN122553968A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of wireless communication technology, and more particularly to a spectrum resource scheduling method for unmanned aerial vehicle (UAV) interaction information lacking conditions. Background Technology
[0002] With the rapid development of UAV and wireless communication technologies, multi-UAV collaborative reconnaissance, monitoring, and information gathering tasks have been widely applied in scenarios such as military reconnaissance, emergency rescue, disaster monitoring, and complex environment perception. In these applications, multiple UAVs typically need to complete data acquisition and transmission under limited spectrum resources. The efficiency of task allocation and spectrum resource scheduling directly impacts the overall system's mission effectiveness and the timeliness of information acquisition. Therefore, task selection and spectrum resource scheduling for multi-UAV collaborative scenarios have become crucial issues in current research and engineering applications.
[0003] In existing technologies, most multi-UAV task allocation and spectrum scheduling methods are based on centralized optimization or the assumption of complete information. For example, a centralized control node obtains the status and task information of all UAVs and uniformly allocates tasks and spectrum resources. However, in practical applications, due to limited communication bandwidth, dynamic changes in network topology, and imperfect information sharing mechanisms between nodes, UAVs often find it difficult to obtain each other's task decision information and communication status in real time and accurately. This makes centralized methods difficult to implement in complex environments and suffers from problems such as high communication overhead and poor robustness.
[0004] To overcome the limitations of centralized methods, some studies have introduced distributed or game-theoretic approaches to model multi-UAV systems, characterizing the competitive relationships and cooperative behaviors among UAVs through non-cooperative or evolutionary game theory. However, most existing game models assume that UAVs can perceive the immediate strategies of other individuals or the complete system state, failing to adequately consider the incompleteness of UAVs' understanding of the system state under conditions of missing interactive information. In scenarios with incomplete information, if deterministic task selection or simple game equilibrium solutions are still used, multiple UAVs are prone to concentrating on selecting high-value tasks, leading to task congestion, concentrated communication load, and spectrum resource competition, thereby reducing the overall system performance.
[0005] On the other hand, in terms of spectrum resource scheduling, existing technologies mostly focus on bandwidth allocation and power control in static or quasi-static scenarios, making it difficult to adapt to complex environments with multi-task parallelism, heterogeneous tasks, and dynamically changing communication needs. Some studies have attempted to introduce reinforcement learning methods to optimize spectrum allocation strategies, but they often directly apply reinforcement learning to the resource decision-making process, ignoring the formation mechanism of the multi-UAV cooperative structure itself, which can easily lead to problems such as an excessively large policy search space, slow convergence speed, or unstable training.
[0006] Furthermore, existing technologies rarely consider task heterogeneity, UAV capability differences, and communication latency constraints from a system-level perspective. The modeling of the coupling relationship between task benefits and communication performance is insufficient, making it difficult to simultaneously improve overall task benefits while ensuring communication fairness and timeliness. Therefore, there is an urgent need for a scheduling method that can effectively guide UAVs to form stable cooperative strategies in environments lacking interactive information and adaptively perform joint optimization of spectrum resources. Summary of the Invention
[0007] In view of this, the purpose of this invention is to propose a spectrum resource scheduling method for the condition of missing UAV interactive information, so as to solve the problems of low efficiency of multi-UAV collaboration, unreasonable spectrum allocation and difficulty in guaranteeing communication latency under the condition of incomplete information.
[0008] The technical means employed in this invention are as follows:
[0009] A spectrum resource scheduling method for conditions where UAV interactive information is lacking includes the following steps: S1. Establish a basic framework for UAV mission selection and spectrum resource scheduling that comprehensively considers differences in mission type and UAV capabilities. S2. Under the basic framework of UAV mission selection and spectrum resource scheduling, UAVs are modeled as decision-making entities under partial information sharing, the mission competition and cooperation relationships between UAVs are characterized, and a mathematical model for spectrum resource scheduling with the goal of maximizing mission benefits and minimizing communication latency is constructed. S3. Based on the mathematical model of spectrum resource scheduling, generate a collaborative strategy for UAV missions through non-cooperative game theory. S4. Based on the UAV mission coordination strategy, an evolutionary strategy update mechanism guided by near-end strategy optimization is introduced to jointly adjust the mission-level spectrum allocation and UAV transmit power configuration, and output the mission-level spectrum allocation scheme and UAV transmit power configuration scheme as the final spectrum resource scheduling result.
[0010] Furthermore, S1 specifically includes the following steps: S11. Construct a multi-UAV collaborative reconnaissance system, where the UAV ID set in the system is... The reconnaissance missions are set as follows During the mission execution phase, each UAV chooses to participate in a specific reconnaissance mission based on its own decision-making. A mission-based clustered cooperative communication architecture is adopted, where the UAVs participating in each mission k form a mission cluster, and the set of mission cluster members is represented as... Within each task cluster, one drone is selected as the cluster head node, and its number is denoted as [node name missing]. It is used to aggregate reconnaissance data from members within the cluster and is responsible for communicating with ground base stations; S12, targeting any dronei With the task k Define drone i Execute the task k The obtained task utility is This is used to describe the heterogeneity at the task utility level; the corresponding amount of reconnaissance data generated is defined as... This is used to describe the heterogeneity at the communication load layer; it introduces tasks. k Maximum acceptable cumulative utility limit Regarding spectrum resources, assume the total available bandwidth of the system is... B The system is uniformly scheduled by ground base stations and divided into multiple orthogonal sub-channels. Each reconnaissance mission is assigned an independent sub-channel, and its bandwidth is denoted as [missing information]. .
[0011] Furthermore, S2 specifically includes the following steps: S21. Introduce a probabilistic task selection strategy and define the UAV. i Select task k The probability is And constitute its strategy vector. The mission selection strategies of all UAVs collectively constitute the global strategy matrix at the system level. It is used to statistically characterize the overall task distribution of multiple UAVs under conditions of partial information sharing. S22. Under the probabilistic task selection mechanism, to characterize the cooperative gains and resource competition relationships among UAVs, define the UAV... i Participate in the task k Expected utility function The specific formula is as follows:
[0012] The utility function consists of three parts: the first term For direct utility, reflecting drones i Execute tasks independently k The intrinsic benefits obtained; the second term is synergistic utility, calculated using weighting coefficients. The modeling of the collaborative benefits brought by other UAVs participating in the same mission reflects the gain effect of multi-UAV cooperative reconnaissance; the third term is the competition penalty term, which is applied when the system targets the mission. k expected utility aggregate value Exceeding the task utility limit At that time, a penalty factor was introduced. Suppress overload to avoid excessive drones concentrating on high-value tasks, which could lead to resource competition and diminishing returns; S23. During the task collaboration process, a task-based cluster communication architecture is adopted, dividing the communication process into two stages: intra-cluster transmission and inter-cluster transmission. In the intra-cluster communication stage, non-cluster leader UAVs transmit reconnaissance data to the cluster leader node. k The allocated communication bandwidth is denoted as drones i The transmission power is ;set up Indicates drone i Intra-cluster channel gain to the cluster head node Given the additive white Gaussian noise power; then the task... k The total intra-cluster communication delay is expressed as:
[0013] During the inter-cluster communication phase, the cluster head node transmits the aggregated data to the ground base station; assuming... Indicates the cluster head node Channel gain with base station The transmit power of the cluster head node. The noise power spectral density is given. During inter-cluster transmission, the cluster head node needs to forward data generated by all UAVs within the task cluster, with a total data volume of [data missing]. Then the task k Inter-cluster communication latency can be modeled as follows:
[0014] Task k The total communication delay is expressed as:
[0015] S24. With minimizing the maximum communication latency of all task clusters in the system as the optimization objective, construct a mathematical model for spectrum resource scheduling. The specific formula is as follows: .
[0016] Furthermore, S3 specifically includes the following steps: S31. Under the condition of missing interactive information, the task selection process among UAVs is modeled as a non-cooperative game, in which each UAV, as an independent rational agent, selects a task participation strategy to maximize its expected payoff function based only on its own available local information and historical observation results. The expected utility function defined in S22 is used to characterize the trade-off between task payoff, cooperative gain and competitive penalty, and to describe the strategy interaction behavior of multiple UAVs under the condition of missing interactive information. S32. Under the non-cooperative game framework, through the adaptive adjustment of individual UAV strategies, the UAV task selection strategy gradually converges to a set of task coordination strategies that satisfy the equilibrium conditions of the non-cooperative game, thereby forming a stable UAV task coordination strategy.
[0017] Furthermore, S4 specifically includes the following steps: S41. Based on the task collaboration structure formed by non-cooperative game theory, an evolutionary policy update algorithm guided by proximal policy optimization is constructed. The evolutionary iteration process of the spectrum allocation scheme is regarded as an interactive environment. Proximal policy optimization dynamically adjusts the execution strategies of the selection, crossover, and mutation operators in the evolutionary algorithm by observing the evolutionary state and outputting control actions. Let the first step be... t During generational evolution, the states observed by PPO are denoted as , Defined as:
[0018] in, This represents the historical best individual of the current generation; This is the sequence of control actions for the evolutionary operators in the last 100 generations; Statistical information representing the fitness of the population over the past 100 generations, including average fitness and its changing trend; For the first t A measure of population diversity used to characterize the breadth of population distribution in the coding space; Average fitness Its change is defined as:
[0019] in, G Indicates population size, For the first t The middle generation i Chromosomal coding of an individual, For fitness evaluation function; Population diversity The metric based on gene differences is defined as follows:
[0020] in, H Chromosome length Represents an individual i In the h The values of each gene locus; In terms of motion space design, PPO outputs a continuous-discrete hybrid motion in each generation:
[0021] in, This indicates the crossover position used by the crossover operator. It is a set of variant gene loci, the size of which does not exceed , This parameter represents the mutation intensity and is used to adjust the magnitude of the search perturbation. This action directly affects the execution mode of the evolution operator, thereby guiding the population's search direction and exploration intensity. In designing the reward function, taking into account both the improvement in scheduling performance and changes in population diversity, the following reward function is constructed:
[0022] in, Indicates the first t The optimal fitness of the generation The best fitness in history, This is the diversity weighting coefficient, used to balance the relationship between performance convergence and search diversity; S42. Based on the evolutionary state and the output results of the near-end strategy optimization, guide the evolutionary algorithm to update the task-level spectrum allocation scheme and UAV launch power configuration, so that the evolutionary process, under the premise of meeting the system constraints, prioritizes the direction of improving the overall mission benefits and reducing communication latency. S43. Construct an environmental reward signal based on task execution benefits and communication latency feedback, and input the reward signal into the Critic network to evaluate the quality of the current evolution update result, thereby providing a basis for subsequent evolution direction adjustments; S44. An Actor-Critic network for updating the PPO based on a policy optimization objective function is used to limit the magnitude of policy changes during evolution, ensuring the stability and convergence of the spectrum allocation and power configuration evolution process; the policy optimization objective function is defined as:
[0023] in, Indicates the state of the old and new strategies. Select action The probability ratio, For the estimation of the advantage function, This is the pruning factor, used to limit the magnitude of policy updates and ensure the stability of the training process; S45. Repeat the evolution strategy update and Actor-Critic network strategy update process from S41 to S44 until the preset convergence condition or iteration termination condition is met, and output the corresponding task-level spectrum allocation scheme and UAV launch power configuration scheme as the final spectrum resource scheduling result.
[0024] The present invention also provides a storage medium comprising a stored program, wherein, when the program is executed, it performs any of the above-described spectrum resource scheduling methods for conditions where UAV interactive information is missing.
[0025] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes any of the above-described spectrum resource scheduling methods for unmanned aerial vehicle (UAV) interactive information missing conditions through the computer program.
[0026] Compared with the prior art, the present invention has the following advantages: This invention models and optimizes drones with only partial interactive information, avoiding reliance on global state and complete information sharing, significantly reducing system communication overhead, and improving scheduling feasibility and practicality in complex dynamic environments.
[0027] This invention integrates the maximization of mission benefits and the minimization of communication latency into the same optimization framework by uniformly modeling the competition and cooperation relationships of UAV missions, thereby achieving a synergistic improvement in spectrum resource allocation and mission execution efficiency, and enhancing the overall system performance and fairness. This invention introduces a proximal policy optimization (PPO)-guided evolutionary update mechanism, enabling the spectrum allocation strategy to continuously adjust and converge under conditions of environmental changes and incomplete information, effectively improving the stability, robustness, and long-term optimization capability of the scheduling strategy. Attached Figure Description
[0028] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0029] Figure 1 This is a basic framework diagram of the present invention.
[0030] Figure 2 This is a flowchart of the algorithm of the present invention.
[0031] Figure 3 The graph shows the change in mission revenue as the number of drones increases when the number of missions is 20.
[0032] Figure 4 The graph shows the change in mission revenue as the number of missions increases when the number of drones is 20.
[0033] Figure 5 The graph shows the algorithm's runtime as a function of the number of drones when the number of tasks is 20.
[0034] Figure 6 The graph shows the algorithm's runtime as a function of the number of tasks when there are 20 drones.
[0035] Figure 7 The chart shows a comparison of the maximum communication latency of the algorithm when the number of tasks and the number of drones are both 20. Detailed Implementation
[0036] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0037] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0038] like Figure 1 and 2 As shown, this invention provides a spectrum resource scheduling method under conditions of missing UAV interactive information, comprising the following steps: S1. Establish a basic framework for UAV mission selection and spectrum resource scheduling that comprehensively considers differences in mission type and UAV capabilities. S11. Construct a multi-UAV collaborative reconnaissance system, where the UAV ID set in the system is... The reconnaissance missions are set as follows During the mission execution phase, each UAV chooses to participate in a specific reconnaissance mission based on its own decision-making. To improve system coordination efficiency, a mission-based clustered cooperative communication architecture is adopted. For each mission... k The drones participating in this mission constitute a mission cluster, whose set of members is represented as follows: Within each task cluster, one drone is selected as the cluster head node, and its number is denoted as [node name missing]. It is used to aggregate reconnaissance data from members within the cluster and is responsible for communicating with ground base stations; S12. For the same task, different drones differ in their perception capabilities, communication conditions, and resource status, resulting in inconsistencies between their task benefits and communication load. For any drone... i With the task k Define drone i Execute the task k The obtained task utility is This is used to describe the heterogeneity at the task utility level; the corresponding amount of reconnaissance data generated is defined as... This is used to describe the heterogeneity at the communication load layer; at the same time, to characterize the reward constraints of tasks at the system level, a task is introduced. k Maximum acceptable cumulative utility limit Regarding spectrum resources, assume the total available bandwidth of the system is... B The system is uniformly scheduled by ground base stations and divided into multiple orthogonal sub-channels. Each reconnaissance mission is assigned an independent sub-channel, and its bandwidth is denoted as [missing information]. Based on the aforementioned descriptions of UAV heterogeneity and mission heterogeneity, a basic framework for UAV mission selection and spectrum resource scheduling is established.
[0039] S2. Under the aforementioned framework, the UAV is modeled as a decision-making entity under partial information sharing, the task competition and cooperation relationship between UAVs is characterized, and a mathematical model for spectrum resource scheduling with the goal of maximizing task benefits and minimizing communication latency is constructed. S21, any drone in the system i Before a task can be executed, a task must be selected from the task set, but real-time task selection information from other drones is unavailable. Therefore, a probabilistic task selection strategy is introduced, defining the drone... i Select task k The probability is And constitute its strategy vector. The mission selection strategies of all UAVs collectively constitute the global strategy matrix at the system level. It is used to statistically characterize the overall task distribution of multiple UAVs under conditions of partial information sharing. S22. Under the probabilistic task selection mechanism, to characterize the cooperative gains and resource competition relationships among UAVs, define the UAV... i Participate in the task k Expected utility function The specific formula is as follows:
[0040] The utility function consists of three parts: the first term For direct utility, reflecting dronesi Execute tasks independently k The intrinsic benefits obtained; the second term is synergistic utility, calculated using weighting coefficients. The modeling of the collaborative benefits brought by other UAVs participating in the same mission reflects the gain effect of multi-UAV cooperative reconnaissance; the third term is the competition penalty term, which is applied when the system targets the mission. k expected utility aggregate value Exceeding the task utility limit At that time, a penalty factor was introduced. Suppress overload to avoid excessive drones concentrating on high-value tasks, which could lead to resource competition and diminishing returns; S23. During the mission collaboration process, a mission-based cluster communication architecture is adopted, dividing the communication process into two stages: intra-cluster transmission and inter-cluster transmission. In the intra-cluster communication stage, non-cluster leader UAVs transmit reconnaissance data to the cluster leader node. k The allocated communication bandwidth is denoted as drones i The transmission power is Considering multiple UAVs within the same task cluster sharing the same spectrum resource, their communication process is inevitably affected by intra-cluster interference. Let... Indicates drone i Intra-cluster channel gain to the cluster head node The power is the additive white Gaussian noise. Then the task... k The total intra-cluster communication delay can be expressed as:
[0041] During the inter-cluster communication phase, the cluster head node transmits the aggregated data to the ground base station. Let... Indicates the cluster head node Channel gain with base station The transmit power of the cluster head node. Let be the noise power spectral density. During inter-cluster transmission, the cluster head node needs to forward data generated by all UAVs within the task cluster, with a total data volume of . Then the task k Inter-cluster communication latency can be modeled as follows:
[0042] Task k The total communication delay can be expressed as:
[0043] S24. Based on the above modeling, and with the optimization objective of minimizing the maximum communication latency of all task clusters in the system, a mathematical model for spectrum resource scheduling is constructed, the specific formula of which is:
[0044] S3. Based on the mathematical model obtained in S2, generate a collaborative strategy for drone missions through non-cooperative game theory. S31. Under conditions of missing interactive information, the task selection process among UAVs is modeled as a non-cooperative game, where each UAV acts as an independent rational agent, selecting a task participation strategy to maximize its expected payoff function based solely on its available local information and historical observations. The individual utility function of the UAVs defined in S22 characterizes the trade-off between task payoffs, collaborative gains, and competitive penalties, thus describing the strategic interaction behavior of multiple UAVs under conditions of missing interactive information. S32. Under the aforementioned non-cooperative game framework, through the adaptive adjustment of individual UAV strategies, the UAV task selection strategies gradually converge to a set of task cooperation strategies that satisfy the equilibrium conditions of the non-cooperative game, thereby forming a stable UAV task cooperation structure.
[0045] S4. Based on the task coordination strategy obtained in S3, an evolutionary strategy update mechanism guided by near-end strategy optimization is introduced to jointly adjust the task-level spectrum allocation and UAV launch power configuration.
[0046] S41. Based on the task collaboration structure formed by non-cooperative game theory, an evolutionary policy update algorithm guided by proximal policy optimization is constructed. The evolutionary iteration process of the spectrum allocation scheme is regarded as an interactive environment. Proximal policy optimization dynamically adjusts the execution strategies of the selection, crossover, and mutation operators in the evolutionary algorithm by observing the evolutionary state and outputting control actions. Let the first step be... t During generational evolution, the states observed by PPO are denoted as , Defined as:
[0047] in, This represents the historical best individual of the current generation; This is the sequence of control actions for the evolutionary operators in the last 100 generations; Statistical information representing the fitness of the population over the past 100 generations, including average fitness and its changing trend; For the first t A measure of population diversity used to characterize the breadth of population distribution in the coding space; Average fitness Its change is defined as:
[0048] in, G Indicates population size, For the first t The middle generationi Chromosomal coding of an individual, For fitness evaluation function; Population diversity The metric based on gene differences is defined as follows:
[0049] in, H Chromosome length Represents an individual i In the h The values of each gene locus; In terms of motion space design, PPO outputs a continuous-discrete hybrid motion in each generation:
[0050] in, This indicates the crossover position used by the crossover operator. It is a set of variant gene loci, the size of which does not exceed , This parameter represents the mutation intensity and is used to adjust the magnitude of the search perturbation. This action directly affects the execution mode of the evolution operator, thereby guiding the population's search direction and exploration intensity. In designing the reward function, taking into account both the improvement in scheduling performance and changes in population diversity, the following reward function is constructed:
[0051] in, Indicates the first t The optimal fitness of the generation The best fitness in history, This is the diversity weighting coefficient, used to balance the relationship between performance convergence and search diversity; S42. Based on the evolutionary state and the output results of the near-end strategy optimization, guide the evolutionary algorithm to update the task-level spectrum allocation scheme and UAV launch power configuration, so that the evolutionary process, under the premise of meeting the system constraints, prioritizes the direction of improving the overall mission benefits and reducing communication latency. S43. Construct an environmental reward signal based on task execution benefits and communication latency feedback, and input the reward signal into the Critic network to evaluate the quality of the current evolution update result, thereby providing a basis for subsequent evolution direction adjustments; S44. An Actor-Critic network for updating the PPO based on a policy optimization objective function is used to limit the magnitude of policy changes during evolution, ensuring the stability and convergence of the spectrum allocation and power configuration evolution process; the policy optimization objective function is defined as:
[0052] in, Indicates the state of the old and new strategies. Select action The probability ratio, For the estimation of the advantage function, This is the pruning factor, used to limit the magnitude of policy updates and ensure the stability of the training process; S45. Repeat the evolution strategy update and Actor-Critic network strategy update process from S41 to S44 until the preset convergence condition or iteration termination condition is met, and output the corresponding task-level spectrum allocation scheme and UAV launch power configuration scheme as the final spectrum resource scheduling result.
[0053] This invention proposes a spectrum resource scheduling method under the condition of missing UAV interactive information. By using non-cooperative game theory to form a UAV mission coordination strategy, an evolutionary strategy update mechanism guided by near-end strategy optimization is introduced to jointly adjust the mission-level spectrum allocation and UAV transmit power configuration, with the optimization objectives of maximizing mission benefits and minimizing maximum communication latency.
[0054] This embodiment conducts experiments in a real-world mission scenario, testing under varying numbers of drones and missions. The comparative algorithms used in this invention employ centralized optimization algorithms such as Particle Swarm Optimization (PSO), Genetic Algorithm (GA), Grey Wolf Algorithm (GWO), Whale Algorithm (WOA), and Kingfisher Algorithm (PKO).
[0055] like Figure 3 As shown, the curves of mission revenue change with the number of drones when the number of missions is 20.
[0056] like Figure 4 The figure shows the curve of mission revenue as a function of the number of missions when the number of drones is 20.
[0057] Depend on Figure 3 and Figure 4 It can be seen that the UAV autonomous decision-making based on non-cooperative game theory consistently achieves greater mission benefits than the PSO algorithm under centralized optimization.
[0058] like Figure 5 The figure shows the algorithm's running time as the number of drones changes when the number of tasks is 20.
[0059] like Figure 6 The figure shows the algorithm's running time as a function of the number of tasks when the number of drones is 20.
[0060] Depend on Figure 5 and Figure 6It can be seen that the UAV autonomous decision-making based on non-cooperative game theory can consistently maintain a lower algorithm running time than the PSO algorithm under centralized optimization.
[0061] like Figure 7 The figure shown is a comparison of the maximum communication latency of the algorithm when the number of tasks is 20 and the number of drones is 20.
[0062] Depend on Figure 7 It can be seen that the spectrum allocation scheme obtained by the evolution update mechanism guided by near-end strategy optimization has the minimum communication latency.
[0063] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A spectrum resource scheduling method for unmanned aerial vehicle (UAV) interactive information deficiency, characterized in that, Includes the following steps: S1. Establish a basic framework for UAV mission selection and spectrum resource scheduling that comprehensively considers differences in mission type and UAV capabilities. S2. Under the basic framework of UAV mission selection and spectrum resource scheduling, UAVs are modeled as decision-making entities under partial information sharing, the mission competition and cooperation relationships between UAVs are characterized, and a mathematical model for spectrum resource scheduling with the goal of maximizing mission benefits and minimizing communication latency is constructed. S3. Based on the mathematical model of spectrum resource scheduling, generate a collaborative strategy for UAV missions through non-cooperative game theory. S4. Based on the UAV mission coordination strategy, an evolutionary strategy update mechanism guided by near-end strategy optimization is introduced to jointly adjust the mission-level spectrum allocation and UAV transmit power configuration, and output the mission-level spectrum allocation scheme and UAV transmit power configuration scheme as the final spectrum resource scheduling result.
2. The spectrum resource scheduling method for unmanned aerial vehicle (UAV) interactive information deficiency conditions as described in claim 1, characterized in that, S1 specifically includes the following steps: S11. Construct a multi-UAV collaborative reconnaissance system, where the UAV ID set in the system is... The reconnaissance missions are set as follows ; During the mission execution phase, each UAV chooses to participate in a specific reconnaissance mission based on its own decision-making. A mission-based clustered cooperative communication architecture is adopted, and for each mission k, the UAVs participating in that mission form a mission cluster. The set of mission cluster members is represented as... Within each task cluster, one drone is selected as the cluster head node, and its number is denoted as [node name missing]. It is used to aggregate reconnaissance data from members within the cluster and is responsible for communicating with ground base stations; S12, targeting any drone i With the task k Define drone i Execute the task k The obtained task utility is This is used to describe the heterogeneity at the task utility level; the corresponding amount of reconnaissance data generated is defined as... This is used to describe the heterogeneity at the communication load layer; it introduces tasks. k Maximum acceptable cumulative utility limit Regarding spectrum resources, assume the total available bandwidth of the system is... B The system is uniformly scheduled by ground base stations and divided into multiple orthogonal sub-channels. Each reconnaissance mission is assigned an independent sub-channel, and its bandwidth is denoted as [missing information]. .
3. The spectrum resource scheduling method for unmanned aerial vehicle (UAV) interactive information deficiency conditions as described in claim 1, characterized in that, S2 specifically includes the following steps: S21. Introduce a probabilistic task selection strategy and define the UAV. i Select task k The probability is And constitute its strategy vector. The mission selection strategies of all UAVs collectively constitute the global strategy matrix at the system level. It is used to statistically characterize the overall task distribution of multiple UAVs under conditions of partial information sharing. S22. Under the probabilistic task selection mechanism, to characterize the cooperative gains and resource competition relationships among UAVs, define the UAV... i Participate in the task k Expected utility function The specific formula is as follows: The utility function consists of three parts: the first term For direct utility, reflecting drones i Execute tasks independently k The intrinsic benefits obtained; the second term is synergistic utility, calculated using weighting coefficients. The modeling of the collaborative benefits brought by other UAVs participating in the same mission reflects the gain effect of multi-UAV cooperative reconnaissance; the third term is the competition penalty term, which is applied when the system targets the mission. k expected utility aggregate value Exceeding the task utility limit At that time, a penalty factor was introduced. Suppress overload to avoid excessive drones concentrating on high-value tasks, which could lead to resource competition and diminishing returns; S23. During the task collaboration process, a task-based cluster communication architecture is adopted, dividing the communication process into two stages: intra-cluster transmission and inter-cluster transmission. In the intra-cluster communication stage, non-cluster leader UAVs transmit reconnaissance data to the cluster leader node. k The allocated communication bandwidth is denoted as drones i The transmission power is ;set up Indicates drone i Intra-cluster channel gain to the cluster head node Given the additive white Gaussian noise power; then the task... k The total intra-cluster communication delay is expressed as: During the inter-cluster communication phase, the cluster head node transmits the aggregated data to the ground base station; assuming... Indicates the cluster head node Channel gain with base station The transmit power of the cluster head node. The noise power spectral density; During inter-cluster transmission, the cluster head node needs to forward data generated by all UAVs within the task cluster, with a total data volume of [data missing]. Then the task k Inter-cluster communication latency can be modeled as follows: Task k The total communication delay is expressed as: S24. With minimizing the maximum communication latency of all task clusters in the system as the optimization objective, construct a mathematical model for spectrum resource scheduling. The specific formula is as follows: 。 4. The spectrum resource scheduling method for unmanned aerial vehicle (UAV) interactive information deficiency conditions as described in claim 3, characterized in that, S3 specifically includes the following steps: S31. Under the condition of missing interactive information, the task selection process among UAVs is modeled as a non-cooperative game, in which each UAV, as an independent rational agent, selects a task participation strategy to maximize its expected payoff function based only on its own available local information and historical observation results. The expected utility function defined in S22 is used to characterize the trade-off between task payoff, cooperative gain and competitive penalty, and to describe the strategy interaction behavior of multiple UAVs under the condition of missing interactive information. S32. Under the non-cooperative game framework, through the adaptive adjustment of individual UAV strategies, the UAV task selection strategy gradually converges to a set of task coordination strategies that satisfy the equilibrium conditions of the non-cooperative game, thereby forming a stable UAV task coordination strategy.
5. The spectrum resource scheduling method for unmanned aerial vehicle (UAV) interactive information deficiency conditions as described in claim 1, characterized in that, S4 specifically includes the following steps: S41. Based on the task collaboration structure formed by non-cooperative game theory, an evolutionary policy update algorithm guided by proximal policy optimization is constructed. The evolutionary iteration process of the spectrum allocation scheme is regarded as an interactive environment. Proximal policy optimization dynamically adjusts the execution strategies of the selection, crossover, and mutation operators in the evolutionary algorithm by observing the evolutionary state and outputting control actions. Let the first step be... t During generational evolution, the states observed by PPO are denoted as , Defined as: in, This represents the historical best individual of the current generation; This is the sequence of control actions for the evolutionary operators in the last 100 generations; Statistical information representing the fitness of the population over the past 100 generations, including average fitness and its changing trend; For the first t A measure of population diversity used to characterize the breadth of population distribution in the coding space; Average fitness Its change is defined as: in, G Indicates population size, For the first t The middle generation i Chromosomal coding of an individual, For fitness evaluation function; Population diversity The metric based on gene differences is defined as follows: in, H Chromosome length Represents an individual i In the h The values of each gene locus; In terms of motion space design, PPO outputs a continuous-discrete hybrid motion in each generation: in, This indicates the crossover position used by the crossover operator. It is a set of variant gene loci, the size of which does not exceed , This parameter represents the mutation intensity and is used to adjust the magnitude of the search perturbation. This action directly affects the execution mode of the evolution operator, thereby guiding the population's search direction and exploration intensity. In designing the reward function, taking into account both the improvement in scheduling performance and changes in population diversity, the following reward function is constructed: in, Indicates the first t The optimal fitness of the generation The best fitness in history, This is the diversity weighting coefficient, used to balance the relationship between performance convergence and search diversity; S42. Based on the evolutionary state and the output results of the near-end strategy optimization, guide the evolutionary algorithm to update the task-level spectrum allocation scheme and UAV launch power configuration, so that the evolutionary process, under the premise of meeting the system constraints, prioritizes the direction of improving the overall mission benefits and reducing communication latency. S43. Construct an environmental reward signal based on task execution benefits and communication latency feedback, and input the reward signal into the Critic network to evaluate the quality of the current evolution update result, thereby providing a basis for subsequent evolution direction adjustments; S44. An Actor-Critic network for updating the PPO based on a policy optimization objective function is used to limit the magnitude of policy changes during evolution, ensuring the stability and convergence of the spectrum allocation and power configuration evolution process; the policy optimization objective function is defined as: in, Indicates the state of the old and new strategies. Select action The probability ratio, For the estimation of the advantage function, This is the pruning factor, used to limit the magnitude of policy updates and ensure the stability of the training process; S45. Repeat the evolution strategy update and Actor-Critic network strategy update process from S41 to S44 until the preset convergence condition or iteration termination condition is met, and output the corresponding task-level spectrum allocation scheme and UAV launch power configuration scheme as the final spectrum resource scheduling result.
6. A storage medium, characterized in that, The storage medium includes a stored program, wherein when the program is executed, it performs the spectrum resource scheduling method for unmanned aerial vehicle (UAV) interactive information missing conditions as described in any one of claims 1 to 5.
7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, The processor executes the spectrum resource scheduling method for unmanned aerial vehicle (UAV) interactive information missing conditions through the computer program according to any one of claims 1 to 5.