Intelligent communication decision method and system for multi-uav cooperative situation assessment
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- NAT UNIV OF DEFENSE TECH
- Filing Date
- 2026-06-29
- Publication Date
- 2026-08-04
AI Technical Summary
[0004]然而,现有多智能体协同方法虽然能够在一定程度上提高多无人机系统的任务回报,但多数方法默认通信可靠,或采用全广播通信、稠密注意力通信等机制
[0009]The aforementioned intelligent communication decision-making method and system for multi-UAV cooperative situation assessment models the communication process as a collaborative, decentralized, partially observable Markov decision process. By introducing an information reliability index into the global state, it can explicitly quantify the impact of interference on information quality, effectively distinguish between genuine task-related information and pseudo-related information caused by interference, and suppress the spread of low-quality information in UAV swarms. By jointly calculating the comprehensive value of links based on task context, candidate sender information reliability, and the potential task assistance of candidate links, and selecting high-value links to generate communication link selection masks, it can break away from the traditional mode of selecting communication based on spatial distance and surface feature similarity, and optimize communication targets to meet the actual task requirements of each UAV. By constructing a joint training objective including communication regularization terms, it can standardize communication density and distribution patterns, reduce redundant communication links, and adapt to bandwidth-constrained operating conditions under dynamic interference scenarios. By configuring a team-shared reward value that integrates target service rewards, safe operation penalties, and communication cost penalties, it can achieve synergistic optimization of task benefits, flight safety, and communication consumption. The embodiments of the present invention can improve the accuracy of collaborative situational assessment of multi-UAV swarms and the continuous service capability of key targets in complex environments with dynamic interference and partial observability, while significantly improving the utilization efficiency of limited communication resources.
Smart Images

Figure CN122513397A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of unmanned aerial vehicle (UAV) communication technology, and in particular to an intelligent communication decision-making method and system for multi-UAV collaborative situation assessment. Background Technology
[0002] Dynamic interference coordination scenarios typically exhibit significant dynamism, uncertainty, and communication constraints. These scenarios can occur in post-disaster areas, post-war areas, or other areas requiring long-term situational awareness assessment for coordinated missions. On one hand, factors such as road blockages, building damage, smoke and dust obscuring the view, terrain masking, residual radiation, or facility destruction can affect the UAV's spatial coverage and perception quality. On the other hand, external interference sources such as temporary base stations, emergency communication vehicles, portable relay nodes, electronic jamming equipment, and residual electromagnetic radiation sources are constantly changing; their location, transmission power, and activity status can simultaneously affect the UAV's local observation quality and wireless link quality.
[0003] For situational assessment tasks involving multiple UAVs, the focus is not only on target discovery, but also on continuous updates to target status, long-term service for high-priority targets, and the establishment of stable information feedback links in key areas. Therefore, when performing such tasks, multi-UAV systems require both reasonable spatial division of labor and effective information sharing under conditions of limited bandwidth, limited links, and partial observability.
[0004] However, while existing multi-agent cooperative methods can improve the mission rewards of multi-UAV systems to some extent, most methods assume reliable communication or employ mechanisms such as full-broadcast communication or dense attention communication. These methods are effective in ideal environments but have significant limitations in dynamic interference cooperative scenarios. In particular, when several UAVs are simultaneously within the range of the same interference source, their observational characteristics may exhibit high similarity, but this similarity may stem from shared noise rather than genuine mission value. If communication is still established directly based on feature similarity or proximity, low-quality information is easily further propagated within the system, making it difficult to distinguish between mission-related information and pseudo-related information caused by shared interference. Furthermore, the selection of communication targets often relies on spatial distance or surface feature matching, making it difficult to guarantee that the selected communication targets can provide effective support for the receiver's mission decisions. In addition, existing methods struggle to effectively control the number of communication links, easily leading to excessively high communication density and scattered distribution, resulting in a large amount of redundant information exchange. This makes them unsuitable for bandwidth and link resource constraints in dynamic interference scenarios and makes it difficult to achieve an optimal balance between mission rewards, critical target service capabilities, and communication resource utilization efficiency.
[0005] Therefore, improving the accuracy of multi-UAV cooperative situation assessment and the efficiency of communication resource utilization in dynamic interference cooperative scenarios has become an urgent technical problem to be solved. Summary of the Invention
[0006] Therefore, it is necessary to provide an intelligent communication decision-making method and system for multi-UAV collaborative situation assessment to address the aforementioned technical problems.
[0007] A smart communication decision-making method for multi-UAV cooperative situation assessment, the method comprising: The multi-UAV cooperative situation assessment communication process in a dynamic interference cooperative scenario is modeled as a cooperative distributed partially observable Markov decision process, resulting in a situation assessment task model. Each time step of the situation assessment task model includes a global state, local observations, joint actions, state transition probabilities, and a team-shared reward value. The global state includes a set of UAV states, a set of target states, and a set of dynamic interference source states, where the UAV states include information reliability indicators. The local observations are the state of a single UAV that it can acquire, the states of its visible teammates, and the states of visible targets. The joint actions include the motion actions and communication selection actions of each UAV, where the communication selection actions are represented by a communication link selection mask. The team-shared reward value includes a target service reward, a safe operation penalty, and a communication cost penalty. A joint training objective is constructed based on the aforementioned situation assessment task model; the joint training objective includes a temporal difference loss, a communication regularization term, a link utility auxiliary term, and a pairing interaction auxiliary term; the communication regularization term is used to constrain communication density and communication distribution pattern. A multi-agent reinforcement learning algorithm with centralized training and decentralized execution is used to solve the joint training objective, resulting in an optimal multi-UAV communication cooperation strategy. This optimal strategy includes calculating the comprehensive value of each candidate communication link for each receiver UAV based on the task context, the reliability of candidate sender information, and the potential task helpness of the candidate links. A predetermined number of candidate communication links with the highest comprehensive value are retained to generate a communication link selection mask. The task context is a feature vector extracted from the receiver UAV's local observations and historical hidden states. The potential task helpness is an estimated link utility value obtained through iterative updates based on historical link activation data and corresponding task benefit changes during training. The optimal multi-UAV communication and coordination strategy is deployed to the onboard processor of each UAV, enabling each UAV to complete communication and coordination decisions and situation assessments during the online execution phase.
[0008] An intelligent communication and decision-making system for collaborative situation assessment of multiple unmanned aerial vehicles (UAVs), the system comprising: The task modeling module is used to model the multi-UAV cooperative situation assessment communication process under dynamic interference cooperative scenarios as a cooperative distributed partially observable Markov decision process, resulting in a situation assessment task model. Each time step of the situation assessment task model includes global state, local observation, joint actions, state transition probability, and team-shared reward value. The global state includes a set of UAV states, a set of target states, and a set of dynamic interference source states, where the UAV states include information reliability indicators. The local observations are the state of a single UAV that can be acquired, the states of its visible teammates, and the states of visible targets. The joint actions include the motion actions and communication selection actions of each UAV, where the communication selection actions are represented by a communication link selection mask. The team-shared reward value includes target service reward, safe operation penalty, and communication cost penalty. The training objective construction module is used to construct a joint training objective based on the situation assessment task model; the joint training objective includes a temporal difference loss, a communication regularization term, a link utility auxiliary term, and a pairing interaction auxiliary term; the communication regularization term is used to constrain the communication density and communication distribution pattern. The strategy solving module is used to solve the joint training objective using a multi-agent reinforcement learning algorithm with centralized training and decentralized execution, to obtain the optimal multi-UAV communication cooperation strategy. The optimal multi-UAV communication cooperation strategy includes, for each receiving UAV, calculating the comprehensive value of each candidate communication link based on the task context, the reliability of the candidate sender's information, and the potential task helpness of the candidate link, and retaining a preset number of candidate communication links with the highest comprehensive value to generate a communication link selection mask. The task context is a feature vector extracted from the local observations and historical hidden states of the receiving UAV; the potential task helpness is an estimated link utility value obtained through iterative updates based on historical link activation data and corresponding task benefit changes during training. The strategy deployment module is used to deploy the optimal multi-UAV communication and coordination strategy to the onboard processor of each UAV, so that each UAV can complete communication and coordination decision-making and situation assessment during the online execution phase.
[0009] The aforementioned intelligent communication decision-making method and system for multi-UAV cooperative situation assessment models the communication process as a collaborative, decentralized, partially observable Markov decision process. By introducing an information reliability index into the global state, it can explicitly quantify the impact of interference on information quality, effectively distinguish between genuine task-related information and pseudo-related information caused by interference, and suppress the spread of low-quality information in UAV swarms. By jointly calculating the comprehensive value of links based on task context, candidate sender information reliability, and the potential task assistance of candidate links, and selecting high-value links to generate communication link selection masks, it can break away from the traditional mode of selecting communication based on spatial distance and surface feature similarity, and optimize communication targets to meet the actual task requirements of each UAV. By constructing a joint training objective including communication regularization terms, it can standardize communication density and distribution patterns, reduce redundant communication links, and adapt to bandwidth-constrained operating conditions under dynamic interference scenarios. By configuring a team-shared reward value that integrates target service rewards, safe operation penalties, and communication cost penalties, it can achieve synergistic optimization of task benefits, flight safety, and communication consumption. The embodiments of the present invention can improve the accuracy of collaborative situational assessment of multi-UAV swarms and the continuous service capability of key targets in complex environments with dynamic interference and partial observability, while significantly improving the utilization efficiency of limited communication resources. Attached Figure Description
[0010] Figure 1 This is a flowchart illustrating an intelligent communication decision-making method for multi-UAV collaborative situation assessment in one embodiment. Figure 2 This is a schematic diagram of a Top-K sparse communication topology in one embodiment; Figure 3 This is a schematic diagram of a multi-UAV communication and coordination method in one embodiment; Figure 4 This is a schematic diagram comparing the task reward performance of the method of the present invention and the comparative scheme in one embodiment, wherein, Figure 4 (a) is a schematic diagram showing the evolution trend of the average test reward of each scheme with the number of training steps, and Figure 4(b) is a schematic diagram showing the final evaluation reward statistics of each scheme after training convergence. Figure 5 Figure 5(a) is a schematic diagram of the communication structure evolution in one embodiment, and Figure 5(b) is a schematic diagram of the communication frequency distribution between UAVs in the early training stage. Figure 6Figure 6(a) shows a comparison of the module comparison results in one embodiment. Figure 6(b) shows a comparison of the test reward changing with the number of training steps. Figure 6(c) shows a comparison of the fairness index changing with the number of training steps. Figure 6(d) shows a comparison of the key target recall rate changing with the number of training steps. Figure 6 (e) is a schematic diagram comparing the number of activated communication links with the number of training steps, and Figure 6(f) is a schematic diagram comparing the single-step communication cost with the number of training steps. Figure 7 This is a schematic diagram of the hierarchical operation architecture of an intelligent communication decision-making system for collaborative situation assessment of multiple UAVs in one embodiment. Detailed Implementation
[0011] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0012] This invention explicitly incorporates the communication selection problem among multiple UAVs into the collaborative decision-making process. Within a multi-agent reinforcement learning framework of centralized training and distributed execution, it constructs a method based on a joint scoring of task context and information reliability. A sparse communication method. This method includes: unified modeling of dynamic interference cooperative environments; explicit quantification of the reliability of UAV information sources; joint scoring of communication partners based on the receiver's task context, the sender's reliability, and the potential assistance of candidate links to the receiver's current task; and utilizing... A sparse selection mechanism limits the number of active communication links; and a communication regularization mechanism constrains communication density and distribution, thereby forming an efficient, controllable, and scalable multi-UAV cooperative communication method in complex interference environments.
[0013] In one embodiment, such as Figure 1 As shown, an intelligent communication decision-making method for multi-UAV cooperative situation assessment is provided, including the following steps: Step 102: The multi-UAV cooperative situation assessment communication process in the dynamic interference cooperative scenario is modeled as a cooperative distributed partially observable Markov decision process, thus obtaining the situation assessment task model.
[0014] Each time step of the situation assessment task model includes the global state, local observations, joint actions, state transition probabilities, and team-shared reward values. The global state includes the set of UAV states, the set of target states, and the set of states of dynamic interference sources, with UAV states including information reliability indicators. Local observations consist of the state of the individual UAV, the states of visible teammates, and the states of visible targets. Joint actions include the motion actions and communication selection actions of each UAV, with communication selection actions represented by a communication link selection mask. Team-shared reward values include target service rewards, safe operation penalties, and communication cost penalties.
[0015] The cooperative, decentralized, partially observable Markov decision process (MOD) is a standard mathematical modeling framework in the field of multi-agent cooperation. Cooperative means that all agents share a completely consistent common goal and a unified team reward; decentralized means that each agent makes decisions independently without a central controller for unified scheduling; and partially observable means that each agent can only acquire local information within its own perception range and cannot grasp the overall global state. The situation assessment task model is a mathematical abstraction of the multi-UAV cooperative problem under dynamic interference scenarios in this invention, based on this framework. It transforms the communication and control processes of multiple UAVs into an optimization problem solvable mathematically. This modeling framework is applicable to various complex cooperative scenarios, such as post-disaster search and rescue scenarios, where each UAV is used to perform personnel search and rescue and environmental reconnaissance tasks. Targets can be trapped personnel, relief supply drop points, dangerous area boundaries, and other objects that need to be discovered and continuously monitored. Cooperative communication manifests as UAVs sharing key information such as the location of trapped personnel and the extent of dangerous areas on demand, rather than indiscriminately broadcasting all perceived data. In the scenario of continuous monitoring of key targets, each UAV performs area patrol and target tracking tasks. Targets refer to key protective facilities, suspicious moving targets, areas of abnormal activity, etc. Cooperative communication is manifested in the UAVs selectively sharing information such as target movement trajectory and abnormal event warnings according to the division of monitoring tasks.
[0016] It is understandable that this modeling approach can integrate communication choices and flight maneuvers into a single decision-making framework, enabling joint optimization of communication and control. It can accurately characterize scenario characteristics such as partial observability caused by rubble obstruction during post-disaster search and rescue, and communication quality fluctuations caused by electromagnetic interference during monitoring of key targets.
[0017] Step 104: Construct a joint training objective based on the situation assessment task model.
[0018] The joint training objective includes temporal difference loss, communication regularization term, link utility auxiliary term, and pairing interaction auxiliary term. The communication regularization term is used to constrain communication density and communication distribution pattern.
[0019] Temporal difference loss is a core loss function in reinforcement learning used to approximate the value function. It updates model parameters by minimizing the error between the current value estimate and the actual cumulative reward in the future. The communication regularization term is a constraint specifically introduced in this invention for the communication cooperation problem, used to suppress redundant communication during training. The link utility auxiliary term is used to supervise the marginal contribution of a single communication link to the task reward. The pairing interaction auxiliary term is used to supervise the interaction effect of pairing communication between UAVs.
[0020] It is understandable that by constructing a joint training objective that includes multiple dimensions, it is possible to effectively constrain the number and distribution of communication links while ensuring the performance of the basic situation assessment task. This can accelerate the convergence speed of model training, guide the model to learn more efficient and targeted communication strategies, and balance task benefits with the efficiency of communication resource utilization.
[0021] Step 106: A multi-agent reinforcement learning algorithm with centralized training and decentralized execution is used to solve the joint training objective and obtain the optimal multi-UAV communication and cooperation strategy.
[0022] The optimal multi-UAV communication cooperation strategy involves calculating the comprehensive value of each candidate communication link for each receiver UAV based on the task context, the reliability of information from candidate senders, and the potential task assistance of candidate links. A predetermined number of candidate communication links with the highest comprehensive value are retained to generate a communication link selection mask. The task context is a feature vector extracted from the receiver UAV's local observations and historical hidden states. The potential task assistance is an estimate of the link utility obtained through iterative updates based on historical link activation data and corresponding task benefit changes during training.
[0023] Multi-agent reinforcement learning (MAL) is a machine learning method that enables multiple agents to automatically learn optimal cooperative behavioral strategies through continuous interaction and trial and error with their environment. Centralized training and distributed execution is the mainstream practical architecture for MML. During the training phase, global state information is used to centrally optimize the policy parameters of all agents, while during the execution phase, each agent makes independent decisions based solely on its own local observations.
[0024] It is understandable that this architecture can fully utilize global information during the training phase to improve the overall quality of the strategy, while ensuring the fully distributed nature and system robustness during the execution phase. It can adapt to the actual deployment requirements of drone swarms without a central controller, effectively solve large-scale multi-drone collaborative problems, and the optimal strategy obtained can achieve efficient communication and collaboration under dynamic interference and communication constraints.
[0025] Step 110: Deploy the optimal multi-UAV communication and coordination strategy to the onboard processor of each UAV, so that each UAV can complete communication and coordination decision-making and situation assessment during the online execution phase.
[0026] The airborne processor is an embedded computing device carried by the drone, characterized by low power consumption, small size, and high real-time performance. It is responsible for running the deployed model and executing real-time task decisions. During the online execution phase, it only performs forward inference calculations of the model and does not perform any reverse updates or optimizations of parameters.
[0027] It is understandable that deploying the trained optimal strategy to the onboard processors of each UAV can enable the fully distributed autonomous operation of multiple UAV systems without relying on real-time command from ground stations or central controllers. This reduces the system's dependence on communication infrastructure, improves the system's survivability and mission execution efficiency in complex interference environments, and enables the real-time output of situation assessment results and control commands under the limited onboard computing power of UAVs, thus meeting the real-time requirements of dynamic mission scenarios.
[0028] The aforementioned intelligent communication decision-making method for multi-UAV collaborative situation assessment models the communication process as a collaborative, decentralized, partially observable Markov decision process. By introducing an information reliability index into the global state, it can explicitly quantify the impact of interference on information quality, effectively distinguish between genuine task-related information and pseudo-related information caused by interference, and suppress the spread of low-quality information in UAV swarms. By jointly calculating the comprehensive value of links based on task context, candidate sender information reliability, and the potential task assistance of candidate links, and by selecting high-value links to generate communication link selection masks, it can move away from the traditional mode of selecting communication based on spatial distance and surface feature similarity, and optimize communication targets to meet the actual task requirements of each UAV. By constructing a joint training objective that includes communication regularization terms, it can standardize communication density and distribution patterns, reduce redundant communication links, and adapt to bandwidth-constrained operating conditions under dynamic interference scenarios. By configuring a team-shared reward value that integrates target service rewards, safe operation penalties, and communication cost penalties, it can achieve synergistic optimization of task benefits, flight safety, and communication consumption. The embodiments of the present invention can improve the accuracy of collaborative situational assessment of multi-UAV swarms and the continuous service capability of key targets in complex environments with dynamic interference and partial observability, while significantly improving the utilization efficiency of limited communication resources.
[0029] In one embodiment, the UAV state set includes the UAV state of each UAV deployed in the environment; the UAV state includes UAV location, information reliability indicators, and at least one additional platform information such as boundary state and collision state; the target state set includes the target state of multiple task targets; the target state includes target location, target discovery status, and target priority; the task target refers to the object that the multi-UAV system needs to discover, continuously track, and provide situational awareness services in a dynamic interference cooperative scenario; the dynamic interference source state set includes the dynamic interference source state of multiple dynamic interference sources; the dynamic interference source state includes the interference source location and interference intensity.
[0030] In this embodiment, the multi-UAV situation assessment task under dynamic interference cooperative scenarios is modeled as a cooperative distributed partially observable Markov decision process (Dec-POMDP):
[0031] in, Indicates a collection of drones. Represents the global state space. Indicates the first The operational space of a drone Indicates its local observation space, Represents the state transition function. Represents the observation function, This indicates that the team shares the rewards. This represents the discount factor.
[0032] At any moment The system's global state is represented as , in Represents the set of drone states. Represents the set of target states. This represents the set of states of dynamic environmental interference sources.
[0033] The state of a drone can be represented as , in Location of the drone. For local reliability metrics, It indicates additional platform information such as boundary status and collision status.
[0034] The target state can be represented as , in For the target location, To determine the status of the target, The current priority of the target.
[0035] The state of dynamic environmental interference sources can be represented as , in Location of the interference source This refers to the interference intensity or operating level.
[0036] In one embodiment, the step of calculating the information reliability index includes: calculating the aggregated interference intensity based on the current location of the UAV and the state of the dynamic interference source, and generating the information reliability index based on the aggregated interference intensity; or, estimating the equivalent interference intensity based on the link quality index measured by the UAV's onboard equipment, and generating the information reliability index based on the equivalent interference intensity.
[0037] In this embodiment, to reflect the combined impact of interference on the observation quality and link quality of the UAV, the present invention introduces a reliability modeling method based on the degree of interference exposure. For the UAV... Define its time at time The polymerization interference intensity is
[0038] in This is used to avoid numerical singularities caused by excessively small distances.
[0039] Based on this aggregated interference intensity, the following reliability indicators for the UAV are constructed:
[0040] in The sensitivity of control to the degradation of reliability as interference exposure increases. In implementations where prior environmental information is available, reliability variables can be directly calculated from the scenario state; in actual deployments, they can be further estimated from link quality indicators such as received signal strength indication, signal-to-interference-plus-noise ratio, packet loss rate, and spectral energy detection.
[0041] In situations where the absolute location and transmission power of the interference source cannot be directly obtained, a single UAV can measure link quality indicators such as received signal strength indication, signal-to-interference-plus-noise ratio, packet loss rate, retransmission count, bit error rate, or spectrum occupancy rate using an onboard communication chip, RF front-end, or link monitoring module. After normalizing each indicator, an equivalent interference strength estimate can be constructed. For example:
[0042] in , , and These represent the standardized metrics obtained by mapping received signal strength indication, signal-to-interference-plus-noise ratio, packet loss rate, and spectrum occupancy rate, respectively. , , and These are the weighting coefficients. The reliability variable can be further obtained through the following mapping:
[0043] in This represents the Sigmoid mapping function. Therefore, even if the physical parameters of external interference sources cannot be directly observed, the airborne processor can still estimate information reliability in real time based on actual link quality indicators.
[0044] In one embodiment, the motion action is selected from a discrete set of actions, which includes forward, backward, left turn, and right turn.
[0045] In this embodiment, each UAV performs a combined motion control and communication selection action at each decision step:
[0046] Among them, motion actions can be derived from discrete action sets. Selected from the middle; communication actions are represented by the link selection mask:
[0047] Local observation is
[0048] These represent the user's own status, the status of visible teammates, and the status of visible targets, respectively. Due to dynamic interference and partial observability, the quality of observations received by different UAVs is not consistent.
[0049] In one embodiment, the step of calculating the comprehensive value includes: calculating the task context matching degree based on the task context of the receiving UAV and the local observation features of the candidate sender; generating an information reliability correction term based on the information reliability index of the candidate sender; generating a link utility correction term based on the potential task helpness of the candidate link; and weighted summing the task context matching degree, the information reliability correction term, and the link utility correction term to obtain the comprehensive value of each candidate communication link.
[0050] In this embodiment, as Figure 2 The diagram illustrates a Top-K sparse communication topology. In the communication selection phase, this invention does not directly establish communication based on the spatial distance or surface feature similarity between UAVs. Instead, it first scores whether candidate senders are truly helpful to the receiver's current task before performing sparse selection. For the receiver UAV... i First, a task context query vector is constructed based on its own observations and historical decision-making states:
[0051] in This represents the historical hidden state of the receiver. This query vector reflects the receiver's current task focus, local situational needs, and decision context.
[0052] For candidate senders Construct the key vector:
[0053] in This represents the characteristics of the sender.
[0054] Therefore, a basic compatibility score is calculated between the receiver's current task context and the candidate sender's information representation:
[0055] Further incorporating the sender reliability term and the link utility term learned during the training phase, an enhanced score is obtained:
[0056] in For reliability sensitivity coefficient, Used to avoid taking the logarithm of zero. For link-assisted utility estimation, The weighting coefficients are used to determine the relative importance of the sender's information within the receiver's current task context. In the above scoring, the basic compatibility term measures whether the sender's information fits the receiver's current task context; the reliability term suppresses candidate senders with strong interference and low information quality; and the link utility term characterizes the marginal contribution of a candidate link to improving the receiver's current task value. Therefore, the communication target selection logic of this invention is not "communicating with whoever is closest" or "communicating with whoever looks similar," but rather "prioritizing communication with whoever's information is both reliable, more relevant to the receiver's current task, and more likely to bring about task benefit improvement."
[0057] To control communication costs and meet bandwidth constraints, this invention reserves no more than [amount missing] for each receiver. There are only a few communication links. That is, after completing the joint scoring, communication is not established with all visible neighbors, but only with a few candidate senders with the highest scores. During the training phase, Gumbel noise is added to the scores, and a differentiable approximately discrete distribution is generated using temperature-scaled softmax. Operator generates hard communication mask:
[0058] in For Gumbel noise, This refers to the communication temperature parameter.
[0059] The retained neighbor information is aggregated to obtain communication messages:
[0060] In one embodiment, the communication regularization term includes a communication density constraint term, a communication distribution entropy constraint term, an average budget constraint term, and a budget distribution entropy constraint term. The communication density constraint term is used to limit the average number of communication links per UAV, and the communication distribution entropy constraint term is used to limit the dispersion of communication links. In the joint training objective, the temporal difference loss is directly added to the communication regularization term, and the link utility auxiliary term and the pairing interaction auxiliary term are adjusted by their respective outer weight coefficients and then added to the joint training objective.
[0061] In this embodiment, to suppress redundant communication and constrain the communication organization pattern, the present invention introduces a communication regularization term in the training objective:
[0062] in Indicates the first Communication density of each receiver, This represents its communication distribution entropy. Indicates the average trigger budget level. Represents the budget distribution entropy. , , and These are the internal weight coefficients for the corresponding items.
[0063] The overall training objective of the system is
[0064] in For temporal difference loss based on joint value hybrid networks, For communication regular expressions, Used for supervising single-link utility estimation Used to monitor the effectiveness of sender pairing interactions. In the current implementation, and They are added directly without setting additional outer weights; The components are weighted by their internal coefficients; only and Through outer layer coefficient and Adjustments were made.
[0065] In one embodiment, the step of calculating the team-shared reward value includes: calculating the target service reward based on the service rate of each target per unit time and the priority of the corresponding target; calculating the collision penalty based on the collision state of the drone, calculating the boundary penalty based on the boundary crossing state of the drone, and summing the collision penalty and the boundary penalty to obtain the safe operation penalty; calculating the communication cost penalty based on the total number of active communication links in the system at the current moment; subtracting the safe operation penalty and the communication cost penalty from the target service reward to obtain the original reward value, and cropping the original reward value within a preset value range to obtain the team-shared reward value.
[0066] In this embodiment, to balance the benefits of the target service with the communication cost, the present invention adopts a shared reward function:
[0067] in To the target Service speed, Prioritize the target. and These are the average collision penalty and the boundary penalty, respectively. This is the cost of communication.
[0068] In one embodiment, the optimal multi-UAV communication cooperation strategy is deployed to the onboard processor of each UAV, enabling each UAV to complete communication cooperation decision-making and situation assessment during the online execution phase. This includes: during the online execution phase, each UAV performs forward inference based on the deployed optimal multi-UAV communication cooperation strategy, selects a mask to receive information from the corresponding sender based on the current communication link, and performs weighted aggregation according to the weights obtained after normalization of the comprehensive value of each retained communication link to obtain the communication message; each UAV fuses the communication message with local observations and outputs the situation assessment result, target service decision, and flight action control command at the current moment.
[0069] In this embodiment, a layered architecture of offline centralized training and online distributed execution is adopted. The computationally intensive model parameter iterative optimization process is completed on ground control computing equipment or edge nodes, thus avoiding the practical deployment shortcomings of UAV onboard processors, which have limited computing power and power consumption and cannot support large-scale training operations. Figure 3 The flowchart of the multi-UAV communication and cooperation method shown is illustrated. The method of the present invention can be implemented according to the following steps: Step 1: Initialize the dynamic interference cooperative scenario by setting the task area, number of UAVs, number of targets, status of dynamic interference sources, communication budget constraints, and task time domain, and establish a multi-UAV cooperative situation assessment task model.
[0070] Step 2: Each UAV collects local observation information during the current decision-making cycle. The local observation information includes at least its own status, the status of visible teammates, the status of visible targets, and historical hidden status.
[0071] Step 3: Based on the current location of the UAV and the status of dynamic interference sources, combined with the received signal strength indication, signal-to-interference-plus-noise ratio, packet loss rate or spectrum detection information, estimate the reliability of each candidate sender's information to obtain the reliability index of the candidate sender.
[0072] Step 4: The receiving UAV extracts the current task context based on its own local observations and historical hidden states, and constructs the receiving query vector; at the same time, it constructs candidate feature representations for the local information of the candidate senders.
[0073] Step 5: Jointly score the receiver's task context, the reliability of the candidate sender's information, and the potential help of the candidate link to the receiver's current task to obtain a comprehensive score for each candidate communication link.
[0074] Step 6: After completing the joint scoring, only the highest-scoring drone from each recipient will be retained. The candidate links are selected, a sparse communication mask is generated, and the information of the retained links is aggregated to form the communication message of the receiving UAV.
[0075] Step 7: Input the communication messages and local observations into the UAV decision module, output the current situation assessment results, target service decisions and flight action control commands, and apply them to the environment to complete the status update.
[0076] Step 8: During the offline training phase, construct the overall training objective based on team shared rewards, communication regularization terms, link utility auxiliary terms, and pairing interaction auxiliary terms, and iteratively optimize the network parameters; after training is completed, deploy the optimized parameters to the UAV's onboard processor, edge nodes, or ground control computing equipment.
[0077] Step 9: In the online decision-making stage, the parameters are no longer updated in reverse. Instead, the offline trained model parameters are called and Steps 2 to 7 are repeated until the task is terminated.
[0078] To verify the effectiveness of the embodiments of the present invention, the method of the present invention can be applied to a specific simulation environment or actual test scenario containing multiple UAVs, multiple mission targets, and multiple dynamic interference sources. The mission area can be a post-disaster area, a post-war area, or other collaborative mission area that requires long-term situational estimation. Each UAV obtains information about itself, visible teammates, and visible targets based on its own local observations; targets have a continuous detection status and service priority; dynamic environmental interference sources continuously affect the perception quality and communication quality.
[0079] In this embodiment of the invention, the entire method flow is executed according to Steps 1 to 9 as described above. Specifically, the offline training phase can be completed at a ground station, a central server, an edge computing center, or other high-computing-power devices. During training, the central training device reads task scenario samples, historical interaction data, and reward feedback, and centrally optimizes the task context extraction parameters, joint scoring parameters, link-assisted utility estimation parameters, and collaborative controller parameters. After training convergence, it generates deployable model weight files, link utility estimation parameters, and communication selection parameters.
[0080] During the online decision-making phase, each UAV loads the pre-trained model parameters into its onboard processor, flight control computing unit, edge nodes, or collaborative control terminal before actual flight or during mission execution. During online execution, each UAV no longer performs backpropagation or centralized parameter updates, but instead relies solely on local observations, historical hidden states, and Top-down parameters. The neighbor messages obtained by sparse filtering are used for forward reasoning, thereby independently completing situation assessment, communication selection and action output.
[0081] In each decision cycle, the receiving UAV first constructs a query vector based on its own state, historical hidden states, and local task context. Then, it calculates a joint score by combining candidate sender features, information reliability, and the potential help of candidate links to the current task. Finally, it uses a Top-down algorithm to evaluate the results. The mechanism selects only a limited number of communication objects with the highest scores, and finally inputs the aggregated messages and local observations into the controller to generate situation assessment results, target service actions and flight actions, and drives the environment into the next moment state.
[0082] It should be noted that the parameters mentioned above, such as the number of drones, the number of targets, the number of interference sources, the size of the mission area, the communication radius, and the number of training batches, can all be adjusted according to the application scenario. The aforementioned parameters are merely specific application examples and do not limit the scope of protection of this invention. This invention is also applicable to drone swarms of other sizes, as well as various complex post-disaster environments, post-war environments, and other collaborative situational assessment scenarios with dynamic interference and limited communication conditions.
[0083] To verify the effectiveness of the embodiments of the present invention, the method of the present invention was applied to a specific simulation environment and compared with various comparative schemes. To facilitate the illustration of the technical effects of the present invention, Table 1 presents representative results of the embodiments of the present invention and the comparative schemes.
[0084]
[0085] As shown in Table 1, the method of this invention outperforms the listed comparative schemes in both the team return and key target recall rates. Specifically, the team return reaches 169.3606. The value of 10.6981 is significantly higher than that of Comparative Solution A (125.8021). 22.6977 and 106.7144 for comparison scheme C. 11.8410; the key target recall rate reached 0.9049. 0.0105, significantly higher than 0.7201 in comparison scheme A. 0.1520 and 0.2505 of the comparison scheme C 0.0478. This indicates that the present invention is more advantageous for the continuous service and situational awareness of key targets under dynamic interference conditions. For example... Figure 4 The schematic diagrams shown below illustrate the results of the embodiments of the present invention and the comparative schemes. The method of the present invention shows superior results in terms of return evolution trend and terminal return statistics.
[0086] Under the same test conditions as the communication-type comparison schemes D and E, the method of the present invention achieved the following results: fairness is 0.8004. 0.0048, key target recall rate 0.8744 0.0210, throughput is This indicates that even under conditions of limited communication, the team can still maintain strong task benefits and team coordination capabilities.
[0087] Based on the module comparison results, the communication density of the system after removing the Top-K sparse communication selection mechanism in Comparative Example 1 is reduced from 0.836. The value increased from 0.020 to 1.000. The value of 0.000 indicates that the sparse communication selection mechanism in this invention can directly control the number of communication links, significantly reducing the burden of redundant communication. In contrast, the system in Example 2, which removes the reliability-aware query module, has a communication entropy of 0.153. The value increased from 0.045 to 0.226. A score of 0.000 indicates that this invention does not rely on distance or surface similarity for mechanical communication, but rather improves the targeting of communication object selection through a joint scoring system of "task context + information reliability + link helpability," thus avoiding overly fragmented communication. Figure 5 The diagram illustrating the evolution of the communication structure shows that, as the execution process progresses, the system's communication structure gradually evolves from an early dense switching mechanism to a more task-specific selective communication structure; for example... Figure 6 The schematic diagram showing the comparison results of the modules indicates that the communication density and communication cost are significantly increased in the first comparative embodiment, which shows that the Top-K sparse communication selection mechanism has a direct effect on the control of communication resources.
[0088] Therefore, this invention can explicitly model the reliability of information sources in dynamic interference cooperative scenarios, reducing the risk of the spread of pseudo-related information and low-quality information. It can also jointly score communication objects based on the receiver's current task context, the sender's reliability, and the link help, avoiding blind communication based on distance or surface feature similarity. At the same time, it can control the number of communication links under bandwidth-constrained conditions through the Top-K sparse selection mechanism, effectively reducing communication overhead. Furthermore, it relies on the communication regularization mechanism to constrain communication density and communication distribution entropy, improving communication organization discipline. Thus, it can improve the unit communication resource utilization efficiency in multi-UAV cooperative situation assessment tasks while maintaining high task rewards and key target recall rates.
[0089] It should be understood that, although Figure 1 The steps in the flowchart are shown sequentially as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order in which these steps are executed, and they can be performed in other orders. Figure 1 At least some of the steps in the process may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least some of the sub-steps or stages of other steps.
[0090] In one embodiment, an intelligent communication decision-making system for multi-UAV cooperative situation assessment is provided, comprising: The task modeling module is used to model the multi-UAV cooperative situation assessment communication process in dynamic interference cooperative scenarios as a cooperative distributed partially observable Markov decision process, resulting in a situation assessment task model. Each time step of the situation assessment task model includes global state, local observations, joint actions, state transition probabilities, and team-shared reward values. The global state includes the UAV state set, the target state set, and the dynamic interference source state set, where the UAV state includes information reliability indicators. Local observations are the state that a single UAV can acquire, the state of its visible teammates, and the state of visible targets. Joint actions include the motion actions and communication selection actions of each UAV, with communication selection actions represented by a communication link selection mask. The team-shared reward values include target service rewards, safe operation penalties, and communication cost penalties. The training objective construction module is used to construct a joint training objective based on the situation assessment task model. The joint training objective includes a temporal difference loss, a communication regularization term, a link utility auxiliary term, and a pairing interaction auxiliary term. The communication regularization term is used to constrain the communication density and communication distribution pattern. The strategy solving module is used to solve the joint training objective using a multi-agent reinforcement learning algorithm with centralized training and decentralized execution, to obtain the optimal multi-UAV communication cooperation strategy. The optimal multi-UAV communication cooperation strategy includes calculating the comprehensive value of each candidate communication link for each receiver UAV based on the task context, the reliability of the candidate sender's information, and the potential task helpness of the candidate link, and retaining a preset number of candidate communication links with the highest comprehensive value to generate a communication link selection mask. The task context is a feature vector extracted from the receiver UAV's local observations and historical hidden states. The potential task helpness is an estimate of the link utility obtained by iteratively updating the link's historical activation data and corresponding task benefit changes during training. The strategy deployment module is used to deploy the optimal multi-UAV communication and coordination strategy to the onboard processor of each UAV, enabling each UAV to complete communication and coordination decision-making and situation assessment during the online execution phase.
[0091] In one specific embodiment, such as Figure 7 The diagram illustrates the layered operational architecture of an intelligent communication decision-making system for multi-UAV collaborative situation assessment. The system comprises five layers from top to bottom: the environmental information layer, the training and deployment layer, the collaborative communication decision-making layer, the UAV swarm layer, and the airborne processing and task execution layer. It adopts a centralized offline training and distributed online execution operation mode. The environmental information layer contains two types of elements: a set of ground targets and a set of dynamic interference sources. These elements respectively carry information on the target's location, detection status, priority, and interference status related to interference exposure and link degradation. Both types of environmental data are synchronously input to the environmental status aggregation module of the collaborative communication decision-making layer. The training and deployment layer uses a ground station or a central training server to complete offline model training and parameter optimization. Through the model parameter distribution process, the converged strategy parameters are transmitted to the communication collaborative decision-making core of the collaborative communication decision-making layer. Within the collaborative communication decision-making layer, the environmental status aggregation module summarizes target status, interference, and link quality information and inputs it into the communication collaborative decision-making core. This core performs task context generation, reliability assessment, joint scoring, and Top-K sparse selection, enabling intelligent optimization decisions for communication targets. Decision commands are issued to each UAV in the UAV swarm layer. All UAVs are connected to the airborne processing and mission execution layer. Through the airborne processing function, they complete local observation and data acquisition, sparse message reception, situation assessment and collaborative control calculations, and finally output flight actions, target service actions and situation assessment results, forming a complete data flow closed loop from environmental perception to decision execution.
[0092] Specific limitations regarding the intelligent communication decision-making system for multi-UAV cooperative situation assessment can be found in the limitations of the intelligent communication decision-making method for multi-UAV cooperative situation assessment mentioned above, and will not be repeated here. Each module in the aforementioned intelligent communication decision-making system for multi-UAV cooperative situation assessment can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the corresponding operations of each module.
[0093] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0094] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. An intelligent communication decision-making method for multi-UAV cooperative situation assessment, characterized in that, The method includes: The multi-UAV cooperative situation assessment communication process in a dynamic interference cooperative scenario is modeled as a cooperative distributed partially observable Markov decision process, resulting in a situation assessment task model. Each time step of the situation assessment task model includes a global state, local observations, joint actions, state transition probabilities, and a team-shared reward value. The global state includes a set of UAV states, a set of target states, and a set of dynamic interference source states, where the UAV states include information reliability indicators. The local observations are the state of a single UAV that it can acquire, the states of its visible teammates, and the states of visible targets. The joint actions include the motion actions and communication selection actions of each UAV, where the communication selection actions are represented by a communication link selection mask. The team-shared reward value includes a target service reward, a safe operation penalty, and a communication cost penalty. A joint training objective is constructed based on the aforementioned situation assessment task model; the joint training objective includes a temporal difference loss, a communication regularization term, a link utility auxiliary term, and a pairing interaction auxiliary term; the communication regularization term is used to constrain communication density and communication distribution pattern. A multi-agent reinforcement learning algorithm with centralized training and decentralized execution is used to solve the joint training objective, resulting in an optimal multi-UAV communication cooperation strategy. This optimal strategy includes calculating the comprehensive value of each candidate communication link for each receiver UAV based on the task context, the reliability of candidate sender information, and the potential task helpness of the candidate links. A predetermined number of candidate communication links with the highest comprehensive value are retained to generate a communication link selection mask. The task context is a feature vector extracted from the receiver UAV's local observations and historical hidden states. The potential task helpness is an estimated link utility value obtained through iterative updates based on historical link activation data and corresponding task benefit changes during training. The optimal multi-UAV communication and coordination strategy is deployed to the onboard processor of each UAV, enabling each UAV to complete communication and coordination decisions and situation assessments during the online execution phase.
2. The method according to claim 1, characterized in that, The drone status set includes the drone status of each drone deployed in the environment; the drone status includes drone location, information reliability indicators, and at least one additional platform information, such as boundary status and collision status. The target state set includes the target states of multiple task targets; the target states include target location, target discovery status, and target priority; the task targets refer to the objects that multiple UAV systems need to discover, continuously track, and provide situational awareness services in a dynamic interference coordination scenario. The dynamic interference source state set includes the dynamic interference source states of multiple dynamic interference sources; the dynamic interference source states include the location of the interference source and the interference intensity.
3. The method according to claim 1, characterized in that, The steps for calculating information reliability indicators include: The aggregated interference intensity is calculated based on the current location of the UAV and the status of dynamic interference sources, and an information reliability index is generated based on the aggregated interference intensity. Alternatively, the equivalent interference intensity is estimated based on the link quality index measured by the UAV's onboard equipment, and an information reliability index is generated based on the equivalent interference intensity.
4. The method according to claim 1, characterized in that, The movement action is selected from a discrete action set, which includes forward, backward, left turn, and right turn.
5. The method according to claim 1, characterized in that, The steps for calculating the overall value include: The task context matching degree is calculated based on the task context of the receiving UAV and the local observation features of the candidate sender; Generate information reliability correction items based on the information reliability indicators of candidate senders; Generate link utility correction terms based on the potential task helpness of candidate links; The comprehensive value of each candidate communication link is obtained by weighted summing of the task context matching degree, information reliability correction term, and link utility correction term.
6. The method according to claim 1, characterized in that, The communication regularization terms include a communication density constraint, a communication distribution entropy constraint, an average budget constraint, and a budget distribution entropy constraint. The communication density constraint is used to limit the average number of communication links per UAV, and the communication distribution entropy constraint is used to limit the dispersion of communication links.
7. The method according to claim 1, characterized in that, In the joint training objective, the temporal difference loss and the communication regularization term are directly added together, while the link utility auxiliary term and the pairing interaction auxiliary term are added to the joint training objective after being adjusted by their respective outer weight coefficients.
8. The method according to claim 1, characterized in that, The steps for calculating the team-shared reward value include: The target service reward is calculated based on the service rate of each target per unit time and the priority of the corresponding target. The collision penalty is calculated based on the collision state of the drone, the boundary penalty is calculated based on the boundary crossing state of the drone, and the collision penalty and the boundary penalty are summed to obtain the safe operation penalty. The communication cost penalty is calculated based on the total number of active communication links in the system at the current moment. Subtract the security operation penalty and communication cost penalty from the target service reward to obtain the original reward value. Then, trim the original reward value to a preset range to obtain the team-shared reward value.
9. The method according to claim 1, characterized in that, The optimal multi-UAV communication and coordination strategy is deployed to the onboard processors of each UAV, enabling each UAV to complete communication and coordination decisions and situation assessments during the online execution phase, including: During the online execution phase, each UAV performs forward inference based on the deployed optimal multi-UAV communication and cooperation strategy. It selects a mask to receive information from the corresponding sender based on the current communication link, and performs weighted aggregation according to the weights obtained after normalization of the comprehensive value of each retained communication link to obtain the communication message. Each UAV integrates communication messages with local observations to output the current situation assessment results, target service decisions, and flight maneuver control commands.
10. An intelligent communication and decision-making system for multi-UAV cooperative situation assessment, characterized in that, The system includes: The task modeling module is used to model the multi-UAV cooperative situation assessment communication process under dynamic interference cooperative scenarios as a cooperative distributed partially observable Markov decision process, resulting in a situation assessment task model. Each time step of the situation assessment task model includes global state, local observation, joint actions, state transition probability, and team-shared reward value. The global state includes a set of UAV states, a set of target states, and a set of dynamic interference source states, where the UAV states include information reliability indicators. The local observations are the state of a single UAV that can be acquired, the states of its visible teammates, and the states of visible targets. The joint actions include the motion actions and communication selection actions of each UAV, where the communication selection actions are represented by a communication link selection mask. The team-shared reward value includes target service reward, safe operation penalty, and communication cost penalty. The training objective construction module is used to construct a joint training objective based on the situation assessment task model; the joint training objective includes a temporal difference loss, a communication regularization term, a link utility auxiliary term, and a pairing interaction auxiliary term; the communication regularization term is used to constrain the communication density and communication distribution pattern. The strategy solving module is used to solve the joint training objective using a multi-agent reinforcement learning algorithm with centralized training and decentralized execution, to obtain the optimal multi-UAV communication cooperation strategy. The optimal multi-UAV communication cooperation strategy includes, for each receiving UAV, calculating the comprehensive value of each candidate communication link based on the task context, the reliability of the candidate sender's information, and the potential task helpness of the candidate link, and retaining a preset number of candidate communication links with the highest comprehensive value to generate a communication link selection mask. The task context is a feature vector extracted from the local observations and historical hidden states of the receiving UAV; the potential task helpness is an estimated link utility value obtained through iterative updates based on historical link activation data and corresponding task benefit changes during training. The strategy deployment module is used to deploy the optimal multi-UAV communication and coordination strategy to the onboard processor of each UAV, so that each UAV can complete communication and coordination decision-making and situation assessment during the online execution phase.