Multi-unmanned aerial vehicle cooperative inspection method and system based on risk perception reinforcement learning
Patent Information
- Application Number
- CN202610669438.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-15
- Publication Date
- 2026-08-21
- Estimated Expiration
- 2046-05-15
AI Technical Summary
[0005]本发明的主要目的在于提供了一种基于风险感知强化学习的多无人机协同巡检方法及系统,旨在解决现有技术在大范围、长周期、复杂地形和动态风险条件下,难以利用多无人机有限的飞行时间和通信资源实现监测广域覆盖与重点区域精细复查相统一、且容易导致局部漏检和协同效率低下的技术问题
[0009] This invention achieves effective mining and state quantification of key features in complex soil and water conservation scenarios through environmental discretization and dynamic risk simulation. The introduction of hierarchical action graphs and benefit prediction mechanisms significantly optimizes the allocation logic of computing resources and flight time, avoiding redundant exploration and computing power bottlenecks. The deep coupling of risk-sensitive policy networks and comprehensive reward criteria ensures that multi-UAV clusters can maintain efficient collaboration and safety boundaries even under dynamic interference. It significantly improves the collaborative operation efficiency of multiple UAVs in large-scale complex terrain, realizing a leap from traditional rule-based static inspection to intelligent risk-based dynamic inspection. Through dynamic risk perception and hierarchical decision-making mechanisms, it can accurately locate hidden danger areas and automatically allocate inspection forces, greatly reducing the cost of manual participation and equipment maintenance. At the same time, the risk-sensitive optimization objectives effectively ensure the safety of the inspection process, avoiding missed inspections and resource misallocation, providing an efficient, intelligent, and low-cost inspection method for long-term, wide-area soil and water conservation monitoring.
Smart Images

Figure CN122219615B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent inspection of unmanned aerial vehicles (UAVs), and in particular to a multi-UAV collaborative inspection method and system based on risk perception reinforcement learning. Background Technology
[0002] Soil and water conservation monitoring projects are typically characterized by wide monitoring areas, highly heterogeneous targets, harsh environmental conditions, significant risk evolution, and long durations. For example, some projects have monitoring responsibilities covering thousands of hectares, with monitoring periods exceeding ten years, and involve multiple key targets such as large excavated slopes, large spoil heaps, drainage structures, and erosion-sensitive areas. Faced with such large-scale, long-term, and complex terrain conditions, there is an urgent need for an effective solution capable of autonomous inspection, risk identification, dynamic revisiting, and long-term time-series monitoring.
[0003] Currently, for such engineering monitoring scenarios, the industry mainly relies on traditional manual inspections, continuous fixed-point observations by installing conventional automatic monitoring equipment in localized areas, or using conventional single drones to patrol according to pre-set rules and routes. In terms of drone applications, existing technologies mostly employ simple area division or static division of labor for path planning and operations.
[0004] However, these existing practices face many limitations in practical applications: First, the regular routes are usually set offline in advance, making it difficult to dynamically change the key inspection areas according to rainfall, construction disturbances, or increased local risks; Second, single UAVs have limited flight time in wide-area scenarios, making it impossible to balance full coverage with high-frequency re-inspection of key areas; Third, if multiple UAVs only adopt simple area division or static division of labor, it is easy to lead to repeated inspections, partial omissions, communication interruptions, and low collaborative efficiency; Fourth, existing multi-robot path planning methods focus more on static target discovery or exploration of unknown areas, while soil and water conservation inspection is a long-term cognitive task oriented towards time-varying risk fields. Its core is not to find a single target, but to continuously estimate the risk status of different areas, update risk cognition, and prioritize the allocation of limited resources to areas that are more worthy of inspection. Summary of the Invention
[0005] The main objective of this invention is to provide a multi-UAV collaborative inspection method and system based on risk perception reinforcement learning. This aims to solve the technical problems of existing technologies, which struggle to achieve both wide-area monitoring coverage and detailed review of key areas using the limited flight time and communication resources of multiple UAVs under conditions of large scale, long cycle, complex terrain, and dynamic risks. These problems also tend to lead to local missed detections and low collaborative efficiency.
[0006] To achieve the above objectives, this invention provides a multi-UAV collaborative inspection method based on risk perception reinforcement learning, the method comprising the following steps: The monitoring area is discretized into a risk voxel grid. Based on the prior risk characteristics and time-varying event field characteristics of the risk voxel grid, dynamic risk prediction is performed to obtain the predicted risk intensity and risk uncertainty values of each voxel unit in the risk voxel grid. Based on the local observation information of each target UAV in the UAV cluster, a coarse patrol action set and a fine patrol action set are constructed. The coarse patrol action set and the fine patrol action set are merged to obtain a hierarchical candidate action map. The UAV cluster contains multiple UAVs. The action revenue of each candidate action in the hierarchical candidate action graph is predicted based on the action revenue prediction model, and an action feature matrix is constructed based on the action revenue, the predicted risk intensity value, and the predicted risk uncertainty value. The action feature matrix is input into a risk-sensitive multi-agent policy network to predict the action probability distribution and output the target collaborative action decision. The risk-sensitive multi-agent policy network is trained based on the inspection reward function. Each target drone in the drone swarm is configured to perform collaborative inspection actions based on the target collaborative action decision.
[0007] Furthermore, to achieve the above objectives, this invention also proposes a multi-UAV collaborative inspection system based on risk perception reinforcement learning, which applies the risk perception reinforcement learning-based multi-UAV collaborative inspection method described above. The risk perception reinforcement learning-based multi-UAV collaborative inspection system includes: The risk prediction module is used to discretize the monitoring area into a risk voxel grid, perform dynamic risk prediction based on the prior risk characteristics and time-varying event field characteristics of the risk voxel grid, and obtain the predicted risk intensity and risk uncertainty values of each voxel unit in the risk voxel grid. The action construction module is used to construct a coarse patrol action set and a fine patrol action set based on the local observation information of each target UAV in the UAV cluster, and merge the coarse patrol action set and the fine patrol action set to obtain a hierarchical candidate action map. The UAV cluster contains multiple UAVs. The feature construction module is used to predict the action revenue of each candidate action in the hierarchical candidate action graph based on the action revenue prediction model, and to construct an action feature matrix based on the action revenue, the risk intensity prediction value, and the risk uncertainty prediction value. The action decision module is used to input the action feature matrix into the risk-sensitive multi-agent policy network to predict the action probability distribution and output the target collaborative action decision. The risk-sensitive multi-agent policy network is trained based on the inspection reward function. Each target drone in the drone cluster is configured to perform collaborative inspection actions based on the target collaborative action decision.
[0008] Furthermore, to achieve the above objectives, this application also proposes a multi-UAV collaborative inspection device based on risk perception reinforcement learning. The device includes: a memory, a processor, and a multi-UAV collaborative inspection program stored in the memory. The processor is used to run the multi-UAV collaborative inspection program, and the computer program is configured to implement the steps of the multi-UAV collaborative inspection method based on risk perception reinforcement learning as described above.
[0009] This invention achieves effective mining and state quantification of key features in complex soil and water conservation scenarios through environmental discretization and dynamic risk simulation. The introduction of hierarchical action graphs and benefit prediction mechanisms significantly optimizes the allocation logic of computing resources and flight time, avoiding redundant exploration and computing power bottlenecks. The deep coupling of risk-sensitive policy networks and comprehensive reward criteria ensures that multi-UAV clusters can maintain efficient collaboration and safety boundaries even under dynamic interference. It significantly improves the collaborative operation efficiency of multiple UAVs in large-scale complex terrain, realizing a leap from traditional rule-based static inspection to intelligent risk-based dynamic inspection. Through dynamic risk perception and hierarchical decision-making mechanisms, it can accurately locate hidden danger areas and automatically allocate inspection forces, greatly reducing the cost of manual participation and equipment maintenance. At the same time, the risk-sensitive optimization objectives effectively ensure the safety of the inspection process, avoiding missed inspections and resource misallocation, providing an efficient, intelligent, and low-cost inspection method for long-term, wide-area soil and water conservation monitoring. Attached Figure Description
[0010] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0011] Figure 1 This is a schematic diagram of the structure of a multi-UAV collaborative inspection device based on risk perception reinforcement learning, which is part of the hardware operating environment of the embodiment of the present invention. Figure 2 This is a flowchart illustrating the first embodiment of the multi-UAV collaborative inspection method based on risk perception reinforcement learning of the present invention. Figure 3 This is a flowchart illustrating the second embodiment of the multi-UAV collaborative inspection method based on risk perception reinforcement learning of the present invention. Figure 4 This is a flowchart illustrating the third embodiment of the multi-UAV collaborative inspection method based on risk perception reinforcement learning of the present invention. Figure 5 This is a flowchart illustrating the fourth embodiment of the multi-UAV collaborative inspection method based on risk perception reinforcement learning of the present invention. Figure 6This is a flowchart illustrating the fifth embodiment of the multi-UAV collaborative inspection method based on risk perception reinforcement learning of the present invention. Figure 7 This is a structural block diagram of the first embodiment of the multi-UAV collaborative inspection system based on risk perception reinforcement learning of the present invention.
[0012] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0013] It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the invention.
[0014] Reference Figure 1 , Figure 1 This is a schematic diagram of the structure of a multi-UAV collaborative inspection device based on risk perception reinforcement learning, which is part of the hardware operating environment of the embodiment of the present invention.
[0015] like Figure 1 As shown, the multi-UAV collaborative inspection device based on risk perception reinforcement learning may include: a processor 1001, such as a central processing unit (CPU), a communication bus 1002, a user interface 1003, a network interface 1004, and a memory 1005. The communication bus 1002 is used to enable communication between these components. The user interface 1003 may include a display screen and an input unit such as a keyboard; the user interface 1003 may also include standard wired and wireless interfaces. The network interface 1004 may optionally include standard wired and wireless interfaces (such as Wireless-Fidelity (Wi-Fi) interfaces). The memory 1005 may be high-speed random access memory (RAM) or stable non-volatile memory (NVM), such as a disk drive. The memory 1005 may also optionally be a storage system independent of the aforementioned processor 1001.
[0016] Those skilled in the art will understand that Figure 1 The structure shown does not constitute a limitation on multi-UAV collaborative inspection equipment based on risk perception reinforcement learning. It may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0017] like Figure 1 As shown, the memory 1005, which is a computer-readable storage medium, may include an operating system, a network communication module, a user interface module, and a multi-UAV collaborative inspection program.
[0018] exist Figure 1 In the multi-UAV collaborative inspection device based on risk perception reinforcement learning shown, the network interface 1004 is mainly used for data communication with the network server; the user interface 1003 is mainly used for data interaction with the user; the processor 1001 and memory 1005 in the multi-UAV collaborative inspection device based on risk perception reinforcement learning of the present invention can be set in the multi-UAV collaborative inspection device based on risk perception reinforcement learning. The multi-UAV collaborative inspection device based on risk perception reinforcement learning calls the multi-UAV collaborative inspection program stored in the memory 1005 through the processor 1001 and executes the multi-UAV collaborative inspection method based on risk perception reinforcement learning provided in the embodiment of the present invention.
[0019] This invention provides a multi-UAV collaborative inspection method based on risk perception reinforcement learning, referring to... Figure 2 , Figure 2 This is a flowchart illustrating the first embodiment of the multi-UAV collaborative inspection method based on risk perception reinforcement learning of the present invention.
[0020] In this embodiment, the multi-UAV collaborative inspection method based on risk perception reinforcement learning includes the following steps: Step S10: Discretize the monitoring area into a risk voxel grid, perform dynamic risk prediction based on the prior risk characteristics and time-varying event field characteristics of the risk voxel grid, and obtain the predicted risk intensity and risk uncertainty of each voxel unit in the risk voxel grid.
[0021] It should be noted that this embodiment is applied to multi-UAV inspection scenarios and is suitable for autonomous inspection, risk identification, dynamic revisit and long-term time series monitoring of large excavated slopes, large spoil heaps, drainage facilities, gully erosion areas, vegetation restoration areas and other units sensitive to soil and water loss.
[0022] To ensure logical consistency and symbol uniformity throughout the text, a unified variable system for this invention is first defined. All subsequent formulas are expanded within this unified symbol system.
[0023] The monitoring area is recorded as: The three-dimensional spatial region to be inspected is represented by the spatial coordinates of any point thereon: in, and Represents planar coordinates, Represents elevation coordinates.
[0024] The number of drones in the system is recorded as follows: Among them, the The index of the drone satisfies: Discrete decision time is denoted as: in, This represents the maximum number of decision steps in a single task or a single training round.
[0025] No. A drone at all times The state vector is defined as: in: Indicates the current location of the drone; Indicates the heading angle; This represents the normalized remaining battery power. It indicates the working status of airborne sensors and can be used to characterize whether the sensors are functioning properly, the status of multi-load switching, or the imaging quality level. It indicates the communication status, which can indicate whether it is connected to a neighboring machine, the link quality, or the current available backhaul bandwidth.
[0026] Discretize the monitoring area into A risk voxel unit, the voxel set is defined as: Each voxel Corresponding to a fixed spatial center point And associate it with a time-varying attribute vector: in: Voxel representation At any moment The risk intensity estimate is used to describe the relative severity of anomalies, instability, erosion, seepage or other soil and water loss problems in the area; This represents the uncertainty of risk, which characterizes the reliability of the system's current risk estimate for the voxel. The higher the uncertainty, the more the voxel needs further observation. Indicates the object category to which the voxel belongs, such as slope, spoil heap, drainage facilities, gully erosion area or vegetation restoration area; This indicates the time interval since the last valid inspection of the voxel, reflecting the freshness of the information in that area; It indicates accessibility or flight difficulty coefficient, used to represent terrain complexity, obstacle density, flight risk, or inspection cost.
[0027] The object category collection is defined as: No. A drone at all times The set of candidate actions is denoted as: in, Indicates the total number of candidate actions. Indicates the first One candidate action.
[0028] The communication radius, safe collision avoidance radius, and effective observation radius are defined as follows: in: This indicates the maximum distance at which drones can stably exchange information. This indicates the minimum safe distance that a drone must maintain during flight; This indicates the effective observation radius of the airborne sensor, and its value is related to the sensor type, flight altitude, and observation resolution.
[0029] The overall goal of this invention is not simply to maximize the spatial coverage area, but to maximize the effective inspection benefits in time-varying risk fields, that is, to prioritize high-quality observations in high-risk, high-uncertainty, long-uninspected, and event-triggered areas under limited resource conditions.
[0030] It should be understood that the executing entity of this embodiment can be a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, or mobile phone, or a terminal electronic device capable of realizing the above functions. The following description uses a multi-UAV collaborative inspection device based on risk perception reinforcement learning (hereinafter referred to as the inspection device) as an example to illustrate this embodiment and the following embodiments.
[0031] It should be noted that prior risk characteristics can be static risk indicators pre-calculated based on topographic slope, historical damage records, engineering object type, and geological structural parameters. Time-varying event field characteristics refer to external environmental disturbance data that dynamically changes over time and excites soil erosion or structural instability.
[0032] The predicted risk intensity value refers to the system's quantitative estimate of the probability of an anomaly or disaster occurring in a single voxel unit at the next decision time. The predicted risk uncertainty value is a measure of the reliability of the system's current risk intensity estimate.
[0033] Understandably, the goal of this step is to transform the real-world engineering scenario into a state-space model that can be processed by the algorithm. Because real-world scenarios are continuous, complex, and contain multiple monitoring objects, directly planning and learning in a continuous space would result in excessively high state dimensionality and decision-making difficulty. Therefore, this invention first discretizes the environment in three dimensions and establishes a dynamically updated digital twin of the risk over time.
[0034] In the specific implementation, the monitoring area will be Discretize into a voxel grid according to a given spatial resolution: In the above formula, It is a continuous three-dimensional spatial region. This is the set of voxels obtained after discretization. The total prime number. The essence of discretization is to project the risk state in continuous space onto a finite number of computable cells, so that subsequent risk estimation, information updates, and action benefit calculations can be performed on a unified grid.
[0035] Each voxel Corresponding center coordinates and attribute vector Spatial resolution is defined as: in: Indicates voxels in Dimensions of direction; Indicates voxels in Dimensions of direction; Indicates voxels in Dimensions of direction.
[0036] The three parameters mentioned above collectively determine the spatial discretization accuracy. If the values are too large, the scene representation is too coarse, making it impossible to distinguish local detail anomalies; if the values are too small, the number of voxels increases dramatically, leading to a significant increase in computational and storage requirements. Therefore, in engineering implementation, finer voxels are typically used for key areas, while coarser voxels are used for low-risk background areas, resulting in a non-uniform resolution voxel representation.
[0037] Step S20: Construct a coarse patrol action set and a fine patrol action set based on the local observation information of each target UAV in the UAV cluster, and merge the coarse patrol action set and the fine patrol action set to obtain a hierarchical candidate action map.
[0038] It should be noted that local observation information refers to the surrounding terrain images, meteorological data, communication status of neighboring drones, and its own flight parameters that a single drone acquires in real time through its onboard sensors.
[0039] Coarse patrol maneuver sets can be long-range, wide-coverage macro-level flight command sequences generated to meet the needs of wide-area census surveys. Fine patrol maneuver sets refer to short-range, multi-view, high-resolution fine-review command sequences generated for high-risk hotspots or areas with initial anomalies.
[0040] A hierarchical candidate action graph is a decision space model that maps macroscopic scanning nodes and microscopic review nodes to the same topological structure and adds trajectory connectivity and flight cost information.
[0041] In practical implementation, each UAV acquires real-time data from its onboard camera, LiDAR point cloud, its own pose, and summary data broadcast by neighboring UAVs during flight. The inspection equipment uses the UAV's current location as the origin and generates macroscopic scanning waypoints within a large search radius at fixed angle intervals or through random sampling, forming a wide-area survey instruction set. Simultaneously, the inspection equipment identifies abnormal hotspots or high-risk object boundaries in local observation data, generating detailed re-examination waypoints within a smaller neighborhood using multi-view surround, low-altitude hovering, or high-frequency revisit methods, forming a key in-depth analysis instruction set. After trajectory feasibility verification and obstacle avoidance pre-screening, both instruction sets are uniformly mapped to the same topology, where nodes represent target waypoints, edges represent connected trajectories, and flight cost parameters such as distance, climb rate, and wind resistance are added. For example, when a UAV cruises over a drainage ditch and identifies local erosion marks, the inspection equipment generates regular cruise waypoints within a 500-meter range and multi-angle close-up shooting waypoints within a 50-meter radius of the erosion point; both together constitute the current decision space.
[0042] Understandably, this step breaks through the limitations of single-scale path planning, taking into account both overall basic coverage and in-depth exploration of key local areas, effectively compressing the dimension of continuous control space, avoiding ineffective exploration and waste of computing power, and significantly improving the task adaptability and decision generation efficiency of multi-machine systems in complex terrain.
[0043] Step S30: Predict the action payoff of each candidate action in the hierarchical candidate action graph based on the action payoff prediction model, and construct an action feature matrix based on the action payoff, the predicted risk intensity value, and the predicted risk uncertainty value.
[0044] It should be noted that a drone swarm comprises multiple drones. The action revenue prediction model refers to a computational module that uses historical inspection data and current environmental conditions to quantitatively evaluate the potential value of candidate actions through regression algorithms or probabilistic models.
[0045] Action benefits refer to the comprehensive quantitative value of the expected improvement in risk perception, probability of anomaly detection, or degree of mission objective achievement after performing a specific flight action.
[0046] The action feature matrix refers to a structured data tensor formed by arranging multi-dimensional indicators such as the predicted benefits of each candidate action, the risk intensity of related voxels, uncertainty, energy consumption cost, communication gain and conflict factor in a unified dimension.
[0047] In some embodiments, the inspection equipment retrieves historical flight records of executed actions and their corresponding actual inspection benefits to construct a feature-benefit mapping sample library. For the currently generated hierarchical candidate action map, the average risk level, non-inspection duration, communication link quality, and expected energy consumption parameters of each candidate waypoint's coverage area are extracted and combined into a multi-dimensional feature vector. Using a Gaussian process regression model or a Bayesian neural network, the feature vector is input into a trained benefit estimator, outputting the expected information gain estimate and confidence interval of the action. Subsequently, the system performs tensor concatenation of the predicted benefit value with the corresponding area's risk intensity estimate and uncertainty estimate, and supplements novelty indicators, event revisit priority, multi-aircraft conflict factors, and energy consumption cost parameters, arranging them in a unified dimension to form a structured data matrix. For example, for a candidate action flying to the top of a high slope, the model combines the benefit performance of similar historical actions with the current high-risk state of the area, outputting a high predicted information gain value, and encapsulating it along with risk intensity, uncertainty, and other parameters into a single row of data in a matrix.
[0048] Understandably, this step explicitly links the long-term value of discrete actions with the state of environmental risk. By fusing high-dimensional features, it eliminates the one-sidedness of single-indicator decision-making and provides the subsequent reinforcement learning network with input representations rich in physical meaning and task orientation, which greatly accelerates the convergence process of the policy model.
[0049] Step S40: Input the action feature matrix into the risk-sensitive multi-agent policy network to predict the action probability distribution and output the target collaborative action decision.
[0050] It should be noted that the risk-sensitive multi-agent policy network is trained based on the inspection reward function, and each target drone in the drone swarm is configured to perform collaborative inspection actions based on the target collaborative action decision.
[0051] It should be noted that risk-sensitive multi-agent policy network refers to a deep neural network model that adopts a centralized training and decentralized execution architecture and introduces conditional risk value or tail omission penalty terms into the objective function.
[0052] Action probability distribution prediction refers to the strategy network outputting a sequence of relative probabilities of each candidate action being selected as the optimal execution instruction based on the input feature matrix.
[0053] Target cooperative action decision-making refers to the allocation of specific flight waypoints, speed commands, and sensor control parameters to each UAV in the cluster after probabilistic sampling or deterministic selection.
[0054] The inspection reward function is a multi-objective reward function that comprehensively considers risk coverage benefits, information novelty gains, event-driven priority, communication connectivity maintenance, energy consumption costs, and multi-machine conflict penalties.
[0055] In some embodiments, the inspection equipment inputs the constructed action feature matrix into a pre-trained multi-agent deep neural network. This network employs a centralized training and distributed execution architecture. During training, it utilizes a global risk map and joint action information to optimize the critic network; during execution, it relies solely on the local feature matrix of each individual drone to drive the policy network. Internally, the network extracts feature interaction relationships through a multilayer perceptron or graph attention mechanism, outputting a sequence of relative probabilities for each candidate action to be selected. Based on the probability distribution, the system uses deterministic maximum selection or random sampling with exploration noise to determine the optimal flight waypoint, speed command, and payload control parameters for the current moment. During training, the system updates network weights according to a comprehensive reward criterion, which simultaneously considers high-risk area coverage gain, novelty compensation for long-uninspected areas, event revisit incentives after heavy rain or construction, communication connectivity maintenance rewards, flight energy consumption deduction, multi-drone aggregation penalties, and severe penalties for intrusion into no-fly zones. After receiving decision commands, each drone autonomously adjusts its flight attitude and sensor pointing, and exchanges local risk maps with neighboring drones when communication distance conditions are met, achieving cluster collaborative operation. For example, under continuous rainfall conditions, the policy network will significantly increase the probability of selecting actions to fly to gully erosion-sensitive areas, while automatically reducing the probability of actions that may cause cross-flight routes between two aircraft, ensuring that the cluster prioritizes response to high-risk areas under the premise of safety.
[0056] Understandably, this step enables autonomous collaboration and dynamic game among multiple UAVs under constraints of limited communication and energy consumption. Through a risk-sensitive mechanism, it effectively suppresses the tail risk of missed detections in key areas, ensuring the system's decision-making robustness and engineering safety in long-term, high-interference environments.
[0057] This embodiment achieves effective mining and state quantification of key features in complex soil and water conservation scenarios through environmental discretization and dynamic risk simulation. The introduction of hierarchical action graphs and benefit prediction mechanisms significantly optimizes the allocation logic of computing resources and flight time, avoiding redundant exploration and computing power bottlenecks. The deep coupling of risk-sensitive policy networks and comprehensive reward criteria ensures that multi-UAV clusters can maintain efficient collaboration and safety boundaries even under dynamic interference. It significantly improves the collaborative operation efficiency of multiple UAVs in large-scale complex terrain, realizing a leap from traditional rule-based static inspection to intelligent risk-based dynamic inspection. Through dynamic risk perception and hierarchical decision-making mechanisms, it can accurately locate hidden danger areas and automatically allocate inspection forces, greatly reducing the cost of manual participation and equipment maintenance. At the same time, the risk-sensitive optimization objectives effectively ensure the safety of the inspection process, avoiding missed inspections and resource misallocation, providing an efficient, intelligent, and low-cost inspection method for long-term, wide-area soil and water conservation monitoring.
[0058] refer to Figure 3 , Figure 3 This is a flowchart illustrating the second embodiment of the multi-UAV collaborative inspection method based on risk perception reinforcement learning of the present invention.
[0059] Based on the first embodiment described above, in this embodiment, step S10 further includes: Step S101: Assign prior risk weights to objects based on their regulatory importance within the monitoring area.
[0060] It should be noted that, in this embodiment, the prior risk characteristics include the object's prior risk weight, slope factor, elevation difference factor, and historical disturbance factor; the time-varying event field characteristics include the event incentive intensity and time-dependent growth characteristics.
[0061] Understandably, given the different risk formation mechanisms and regulatory importance of different types of objects, relying solely on subsequent real-time observations would lead to a lack of direction in initial task allocation. Therefore, this invention introduces object-level prior risk weights to assign different initial levels of attention to different monitoring objects.
[0062] Define the prior risk weight of the object as follows: in, Indicates category The corresponding prior importance. For example, large excavated slopes and large spoil heaps typically have more severe instability consequences, therefore their weights are generally higher than those of ordinary vegetation restoration areas. A typical relationship can be expressed as: Furthermore, voxels The initial risk intensity is defined as: in: Voxel representation The intensity of risk at the initial moment; Indicates the object category to which the voxel belongs; This represents the prior weight corresponding to the category; Represents the prior feature vector related to voxels; This represents a mapping function that calculates the initial risk level based on prior features.
[0063] To enhance interpretability, the following definition is provided: in: The slope factor is used to represent the steepness of the surface slope. The steeper the slope, the higher the risk of slope or surface erosion. The elevation difference factor is used to represent the characteristics of local elevation fluctuations and is often related to potential energy and local instability risk. The distance factor to the drainage channel is used to reflect the spatial relationship between the voxel and the ditch, interception and drainage facilities or confluence path; This is the historical disturbance factor, used to represent the intensity of historical construction disturbances, historical defects, or historical anomaly records; The weighting parameter is used to adjust the contribution of different prior factors to the initial risk.
[0064] The logical significance of this formula is that the initial risk is not set out of thin air, but is jointly determined by the object category and terrain engineering characteristics. The result will serve as the starting point for subsequent dynamic updates to the risk.
[0065] Step S102: Extract the slope factor, elevation difference factor and historical disturbance factor of the risk voxel grid, and calculate the initial risk intensity of each voxel unit in combination with the prior risk weight of the object.
[0066] Understandably, further, voxels The initial risk intensity is defined as: in: Voxel representation The intensity of risk at the initial moment; Indicates the object category to which the voxel belongs; This represents the prior weight corresponding to the category; Represents the prior feature vector related to voxels; This represents a mapping function that calculates the initial risk level based on prior features.
[0067] To enhance interpretability, the following definition is provided: in: The slope factor is used to represent the steepness of the surface slope. The steeper the slope, the higher the risk of slope or surface erosion. The elevation difference factor is used to represent the characteristics of local elevation fluctuations and is often related to potential energy and local instability risk. The distance factor to the drainage channel is used to reflect the spatial relationship between the voxel and the ditch, interception and drainage facilities or confluence path; This is the historical disturbance factor, used to represent the intensity of historical construction disturbances, historical defects, or historical anomaly records; The weighting parameter is used to adjust the contribution of different prior factors to the initial risk.
[0068] The logical significance of this formula is that the initial risk is not set out of thin air, but is jointly determined by the object category and terrain engineering characteristics. The result will serve as the starting point for subsequent dynamic updates to the risk.
[0069] Step S103: Obtain the event excitation intensity of the external time-varying event field, and calculate the time-dependent growth characteristics of each voxel unit in the risk voxel grid based on the time interval driving function.
[0070] It should be noted that soil and water conservation risks are not static and unchanging, but are affected by external events such as rainfall, construction, freeze-thaw cycles, wind erosion, and localized landslides. Therefore, using only static priors is insufficient to reflect the actual risk evolution process. For this reason, an external event field is introduced: Used to indicate at time ,Location The intensity of external disturbances at the location.
[0071] voxels Its event stimulus term is defined as: in: Indicates the event's effect on voxels The intensity of the incentive from risk; This indicates the value of the event field at the center point of the voxel; This is an event mapping function used to convert raw event quantities into risk gain quantities.
[0072] For example, if the event field is rainfall intensity, then Hourly rainfall intensity, cumulative rainfall, or rainfall level can be mapped to risk increments; if the event site is a construction disturbance, then... The extent of local risk escalation can be mapped based on the construction location and construction level.
[0073] Step S104: Construct a risk prediction model based on the initial risk intensity, the event incentive intensity, and the time-dependent growth characteristics.
[0074] Based on this, a risk prediction model is defined: in: Indicates at time Risks predicted before observation updates; Indicates time Updated risk intensity; This is the risk memory coefficient, used to retain historical risk status. This is the event impact coefficient, used to adjust the intensity of the current event's impact on risk growth; This is a timeliness growth coefficient, used to reflect that "the longer the period without inspection, the more attention should be paid to its potential risks"; The time interval driving function is used as the calculation result, and the product of the calculation result of the time interval driving function and the failure growth coefficient is the time-dependent growth characteristic.
[0075] This formula reflects three sources of risk evolution: inheritance of historical risks, growth triggered by events, and accumulation of uncertainty due to long-term lack of inspection. Among these, The larger the value, the more the system values the continuity of historical risks; The larger the value, the more sensitive the system is to unexpected events; The larger the value, the more the system emphasizes the need for supplementary inspections of areas that have not been inspected for a long time.
[0076] To characterize the feature that "the priority of areas that have not been inspected for a long time gradually increases," the following definition is made: in: This indicates the time interval since the last valid inspection of the voxel; For time-dependent growth parameters; This is the normalized time-driven value.
[0077] This function follows Monotonically increasing, and at a relatively large level The timeframe should gradually saturate to avoid unlimited growth that could lead to time-sensitive items dominating the overall risk.
[0078] Step S105: The risk voxel grid is observed by the target UAV, and the output of the risk prediction model is corrected a posteriori based on the Bayesian update formula to obtain the risk intensity prediction value and risk uncertainty prediction value of each voxel unit.
[0079] In practical implementation, the inspection equipment defines the observation model: [The text abruptly ends here, likely due to an incomplete sentence or a formatting error.] A drone at all times voxels The observation is defined as: in: Indicates the drone's voxel The observation results can be image discrimination scores, crack probability, gully intensity estimation, vegetation cover change, or fused observation features; This represents an ideal observation model, which is determined by the UAV's state, observation angle, distance, payload type, and the actual state of the voxels. To represent observation noise, satisfying: in, To observe noise variance, which reflects factors such as image quality, illumination, occlusion, and sensor error.
[0080] To determine whether a voxel can be effectively observed, a visibility indicator function is defined: in: This represents the Euclidean distance from the drone to the center of the voxel; The effective observation radius; "unobstructed" means that there is no terrain, building or other obstacle obstructing the line of sight.
[0081] This function filters out voxels that, although geometrically similar, cannot actually form an effective image; it only filters out voxels when... Only then is it considered that the observation of this voxel can be used for risk updates.
[0082] This invention does not model the inspection task as a simple target counting problem, but instead treats each voxel as a random unit with a "high-risk / non-high-risk" probability. The posterior probability of voxel risk is defined as: in: Voxel representation At any moment In an abnormal or high-risk state; This indicates that the voxel is in a normal or low-risk state; Represents from the initial time to time 1 The complete set of observation information.
[0083] At any moment After the new observations arrive, the Bayesian update formula is used: in: This indicates the prior high-risk probability before the observation update; This represents the likelihood of the current observation under a high-risk voxel condition; This represents the likelihood of the current observation obtained under the condition that the voxel is not at high risk.
[0084] This formula embodies a core logic: if the current observation is more consistent with a high-risk pattern, the posterior probability increases; if the current observation is more consistent with a normal pattern, the posterior probability decreases. In this way, the system can continuously revise its judgment of the risk status of each region based on observations.
[0085] To align with the subsequent inspection reward function, the risk probability is further mapped to risk intensity: in, This is a mapping function from probability to risk intensity. Common possible forms include: Linear form: Logarithmic odds form: If a linear form is used, the risk intensity remains at [a certain level]. Using intervals facilitates normalized calculations; if logarithmic odds are used, it is more conducive to amplifying the differences between high-risk and low-risk areas.
[0086] In addition to the risk intensity itself, the system also needs to consider whether the current estimate is reliable; therefore, the risk uncertainty is defined as: in, exist When it reaches its maximum, it indicates that the system is at its most uncertain; when near or When the uncertainty decreases, it indicates that the system's judgment is relatively clear. This uncertainty will be directly used for subsequent novelty payoff calculations and action value estimations.
[0087] This embodiment transforms the traditional risk assessment model, which relies on a single threshold or fixed rules, into a dynamic cognitive process involving the fusion of multi-source heterogeneous data and probabilistic recursive evolution. This effectively overcomes technical bottlenecks such as strong environmental interference, rapid risk evolution, and sparse observational data in complex soil and water conservation scenarios. Through prior weight guidance and extraction of historical terrain features, it achieves deep alignment between engineering supervision logic and geospatial attributes. Utilizing event incentives and time-varying mechanisms, it possesses the ability to respond quickly to emergencies and automatically fill long-term blind spots. Combining a recursive prediction model with probabilistic posterior correction, it significantly improves noise robustness and cognitive accuracy while ensuring the continuity of state estimation. Overall, it significantly enhances the environmental perception depth and state prediction accuracy of multi-UAV swarms in wide-area time-varying risk fields, providing a high-fidelity, quantifiable, and physically meaningful state-space benchmark for subsequent hierarchical action planning and reinforcement learning strategy optimization.
[0088] refer to Figure 4 , Figure 4 This is a flowchart illustrating the third embodiment of the multi-UAV collaborative inspection method based on risk perception reinforcement learning of the present invention.
[0089] Based on the above embodiments, in this embodiment, step S20 further includes: Step S201: Extract the location information and abnormal hotspot information of each target UAV based on the local observation information of each target UAV in the UAV cluster.
[0090] It should be noted that location information refers to the real-time coordinates, heading angle, and flight altitude data of the target UAV in three-dimensional space. Abnormal hotspot information can be the spatial distribution characteristics of areas of surface deformation, erosion marks, or structural anomalies identified in local observation data that exceed a preset risk threshold and require focused review.
[0091] Understandably, relying solely on continuous control output would make training multi-UAV reinforcement learning in large-scale scenarios difficult. Therefore, this invention first discretizes the actionable actions into candidate action graphs, from which the policy network then selects the appropriate action. This embodiment further considers that soil and water conservation inspections simultaneously require both wide-area surveys and detailed local surveys, thus designing a hierarchical candidate action graph.
[0092] Step S202: Perform motion sampling within the first preset sampling radius corresponding to the location information to generate a coarse patrol motion set.
[0093] In practical implementation, the inspection equipment operates within a large search radius. A coarse-grained set of candidate actions is generated within the first preset sampling radius. in: Indicates the first A drone at all times A set of rough patrol actions; Indicates the sampling radius of the coarse survey layer; This indicates the number of actions performed in the coarse patrol layer.
[0094] The function of the coarse survey layer is to quickly cover a large area and prioritize the discovery of potential new risk areas. It is suitable for macro-scanning and area redistribution in long-cycle tasks.
[0095] Step S203: Perform action sampling within the second preset sampling radius corresponding to the abnormal hotspot information to generate a fine patrol action set.
[0096] In practice, the inspection equipment generates a set of detailed inspection actions around key objects or in the vicinity of abnormal hotspots: in: This represents the set of actions for the precision patrol layer; The sampling radius of the fine-tuning layer (i.e., the second preset sampling radius) usually satisfies ; This indicates the number of actions performed during the fine patrol layer.
[0097] The purpose of the fine-tuning layer is to conduct multi-angle, high-resolution, and short-interval re-inspections around high-risk hotspots. It is applicable to local key phenomena such as cracks, slope gullies, abnormal slope toe accumulation, and abnormal seepage.
[0098] Step S204: Construct a trajectory cost function that includes trajectory length, elevation difference cost, wind field disturbance cost, and obstacle avoidance cost.
[0099] In practical implementation, to evaluate the cost of action execution, for either side... Define trajectory cost: in: Indicates the length of the trajectory; This indicates the elevation difference cost corresponding to climbing or descending; This indicates the additional flight cost caused by wind field disturbances; This indicates the cost of obstacle avoidance or traversing complex terrain; These are the cost weights for each of the above items.
[0100] This cost is not directly equivalent to the energy penalty in the final reward, but is used for action graph generation and candidate action pre-screening, so that obviously undesirable or unreachable actions are eliminated before entering the policy network.
[0101] Step S205: Perform safety screening on the coarse patrol action set and the fine patrol action set based on the trajectory cost function to eliminate unreachable actions.
[0102] In practice, the inspection equipment substitutes each candidate waypoint in the coarse and fine inspection action sets into the constructed trajectory cost function for calculation one by one to obtain the comprehensive flight cost value corresponding to each action. The calculation results are compared with the preset safety energy consumption threshold and terrain access limit. At the same time, combined with hard safety distance constraints, it checks whether the trajectory has spatial interference with known obstacles or no-fly zones, automatically filters out instructions that exceed cost limits or have collision risks, and retains a subset of feasible actions that comply with flight safety regulations.
[0103] Step S206: Aggregate the nodes of the coarse patrol action set and the fine patrol action set after safety screening to obtain a hierarchical candidate action graph.
[0104] In practical implementation, the inspection equipment combines coarse and fine inspection actions into a unified motion diagram: The node set is as follows: edge set This represents the reachable trajectory from the current state of the drone to the endpoint corresponding to each candidate action.
[0105] This embodiment fundamentally changes the problems of dimensionality curse and convergence difficulties caused by direct optimization of the traditional continuous control space. It achieves an organic unity of wide-area general survey and focused detailed survey through a coarse-fine dual-radius sampling mechanism, enabling the system to balance macroscopic coverage breadth and microscopic observation depth under limited onboard computing power. The introduction of the trajectory cost function and safety screening mechanism brings the complex three-dimensional environmental constraints and physical flight limitations to the action generation stage, effectively filtering out invalid and dangerous commands, significantly compressing the search space of the policy network, and improving training stability.
[0106] refer to Figure 5 , Figure 5 This is a flowchart illustrating the fourth embodiment of the multi-UAV collaborative inspection method based on risk perception reinforcement learning of the present invention.
[0107] Based on the above embodiments, in this embodiment, step S30 further includes: Step S301: Construct a training sample set based on historical data, the training sample set including historical action features and historical inspection benefit labels.
[0108] It should be noted that although the inspection equipment can directly decide to fly to a high-risk area based on the current risk map, relying solely on local instantaneous risk is insufficient to assess "how much long-term benefit a certain action can bring." This is because the benefit of an action is simultaneously affected by risk density, uncertainty distribution, communication environment, energy consumption cost, and historical inspection status. Therefore, this embodiment introduces a Gaussian process model to predict the potential inspection benefits of candidate actions and simultaneously provide the prediction uncertainty.
[0109] In practical implementation, the inspection equipment tracks the historical actions performed by the drone. The corresponding real-time return sample is defined as: in: Indicates a specific historical decision-making moment; Indicates the first A drone at all times The actual actions performed; This indicates that the immediate inspection benefit after the action is performed can be calculated by the subsequent unified reward function.
[0110] Construct a training sample set based on historical data: in: Indicates the first Deadline for drone deployment The accumulated action-reward sample set; Indicates the total number of samples; Feature vectors representing actions; This indicates the reward label corresponding to the action.
[0111] Step S302: Define the action feature vector of the UAV, construct an action revenue prediction model based on the action feature vector and Gaussian process regression, and train the parameters of the action revenue prediction model based on the training sample set.
[0112] It should be noted that the motion vector of a drone is defined as follows: in: Indicates the starting position of the action; Indicates the starting heading; This indicates the average risk intensity of the area covered by the action; This represents the average uncertainty of the area covered by the action; This indicates the average uninspected duration of the area covered by the action. It represents communication-related characteristics, such as the average connectivity with neighboring machines or base stations after an action is performed; This indicates energy consumption-related characteristics, such as action path length, elevation difference changes, and expected power consumption levels.
[0113] The purpose of this vector is to encode the environmental context, risk context, and execution cost information of "an action" in a unified way, so that it can be used by Gaussian processes for regression learning.
[0114] It should be noted that the action revenue prediction model is composed of action revenue functions, and the unknown action revenue function is defined as follows: in: The characteristic is The potential benefits of the action; Represent a Gaussian process; This is a mean function used to give the prior mean of the action's reward; This is a kernel function used to measure the similarity between two action features.
[0115] Step S303: Input the candidate action features in the hierarchical candidate action map into the trained action revenue prediction model, and output the action revenue prediction value and prediction revenue uncertainty of each candidate action.
[0116] It is understandable that for the candidate action feature set: Its posterior mean is: The posterior covariance is: in: This is a matrix of historical action features; This is a vector of historical return labels; Represents the kernel matrix among historical samples; The kernel matrix represents the relationship between candidate samples and historical samples; Represents the kernel matrix between candidate samples; It is the identity matrix; To observe the noise variance, this is used to improve numerical stability and reflect random perturbations in the gain label.
[0117] For the For each candidate action, its predicted return and prediction uncertainty are defined as follows: in: Indicates the system's response to the first... Estimation of the potential revenue of each candidate action; This indicates the degree of uncertainty in the estimate.
[0118] In subsequent action selection, not only should we consider The size is also taken into consideration. This allows the system to either favor high-yield areas or proactively explore areas where current returns are uncertain but may be valuable.
[0119] Step S304: Calculate the information novelty of each voxel unit based on the predicted risk uncertainty value and the timeliness growth characteristics to obtain the novelty gain of each candidate action.
[0120] It should be noted that the risk intensity and uncertainty defined above reflect "where danger may occur" and "where uncertainty remains," but in long-term inspections, it is also necessary to answer "which areas have outdated information and need to be reviewed again." Therefore, this embodiment further defines information novelty and establishes event-driven revisit priorities.
[0121] In the specific implementation, voxels are defined. At any moment The novelty of information is: in: Indicates the novelty of information in voxels; Indicates the degree of risk uncertainty; Indicates the time-sensitive driving value; These are the weighting coefficients.
[0122] This formula indicates that an area may be worth inspecting either because the system's understanding of it is unclear, or because although the area has been observed before, it has not been re-observed for a long time, and the information is outdated. The former is derived from... The latter is described by describe.
[0123] For candidate actions The total novelty gain of its covered area is defined as: in, This represents the set of voxels expected to be covered after performing the action. This formula allows the novelty of the candidate action to be obtained by summing the novelty of the covered voxels.
[0124] In practical engineering, certain areas should be temporarily assigned higher priority after rainfall, construction, or when anomalies are identified during initial inspection. Therefore, the event-triggered areas should be... Define the event priority field: in: Indicates the event revisit priority of a voxel; Indicates the set of regions where the current event is triggered; This is the base priority assigned after the event is triggered; This is the amplification factor for the impact of the event; The event stimulus intensity is defined in step 1.
[0125] Therefore, the event revisit benefit for a candidate action is defined as: The greater the benefit of revisiting an event, the more effectively the action can cover the current hotspots of events that should be prioritized for review.
[0126] Step S305: Determine the conflict factor based on the distance between the predicted position of the target UAV after performing the candidate action and the predicted position of the neighboring UAV.
[0127] It should be noted that, since the system consists of multiple drones, each drone must share risk map information locally, while avoiding multiple drones repeatedly flying to the same area. Therefore, this embodiment can also establish a communication map, a local map fusion mechanism, and a conflict suppression model.
[0128] In the specific implementation, the time is defined. The communication diagram is as follows: in: Represents a set of drone nodes; Indicates at time There is a set of edges for communication links.
[0129] When any two drones satisfy: Then we have: That is, when the distance between the two machines does not exceed the communication radius At that time, it was believed that the two could exchange information reliably.
[0130] UAV-based local risk map fusion: Part 1 The drone received the neighbor set After obtaining the local risk map, perform weighted fusion on the shared voxels: in: Indicates that after fusion, it will be controlled by drones The risk intensity of the voxels held; Indicates it comes from a drone Voxel risk estimation; This represents the confidence weight of the corresponding estimate.
[0131] Credibility weight The following factors can be considered comprehensively: observation clarity, observation time recentity, observation angle quality, sensor quality level, and communication latency. Essentially, this formula is a weighted average of the local perceptions of different UAVs to form a more accurate and consistent risk map.
[0132] To reduce redundant inspections by multiple machines and potential collisions, the first... The conflict factor for each candidate action is: in: Indicates drone Execute candidate actions The predicted position after; Indicates neighboring machines The predicted location; The conflict sensitivity coefficient determines how quickly the distance changes affect the decay of conflict terms.
[0133] The formula has the following characteristics: when the predicted positions of the two machines are close, the exponential term is large and the conflict factor increases; when the distance is far enough, the exponential term decays rapidly and the impact of the conflict weakens.
[0134] Furthermore, if the following conditions are met: Then the candidate action is directly determined to be an unenforceable action. Here... It is a hard safety constraint, and These are soft conflict constraints, and together they ensure system security.
[0135] Step S306: Extract the expected communication bandwidth gain and expected energy consumption indicators corresponding to each candidate action.
[0136] It should be noted that the expected communication bandwidth gain refers to the combined increase in the number of neighboring machine communication links and data backhaul capacity that are expected to be restored or enhanced after the candidate action is executed. The expected energy consumption index can be the estimated battery discharge amount obtained by combining the trajectory geometric length, vertical climb range, wind field drag, and load power consumption curve through physical equations.
[0137] In some embodiments, the inspection equipment can retrieve a three-dimensional terrain occlusion model and a real-time radio signal propagation map to simulate the changes in link connectivity during the flight of the target UAV to each candidate waypoint. It can calculate the number of neighboring aircraft communication links that can be restored or enhanced after the action is executed and the data backhaul capacity, generate the expected communication bandwidth gain, and combine the geometric length of the candidate trajectory, vertical climb amplitude, wind drag coefficient and airborne load power consumption curve to estimate the amount of battery discharge required to complete the action through the physical energy consumption equation, and output the expected energy consumption index.
[0138] Step S307: Aggregate the predicted action revenue, predicted revenue uncertainty, novelty revenue, conflict factor, expected communication bandwidth gain, expected energy consumption index and risk intensity prediction to construct an action feature matrix.
[0139] In the specific implementation, a unified feature vector is constructed for each candidate action node: in: A geometric description representing a candidate action, such as target position, heading, or velocity command; This represents the predicted value of the action's revenue; This indicates the uncertainty of the predicted value; This indicates the average risk intensity of the area covered by the action; Indicates the average uncertainty of the area covered by the action; This indicates the average uninspected time in the area covered by the action; This indicates the novelty of the information corresponding to the action; Indicates communication gain; Indicates the energy consumption cost index; Indicates the first Conflict factors of candidate actions.
[0140] These features together constitute the action description input of the policy network, enabling the network to learn a balance between multidimensional indicators, rather than making greedy choices based solely on a single risk value.
[0141] This embodiment achieves simultaneous output of long-term action value and assessment confidence through Gaussian process regression, enabling the system to achieve a dynamic balance between high-yield utilization and high-uncertainty exploration. The introduction of information novelty and conflict factors effectively addresses both global cognitive blind spot supplementation and multi-machine spatial resource allocation, avoiding redundant inspections and local missed inspections. Explicit modeling of communication gain and energy consumption indicators deeply embeds engineering physical constraints and network connectivity requirements into the decision feature space. By constructing an action feature matrix, multi-source heterogeneous indicators are transformed into standardized high-dimensional tensors, providing a complete, physically meaningful, and computationally efficient input benchmark for reinforcement learning policy networks.
[0142] refer to Figure 6 , Figure 6 This is a flowchart illustrating the fifth embodiment of the multi-UAV collaborative inspection method based on risk perception reinforcement learning of the present invention.
[0143] Based on the above embodiments, in this embodiment, step S40 further includes: Step S401: Obtain the local observation features corresponding to the action feature matrix, input the local observation features into the initial multi-agent policy network, and output the initial action probability distribution.
[0144] In its implementation, the inspection equipment models the multi-UAV inspection task as a Markov game: in: Represents the global state space, including the states of all drones and the global risk graph state; Represents the joint action space; Indicates the state transition probability; Represents the joint reward function; This is a discount factor used to balance immediate returns with long-term returns.
[0145] The reason for modeling the task as a Markov game is that each drone's decision not only affects its own subsequent state, but also the visible area, communication relationships and risk graph updates of other drones. Therefore, it must be regarded as a multi-agent coupling problem.
[0146] While more complete information can be utilized during the training phase, each drone can only see a localized portion of the environment during actual deployment. Therefore, the first... The local observations from the drone are as follows: in: This refers to the drone's own status. For local digital twin risk maps; The candidate action feature matrix; Provide summary information about neighboring devices, such as their location, remaining battery power, and declared intent.
[0147] The above observation definition ensures the feasibility of decentralized decision-making during the execution phase.
[0148] Define a shared parameter policy network: in: Indicates the policy network parameters; Indicates local observation Select action The probability of.
[0149] Simultaneously define the value network: in, Represents the parameters of the value network. Used to estimate the expected cumulative return under the current local observation.
[0150] Within a centralized training framework, centralized commentators can also be defined: This is used to improve the accuracy of value assessment by utilizing more global information during training.
[0151] Step S402: Obtain the missed detection loss features corresponding to the initial action probability distribution, and calculate the conditional value of risk features based on the missed detection loss features.
[0152] It should be noted that the missed detection loss feature refers to the number and duration of high-risk voxels not effectively observed within the statistical decision-making period, and is a tail risk exposure indicator quantified by combining a high-risk threshold. The conditional value at risk feature can be the conditional expected value calculated for the high quantile interval of the missed detection loss distribution, used to characterize the severity of extreme missed detection situations.
[0153] In some embodiments, the inspection equipment traverses the high-risk voxel units within the current decision-making cycle, counts the number and duration of high-risk areas not effectively observed and covered by any UAV, generates a missed detection loss feature sequence in combination with a preset high-risk threshold, sorts the sequence in ascending order according to the loss value, extracts the tail high quantile interval corresponding to a preset confidence level, calculates the conditional expected value of missed detection loss within the interval, and outputs the conditional risk value feature.
[0154] Step S403: Add the conditional risk value feature as a tail constraint penalty term to the preset strategy optimization function to construct a risk-sensitive objective function.
[0155] It is understandable that the definition of the first The cumulative discount return for each drone is: If only optimization The strategy may favor high returns on average, but may be insufficiently sensitive to missed detections in a few key high-risk areas. Therefore, this invention introduces a conditional value-at-risk term and defines the risk-sensitive objective function as follows: in: This indicates the expected cumulative return; This indicates the loss due to missed detection; Indicates confidence level Conditional value at risk of loss due to missed detection; This represents the risk penalty coefficient.
[0156] Introducing Conditional Value at Risk (VaR) The significance of this is that it not only reduces the average number of missed detections, but also focuses on suppressing "extreme missed detection situations", thus making it more suitable for engineering safety supervision scenarios.
[0157] Furthermore, in order to enable multi-drone swarms to automatically identify and prioritize the filling of the most threatening regulatory vulnerabilities during collaborative operations, step S402 above may include: Step S4021: Extract the voxel high-risk state indication features of each voxel unit.
[0158] It should be noted that the voxel high-risk status indication feature refers to a binary tag or numerical identifier used at the logical level to identify whether a specific spatial unit is above the safety warning line, and is used to indicate whether the predicted risk intensity value of the voxel unit is greater than a preset high-risk intensity threshold.
[0159] Step S4022: Extract the voxel missed detection status indication features of each voxel unit.
[0160] It should be noted that the voxel missed detection status indication feature refers to the logical judgment result reflecting whether a specific monitoring unit is effectively covered by the UAV payload within a predetermined task time, and is used to indicate whether the voxel unit has not been effectively inspected within a preset time window.
[0161] Step S4023: Perform a product operation on the voxel high-risk state indication feature and the voxel missed detection state indication feature to obtain the high-risk missed detection feature; Step S4024: Sum the high-risk missed detection features of all voxel units within a preset decision period to obtain an initial set of missed parameters; Step S4025: Calculate the tail estimate of the distribution of the initial omission parameter set under the preset confidence level probability, and obtain the missed detection loss feature corresponding to the initial action probability distribution.
[0162] In practical implementation, the formula for missed detection loss is defined as: in, This is a high-risk threshold; The above formula for missed detection loss is used to calculate the degree to which high-risk areas are missed. The first term indicates whether a voxel belongs to a high-risk area, and the second term indicates whether the high-risk voxel was not effectively inspected within the specified time window.
[0163] Step S404: Based on the risk-sensitive objective function, perform iterative parameter updates on the initial multi-agent policy network to obtain the risk-sensitive multi-agent policy network.
[0164] It is understandable that the advantage function is defined as: in: This represents an estimate of the cumulative return at the current moment; This represents the value network's predicted value; It indicates the degree of superiority or inferiority of the current action relative to the average level.
[0165] Optimize the objective using a shearing strategy: in, Update the shearing parameters for the strategy.
[0166] The purpose of this objective function is to limit the magnitude of each update, prevent excessive policy oscillations, and improve training stability.
[0167] Value network loss is defined as: This formula requires the value network to fit the actual returns as accurately as possible.
[0168] The entropy regularization term is defined as: This item is intended to encourage a certain degree of policy exploration and prevent the policy from converging to a local optimum too early.
[0169] The final overall optimization objective is written as: Among them: the first term corresponds to strategy optimization; the second term corresponds to value fitting; the third term corresponds to exploration regularization; and the fourth term corresponds to risk-sensitive tail constraint.
[0170] Thus, the reinforcement learning part achieves a complete training logic that is "based on a unified reward function, with long-term cumulative returns as the goal, and additional control over high-risk missed detection tail risks".
[0171] Step S405: Input the action feature matrix into the risk-sensitive multi-agent policy network to predict the action probability distribution, and perform sampling processing on the prediction results to output the target collaborative action decision.
[0172] In some embodiments, the inspection device inputs the real-time constructed action feature matrix into the trained risk-sensitive multi-agent policy network to perform forward inference, obtain the updated action probability distribution sequence, selects deterministic maximum extraction or random sampling strategy with temperature coefficient according to the degree of environmental uncertainty in the current task stage, extracts the optimal flight waypoint, speed command and sensor control parameters from the probability distribution, and packages them to generate target cooperative action decision.
[0173] Furthermore, in order to reduce the risk of equipment loss due to improper battery management, step S405 above may include: Step S4051: Input the action feature matrix into the policy probability evaluation layer of the risk-sensitive multi-agent policy network, and output the candidate action probability distribution of each candidate action; Step S4052: Extract the maximum value feature from the candidate action probability distribution to obtain the initial action decision; Step S4053: Obtain the predicted remaining power information corresponding to the execution of the initial action decision, and determine whether the predicted remaining power information is less than a preset safe power threshold; Step S4054: If the predicted remaining battery power is less than the preset safe battery power threshold, generate a forced return command; Step S4055: If the predicted remaining power information is not less than the preset safe power threshold, the action feature matrix is input into the risk-sensitive multi-agent policy network to predict the action probability distribution, and sampling processing is performed on the prediction results to output the target collaborative action decision.
[0174] In its implementation, to ensure flight safety and mission closure, this invention incorporates autonomous return-to-home conditions. A drone will trigger a return-to-home event if it meets any of the following conditions: or, or, in: This represents the remaining battery power of the i-th drone at time t; This is the minimum safe power threshold; The maximum allowed duration for a single task; This is the location of the return point; This indicates the cost of returning from the current location to the return point; This represents the cost of security redundancy; This represents the estimated remaining available return capacity based on the current state.
[0175] The physical meaning of the third condition is: if the current remaining energy is insufficient to support the drone's safe return and retain the necessary redundancy, then it must immediately switch to return mode.
[0176] Step S406: Send the target cooperative action decision to each target drone in the drone cluster, so that each target drone performs a cooperative inspection action based on the target cooperative action decision.
[0177] In practice, after training is complete, the system enters the deployment phase. During deployment, the system no longer relies on global information; instead, each drone makes independent decisions based on its own local observations.
[0178] For the The drone performs the following online operation: 1. Obtain the current local digital twin map tiles, neighboring machine summary information, and its own status; 2. Generate coarse and fine patrol actions based on the current location and key areas; 3. Calculate the predicted benefit, prediction uncertainty, novelty, event revisit value, communication gain, energy cost, and conflict factor for each candidate action; 4. Input the action feature matrix into the policy network to obtain the action probability distribution; 5. Determine the current action to be performed based on either the highest probability selection method or random sampling. 6. Perform the action and complete a new round of observations; 7. Update the local risk map and uncertainty map based on the observation results; 8. When communication conditions are met, exchange local maps and anomaly labels with neighboring machines.
[0179] Its online decision-making is represented as: Alternatively, a random sampling strategy can be adopted: The former is more suitable for the stable execution phase, while the latter is more suitable for maintaining appropriate exploratory nature in areas with higher uncertainty.
[0180] This embodiment quantifies the loss from missed inspections by introducing conditional value of risk (VoV) and incorporates it as a tail constraint penalty term into the policy optimization process, successfully constructing a risk-sensitive objective function. This mechanism overcomes the limitation of traditional reinforcement learning models that only focus on optimizing "average expected return," enabling the multi-agent policy network to force a focus on and suppress the "extreme tail risk" of overlooked key high-risk areas during parameter iteration. The resulting collaborative action decisions not only take into account the overall inspection efficiency of the multi-UAV swarm but also significantly reduce the probability of missing major hidden danger points from the algorithmic level, significantly improving the safety lower limit and decision robustness of the entire collaborative inspection system in actual engineering safety supervision scenarios.
[0181] Furthermore, the aforementioned steps have defined risk intensity, uncertainty, novelty, event priority, communication relationships, and action costs. To ensure that the reinforcement learning training objective aligns with the actual inspection objective, these factors need to be uniformly incorporated into the immediate reward function.
[0182] Based on the definition of the inspection reward function, the first A drone at all times The immediate reward after an action is performed, the inspection reward function is defined as follows: in, To provide comprehensive instant rewards, For risk and return items, The weighting coefficients for the risk-return term. For novelty-related benefits, The weighting coefficient for the novelty benefit term. For the benefit of revisiting the event, The weighting coefficients for the event revisit benefit term. For communication value items, These are the weighting coefficients for the communication value item. Energy consumption cost item The weighting coefficient for the energy consumption cost item. For conflict penalty items, The weighting coefficients for the conflict penalty term. To constrain violations of penalties, To constrain the weighting coefficients of the violation penalty items, Indicates the drone index. Indicates a time index; The design philosophy of this unified reward function is to encourage the system to fly to worthwhile places and to encourage the system to maintain information flow, while punishing high-cost, conflict-ridden, and unsafe behaviors.
[0183] Define the effective inspection benefits of actions in high-risk areas as: in, This indicates whether a voxel was effectively observed. This indicates the current risk level of the voxel. This indicates the occupancy coefficient of the voxel, which has recently been observed with high quality. Indicates the voxel unit index. Indicates the execution of candidate actions The set of voxels that are expected to be covered.
[0184] The novelty benefit term is defined as: in, This indicates the novelty of voxel information; this term directly adds the novelty of step voxels to the action level. Its purpose is to encourage UAVs to conduct supplementary observations of areas with high uncertainty and those that have not been inspected for a long time.
[0185] The event revisit benefit term is defined as: in, This indicates the event revisit priority of the voxel; this item explicitly writes the high priority of the event triggering area into the reward, so that the strategy will automatically favor hot spot revisits after heavy rain, construction and initial anomaly detection.
[0186] An action should be given additional positive incentives when it helps enhance information sharing or maintain local network connectivity. The communication value term is defined as: in, This indicates the increase in the amount of shared information resulting from the action. This indicates the improvement in network connectivity; for example, if a drone flies to a relay location and, although it does not directly cover the highest-risk area, it can restore communication between two groups of drones, then this action can still yield some benefits.
[0187] The energy consumption cost term is defined as: in, This indicates the flight distance corresponding to the candidate action. Indicates the change in flight altitude. Indicates the magnitude of the change in heading. This represents the estimated value of wind field disturbance. and These represent the energy consumption coefficients of each component; this formula is used to approximate the overall flight cost required to perform the maneuver. The longer the distance, the greater the climb, the more frequent the sharp turns, and the stronger the wind resistance, the greater the energy consumption cost.
[0188] The conflict penalty term is defined as follows: in, This indicates the conflict factor for candidate actions; this factor is used to penalize multi-machine aggregation, repetitive inspections, and potential collision trends.
[0189] The constraint violation penalty item is used to define uniform penalties for serious violations such as no-fly zone incursion, low battery, timeout, and collision. in, Indicate whether to enter the no-fly zone. This indicates whether the battery level is below the safe threshold. Indicates whether the task time limit has been exceeded. Indicators that indicate whether a collision or irreversible failure has occurred are typically taken as... or If this occurs, a severe penalty is imposed, causing the strategy to automatically avoid dangerous behaviors during training.
[0190] Furthermore, to facilitate engineering implementation, the key parameters involved in the embodiments of the present invention will be further explained.
[0191] N: Number of drones, usually 3 to 20; Too small a size will result in insufficient wide-area coverage. If it is too large, it will increase the difficulty of communication coordination and conflict control.
[0192] Communication radius, typically taken as 100~1000m; It depends on the type of communication module, terrain obstruction, and backhaul bandwidth requirements. The larger the size, the easier it is for multiple machines to work together, but the equipment cost and power consumption may be higher.
[0193] The safe collision avoidance radius is typically taken as 5~30m; It depends on flight speed, positioning error, and control precision. The higher the value, the higher the safety, but the stronger the constraints on the maneuver space.
[0194] For effective observation of a class, a distance of 20-300 meters is typically taken. It is related to payload resolution, flight altitude, observation angle, and target size. The larger the value, the larger the coverage area of a single observation, but the local detail resolution may decrease.
[0195] The number of candidate actions for coarse inspection is usually set to 20-80. The larger the value, the more comprehensive the macroscopic search, but the computational cost increases.
[0196] The number of candidate actions for fine-tuning is usually 10 to 50. The larger the scale, the more detailed the review, but the complexity of local decision-making increases.
[0197] Risk memory factor, typically taken as 0.7 to 0.98; The larger the value, the more the system tends to retain historical risk states; the smaller the value, the more the system relies on current observations and events.
[0198] Event impact coefficient; This parameter is used to adjust the impact of external events such as rainfall and construction on risk growth. It can be calibrated based on historical data during engineering projects.
[0199] : Time-sensitivity growth coefficient; The larger this parameter is, the more the system prioritizes the need for re-covering areas that have not been inspected for a long time.
[0200] The time-dependent growth parameter is usually taken as... ; Its determining function The growth rate. If the monitoring task has a long time scale, a smaller value should be selected to avoid premature saturation.
[0201] : Variance of observation noise in Gaussian processes; The larger this parameter is, the higher the tolerance for noise in historical return samples.
[0202] Discount factor, usually taken as ; The larger the value, the more the system prioritizes long-term inspection benefits; the smaller the value, the more it focuses on immediate benefits.
[0203] : Strategy clipping parameters, typically taken as ; It determines the maximum range of changes allowed for each policy update.
[0204] Tail risk penalty coefficient, usually taken as ; The larger this parameter is, the more the system prioritizes avoiding high-risk missed detections.
[0205] Reward weighting is satisfied: in: Prioritize controlling high-risk areas; The degree to which a control system prefers novel information; The importance of controlling event-driven review; Control the degree of preference for maintaining communication and information sharing; Control the tendency to conserve energy; Controlling the degree of multi-machine dispersion and the intensity of conflict suppression; Control the severity of penalties for violations.
[0206] Generally speaking, if the project places greater emphasis on the rapid detection of anomalies, then the time required for detection will increase. and If a greater emphasis is placed on battery life and large-scale surveys, then the need for increased [capacity / capacity] will be greater. If the number of drones is large and the density is high, the intensity should be increased appropriately. .
[0207] Furthermore, this embodiment of the invention also proposes a computer-readable storage medium storing a multi-UAV collaborative inspection program. When the multi-UAV collaborative inspection program is executed by a processor, it implements the steps of the multi-UAV collaborative inspection method based on risk perception reinforcement learning as described above.
[0208] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.
[0209] The aforementioned computer-readable storage medium may be included in a multi-UAV collaborative inspection device based on risk perception reinforcement learning; or it may exist independently and not be assembled into a multi-UAV collaborative inspection device based on risk perception reinforcement learning.
[0210] Furthermore, this invention also proposes a computer program product, including a multi-UAV collaborative inspection program, which, when executed by a processor, implements the steps of the multi-UAV collaborative inspection method based on risk perception reinforcement learning as described above.
[0211] The specific implementation of the computer program product of the present invention is basically the same as the embodiments of the multi-UAV collaborative inspection method based on risk perception reinforcement learning described above, and will not be repeated here.
[0212] Reference Figure 7 , Figure 7 This is a structural block diagram of the first embodiment of the multi-UAV collaborative inspection system based on risk perception reinforcement learning of the present invention.
[0213] like Figure 7 As shown, the multi-UAV collaborative inspection system based on risk perception reinforcement learning proposed in this embodiment of the invention includes: Risk prediction module 10 is used to discretize the monitoring area into a risk voxel grid, perform dynamic risk prediction based on the prior risk characteristics and time-varying event field characteristics of the risk voxel grid, and obtain the risk intensity prediction value and risk uncertainty prediction value of each voxel unit in the risk voxel grid. Action construction module 20 is used to construct a coarse patrol action set and a fine patrol action set based on the local observation information of each target UAV in the UAV cluster, and merge the coarse patrol action set and the fine patrol action set to obtain a hierarchical candidate action map. The UAV cluster contains multiple UAVs. The feature construction module 30 is used to predict the action revenue of each candidate action in the hierarchical candidate action map based on the action revenue prediction model, and to construct an action feature matrix based on the action revenue, the risk intensity prediction value and the risk uncertainty prediction value. The action decision module 40 is used to input the action feature matrix into the risk-sensitive multi-agent policy network to predict the action probability distribution and output the target collaborative action decision. The risk-sensitive multi-agent policy network is trained based on the inspection reward function. Each target drone in the drone cluster is configured to perform collaborative inspection actions based on the target collaborative action decision.
[0214] This embodiment achieves effective mining and state quantification of key features in complex soil and water conservation scenarios through environmental discretization and dynamic risk simulation. The introduction of hierarchical action graphs and benefit prediction mechanisms significantly optimizes the allocation logic of computing resources and flight time, avoiding redundant exploration and computing power bottlenecks. The deep coupling of risk-sensitive policy networks and comprehensive reward criteria ensures that multi-UAV clusters can maintain efficient collaboration and safety boundaries even under dynamic interference. It significantly improves the collaborative operation efficiency of multiple UAVs in large-scale complex terrain, realizing a leap from traditional rule-based static inspection to intelligent risk-based dynamic inspection. Through dynamic risk perception and hierarchical decision-making mechanisms, it can accurately locate hidden danger areas and automatically allocate inspection forces, greatly reducing the cost of manual participation and equipment maintenance. At the same time, the risk-sensitive optimization objectives effectively ensure the safety of the inspection process, avoiding missed inspections and resource misallocation, providing an efficient, intelligent, and low-cost inspection method for long-term, wide-area soil and water conservation monitoring.
[0215] In some embodiments, a multi-UAV collaborative inspection system based on risk perception reinforcement learning is configured to perform the following steps: S1: Collect topographic data, engineering layout data, historical monitoring data, meteorological data, and construction disturbance data; S2: Discretize the monitoring area into a risk voxel grid and construct a digital twin memory map; S3: Initialize voxel risk intensity based on object category and prior features; S4: Predict risk evolution trends based on external event fields and time-dependent driving functions; S5: Multiple drones take off and acquire local observations; S6: Update the risk posterior and uncertainty of each voxel based on Bayesian methods; S7: Construct candidate action graphs for the coarse-pattern and fine-pattern layers; S8: Predicting the payoff and payoff uncertainty of candidate actions using Gaussian processes; S9: Construct action features by combining novelty, event priority, communication value, energy consumption cost, and conflict penalty; S10: Output action decisions through a risk-sensitive multi-agent reinforcement learning strategy; S11: Perform actions and update the local risk map, while fusing the local map under nearby communication conditions; S12: Trigger dynamic revisits for hotspot areas after heavy rain, after initial abnormal inspections, and after construction disturbances; S13: When the return conditions or mission completion conditions are met, perform autonomous return and output inspection results.
[0216] The multi-UAV collaborative inspection system based on risk perception reinforcement learning provided in this application employs the multi-UAV collaborative inspection method based on risk perception reinforcement learning in the above embodiments, and can solve the technical problems of multi-UAV collaborative inspection based on risk perception reinforcement learning. Compared with the prior art, the beneficial effects of the multi-UAV collaborative inspection system based on risk perception reinforcement learning provided in this application are the same as the beneficial effects of the multi-UAV collaborative inspection method based on risk perception reinforcement learning provided in the above embodiments, and other technical features of the multi-UAV collaborative inspection system based on risk perception reinforcement learning are the same as the features disclosed in the methods of the above embodiments, and will not be repeated here.
[0217] It should be understood that the above are merely illustrative examples and do not constitute any limitation on the technical solutions of the present invention. In specific applications, those skilled in the art can make settings as needed, and the present invention does not impose any restrictions on this.
[0218] It should be noted that the workflow described above is merely illustrative and does not limit the scope of protection of this invention. In practical applications, those skilled in the art can select some or all of the workflow to achieve the purpose of this embodiment according to actual needs, and no restrictions are imposed here.
[0219] In addition, for technical details not described in detail in this embodiment, please refer to the multi-UAV collaborative inspection method based on risk perception reinforcement learning provided in any embodiment of the present invention, which will not be repeated here.
[0220] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or system that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or system. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or system that includes that element.
[0221] It should be noted that the user information (including but not limited to user device information, user personal information, user location information, user behavior information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.
[0222] The sequence numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0223] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as read-only memory / random access memory, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of the present invention.
[0224] The above are merely preferred embodiments of the present invention and do not limit the scope of the patent. Any equivalent structural or procedural transformations made based on the description and drawings of the present invention, or direct or indirect applications in other related technical fields, are similarly included within the scope of patent protection of the present invention.
Claims
1. A multi-UAV collaborative inspection method based on risk perception reinforcement learning, characterized in that, The method includes: The monitoring area is discretized into a risk voxel grid. Based on the prior risk characteristics and time-varying event field characteristics of the risk voxel grid, dynamic risk prediction is performed to obtain the predicted risk intensity and risk uncertainty of each voxel unit in the risk voxel grid. The prior risk characteristics include object prior risk weight, slope factor, elevation difference factor and historical disturbance factor. The time-varying event field characteristics include event excitation intensity and time-dependent growth characteristics. Based on the local observation information of each target UAV in the UAV cluster, a coarse patrol action set and a fine patrol action set are constructed. The coarse patrol action set and the fine patrol action set are merged to obtain a hierarchical candidate action map. The UAV cluster contains multiple UAVs. The action revenue of each candidate action in the hierarchical candidate action graph is predicted based on the action revenue prediction model, and an action feature matrix is constructed based on the action revenue, the predicted risk intensity value, and the predicted risk uncertainty value. The action feature matrix is input into a risk-sensitive multi-agent policy network to predict the action probability distribution and output the target collaborative action decision. The risk-sensitive multi-agent policy network is trained based on the inspection reward function. Each target drone in the drone swarm is configured to perform collaborative inspection actions based on the target collaborative action decision.
2. The multi-UAV collaborative inspection method based on risk perception reinforcement learning as described in claim 1, characterized in that, The process of dynamically predicting risk based on the prior risk characteristics and time-varying event field characteristics of the risk voxel grid, and obtaining the predicted risk intensity and risk uncertainty values for each voxel cell in the risk voxel grid, includes: Prior risk weights for objects are assigned based on their regulatory importance within the monitoring area. Extract the slope factor, elevation difference factor, and historical disturbance factor of the risk voxel grid, and calculate the initial risk intensity of each voxel unit in combination with the prior risk weight of the object. The event excitation intensity of the external time-varying event field is obtained, and the time-dependent growth characteristics of each voxel unit in the risk voxel grid are calculated based on the time interval driving function. A risk prediction model is constructed based on the initial risk intensity, the event incentive intensity, and the time-dependent growth characteristics. The risk voxel grid is observed by the target UAV, and the output of the risk prediction model is corrected a posteriori based on the Bayesian update formula to obtain the risk intensity prediction value and risk uncertainty prediction value of each voxel unit.
3. The multi-UAV collaborative inspection method based on risk perception reinforcement learning as described in claim 1, characterized in that, The process involves constructing a coarse-survey action set and a fine-survey action set based on local observation information of each target UAV in the UAV swarm. Merging the coarse-survey action set and the fine-survey action set yields a hierarchical candidate action map, including: Based on local observation information of each target drone in the drone swarm, the location information and abnormal hotspot information of each target drone are extracted; Action sampling is performed within the first preset sampling radius corresponding to the location information to generate a coarse patrol action set; Action sampling is performed within the second preset sampling radius corresponding to the abnormal hotspot information to generate a set of fine patrol actions; Construct a trajectory cost function that includes trajectory length, elevation difference cost, wind field disturbance cost, and obstacle avoidance cost; Based on the trajectory cost function, a safety screening is performed on the coarse patrol action set and the fine patrol action set to eliminate unreachable actions; The coarse patrol action set and fine patrol action set after safety screening are aggregated into nodes to obtain a hierarchical candidate action graph.
4. The multi-UAV collaborative inspection method based on risk perception reinforcement learning as described in claim 1, characterized in that, The step of predicting the action payoff of each candidate action in the hierarchical candidate action graph based on the action payoff prediction model, and constructing an action feature matrix based on the action payoff, the predicted risk intensity value, and the predicted risk uncertainty value, includes: A training sample set is constructed based on historical data, which includes historical action features and historical inspection benefit labels; Define the action feature vector of the UAV, construct an action revenue prediction model based on the action feature vector and Gaussian process regression, and train the parameters of the action revenue prediction model based on the training sample set. The candidate action features in the hierarchical candidate action map are input into the trained action benefit prediction model, and the action benefit prediction value and prediction benefit uncertainty of each candidate action are output. Based on the predicted risk uncertainty and the time-dependent growth characteristics, the information novelty of each voxel unit is calculated to obtain the novelty gain of each candidate action. The conflict factor is determined based on the distance between the predicted position of the target UAV after performing candidate actions and the predicted position of neighboring UAVs; Extract the expected communication bandwidth gain and expected energy consumption indicators for each candidate action; The predicted action revenue, the uncertainty of the predicted revenue, the novelty revenue, the conflict factor, the expected communication bandwidth gain, the expected energy consumption index, and the predicted risk intensity are aggregated to construct an action feature matrix.
5. The multi-UAV collaborative inspection method based on risk perception reinforcement learning as described in claim 1, characterized in that, The step of inputting the action feature matrix into a risk-sensitive multi-agent policy network to predict the action probability distribution and outputting a target collaborative action decision includes: Obtain the local observation features corresponding to the action feature matrix, input the local observation features into the initial multi-agent policy network, and output the initial action probability distribution; Obtain the missed detection loss features corresponding to the initial action probability distribution, and calculate the conditional value of risk features based on the missed detection loss features; The conditional value-at-risk feature is added as a tail constraint penalty term to a preset strategy optimization function to construct a risk-sensitive objective function; Based on the risk-sensitive objective function, the parameters of the initial multi-agent policy network are iteratively updated to obtain the risk-sensitive multi-agent policy network. The action feature matrix is input into the risk-sensitive multi-agent policy network to predict the action probability distribution, and the prediction results are sampled and processed to output the target collaborative action decision. The target cooperative action decision is sent to each target drone in the drone cluster, so that each target drone performs a cooperative inspection action based on the target cooperative action decision.
6. The multi-UAV collaborative inspection method based on risk perception reinforcement learning as described in claim 5, characterized in that, The step of obtaining the missed detection loss features corresponding to the initial action probability distribution includes: Extract the voxel high-risk state indication features of each voxel unit. The voxel high-risk state indication features are used to indicate whether the predicted risk intensity value of the voxel unit is greater than a preset high-risk intensity threshold. Extract the voxel missed detection status indication features of each voxel unit. The voxel missed detection status indication features are used to indicate whether the voxel unit has not been effectively inspected within a preset time window. The high-risk missed detection feature is obtained by multiplying the voxel high-risk state indication feature and the voxel missed detection state indication feature. The high-risk missed detection features of all voxel units are summed and accumulated within a preset decision period to obtain an initial set of missed parameters. Calculate the tail estimate of the distribution of the initial set of missing parameters under a preset confidence level probability, and obtain the missed detection loss feature corresponding to the initial action probability distribution.
7. The multi-UAV collaborative inspection method based on risk perception reinforcement learning as described in claim 5, characterized in that, The step of inputting the action feature matrix into the risk-sensitive multi-agent policy network to predict the action probability distribution, and sampling the prediction results to output the target collaborative action decision includes: The action feature matrix is input into the policy probability evaluation layer of the risk-sensitive multi-agent policy network, and the candidate action probability distribution of each candidate action is output. The initial action decision is obtained by extracting the maximum value feature from the candidate action probability distribution; Obtain the predicted remaining battery power information corresponding to the initial action decision, and determine whether the predicted remaining battery power information is less than a preset safe battery power threshold; If the predicted remaining battery power is less than the preset safe battery power threshold, a forced return command is generated. If the predicted remaining battery power is not less than the preset safe battery power threshold, the action feature matrix is input into the risk-sensitive multi-agent policy network to predict the action probability distribution, and the prediction results are sampled and processed to output the target collaborative action decision.
8. The multi-UAV collaborative inspection method based on risk perception reinforcement learning as described in claim 1, characterized in that, The inspection reward function includes risk reward, novelty reward, event revisit reward, communication value, energy cost, conflict penalty, and constraint violation penalty, as shown in the following formula: in, To provide comprehensive instant rewards, For risk and return items, The weighting coefficients for the risk-return term. For novelty-related benefits, The weighting coefficient for the novelty benefit term. For the benefit of revisiting the event, The weighting coefficients for the event revisit benefit term. For communication value items, These are the weighting coefficients for the communication value item. Energy consumption cost item The weighting coefficient for the energy consumption cost item. For conflict penalty items, The weighting coefficients for the conflict penalty term. To constrain violations of penalties, To constrain the weighting coefficients of the violation penalty terms, Indicates the drone index. Indicates a time index; in, This indicates whether a voxel was effectively observed. This indicates the current risk level of the voxel. This indicates the occupancy coefficient of the voxel, which has recently been observed with high quality. Indicates the voxel unit index. Indicates the execution of candidate actions The expected set of voxels to be covered; in, Indicates the novelty of information in voxels; in, Indicates the event revisit priority of a voxel; in, This indicates the increase in the amount of shared information resulting from the action. This indicates the amount of improvement in network connectivity. in, This indicates the flight distance corresponding to the candidate action. Indicates the change in flight altitude. Indicates the magnitude of the change in heading. This represents the estimated value of wind field disturbance. and These represent the energy consumption coefficients of each part; in, Indicates the conflict factor of the candidate action; in, Indicate whether to enter the no-fly zone. This indicates whether the battery level is below the safe threshold. Indicates whether the task time limit has been exceeded. Indicates whether a collision or irreversible failure has occurred.
9. A multi-UAV collaborative inspection system based on risk perception reinforcement learning, applying the multi-UAV collaborative inspection method based on risk perception reinforcement learning as described in any one of claims 1 to 8, characterized in that, The system includes: The risk prediction module is used to discretize the monitoring area into a risk voxel grid, and perform dynamic risk prediction based on the prior risk characteristics and time-varying event field characteristics of the risk voxel grid. It obtains the predicted risk intensity and risk uncertainty of each voxel unit in the risk voxel grid. The prior risk characteristics include object prior risk weight, slope factor, elevation difference factor and historical disturbance factor. The time-varying event field characteristics include event excitation intensity and time-effect growth characteristics. The action construction module is used to construct a coarse patrol action set and a fine patrol action set based on the local observation information of each target UAV in the UAV cluster, and merge the coarse patrol action set and the fine patrol action set to obtain a hierarchical candidate action map. The UAV cluster contains multiple UAVs. The feature construction module is used to predict the action revenue of each candidate action in the hierarchical candidate action graph based on the action revenue prediction model, and to construct an action feature matrix based on the action revenue, the risk intensity prediction value, and the risk uncertainty prediction value. The action decision module is used to input the action feature matrix into the risk-sensitive multi-agent policy network to predict the action probability distribution and output the target collaborative action decision. The risk-sensitive multi-agent policy network is trained based on the inspection reward function. Each target drone in the drone cluster is configured to perform collaborative inspection actions based on the target collaborative action decision.
10. A multi-UAV collaborative inspection device based on risk perception reinforcement learning, characterized in that, The multi-UAV collaborative inspection device based on risk perception reinforcement learning includes: a memory, a processor, and a multi-UAV collaborative inspection program stored in the memory. The processor is used to run the multi-UAV collaborative inspection program, which is configured to implement the multi-UAV collaborative inspection method based on risk perception reinforcement learning as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Automatic driving behavior prediction method and system fusing scene risk perception and interactive collaborative decision
CN121278534A