Safety decision-making method for patrol scheduling of multiple unmanned vehicles in confrontation environment
Through the security decision-making method of hybrid architecture and the dual-layer design of the scenario evaluation layer and the decision-making layer, the task allocation and coordination problems of multiple unmanned vehicles in the confrontational environment are solved, and efficient and robust task completion is achieved.
Patent Information
- Application Number
- CN202510434102.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-07-22
AI Technical Summary
The existing multi-unmanned vehicle collaborative control methods are difficult to achieve efficient and robust task allocation and coordination in complex dynamic environments, especially in the confrontation environment, which lacks a safety decision-making level, resulting in high difficulty in algorithm convergence and poor generalization capabilities.
The security decision-making method adopts a hybrid architecture, including the scenario evaluation layer, the decision-making layer and the performance monitoring layer. Through the combination of layered decision-making mechanisms and reinforcement learning, the decision-making execution quality is monitored in real time and automatic recovery and retry are carried out. A two-layer architecture is designed to switch decision paradigms and task allocation is combined with the improved greedy joint auction algorithm.
It realizes efficient and coordinated operation of multiple unmanned vehicle systems in complex environments, improves the robustness and adaptability of the system, and ensures the efficiency and safety of task completion in dynamic environments.
Smart Images

Figure CN120355259A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of cooperative control of unmanned vehicles, and particularly relates to a safety decision-making method for multi-unmanned vehicle inspection and scheduling in an adversarial environment. Background Art
[0002] With the rapid development of unmanned driving technology, the cooperative execution of tasks by multiple unmanned platforms has become an important development direction of intelligent transportation systems, and has broad application prospects in scenarios such as smart cities and emergency rescue. Especially in an adversarial environment, the ability of multiple unmanned vehicles to efficiently cooperate as carriers of unmanned platforms to complete complex tasks is particularly important. However, in a complex unstructured dynamic environment, there are still many challenges in realizing the intelligent scheduling decision-making of multiple unmanned vehicles to autonomously allocate tasks and coordinate behaviors according to task requirements and environmental changes. The inspection and scheduling of multiple unmanned vehicles in a complex environment is a comprehensive problem, which requires maximizing the efficiency of the multiple unmanned vehicle system to search and achieve the expected indicators at the decision-making level, and relying on the autonomy of individual vehicles to complete local path planning, environmental perception and data collection, and finally realizing the macro closed-loop control from perception to decision-making.
[0003] In the field of inspection and scheduling in recent years, the multi-agent reinforcement learning (MARL) method has received extensive attention in the academic community because of its ability to autonomously learn strategies from the environment to well solve complex multi-agent cooperative control problems, and is divided into three major learning paradigms: uncorrelated type, communication rule type, and mutual cooperation type. Especially for the cooperative learning paradigm that does not rely on explicit communication represented by QMIX proposed by Rashid et al. and COMA proposed by Foerster et al., multi-agent reinforcement learning enables each agent to learn a local strategy through the interaction between the agent and the environment, and the individual strategy learning processes can imitate and restrict each other, so as to achieve the cooperation of the entire cluster. In terms of solving practical problems, Deng et al. directly applied the deep reinforcement learning algorithm represented by MADDPG to the search of unknown areas by multiple robots, and directly associated the action space of the agent with the search behavior. However, most of the existing MARL methods assume that the agents are homogeneous, the environmental information is completely observable, and the situation of the higher-dimensional action space of individuals is not considered, which has a large gap with practical applications. In addition, MARL is prone to falling into local optima. In the case of the separation of the decision-making layer and the execution layer and the decentralized execution of individual actions by multiple unmanned vehicles, it is difficult for such methods to solve the environmental non-stationarity problem caused by the decision-making-actions of multiple agents, which will significantly increase the difficulty of algorithm convergence. At the same time, lacking the backup of a safety decision-making layer, it will be difficult for the reinforcement learning method to solve the strategy fine-tuning for a dynamic environment and ensure the generalization ability of the algorithm.
[0004] Traditional multi-unmanned vehicle collaborative control methods are the core algorithms of the common safety decision-making layer. Sarkar et al. believe that its basic idea is based on distributed path planning combined with mutual communication, using heuristic dynamic adjustment of priorities to avoid conflicts, and finding the optimal solution by adjusting subsets. It mainly includes two categories: centralized and distributed. Common centralized methods such as the Hungarian multi-objective allocation algorithm uniformly allocate tasks after the central node obtains global information. Such methods have strong global optimization capabilities, but lack flexibility and robustness in dynamic environments, and have high computational complexity, making it difficult to handle large-scale clusters. Distributed methods such as the market bidding mechanism, where each unmanned vehicle bids for tasks according to its own utility function. A typical example is that Khodayi-mehr et al. proposed dividing the swarm of robots into multiple teams, and each team reduces the gain overlap between robots by sharing information and updating belief coefficients. These two types of methods have good scalability, but it is difficult to obtain the global optimal solution, and it is difficult to effectively handle the coordination and collision avoidance problems between unmanned vehicles. Summary of the Invention
[0005] The technical problem to be solved by the present invention is to provide a safety decision-making method for multi-unmanned vehicle inspection and scheduling in an adversarial environment.
[0006] To achieve the above object, the present invention adopts the following technical solutions:
[0007] A safety decision-making method for multi-unmanned vehicle inspection and scheduling in an adversarial environment, including:
[0008] Step S1, obtaining environmental perception information through the scenario evaluation layer;
[0009] Step S2, inputting the environmental perception information into the decision-making layer for inspection and scheduling planning based on reinforcement learning under a hierarchical decision-making execution mechanism;
[0010] Step S3, the decision-making layer monitors the decision execution quality in real time and performs automatic recovery and retry.
[0011] Preferably, in step S1, the scenario evaluation layer conducts a comprehensive analysis of the task environment through scenario complexity evaluation, environmental complexity analysis, and communication status evaluation.
[0012] Preferably, in step S2, the decision-making layer adopts a two-layer architecture design. The upper layer realizes utility calculation, bidding iteration, and task exchange optimization based on an improved greedy joint auction algorithm, and the lower layer realizes inspection and scheduling based on a reinforcement learning framework.
[0013] Preferably, in step S3, the decision-making layer monitors the decision execution quality in real time and performs automatic recovery and retry through performance index monitoring and failure recovery mechanisms.
[0014] Based on a complex unstructured scenario without a priori maps and with partial known information, and considering heterogeneous constraints among unmanned vehicles, it is assumed that the reinforcement learning layer is based on a multi-agent reinforcement learning algorithm, which is used for the ground station to dynamically publish the inspection target points of each unmanned vehicle; a distributed safety decision-making mechanism is constructed at the unmanned vehicle node layer as a supplementary extension of the reinforcement learning method in special cases. When the scenario undergoes dynamic changes or abnormal situations, the unmanned vehicle node fuses multi-source information about the environment and transmits the information to the decision-making layer, thereby realizing scheduling replanning under the hierarchical decision-making mechanism. On this basis, the unmanned vehicle uses its own local planner to execute end-point planning control, achieving efficient collaborative execution of the inspection task, enabling the multi-unmanned vehicle system to better adapt to environmental changes in a large-scale task dynamic scenario, overcoming the deficiencies in the robustness of traditional multi-agent reinforcement learning algorithms, and achieving the efficient completion of preset inspection performance indicators. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on the provided drawings without creative efforts.
[0016] Figure 1 It is a flowchart of a safety decision-making method for multi-unmanned vehicle inspection scheduling in an adversarial environment according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0017] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0018] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the drawings and specific embodiments.
[0019] Embodiment 1:
[0020] As Figure 1 shown, an embodiment of the present invention provides a safety decision-making method for multi-unmanned vehicle inspection scheduling in an adversarial environment, including:
[0021] Step S1, obtaining environmental perception information through the scenario evaluation layer;
[0022] Step S2: Input the environmental perception information into the decision-making layer for inspection scheduling planning based on reinforcement learning under a hierarchical decision-making mechanism.
[0023] Step S3: The decision-making layer monitors the decision execution quality in real time and performs automatic recovery and retry through performance metric monitoring and failure recovery mechanisms.
[0024] Considering that in actual deployment, the environmental functions of the model can be extended according to specific requirements, such as adding a more complex obstacle generation mechanism or introducing characteristics such as dynamic environmental changes to test the adaptability and robustness of the algorithm. However, relying solely on reinforcement learning methods to face all inspection scheduling problems will inevitably face problems such as unreliable decision-making layers, poor environmental adaptability, and disconnection in decision execution. Therefore, a mobile decision-making module with a hybrid architecture is designed:
[0025] As Figure 1 shown, in order to improve the reliability and robustness of the system in complex scenarios, a hybrid architecture decision-making system is designed, which can adaptively switch decision-making paradigms according to scene characteristics. When the ground multi-unmanned vehicle system realizes distributed perception with the help of an edge computing unit and inputs the environmental perception into the decision-making layer, a scene evaluation layer and a decision-making layer are selected to organically adjust and allocate the scheduling strategy. Next, these three conceptual layers will be introduced in turn:
[0026] The scene evaluation layer serves as the perception entrance of the system and conducts a comprehensive analysis of the task environment through three key modules: scene complexity evaluation, environmental complexity analysis, and communication status evaluation. This layer not only continuously monitors the changes in the environmental situation but also provides a reliable environmental perception basis for the decision-making layer through real-time data collection and analysis to ensure the scientificity and accuracy of subsequent decisions.
[0027] The decision-making layer adopts a two-layer architecture design. The safety decision-making layer is responsible for utility calculation, competitive iteration, and task exchange optimization, and the lower layer realizes the preset inspection strategy based on the reinforcement learning framework. It should be noted that the switching basis largely depends on the grasp of the regional situation by the environmental evaluation layer, which is quantified as shown in Table 1 in this invention.
[0028] Table 1
[0029]
[0030] Considering the complex and changeable inspection scheduling in the adversarial environment and fully taking into account the fields where the two algorithms can play their maximum roles, a two-layer architecture is designed. When the scenario meets one of the situations in the above table, the safety decision-making layer, relying on its advantages in certainty and computational efficiency, can provide a reliable task allocation scheme in these extreme situations. At the same time, the system continuously monitors the scenario status and automatically switches back to the reinforcement learning paradigm to obtain better decision-making performance when the conditions return to normal. This hybrid architecture not only ensures the optimization effect of the system in normal scenarios but also guarantees the reliability and robustness in extreme situations. This hybrid decision-making mechanism adopts a smooth switching mechanism with transitional weighting, and for this purpose, a decision fusion buffer is introduced: a weighted combination of the two decisions is briefly used before and after the switch, which not only ensures the stability of the decision but also improves the system's ability to handle complex scenarios. Finally, the calculated target points are sent to the ground multi-unmanned vehicle system in real time, and the local perception and path planner of the unmanned vehicle are used to avoid dynamic and static obstacles.
[0031] The decision-making layer monitors the quality of decision execution in real time and performs automatic recovery and retry through two functional modules: performance metric monitoring and failure recovery mechanism. This layer is not only responsible for the real-time supervision of the task execution process but also establishes a feedback mechanism with the scenario evaluation layer. By correcting the evaluation results in real time and optimizing the evaluation accuracy, a closed-loop optimization mechanism of the system is formed.
[0032] These three layers of architecture cooperate with each other through an organically unified information flow and control flow to build an adaptive and robust task allocation system. The environmental perception of the scenario evaluation layer provides the basis for decision-making. The hybrid architecture of the decision-making layer ensures the reliability and adaptability of the decision-making, while the decision-making layer guarantees the closed-loop optimization of the entire system through the feedback mechanism, ultimately realizing the efficient cooperative operation of the multi-unmanned vehicle system in a complex environment.
[0033] As an implementation method of the embodiment of the present invention, the improved greedy joint auction algorithm proposed in the safety decision-making layer is essentially a multi-objective allocation decision algorithm. Under this algorithm, at fixed-step moments, an arbitrary number of map inspection points to be inspected (the number does not have to be strictly matched with the current number of unmanned vehicles) are relatively evenly generated in the area without prior information. These inspection points are approximately formed into several regional clusters and are used as inspection points for subsequent multi-objective allocation. Considering the adversarial environment, the multi-unmanned vehicle system needs to track and intercept abnormal targets found during the inspection. Therefore, once such targets are found, they will also be used as a kind of dynamic inspection point. So far, the problem faced by the safety decision-making layer is the task allocation problem of unmanned vehicles in a complex scenario with dynamic and static inspection targets.
[0034] Secondly, it is necessary to preliminarily calculate the cost matrix as the input of the subsequent allocation algorithm. In this stage, only the relative distances between the unmanned vehicles and each cluster are considered, and the distance cost matrix is preliminarily calculated. The specific calculation method is as follows:
[0035]
[0036] Among them, C i,j represents the element in the i-th row and j-th column of the cost matrix, pos i and pos j represent the current coordinates of the unmanned vehicle and the cluster respectively, represents the linear prediction of the coordinates of the unmanned vehicle at the next decision moment. Based on this, k1 and k2 are set as item weights, and generally, the Euclidean distance between the current and predicted coordinates is used as the element of the cost matrix C.
[0037] The core process of the improved greedy joint auction algorithm can be divided into three key stages: initialization, iterative bidding, and post-processing optimization. In the initialization stage, the cost matrix is defined and obtained artificially. The algorithm first converts the cost matrix into utility values, and two key parameters, position persistence reward and distance penalty, are introduced in the utility calculation to balance the stability and efficiency of task allocation. The utility calculation comprehensively considers the basic utility, distance cost, and persistence reward, and this formula will change the weights ω1 and ω2 with the accumulation of the historical data of a single vehicle, thus forming a more comprehensive decision basis. The following is the utility calculation formula:
[0038]
[0039] Among them: U ij represents the comprehensive utility of the unmanned vehicle i for the task j, β is the persistence reward weight, I is the persistence indicator function, which is used to indicate whether the current unmanned vehicle has a task assignment, C ij is the basic cost, α is the additional distance penalty factor, d ij is the distance cost, is the passable coefficient from the unmanned vehicle i to the inspection point j, which will increase with the increase of the historical traversal degree along the straight line from the unmanned vehicle to the inspection point and the decrease of the obstacle density, thereby reducing the additional distance penalty term, ε is a small positive number to avoid division by zero, and ω1 and ω2 are the dynamic weights of the historical single-vehicle data;
[0040] Under the action of these two items, the execution parameters of the single vehicle during historical iterative allocation will be recorded. By fine-tuning the weights of this item, individuals with stronger mobility are encouraged to relatively ignore the distance cost, frequently change the search interval, and individuals with low completion rates relatively attach importance to persistence.
[0041] In the iterative bidding stage, the algorithm adopts a distributed bidding mechanism, and each unmanned vehicle selects the optimal task to bid based on the current utility matrix. By maintaining local bidding information and global allocation results, the algorithm can converge to a stable allocation scheme within a limited number of iterations. When using the softmax function to guide the unmanned vehicle to iteratively bid for inspection points, the retention probability of the sub-optimal solution is ensured, so as to be able to jump out of the local optimum. This process is similar to the auction mechanism in economics, where each unmanned vehicle tries to maximize its own utility and realizes the reasonable allocation of global tasks through the bidding process. The following is the iterative bidding function:
[0042] U choose =[b ij t ,U ij ,
[0043]
[0044] [P1,P2]=softmax(U choose / T tem )
[0045]
[0046] where: b ij (t) is the final bidding price of unmanned vehicle i for task j in the t-th iteration, U ij is the comprehensive utility calculated in this round, and b local is the local bidding winning price for a single inspection point.
[0047]
[0048] The post-processing optimization stage focuses on solving the problem of unallocated unmanned vehicles generated after the bidding iteration in the previous stage, and optimizes the task exchange through a comprehensive scoring mechanism (combining distance score and persistence score). When a better allocation scheme is found, the algorithm will dynamically adjust the task allocation result to ensure the local optimality of the final scheme. This optimization method based on the greedy strategy not only ensures the convergence of the algorithm, but also achieves a good balance between computational efficiency and scheme quality. The task exchange scoring function:
[0049]
[0050] where: S ij is the exchange score, d ij is the distance cost, γ is the interference weight between tasks, I ij is the persistence indicator function, and a ik is the current allocation status, which takes 1 when task j is currently allocated to i, otherwise 0
[0051] In the scenario of the multi-unmanned vehicle inspection and scheduling decision-making area under adversarial conditions, the safety decision-making layer can efficiently solve complex task allocation problems. Suppose N unmanned vehicles are deployed in a certain area and tasks need to be allocated among M potential inspection points. Since it is in an adversarial environment, assume that the multi-unmanned vehicle system needs to complete a relatively large-scale inspection of the area within a limited time and continuously track and finally intercept potential dynamic targets. For this reason, these inspection points include detected dynamic target points and static inspection points. The algorithm quantifies factors such as the path planning cost of the unmanned vehicle to the target point (considering terrain obstacles), the number of historical allocated inspection points of each unmanned vehicle, the effective arrival number of each unmanned vehicle, and the task persistence requirement (to avoid efficiency loss caused by frequent switching) into numerical values by constructing a comprehensive utility function. During the iterative bidding process, each unmanned vehicle makes autonomous decisions based on the current situation, selects the optimal target point for bidding, and gradually converges to a stable allocation plan through local information interaction. For example, when a new dynamic target is discovered, the algorithm weighs the task urgency and the current deployment cost, which may trigger task reallocation, adjusts the nearest or most suitable unmanned vehicle to the interception position, and at the same time considers the cooperative replacement of other unmanned vehicles, and finally completes the dynamic task adjustment while maintaining the overall decision-making efficiency.
[0052]
[0053]
[0054]
[0055]
[0056] The embodiments of the present invention are based on the idea of greedy joint auction, providing support for the decision relay of the unmanned vehicle in the face of specific scenarios as the security decision layer, calculating each cost matrix distributively, introducing the dynamic weight of historical data, and generating a utility matrix based on the adaptive utility calculation formula. An evaluation function is designed based on the comprehensive utility value of the task, distance factor, etc., to accelerate the process of finding the optimal solution. A randomness strategy is introduced in the optimal task selection, and the sub-optimal task is selected with a certain probability to enhance the exploration ability of the algorithm and jump out of the local optimal solution. By setting the softmax random sampling mechanism, when the utility values of the optimal task and the sub-optimal task are close, the sub-optimal task is selected with a certain probability to improve the global optimization ability. A dynamic adjustment mechanism is introduced to adapt to environmental changes and the emergence of threats. After each iteration, the utility matrix is updated according to the situation change, and the task exchange logic is re-executed to re-evaluate and dynamically adjust the assigned tasks. A composite maneuver decision mechanism is constructed. In a regular and stable environment, the reinforcement learning results are used to dynamically allocate exploration areas for individuals, and the states of multiple unmanned vehicles and environmental changes are dynamically and periodically evaluated. When specific conditions are met, a smooth transition to the security decision layer will occur. This transition structure will switch the decision layer with dynamically changed weights, and at the same time, the execution of the end planning will be handed over to the individual edge computing nodes throughout the process, separating the decision layer and the individual action space, and retaining the end decision autonomy of the single vehicle on the premise of taking into account the global coordination.
[0057] The above-described embodiments are only descriptions of the preferred embodiments of the present invention, and do not limit the scope of the present invention. Without departing from the design spirit of the present invention, various deformations and improvements made by those of ordinary skill in the art to the technical solutions of the present invention shall fall within the protection scope determined by the claims of the present invention.
Claims
1. A safety decision-making method for multi-unmanned vehicle inspection and scheduling in an adversarial environment, characterized in that Including: Step S1: Obtain environmental perception information through the scenario evaluation layer; Step S2: Input the environmental perception information into the decision-making layer to perform patrol scheduling planning based on reinforcement learning under a hierarchical decision-making execution mechanism; Step S3: The decision-making layer monitors the decision execution quality in real time and performs automatic recovery and retry.
2. The safety decision-making method for multi-unmanned vehicle inspection and scheduling in an adversarial environment according to claim 1, characterized in that, In Step S1, the scenario evaluation layer conducts a comprehensive analysis of the task environment through scenario complexity evaluation, environmental complexity analysis, and communication status evaluation.
3. The safety decision-making method for multi-unmanned vehicle inspection and scheduling in an adversarial environment according to claim 1, wherein In Step S2, the decision-making layer adopts a two-layer architecture design. The upper layer realizes utility calculation, bidding iteration, and task exchange optimization based on an improved greedy joint auction algorithm, and the lower layer realizes patrol scheduling based on a reinforcement learning framework. The safety decision-making method for multi-unmanned vehicle inspection scheduling in an adversarial environment according to claim 3, wherein In Step S3, the decision-making layer monitors the decision execution quality in real time through performance metric monitoring and failure recovery mechanisms and performs automatic recovery and retry.
Citation Information
Cited By
Production environment autonomous inspection method based on large model
CN120599714A
A large model-based production environment autonomous inspection method
CN120599714B