A cooperative inspection task allocation method and device based on multi-vehicle state perception and a storage medium
By constructing the target cost spectrum matrix through the central coordination unit and learning the strategy of the autonomous inspection unit, dynamic scheduling optimization of the multi-device collaborative inspection system is realized, which solves the efficiency problem of task scheduling and path optimization in the existing technology and improves the adaptability and efficiency of the inspection system.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHENZHEN TIANYOU SATELLITE APPL TECH CO LTD
- Filing Date
- 2025-12-17
- Publication Date
- 2026-04-17
AI Technical Summary
Existing multi-device collaborative inspection systems struggle to achieve efficient task scheduling and path optimization when handling data from different sources and implementing strategy management, especially when the environment changes and equipment status fluctuates, making it difficult to maintain the optimization of task scheduling.
The central coordination unit constructs a target cost spectrum matrix, combines the operating status and environmental characteristics of the autonomous inspection unit, generates a priority index table, and performs policy learning through a composite reward function. The autonomous inspection unit and the central coordination unit then coordinate policies cyclically to achieve dynamic scheduling optimization.
It improves inspection efficiency, maintains optimized task scheduling when the environment changes and equipment status fluctuates, reduces unnecessary movement and resource waste, and enhances overall inspection efficiency and the adaptability of strategies.
Smart Images

Figure CN121329094B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of automation control technology, and in particular to a collaborative inspection task allocation method, device and storage medium based on multi-vehicle status perception. Background Technology
[0002] With the increasing demand for automated and intelligent inspection in industrial parks, power transmission lines, energy storage stations, oil and gas facilities, mobile inspection equipment is gradually being used to replace or assist manual inspections. Mobile autonomous inspection units typically collect environmental information through vehicle-mounted sensors and perform tasks such as point inspection, data recording, and anomaly reporting.
[0003] In many application scenarios, multiple inspection units need to be deployed to cover a larger spatial area. To improve inspection efficiency, a scheduling center or management system is usually introduced to manage the operating status, task requirements, and environmental information of the inspection units, and to coordinate the task execution process of multiple inspection units.
[0004] Current multi-device collaborative execution methods typically rely on existing task planning, path arrangement, or strategy management mechanisms. The system assigns tasks and controls execution of inspection units based on preset objectives. During inspection execution, the system may adjust tasks or execution strategies according to actual conditions or management needs. Related technologies involve task planning, scheduling management, and operational feedback processing.
[0005] With improved data acquisition capabilities and increased system complexity, multi-device collaborative inspection systems may involve various operational information, such as equipment status, environmental characteristics, task information, and scheduling strategies. This information can influence the management methods, execution decisions, and collaborative strategies during the inspection process. How to comprehensively process data from different sources within the system to achieve multi-device inspection execution and strategy management has become a key area of consideration in related fields. Summary of the Invention
[0006] To address the aforementioned technical issues, this application provides a collaborative inspection task allocation method, apparatus, and storage medium based on multi-vehicle status perception.
[0007] The technical solution provided in this application is described below:
[0008] The first aspect of this application provides a collaborative inspection task allocation method based on multi-vehicle state perception, the method comprising:
[0009] The central coordination unit calculates the risk indicators of each inspection target node based on the environmental feature sets and historical event sequences uploaded by several autonomous inspection units, and obtains a priority index table.
[0010] The central coordination unit determines the target's original parameters based on the priority index table and the operating status vector of its respective main inspection unit.
[0011] The central coordination unit constructs a target cost spectrum matrix based on the original target parameters. The matrix elements in the target cost spectrum matrix reflect the comprehensive correlation cost between the autonomous inspection unit and the inspection target node.
[0012] In the central coordination unit, a set of operational constraints is applied to the target cost spectrum matrix, and constraint solving is performed to obtain an initial scheduling scheme;
[0013] Each main inspection unit takes its own operating state vector as input and the executable candidate migration direction in the action output set as the output behavior space, and constructs a composite reward function based on the task completion rate component, energy consumption penalty component and path passage safety component.
[0014] Within the unit task execution cycle, each main inspection unit updates the preset strategy value function according to the composite reward function to obtain a local strategy parameterized sequence;
[0015] Each primary inspection unit sends the local policy parameterization sequence to the central coordination unit, which then performs aggregation calculations on all received policy parameterization sequences to obtain global policy tuning parameters.
[0016] The central coordination unit distributes the global strategy tuning parameters to its respective main inspection units, enabling each main inspection unit to modify its local strategy parameterized sequence and composite reward function based on the global strategy tuning parameters.
[0017] Optionally, the central coordination unit constructs a target cost spectrum matrix based on the original target parameters. The matrix elements in the target cost spectrum matrix reflect the comprehensive correlation cost between the autonomous inspection unit and the inspection target node, including:
[0018] The original parameters of the target are normalized to form a set of feature vectors that describe the status of each main inspection unit and the characteristics of the inspection nodes.
[0019] The geometric spatial distance between each main inspection unit and each inspection target node is calculated based on the feature vector set, and the geometric spatial distance is corrected based on the path reachability parameter and the environmental obstacle density to obtain the path impedance factor.
[0020] Determine the operating state vector of each of the autonomous inspection units;
[0021] Calculate the task requirement tensor based on the risk and urgency indicators of each inspection target node obtained in advance;
[0022] The path impedance factor, running state vector and task requirement tensor are fused by the mapping function to generate the comprehensive cost element between each main inspection unit and each inspection target node.
[0023] The comprehensive cost elements are arranged according to the correspondence between their respective main inspection units and inspection target nodes to form a target cost spectrum matrix.
[0024] Optionally, in the central coordination unit, a set of operational constraints is applied to the target cost spectrum matrix, and constraint solving is performed to obtain an initial scheduling scheme, including:
[0025] The first constraint subset is generated based on the remaining energy value, load capacity and communication reachability parameters of each main inspection unit;
[0026] A second constraint subset is generated based on the risk indicators, inspection time windows, and geographical connectivity indicators of each inspection target node.
[0027] Construct a set of runtime constraint conditions based on the first constraint subset and the second constraint subset;
[0028] The constraint parameters in the set of operating constraints are normalized, and the matrix elements of the target cost spectrum matrix are used as the initial search weights.
[0029] The target cost spectrum matrix is mapped layer by layer by a preset multidimensional constraint projection operator to obtain the constraint clipping matrix;
[0030] Based on the constraint clipping matrix, a solution is obtained to obtain a set of locally feasible optimal solutions;
[0031] The optimal set of solution elements is determined from the set of locally feasible optimal solutions, and an initial scheduling scheme is generated based on the set of solution elements.
[0032] Optionally, each primary inspection unit uses its own operating state vector as input, the executable candidate migration directions in the action output set as the output behavior space, and constructs a composite reward function based on the task completion rate component, energy consumption penalty component, and path passage safety component, including:
[0033] The current operating state vector is used as input, and the operating state vector includes position parameters, velocity parameters, remaining energy value, and load capacity;
[0034] The executable candidate migration directions are defined as action output sets from the action output set, and the running state vector and the action output set are mapped to the behavior space;
[0035] In the behavior space, for each candidate action, the corresponding task completion rate component, energy consumption penalty component, and path passage safety component are calculated;
[0036] Based on the preset weighting coefficients, the task completion rate component, energy consumption penalty component, and path passage safety component are linearly fused to obtain a composite reward function.
[0037] Optionally, the step of updating the preset strategy value function according to the composite reward function within a unit task execution cycle to obtain a local strategy parameterized sequence includes:
[0038] The autonomous inspection unit acquires the running status vector S_t in the current task execution cycle and determines the set of executable actions A_t from the action output set;
[0039] For each candidate action a_i in the set of executable actions A_t, calculate the corresponding composite reward value R_t(a_i);
[0040] Based on the current running state vector S_t, candidate action a_i, and the composite reward value R_t(a_i), the preset strategy value function is updated using the following formula:
[0041] Q_new(S_t, a_i)=(1-η)·Q_old(S_t, a_i)+η·(R_t(a_i)+λ·maxQ_old(S_t+1, a′));
[0042] Where η is the learning rate, λ is the discount factor, a′ is any action in the set of actions that can be executed at the next time step, Q_old represents the value of the original preset policy value function, and Q_new represents the value of the preset policy value function obtained by updating.
[0043] Optionally, each primary inspection unit sends the local policy parameterized sequence to the central coordination unit, which then performs aggregation calculations on all received policy parameterized sequences to obtain global policy tuning parameters, including:
[0044] Each primary inspection unit transmits its local strategy parameterization sequence to the central coordination unit via a wireless communication link.
[0045] The central coordination unit indexes and categorizes the local policy parameterization sequences and constructs a policy parameter set matrix, wherein different local policy parameterization sequences corresponding to the same running state vector are stored in the same index position;
[0046] For different local strategy parameterization sequences corresponding to the same running state vector in the strategy parameter set matrix, the corresponding aggregate weighting coefficient is calculated according to the task completion rate component of each main inspection unit.
[0047] Based on the aggregation weighting coefficients, the central coordination unit performs aggregation calculations on different local policy parameterization sequences corresponding to the same running state vector to obtain global policy tuning parameters.
[0048] Optionally, the central coordination unit distributes the global policy tuning parameters to their respective main inspection units, enabling each autonomous inspection unit to modify its local policy parameterized sequence and composite reward function based on the global policy tuning parameters, including:
[0049] The central coordination unit encapsulates the global policy tuning parameters into policy update instructions and sends them to the autonomous inspection unit via a wireless communication link.
[0050] The autonomous inspection unit extracts global policy tuning parameters from the policy update instruction;
[0051] The local policy parameterization sequence corresponding to the autonomous inspection unit is corrected and calculated with the global policy tuning parameters. The correction calculation is performed using the following formula:
[0052] Q_revised(s,a)=δ·Q_local(s,a)+(1-δ)·Q_global(s,a);
[0053] Where Q_revised represents the modified local policy parameterization sequence corresponding to the running state vector s and candidate action a; Q_global represents the global policy tuning parameters corresponding to the running state vector s and candidate action a; and δ is the self-preservation weight coefficient.
[0054] The composite reward function is recalculated based on the modified local policy parameterized sequence.
[0055] A second aspect of this application provides a collaborative inspection task allocation device based on multi-vehicle state perception, the device comprising:
[0056] The data acquisition unit is used to calculate the risk indicators of each inspection target node and obtain a priority index table based on the environmental feature sets and historical event sequences uploaded by several autonomous inspection units.
[0057] The parameter determination unit is used to determine the target original parameters based on the priority index table and the running status vector of each main inspection unit.
[0058] The matrix construction unit is used to construct a target cost spectrum matrix based on the original target parameters. The matrix elements in the target cost spectrum matrix reflect the comprehensive correlation cost between the autonomous inspection unit and the inspection target node.
[0059] The constraint application unit is used to apply a set of running constraints to the target cost spectrum matrix and perform constraint solving to obtain an initial scheduling scheme.
[0060] The function building unit is used to construct a composite reward function based on the task completion rate component, energy consumption penalty component, and path passage safety component, with its own running state vector as input and the executable candidate migration direction in the action output set as the output behavior space.
[0061] The function update unit is used to update the preset strategy value function according to the composite reward function within a unit task execution cycle to obtain a local strategy parameterized sequence.
[0062] An aggregation calculation unit is used to send the local policy parameterization sequence to the central coordination unit, which then performs aggregation calculations on all received policy parameterization sequences to obtain global policy tuning parameters.
[0063] The function correction unit is used to distribute the global policy tuning parameters to their respective main inspection units, so that each main inspection unit can correct its own local policy parameterized sequence and composite reward function according to the global policy tuning parameters.
[0064] A third aspect of this application provides a collaborative inspection task allocation device based on multi-vehicle state perception, the device comprising:
[0065] Processor, memory, input / output units, and bus;
[0066] The processor is connected to the memory, the input / output unit, and the bus;
[0067] The memory stores a program, which the processor invokes to execute the first aspect and any one of the optional methods in the first aspect.
[0068] A fourth aspect of this application provides a computer-readable storage medium on which a program is stored, which, when executed on a computer, performs the methods of the first aspect and any one of the first aspects.
[0069] As can be seen from the above technical solutions, this application has the following beneficial effects:
[0070] 1. The central coordination unit determines inspection priorities based on environmental characteristics and operational status, and constructs a target cost spectrum matrix for task scheduling, while the autonomous inspection unit learns local strategies through a composite reward function. A cyclical coordination relationship is formed between the two, so that scheduling no longer depends on fixed plans, but evolves gradually with the operation process.
[0071] 2. The autonomous inspection unit updates its local policy parameterization sequence based on task execution feedback and uploads it to the central coordination unit. The central coordination unit performs aggregate calculations on the policy information from multiple inspection units, extracts global policy optimization parameters, and distributes them back to their respective master inspection units. Through this mechanism, each inspection unit can obtain global perspective compensation under local view limitations, significantly improving overall inspection efficiency.
[0072] 3. Through iterative learning and strategy correction, continuous optimization can be achieved as the environment changes. For example, when the risk of an inspection scenario changes or the operating status of equipment fluctuates, the strategy updates of each main inspection unit will be integrated in real time by the central coordination unit, thereby ensuring that task scheduling and path strategies remain optimal or suboptimal.
[0073] 4. Because the composite reward function takes into account the task completion rate, energy consumption and path safety, the autonomous inspection unit will actively select the inspection path with higher cost performance, rather than just executing according to distance or task order, thus avoiding ineffective movement and waste of resources. Attached Figure Description
[0074] To more clearly illustrate the technical solutions in this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0075] Figure 1 This is a flowchart illustrating an embodiment of the collaborative inspection task allocation method based on multi-vehicle state awareness provided in this application.
[0076] Figure 2 This is a flowchart illustrating an implementation of step S103 in the collaborative inspection task allocation method based on multi-vehicle state awareness provided in this application.
[0077] Figure 3 This is a flowchart illustrating an implementation of step S104 in the collaborative inspection task allocation method based on multi-vehicle state awareness provided in this application.
[0078] Figure 4 This is a flowchart illustrating an implementation of step S105 in the collaborative inspection task allocation method based on multi-vehicle state awareness provided in this application.
[0079] Figure 5 This is a flowchart illustrating an implementation of step S107 in the collaborative inspection task allocation method based on multi-vehicle state awareness provided in this application.
[0080] Figure 6This is a schematic diagram of an embodiment of the collaborative inspection task allocation device based on multi-vehicle state perception provided in this application;
[0081] Figure 7 This is a schematic diagram of an embodiment of a collaborative inspection task allocation device based on multi-vehicle state perception provided in this application. Detailed Implementation
[0082] The various steps involved in the embodiments of the present invention can be implemented by software, hardware, or a combination of both. It should be noted that the embodiments of the present invention do not limit the subject performing the task.
[0083] The “central coordination unit” can be a server, an edge computing node, a cloud computing module, or any processing device with task coordination capabilities.
[0084] The "autonomous inspection unit" can be a mobile robot, intelligent inspection vehicle, unmanned vehicle, unmanned boat, or automated equipment with mobility.
[0085] In different specific application environments, the form, number, and functional module division of execution units may vary, but as long as they can complete the data acquisition, task allocation, policy learning, and policy synchronization processes described in this invention, they are considered to fall within the protection scope of this invention. This invention does not limit the physical structure, network topology, or installation location of the execution entity.
[0086] Unless otherwise specified, the operations such as "processing," "computing," "executing," "updating," and "merging" mentioned in the embodiments of this invention refer to the actions of a computing program running on a processor to process data. The processor mentioned above can be a central processing unit (CPU), graphics processing unit (GPU), programmable logic device (FPGA), application-specific integrated circuit (ASIC), or other chips with computing capabilities. This invention does not limit the specific processing platform used.
[0087] Therefore, this invention should be understood as a method invention, and the form of its implementation carrier does not constitute a limitation on the scope of protection of this invention.
[0088] To facilitate understanding of the embodiments of the present invention by those skilled in the art, some terms used in this specification are explained below. These explanations of terminology are for illustrative purposes only and do not constitute a limitation on the terminology used.
[0089] 1. A Central Coordination Unit (CCU) is a functional entity that receives data, performs calculations, and issues control policies. This unit can be implemented by a server, edge computing device, or a terminal with processing capabilities, or by a collaborative cluster of multiple processing entities.
[0090] 2. An Autonomous Inspection Unit (AMU) refers to an intelligent terminal capable of autonomous movement, performing inspection actions, and collecting environmental data. This terminal can be a mobile robot, unmanned vehicle, drone, or other device capable of performing inspection tasks.
[0091] 3. The Environmental Feature Set refers to the set of data related to the environment of the inspection area collected and reported by the autonomous inspection unit, such as temperature characteristics, visible light image characteristics, obstacle information, etc., which are used to represent the current status of the inspection environment.
[0092] 4. Historical Event Sequence refers to historical data related to the inspection target recorded by the system, including anomaly occurrence records, alarm time series, fault event records, etc.
[0093] 5. The Priority Index Table refers to a list of target nodes sorted according to their risk indicators or importance, used to indicate the processing priority of different inspection targets.
[0094] 6. The operating state vector refers to the set of values used to describe the current operating state of the autonomous inspection unit, including but not limited to state variables such as position, velocity, remaining energy, and current load.
[0095] 7. Primitive Object Metrics refer to the basic values used to build the scheduling optimization model, which are calculated by the central coordination unit based on the inspection target and the status of the inspection unit. Examples include inspection distance estimation, resource availability, and risk exposure.
[0096] 8. The Objective Cost Spectrum Matrix is a matrix used to describe the comprehensive cost relationship between inspection units and inspection targets. The elements in the matrix reflect the cost required to execute a certain task node.
[0097] 9. The Operational Constraint Set refers to the set of constraints on the scheduling results, such as energy constraints, task time constraints, and resource availability constraints.
[0098] 10. Action Set refers to the set of actions that the autonomous inspection unit can perform in its current state, such as the direction it can move in next or the actions it can take.
[0099] 11. The composite reward function is a function used to evaluate the comprehensive reward value obtained by the inspection unit after performing the action. Its reward value is composed of multiple reward components, such as task completion rate reward, energy consumption penalty, safety reward, etc.
[0100] 12. The policy value function / Q-value (Action Value Function) is used to evaluate the expected long-term reward obtained by choosing a certain action in a certain state. It is the core computational resource in policy learning algorithms (such as Q-learning).
[0101] 13. The Local Parameterized Strategy Vector refers to the set of strategy parameters updated by the autonomous inspection unit based on local execution feedback, which consists of its strategy value function or the corresponding parameter expression structure.
[0102] 14. Global Strategy Adjustment Parameters refer to the strategy parameters used for global optimization obtained by the central coordination unit after fusing and calculating the local strategy parameter sets uploaded by multiple autonomous inspection units.
[0103] Please see Figure 1 This application first provides an embodiment of a collaborative inspection task allocation method based on multi-vehicle state perception, which includes:
[0104] S101. The central coordination unit calculates the risk indicators of each inspection target node based on the environmental feature set and historical event sequence uploaded by several autonomous inspection units, and obtains a priority index table.
[0105] In this embodiment, the central coordination unit periodically or in real time receives environmental information from multiple autonomous inspection units, including but not limited to: sensor measurements of temperature, humidity, vibration, current, gas concentration, etc. near the node; the node's recent alarm events, maintenance records, and abnormal trend changes; and historical probability values or predicted outputs of node risks.
[0106] The central coordination unit analyzes the above data to determine the risk indicators for each inspection target node. For example, it can statistically analyze the gas concentration change trend at a certain node and combine this with the number of historical abnormal events to generate a risk score. Risk indicators can be numerical, for example:
[0107] Node A: Risk value = 0.82;
[0108] Node B: Risk value = 0.34;
[0109] The central coordination unit sorts the inspection nodes according to their risk values, resulting in a priority index table. This priority index table is used to determine which node should perform the inspection task first.
[0110] In one embodiment, the environmental feature set can be encapsulated in vector form, for example:
[0111] Env_Feature = [T, H, Gas, Vib, Current, ...];
[0112] Historical event sequences may include: whether the node has ever issued an abnormal alarm; whether emergency maintenance has occurred;
[0113] The node exhibits abnormal timing patterns (e.g., triggering an alarm at fixed times every week).
[0114] The sequence of historical events can be represented as:
[0115] Event_Seq = { (t1, type1), (t2, type2), ...};
[0116] Ultimately, risk indicators are formed, such as:
[0117] Risk_Index(Node_i) = f(current features, historical event statistics, node importance);
[0118] The nodes are sorted according to the risk rules to obtain a priority index table:
[0119] Priority = [Node_A, Node_C, Node_D, Node_B, ...].
[0120] S102. The central coordination unit determines the target original parameters based on the priority index table and the running status vector of its respective main inspection unit.
[0121] In step S102, the central coordination unit determines the target raw parameters for subsequent scheduling calculations based on the priority index table and the operating status vectors of each autonomous inspection unit. Specifically, the central coordination unit first receives real-time operating status vectors from multiple autonomous inspection units. These operating status vectors are a set of data continuously updated by the inspection units during execution, typically including parameters such as current physical location coordinates, remaining power or fuel, current task load level, and equipment hardware and software operating status. The operating status vectors reflect the current executable capabilities and physical resource status of the inspection units; for example, remaining energy characterizes the vehicle's range, and the current task load reflects its schedulable resource reserves.
[0122] The central coordination unit integrates the aforementioned operational status with the priority index table. The priority index table sorts the inspection target nodes according to risk indicators. During scheduling, the central coordination unit calculates spatial parameters such as spatial distance and path reachability between the autonomous inspection unit and each node based on the geographical or path topology relationship between the operational status vector and the inspection target nodes. Simultaneously, the central coordination unit assesses the matching degree between the inspection difficulty of the node and the vehicle's execution capability by combining operational capability indicators such as the remaining energy and real-time task load of the autonomous inspection unit. For example, if node A has a high inspection priority, and if inspection unit 1 has a short straight-line distance to node A and sufficient current energy, while inspection unit 2 has idle resources but faces obstacles or path restrictions with node A, then the central coordination unit sets the initial execution tendency of node A to inspection unit 1. Through the above calculations, the central coordination unit obtains the target raw parameters representing the correlation between the inspection unit and the target node, including but not limited to numerical parameters characterizing execution costs such as location distance value, energy availability coefficient, estimated inspection time, and path accessibility level.
[0123] S103. The central coordination unit constructs a target cost spectrum matrix based on the original target parameters. The matrix elements in the target cost spectrum matrix reflect the comprehensive correlation cost between the autonomous inspection unit and the inspection target node.
[0124] Based on the original target parameters, the central coordination unit constructs a matrix, where each matrix element corresponds to "the cost of an autonomous inspection unit performing an inspection task at a certain node". The smaller the value, the lower the execution cost.
[0125] The cost can be calculated based on weights such as the running state vector and node risk. There are no restrictions on the calculation method or the numerical form of the cost value.
[0126] Please see Figure 2 In a specific embodiment, step S103 can be implemented in the following way:
[0127] S1031. Normalize the original parameters of the target to form a feature vector set for describing the status of each main inspection unit and the characteristics of the inspection node.
[0128] In this embodiment, the central coordination unit preprocesses the priority index table obtained in step S101 and the operational status data reported by each of the main inspection units. Data from different sources (such as energy percentage, motor temperature, node risk value, geographical coordinate distance, etc.) have different measurement dimensions and orders of magnitude. If directly used for calculation, some high-order data will dominate the overall evaluation. Therefore, the central coordination unit performs normalization processing on all raw parameters, such as using Min-Max normalization or Z-score standardization, so that all feature dimensions are linearly scaled to the same dimension range. The normalized dataset forms a feature vector set used to describe the status of autonomous inspection units and the characteristics of inspection nodes. This feature vector set characterizes the vehicle's executability and the intensity of node inspection demand.
[0129] S1032. Calculate the geometric spatial distance between each main inspection unit and each inspection target node based on the feature vector set, and correct the geometric spatial distance based on the path reachability parameter and the environmental obstacle density to obtain the path impedance factor.
[0130] The central coordination unit calculates the geometric spatial distance from its main inspection unit to each inspection target node based on the location coordinate parameters in the feature vector set. This distance can be obtained, for example, through Euclidean distance or geographic network distance. In some inspection areas, there may be obstacles such as pipe corridors, fences, and restricted areas. The central coordination unit corrects the geometric distance based on environmental data such as regional obstacle density and path congestion, so that the path distance can reflect accessibility and execution difficulty.
[0131] The corrected distance data is defined as the path impedance factor, representing the path cost for a vehicle to perform an inspection from its current location to the target node. For example, the path impedance factor will increase significantly when there is a narrow passage between the inspection vehicle and the node.
[0132] S1033. Determine the operating state vector of each of the autonomous inspection units;
[0133] The central coordination unit generates an operational status vector based on the operational status vectors of the inspection units (remaining energy, load status, sensor availability, etc.). The operational status vector describes the vehicle's task execution capability, for example:
[0134] Remaining battery power indicates the vehicle's ability to conduct continuous inspections.
[0135] Load capacity indicates whether a vehicle is able to continue performing inspection tasks.
[0136] The operational status of sensors affects the availability of task execution.
[0137] The operating state vector can be understood as a combination of vectors describing the vehicle's current effective execution capability.
[0138] S1034. Calculate the task requirement tensor based on the risk indicators and urgency indicators of each inspection target node obtained in advance.
[0139] The central coordination unit quantifies the inspection needs of each inspection node based on risk indicators, historical alarm density, and real-time temperature anomalies or sound vibrations, thus forming a task requirement tensor. The task requirement tensor is used to characterize the importance and difficulty of the inspection target, for example:
[0140] High-risk nodes have higher demand intensity;
[0141] Nodes exhibiting an alarm trend have a higher urgency indicator.
[0142] S1035. The path impedance factor, running state vector and task requirement tensor are fused by the mapping function to generate the comprehensive cost element between each main inspection unit and each inspection target node.
[0143] The central coordination unit uses a pre-defined mapping function (such as a multi-index weighting function or a fuzzy decision function) to fuse the path impedance factor, the operational state vector, and the task requirement tensor. The fusion result outputs a value representing the "comprehensive cost of a vehicle performing a task at a certain inspection node," which is the comprehensive cost element between the inspection vehicle and the target node. The smaller the comprehensive cost, the more suitable the inspection vehicle is for performing the inspection task at the corresponding node.
[0144] For example, if vehicle 1 is closer to node A, but has low remaining battery or a congested route, the overall cost may be greater than that of another vehicle that is slightly farther away but has more resources.
[0145] S1036. Arrange the comprehensive cost elements according to the correspondence between their respective main inspection units and inspection target nodes to form a target cost spectrum matrix.
[0146] Finally, the central coordination unit organizes all comprehensive cost elements into a matrix according to the two-way mapping relationship between "vehicle and inspection node". Each row of the matrix corresponds to an autonomous inspection unit, each column corresponds to an inspection target node, and each unit value represents the comprehensive cost of the vehicle performing the task at that node.
[0147] To make the technical solution of the present invention clearer, the data structure involved in this embodiment will now be described.
[0148] The central coordination unit constructs an Objective Cost Spectrum Matrix based on the set of inspection target nodes and the set of autonomous inspection units. This matrix represents the "comprehensive cost of a certain inspection vehicle performing a task at a certain node." Its data structure can be represented as follows:
[0149] ;
[0150] Where m represents the number of autonomous inspection units, and n represents the number of inspection target nodes. This represents the comprehensive associated cost of autonomous inspection unit i performing inspection target j.
[0151] The value of a product has the following structure:
[0152] ;
[0153] in, This represents the estimated path distance from the autonomous inspection unit to the inspection target node. This indicates the remaining energy of the autonomous inspection unit. Indicates the risk level or priority of the target node to be inspected. This indicates the current load, status, and other operating parameters of the autonomous inspection unit.
[0154] Each autonomous inspection unit maintains a dynamically updated operating state vector:
[0155] ;
[0156] Where, pos i This represents a position parameter, which can be a two-dimensional or three-dimensional coordinate numerical vector, vel. i This represents a velocity parameter, which can include direction and magnitude; energy i This represents the remaining energy parameter, which can be a percentage or an absolute amount of electricity. (load) i A numerical representation indicating the current load or task backlog.
[0157] S104. In the central coordination unit, a set of operational constraints is applied to the target cost spectrum matrix, and constraint solving is performed to obtain an initial scheduling scheme.
[0158] In this embodiment, the central coordination unit may introduce constraints, including but not limited to:
[0159] Each inspection vehicle is assigned only one task at a time;
[0160] High-risk nodes must be inspected within a limited time.
[0161] Avoid inspection vehicles continuously running on high-energy-consuming routes.
[0162] The central coordination unit can obtain the task allocation results by solving this constrained optimization problem (without limiting the solution algorithm).
[0163] Please see Figure 3 In one specific embodiment, step S104 can be implemented through the following specific implementation:
[0164] S1041. Generate the first constraint subset based on the remaining energy value, load capacity and communication reachability parameters of each main inspection unit;
[0165] The central coordination unit first reads the operational status vector of each autonomous inspection unit and extracts parameters directly related to the vehicle's execution capabilities, such as remaining energy value, current load capacity, and communication link reachability. The central coordination unit then filters this data to determine whether the vehicle possesses the basic conditions to execute the target inspection node task, for example:
[0166] If the remaining energy is insufficient to complete the round trip, the vehicle marks the inspection node as unreachable.
[0167] If the load has reached its maximum capacity, the vehicle will not participate in new task assignments.
[0168] If a communication blind spot exists between a vehicle and the central coordination unit, the vehicle cannot perform inspection tasks in real time. Based on these judgment rules, the central coordination unit constructs a first subset of constraints to restrict the eligibility conditions for vehicles to participate in task scheduling.
[0169] S1042. Generate a second constraint subset based on the risk indicators, inspection time windows, and geographical connectivity indicators of each inspection target node.
[0170] The central coordination unit identifies task-related attribute parameters from the inspection node data associated with the target cost spectrum matrix, including: risk indicators, inspection time windows (e.g., must be completed within a specified time period), and the geographical connectivity of the inspection path. For example:
[0171] Some nodes require completion during device runtime;
[0172] Some nodes must be accessed via specific areas due to path obstacles;
[0173] The central coordination unit constructs a second constraint subset based on the aforementioned node attributes to restrict which inspection nodes a vehicle can perform.
[0174] S1043. Construct a set of runtime constraint conditions based on the first constraint subset and the second constraint subset;
[0175] The central coordination unit merges the two constraint subsets generated in steps S1041 and S1042 to obtain a unified set of operational constraints. Since the parameters from different constraint sources are of different orders of magnitude (e.g., communication signal strength and energy percentage differ significantly), the central coordination unit performs normalization processing on all parameters in the constraint set to ensure that all constraints are within the same order of magnitude during calculation.
[0176] The normalized constraint values are used to adjust the weights of the values in the target cost spectrum matrix, so that the target cost spectrum matrix not only expresses the cost of the vehicle performing the task, but also whether the constraint conditions are met.
[0177] The constraint parameters in the set of operating constraints are normalized, and the matrix elements of the target cost spectrum matrix are used as the initial search weights.
[0178] S1044. The target cost spectrum matrix is mapped layer by layer using a preset multidimensional constraint projection operator to obtain the constraint clipping matrix;
[0179] The central coordination unit performs a multidimensional constrained projection operation on the target cost spectrum matrix.
[0180] The projection process can be understood as:
[0181] Multiple constraints are mapped to the corresponding matrix elements of the target cost spectrum matrix, and elements that do not meet the constraints are penalized or directly masked. The projection operation outputs a pruned matrix, called the constraint pruning matrix. This constraint pruning matrix eliminates the possible task assignments that do not meet the constraints, retaining only the vehicle-node combinations that satisfy the execution conditions.
[0182] S1045. Based on the constraint clipping matrix, perform the solution to obtain a locally feasible optimal solution set;
[0183] The central coordination unit takes the constraint pruning matrix as input and calls the internal solution engine to perform optimization. During the solution process, it outputs several candidate scheduling schemes that satisfy all constraints, which constitute a locally feasible optimal solution set. For example, when the number of vehicles exceeds the number of inspection nodes, the solution engine may generate multiple equivalent execution schemes.
[0184] S1046. Determine the optimal solution set from the local feasible optimal solution set, and generate an initial scheduling scheme based on the solution set.
[0185] The central coordination unit selects the solution with the lowest overall cost from the set of locally feasible optimal solutions, i.e., the optimal tuple. Based on the vehicle-node mapping relationship of each vehicle in this tuple, the central coordination unit generates an initial scheduling plan and distributes it to each vehicle. At this point, each inspection unit obtains its first target inspection node and enters the next stage of reinforcement learning path decision-making.
[0186] S105. Each main inspection unit takes its own operating state vector as input, takes the executable candidate migration direction in the action output set as the output behavior space, and constructs a composite reward function based on the task completion rate component, energy consumption penalty component and path passage safety component.
[0187] The autonomous inspection unit selects a feasible direction of movement from its set of actions based on its own status (position, energy, speed, etc.), for example:
[0188] "Go straight", "Turn left", "Turn right", "Continue to execute node task".
[0189] The autonomous inspection unit calculates the reward based on the effect of the action:
[0190] Task completion rate component: Successfully executing the planned task will result in a positive reward;
[0191] Energy penalty component: Prioritize low-energy paths to reduce negative rewards;
[0192] Path safety component: Gain additional rewards by avoiding high-risk areas.
[0193] Specifically, in one embodiment, step S105 can be implemented using a reinforcement learning decision model. This model takes the real-time operating status of the autonomous inspection unit as input and outputs action decisions for performing path migration. Please refer to [link to relevant documentation]. Figure 4 One embodiment of step S105 includes:
[0194] S1051. The current operating state vector is used as input, and the operating state vector includes position parameters, velocity parameters, remaining energy value and load capacity.
[0195] S1052. Define the executable candidate migration directions from the action output set as the action output set, and map the running state vector and the action output set to the behavior space;
[0196] At time t, the central coordination unit acquires the current inspection unit's operational state vector, encapsulating parameters such as current position coordinates, current speed, remaining energy, and current load capacity as state inputs. These state inputs characterize the inspection unit's real-time physical state and operational capabilities. Subsequently, the central coordination unit selects executable migration directions from the set of action outputs generated for this inspection node, such as left / right / forward / backward / around, ensuring that each action satisfies the vehicle's physical constraints and task framework requirements. The central coordination unit maps the operational state vector and the candidate action set together to the reinforcement learning behavior space, forming state-action pairs for decision-making and reasoning.
[0197] S1053. In the behavior space, calculate the corresponding task completion rate component, energy consumption penalty component and path passage safety component for each candidate action.
[0198] In the behavior space, for each candidate action to be evaluated, the central coordination unit calculates three evaluation components: task completion rate, energy consumption penalty, and path safety. The task completion rate component measures whether the action enables the inspection unit to move towards the target node and whether it helps shorten the time to reach the task objective. The energy consumption penalty component estimates the energy consumed in performing the action and serves as a negative incentive; it generates a penalty value when the action leads to a decrease in remaining energy or poor energy efficiency. The path safety component quantifies the obstacle density, risk value, or traffic feasibility of the path involved in the current action; it decreases when the action causes vehicles to enter obstacle areas or high-risk areas.
[0199] S1054. Based on the preset weighting coefficients, the task completion rate component, energy consumption penalty component, and path passage safety component are linearly fused to obtain a composite reward function.
[0200] The central coordination unit linearly fuses the three evaluation components based on preset weighting coefficients to form a composite reward function. This composite reward function comprehensively reflects the combined benefits of candidate actions in terms of task advancement, energy utilization, and execution safety. The reinforcement learning model ranks or samples candidate actions based on the composite reward function to output the optimal action, thereby enabling the autonomous inspection unit to make real-time path migration decisions.
[0201] S106. Within the unit task execution cycle, each main inspection unit updates the preset strategy value function according to the composite reward function to obtain a local strategy parameterized sequence.
[0202] During execution, the inspection vehicle updates the Q-value table or other strategy model structures in the strategy model based on the reward values obtained.
[0203] If turning left brings high rewards, then the strategy value of the "turn left" action is increased.
[0204] The updated policy model is compressed into a local policy parameterized sequence for later uploading.
[0205] Specifically, the implementation of step S106 includes: the autonomous inspection unit obtains the running state vector S_t in the current task execution cycle, and determines the executable action set A_t from the action output set; for each candidate action a_i in the executable action set A_t, calculates the corresponding composite reward value R_t(a_i); and updates the preset strategy value function according to the current running state vector S_t, the candidate action a_i and the composite reward value R_t(a_i).
[0206] Specifically, in this embodiment, within a unit task execution cycle, each autonomous inspection unit dynamically updates the preset strategy value function based on the composite reward function constructed in step S105 to continuously improve the decision-making strategy. Specifically, the autonomous inspection unit first obtains a state vector S_t from the currently collected operating state. This state vector includes quantitative indicators describing the vehicle's current operating state, such as position parameters, speed parameters, remaining energy value, and load capacity. Subsequently, based on the strategy model, the autonomous inspection unit selects the set of executable actions A_t from the action output set at the current moment.
[0207] For each candidate action a_i in the action set A_t, the autonomous inspection unit calls the composite reward function to calculate the composite reward value R_t(a_i) obtained by executing the action. This reward value reflects the comprehensive performance of the action in terms of task completion rate improvement, energy consumption penalty, and path safety. Based on the calculated reward value, the autonomous inspection unit further updates the policy value of each candidate action using the policy value function. In a preferred embodiment, the policy value function can be updated using a Q-learning-based update method, through the following formula:
[0208] The preset policy value function is updated by the following formula: Q_new(S_t, a_i)=(1-η)·Q_old(S_t, a_i)+η·(R_t(a_i)+λ·maxQ_old(S_t+1, a′));
[0209] Where η is the learning rate, λ is the discount factor, a′ is any action in the set of actions that can be executed at the next time step, Q_old represents the value of the original preset policy value function, and Q_new represents the value of the preset policy value function obtained by updating.
[0210] The above formula indicates that if a candidate action a_i obtains a higher composite reward value in the current state and can lead to a better choice of subsequent actions, then the policy value corresponding to the action will be increased, making the action more likely to be selected in subsequent states; otherwise, the policy value of the action will be decreased.
[0211] As multiple execution cycles iterate, the policy values of the autonomous inspection unit under different operating states will gradually converge, thereby enabling the policy model to gradually acquire better action selection capabilities. After completing the current task execution cycle, the updated policy value function will be extracted into a local policy parameterized sequence through parameter compression and uploaded to the central coordination unit for global policy fusion.
[0212] S107. Each primary inspection unit sends the local policy parameterization sequence to the central coordination unit, which then performs aggregation calculations on all received policy parameterization sequences to obtain global policy tuning parameters.
[0213] In this embodiment, the central coordination unit merges the local strategies submitted by all inspection vehicles, for example:
[0214] It can average or weightedly combine the strategies under the same state to output global strategy tuning parameters.
[0215] Please see Figure 5 In one specific embodiment, step S107 is implemented as follows:
[0216] S1071, Each independent inspection unit sends the local strategy parameterization sequence to the central coordination unit through the wireless communication link;
[0217] In this embodiment, after updating their local policy parameterization sequence, each autonomous inspection unit uploads the policy parameter data to the central coordination unit for global policy fusion processing. Specifically, each autonomous inspection unit first sends the local policy parameterization sequence generated in the current task cycle to the central coordination unit via a wireless communication link (e.g., based on 5G, Wi-Fi Mesh, or a dedicated wireless communication protocol). This policy parameter sequence contains policy value update information corresponding to different operating state vectors and represents the strategic knowledge expression of the autonomous inspection unit after autonomous learning.
[0218] S1072. The central coordination unit indexes and classifies the local strategy parameterized sequence and constructs a strategy parameter set matrix, wherein different local strategy parameterized sequences corresponding to the same running state vector are stored in the same index position.
[0219] After receiving local policy parameterization sequences from multiple autonomous inspection units, the central coordination unit categorizes and organizes all received data according to preset indexing rules. Specifically, the central coordination unit constructs a policy parameter set matrix using the running state vector as the index key, ensuring that policy parameterization sequences from different inspection units with the same running state vector (corresponding to the same decision scenario in the same state space) are stored at the same index position in the matrix. This results in a structured policy parameter set with the state vector as the row index and the inspection unit number as the column index.
[0220] S1073. For different local strategy parameterization sequences corresponding to the same running state vector in the strategy parameter set matrix, calculate the corresponding aggregate weighting coefficient according to the task completion rate component of each main inspection unit.
[0221] In this step, the central coordination unit assigns an aggregation weighting coefficient to each local strategy parameterized sequence based on the task completion rate, energy efficiency, or other performance indicators of different inspection units within the current task cycle. The weighting coefficient represents the degree of contribution of the inspection unit's local strategy to the global strategy fusion process. For example, inspection units with higher task execution efficiency and better path selection will have a higher aggregation weighting coefficient.
[0222] S1074. Based on the aggregation weighting coefficient, the central coordination unit performs aggregation calculation on different local policy parameterization sequences corresponding to the same running state vector to obtain global policy tuning parameters.
[0223] Based on the aforementioned aggregation weighting coefficients, the central coordination unit performs aggregation operations on multiple local policy parameterized sequences located under the same running state vector in the policy parameter set matrix to obtain a global policy tuning parameter. This global policy tuning parameter represents the fusion of the learning results of multiple inspection units and serves as the policy basis for the entire system to update policies in the next round or to be distributed to all inspection units for autonomous inspection.
[0224] S108. The central coordination unit distributes the global strategy tuning parameters to its respective main inspection units, so that each main inspection unit can modify its local strategy parameterized sequence and composite reward function according to the global strategy tuning parameters.
[0225] In step S108, the central coordination unit encapsulates the global policy tuning parameters obtained through aggregation calculation into a policy update package and distributes it to each of its respective master inspection units, enabling each master inspection unit to correct and adjust its local policy parameterized sequence and composite reward function accordingly. Specifically, after generating the global policy tuning parameters, the central coordination unit encapsulates the parameters along with corresponding metadata (including parameter version number, generation timestamp, applicable state space index range, effective policy, and rollback indication, etc.) and sends it to each of its respective master inspection units through a reliable wireless communication channel.
[0226] Upon receiving the policy update package, each primary inspection unit first verifies the integrity and version information of the parameter package, confirming that the update package is applicable to the current state space index range of its unit. Subsequently, for the policy parameters at the same state-action index position, each primary inspection unit merges and corrects the locally stored local policy parameterization sequence.
[0227] In an optional embodiment,
[0228] The local policy parameterization sequence corresponding to the autonomous inspection unit is corrected and calculated with the global policy tuning parameters. The correction calculation is performed using the following formula:
[0229] Q_revised(s,a)=δ·Q_local(s,a)+(1-δ)·Q_global(s,a);
[0230] Where Q_revised represents the modified local policy parameterization sequence corresponding to the running state vector s and candidate action a; Q_global represents the global policy tuning parameters corresponding to the running state vector s and candidate action a; and δ is the self-preservation weight coefficient.
[0231] In addition to directly fusing the Q-value, the global policy tuning parameters issued by the central coordination unit can also include suggested adjustments to the weights of each reward component of the composite reward function. Based on these suggestions, the receiving end adaptively adjusts the weight coefficients of its local composite reward function (e.g., task completion rate weight α, energy consumption penalty weight β, and safety weight γ), thereby changing the target preferences for future learning. For example, when the global policy indicates that the overall task progress lags behind expectations, the local system can appropriately increase the weight of α to emphasize task completion efficiency; when an increase in overall environmental risk is detected, γ can be increased to prioritize safety. Weight adjustments can be implemented using a simple weighted update or a closed-loop adjustment formula based on error feedback, ensuring convergence and stability in weight updates.
[0232] Each independent inspection unit should verify the correction results locally, including but not limited to: performing several decision-making steps through simulation to evaluate the immediate effect of the new strategy in the local environment, or triggering trial operation in low-risk scenarios; if the verification fails, the previous parameters can be restored according to the rollback instructions in the strategy package, or an anomaly can be reported to the central coordination unit for review. To reduce communication overhead and ensure real-time performance, strategy updates can be applied on demand or synchronized periodically: for urgent adjustments (such as an increase in security weight), the strategy should be prioritized for immediate distribution and application; for non-urgent global optimization parameters, batch updates can be performed when the vehicle is idle or returns to the base station.
[0233] After fusion and weight adjustment are completed, each independent inspection unit stores the revised local policy parameterized sequence and the revised composite reward function into its local policy cache, and makes decisions and selects actions based on the revised policy in the next task execution cycle. At the same time, each independent inspection unit will send the revised performance indicators (such as local task completion rate components, energy consumption changes, and number of security events) back to the central coordination unit according to the reporting policy set by the system.
[0234] The foregoing embodiments describe the collaborative inspection task allocation method based on multi-vehicle state perception provided in this application. The embodiments of the device provided in this application are described below.
[0235] See Figure 6 This application provides an embodiment of a collaborative inspection task allocation device based on multi-vehicle state perception, which includes:
[0236] The data acquisition unit 601 is used to calculate the risk indicators of each inspection target node and obtain a priority index table based on the environmental feature set and historical event sequence uploaded by several autonomous inspection units.
[0237] The parameter determination unit 602 is used to determine the target original parameters based on the priority index table and the running status vector of each main inspection unit.
[0238] The matrix construction unit 603 is used to construct a target cost spectrum matrix based on the target original parameters. The matrix elements in the target cost spectrum matrix reflect the comprehensive correlation cost between the autonomous inspection unit and the inspection target node.
[0239] The constraint application unit 604 is used to apply a set of running constraints to the target cost spectrum matrix and perform constraint solving to obtain an initial scheduling scheme;
[0240] The function construction unit 605 is used to construct a composite reward function based on the task completion rate component, energy consumption penalty component and path passage safety component, by taking its own running state vector as input and the executable candidate migration direction in the action output set as the output behavior space.
[0241] The function update unit 606 is used to update the preset strategy value function according to the composite reward function within a unit task execution cycle to obtain a local strategy parameterized sequence.
[0242] The aggregation calculation unit 607 is used to send the local policy parameterization sequence to the central coordination unit, and the central coordination unit performs aggregation calculation on all received policy parameterization sequences to obtain global policy tuning parameters.
[0243] The function correction unit 608 is used to distribute the global policy tuning parameters to their respective main inspection units, so that each main inspection unit can correct its own local policy parameterized sequence and composite reward function according to the global policy tuning parameters.
[0244] Optionally, the matrix building unit 603 is specifically used for:
[0245] The original parameters of the target are normalized to form a set of feature vectors that describe the status of each main inspection unit and the characteristics of the inspection nodes.
[0246] The geometric spatial distance between each main inspection unit and each inspection target node is calculated based on the feature vector set, and the geometric spatial distance is corrected based on the path reachability parameter and the environmental obstacle density to obtain the path impedance factor.
[0247] Determine the operating state vector of each of the autonomous inspection units;
[0248] Calculate the task requirement tensor based on the risk and urgency indicators of each inspection target node obtained in advance;
[0249] The path impedance factor, running state vector and task requirement tensor are fused by the mapping function to generate the comprehensive cost element between each main inspection unit and each inspection target node.
[0250] The comprehensive cost elements are arranged according to the correspondence between their respective main inspection units and inspection target nodes to form a target cost spectrum matrix.
[0251] Optionally, the constraint applying unit 604 is specifically used for:
[0252] The first constraint subset is generated based on the remaining energy value, load capacity and communication reachability parameters of each main inspection unit;
[0253] A second constraint subset is generated based on the risk indicators, inspection time windows, and geographical connectivity indicators of each inspection target node.
[0254] Construct a set of runtime constraint conditions based on the first constraint subset and the second constraint subset;
[0255] The constraint parameters in the set of operating constraints are normalized, and the matrix elements of the target cost spectrum matrix are used as the initial search weights.
[0256] The target cost spectrum matrix is mapped layer by layer by a preset multidimensional constraint projection operator to obtain the constraint clipping matrix;
[0257] Based on the constraint clipping matrix, a solution is obtained to obtain a set of locally feasible optimal solutions;
[0258] The optimal set of solution elements is determined from the set of locally feasible optimal solutions, and an initial scheduling scheme is generated based on the set of solution elements.
[0259] Optionally, function building unit 605 is specifically used for:
[0260] The current operating state vector is used as input, and the operating state vector includes position parameters, velocity parameters, remaining energy value, and load capacity;
[0261] The executable candidate migration directions are defined as action output sets from the action output set, and the running state vector and the action output set are mapped to the behavior space;
[0262] In the behavior space, for each candidate action, the corresponding task completion rate component, energy consumption penalty component, and path passage safety component are calculated;
[0263] Based on the preset weighting coefficients, the task completion rate component, energy consumption penalty component, and path passage safety component are linearly fused to obtain a composite reward function.
[0264] Optionally, the function update unit 606 is specifically used for:
[0265] In the current task execution cycle, obtain the running state vector S_t, and determine the set of executable actions A_t from the action output set;
[0266] For each candidate action a_i in the set of executable actions A_t, calculate the corresponding composite reward value R_t(a_i);
[0267] Based on the current running state vector S_t, candidate action a_i, and the composite reward value R_t(a_i), the preset strategy value function is updated using the following formula:
[0268] Q_new(S_t, a_i)=(1-η)·Q_old(S_t, a_i)+η·(R_t(a_i)+λ·maxQ_old(S_t+1, a′));
[0269] Where η is the learning rate, λ is the discount factor, a′ is any action in the set of actions that can be executed at the next time step, Q_old represents the value of the original preset policy value function, and Q_new represents the value of the preset policy value function obtained by updating.
[0270] Optionally, the aggregation calculation unit 607 is specifically used for:
[0271] The local policy parameterization sequence is transmitted to the central coordination unit via a wireless communication link;
[0272] The local policy parameterization sequence is indexed and classified, and a policy parameter set matrix is constructed, wherein different local policy parameterization sequences corresponding to the same running state vector are stored in the same index position;
[0273] For different local strategy parameterization sequences corresponding to the same running state vector in the strategy parameter set matrix, the corresponding aggregate weighting coefficient is calculated according to the task completion rate component of each main inspection unit.
[0274] Based on the aggregated weighting coefficients, aggregate calculations are performed on different local policy parameterized sequences corresponding to the same running state vector to obtain global policy tuning parameters.
[0275] Optionally, the function correction unit 608 is specifically used for:
[0276] The central coordination unit encapsulates the global policy tuning parameters into policy update instructions and sends them to the autonomous inspection unit via a wireless communication link.
[0277] The autonomous inspection unit extracts global policy tuning parameters from the policy update instruction;
[0278] The local policy parameterization sequence corresponding to the autonomous inspection unit is corrected and calculated with the global policy tuning parameters. The correction calculation is performed using the following formula:
[0279] Q_revised(s,a)=δ·Q_local(s,a)+(1-δ)·Q_global(s,a);
[0280] Where Q_revised represents the modified local policy parameterization sequence corresponding to the running state vector s and candidate action a; Q_global represents the global policy tuning parameters corresponding to the running state vector s and candidate action a; and δ is the self-preservation weight coefficient.
[0281] The composite reward function is recalculated based on the modified local policy parameterized sequence.
[0282] Please see Figure 7This application also provides a collaborative inspection task allocation device based on multi-vehicle state perception, including:
[0283] Processor 701, memory 702, input / output unit 703, bus 704;
[0284] The processor 701 is connected to the memory 702, the input / output unit 703, and the bus 704;
[0285] The memory 702 stores a program, and the processor 701 calls the program to execute any of the methods described above.
[0286] This application also relates to a computer-readable storage medium on which a program is stored, which, when run on a computer, causes the computer to perform any of the methods described above.
[0287] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0288] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection between apparatuses or units through some interfaces, and may be electrical, mechanical, or other forms.
[0289] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0290] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0291] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
Claims
1. A collaborative inspection task allocation method based on multi-vehicle state perception, characterized in that, The method includes: The central coordination unit calculates the risk indicators of each inspection target node based on the environmental feature sets and historical event sequences uploaded by several autonomous inspection units, and obtains a priority index table. The central coordination unit determines the target's original parameters based on the priority index table and the operating status vector of its respective main inspection unit. The central coordination unit constructs a target cost spectrum matrix based on the original target parameters. The matrix elements in the target cost spectrum matrix reflect the comprehensive correlation cost between the autonomous inspection unit and the inspection target node. In the central coordination unit, a set of operational constraints is applied to the target cost spectrum matrix, and constraint solving is performed to obtain an initial scheduling scheme; Each main inspection unit takes its own operating state vector as input and the executable candidate migration direction in the action output set as the output behavior space, and constructs a composite reward function based on the task completion rate component, energy consumption penalty component and path passage safety component. Within the unit task execution cycle, each main inspection unit updates the preset strategy value function according to the composite reward function to obtain a local strategy parameterized sequence; Each primary inspection unit sends the local policy parameterization sequence to the central coordination unit, which then performs aggregation calculations on all received policy parameterization sequences to obtain global policy tuning parameters. The central coordination unit distributes the global strategy tuning parameters to its respective main inspection units, enabling each main inspection unit to modify its local strategy parameterized sequence and composite reward function based on the global strategy tuning parameters.
2. The collaborative inspection task allocation method based on multi-vehicle state perception as described in claim 1, characterized in that, The central coordination unit constructs a target cost spectrum matrix based on the original target parameters. The matrix elements in the target cost spectrum matrix reflect the comprehensive correlation cost between the autonomous inspection unit and the inspection target node, including: The original parameters of the target are normalized to form a set of feature vectors that describe the status of each main inspection unit and the characteristics of the inspection nodes. The geometric spatial distance between each main inspection unit and each inspection target node is calculated based on the feature vector set, and the geometric spatial distance is corrected based on the path reachability parameter and the environmental obstacle density to obtain the path impedance factor. Determine the operating state vector of each of the autonomous inspection units; Calculate the task requirement tensor based on the risk and urgency indicators of each inspection target node obtained in advance; The path impedance factor, running state vector and task requirement tensor are fused by the mapping function to generate the comprehensive cost element between each main inspection unit and each inspection target node. The comprehensive cost elements are arranged according to the correspondence between their respective main inspection units and inspection target nodes to form a target cost spectrum matrix.
3. The collaborative inspection task allocation method based on multi-vehicle state perception as described in claim 1, characterized in that, In the central coordination unit, a set of operational constraints is applied to the target cost spectrum matrix, and constraint solving is performed to obtain an initial scheduling scheme, including: The first constraint subset is generated based on the remaining energy value, load capacity and communication reachability parameters of each main inspection unit; A second constraint subset is generated based on the risk indicators, inspection time windows, and geographical connectivity indicators of each inspection target node. Construct a set of runtime constraint conditions based on the first constraint subset and the second constraint subset; The constraint parameters in the set of operating constraints are normalized, and the matrix elements of the target cost spectrum matrix are used as the initial search weights. The target cost spectrum matrix is mapped layer by layer by a preset multidimensional constraint projection operator to obtain the constraint clipping matrix; Based on the constraint clipping matrix, a solution is obtained to obtain a set of locally feasible optimal solutions; The optimal set of solution elements is determined from the set of locally feasible optimal solutions, and an initial scheduling scheme is generated based on the set of solution elements.
4. The collaborative inspection task allocation method based on multi-vehicle state perception as described in claim 1, characterized in that, Each primary inspection unit uses its own operational state vector as input and the executable candidate migration directions in the action output set as its output behavior space. It constructs a composite reward function based on the task completion rate component, energy consumption penalty component, and path passage safety component, including: The current operating state vector is used as input, and the operating state vector includes position parameters, velocity parameters, remaining energy value, and load capacity; The executable candidate migration directions are defined as action output sets from the action output set, and the running state vector and the action output set are mapped to the behavior space; In the behavior space, for each candidate action, the corresponding task completion rate component, energy consumption penalty component, and path passage safety component are calculated; Based on the preset weighting coefficients, the task completion rate component, energy consumption penalty component, and path passage safety component are linearly fused to obtain a composite reward function.
5. The collaborative inspection task allocation method based on multi-vehicle state perception as described in claim 4, characterized in that, Within a unit task execution cycle, each primary inspection unit updates the preset strategy value function according to the composite reward function to obtain a local strategy parameterized sequence, including: The autonomous inspection unit acquires the running status vector S_t in the current task execution cycle and determines the set of executable actions A_t from the action output set; For each candidate action a_i in the set of executable actions A_t, calculate the corresponding composite reward value R_t(a_i); Based on the current running state vector S_t, candidate action a_i, and the composite reward value R_t(a_i), the preset strategy value function is updated using the following formula: Q_new(S_t, a_i)=(1-η)·Q_old(S_t, a_i)+η·(R_t(a_i)+λ·maxQ_old(S_t+1, a′)); Where η is the learning rate, λ is the discount factor, a′ is any action in the set of actions that can be executed at the next time step, Q_old represents the value of the original preset policy value function, and Q_new represents the value of the preset policy value function obtained by updating.
6. The collaborative inspection task allocation method based on multi-vehicle state perception as described in claim 1, characterized in that, Each primary inspection unit sends its local policy parameterized data sequence to the central coordination unit, which then performs aggregation calculations on all received policy parameterized data sequences to obtain global policy tuning parameters, including: Each primary inspection unit transmits its local strategy parameterization sequence to the central coordination unit via a wireless communication link. The central coordination unit indexes and categorizes the local policy parameterization sequences and constructs a policy parameter set matrix, wherein different local policy parameterization sequences corresponding to the same running state vector are stored in the same index position; For different local strategy parameterization sequences corresponding to the same running state vector in the strategy parameter set matrix, the corresponding aggregate weighting coefficient is calculated according to the task completion rate component of each main inspection unit. Based on the aggregation weighting coefficients, the central coordination unit performs aggregation calculations on different local policy parameterization sequences corresponding to the same running state vector to obtain global policy tuning parameters.
7. The collaborative inspection task allocation method based on multi-vehicle state perception as described in claim 1, characterized in that, The central coordination unit distributes the global policy tuning parameters to their respective main inspection units, enabling each autonomous inspection unit to modify its local policy parameterized sequence and composite reward function based on the global policy tuning parameters, including: The central coordination unit encapsulates the global policy tuning parameters into policy update instructions and sends them to the autonomous inspection unit via a wireless communication link. The autonomous inspection unit extracts global policy tuning parameters from the policy update instruction; The autonomous inspection unit performs correction calculations on the corresponding local policy parameterization sequence and the global policy tuning parameters; The correction calculation is performed using the following formula: Q_revised(s,a)=δ·Q_local(s,a)+(1-δ)·Q_global(s,a); Where Q_revised represents the corrected local policy parameterization sequence corresponding to the running state vector s and candidate action a; Q_global represents the global policy tuning parameters corresponding to the running state vector s and candidate action a; δ is the self-preservation weight coefficient; and Q_local represents the local policy parameterization sequence. The composite reward function is recalculated based on the modified local policy parameterized sequence.
8. A collaborative inspection task allocation device based on multi-vehicle status perception, characterized in that, The device includes: The data acquisition unit is used to calculate the risk indicators of each inspection target node and obtain a priority index table based on the environmental feature sets and historical event sequences uploaded by several autonomous inspection units. The parameter determination unit is used to determine the target original parameters based on the priority index table and the running status vector of each main inspection unit. The matrix construction unit is used to construct a target cost spectrum matrix based on the original target parameters. The matrix elements in the target cost spectrum matrix reflect the comprehensive correlation cost between the autonomous inspection unit and the inspection target node. The constraint application unit is used to apply a set of running constraints to the target cost spectrum matrix and perform constraint solving to obtain an initial scheduling scheme. The function building unit is used to construct a composite reward function based on the task completion rate component, energy consumption penalty component, and path passage safety component, with its own running state vector as input and the executable candidate migration direction in the action output set as the output behavior space. The function update unit is used to update the preset strategy value function according to the composite reward function within a unit task execution cycle to obtain a local strategy parameterized sequence. An aggregation calculation unit is used to send the local policy parameterization sequence to the central coordination unit, which then performs aggregation calculations on all received policy parameterization sequences to obtain global policy tuning parameters. The function correction unit is used to distribute the global policy tuning parameters to their respective main inspection units, so that each main inspection unit can correct its own local policy parameterized sequence and composite reward function according to the global policy tuning parameters.
9. A collaborative inspection task allocation device based on multi-vehicle state perception, characterized in that, The device includes: Processor, memory, input / output units, and bus; The processor is connected to the memory, the input / output unit, and the bus; The memory stores a program, which the processor invokes to perform the method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium contains a program that, when executed on a computer, performs the method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Unmanned intelligent inspection equipment cooperative scheduling method and system in photovoltaic power generation scene
CN119358998A
Oil depot tank field inspection robot task allocation method and system based on Internet of Things
CN119417192A