Task allocation method and system for inspection robots in oil depot tank area based on Internet of Things

By constructing a multi-objective collaborative scheduling model based on deep neural networks and introducing reinforcement learning algorithms to inspect robots in the IoT oil tank area, the problem of lack of targeted and flexible task allocation in the existing technology is solved, efficient and intelligent task allocation and emergency treatment are achieved, and the continuity and reliability of inspection work are improved.

CN119417192BActive Publication Date: 2025-05-06北京网藤科技有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510019788.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-07
Publication Date
2025-05-06
Estimated Expiration
2045-01-07

AI Technical Summary

Technical Problem

The existing oil tank area inspection robot task allocation methods lack targeted and flexible, and cannot take into account multiple goals such as task priority, path optimization and coordinated obstacle avoidance at the same time, and lack an effective emergency response mechanism, which affects the continuity and reliability of inspection work.

Method used

The task allocation method of the oil tank area patrol robot based on the Internet of Things is adopted. By receiving the environmental data collected by the Internet of Things sensing nodes and real-time state information uploaded by the robot, a multi-objective collaborative scheduling model based on the deep neural network is constructed, and the spatial and fusion characteristics are extracted in combination with the spatiotemporal attention mechanism, a task priority evaluation sub-model and path planning sub-model are established to generate the optimal task allocation strategy, and when task conflicts or execution exceptions are detected, optimization and adjustment are made through reinforcement learning algorithms.

Benefits of technology

It improves the accuracy and efficiency of task allocation, realizes intelligent allocation of robot inspection tasks, enhances the robustness of the system and the ability to deal with emergencies, ensures the executability of task allocation, and extends the service life of the robot.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119417192B_ABST
    Figure CN119417192B_ABST
Patent Text Reader

Abstract

The present invention provides a method and system for allocating tasks of inspection robots in oil depot tank areas based on the Internet of Things, which relates to the technical field of the Internet of Things, including receiving environmental data and robot status information; constructing a multi-objective collaborative scheduling model, using a spatiotemporal attention mechanism to extract features, establishing a task priority evaluation and path planning sub-model, and generating an optimal task allocation strategy; issuing inspection task instructions, monitoring the execution status in real time, and starting an emergency module for optimization and adjustment when an abnormality is detected. The present invention realizes intelligent and collaborative inspection task allocation, and improves inspection efficiency and safety.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to Internet of Things technology, and in particular to an Internet of Things-based oil depot tank area inspection robot task allocation method and system. Background Art

[0002] The existing methods for allocating tasks for inspection robots in oil depot tank areas still have some shortcomings. First, most methods fail to make full use of the real-time environmental data and robot status information collected by the Internet of Things, resulting in a lack of pertinence and flexibility in task allocation. Second, existing methods often use a single task allocation algorithm, which cannot take into account multiple goals at the same time, such as task priority, path optimization, and collaborative obstacle avoidance, making it difficult to achieve the global optimal task allocation effect. Finally, when task conflicts or execution anomalies occur, existing methods lack an effective emergency response mechanism and cannot adjust the task allocation strategy in a timely manner, affecting the continuity and reliability of inspection work. Summary of the invention

[0003] The embodiments of the present invention provide a method and system for allocating tasks of oil depot tank area inspection robots based on the Internet of Things, which can solve the problems in the prior art.

[0004] According to a first aspect of the embodiments of the present invention,

[0005] Provides a method for allocating tasks for inspection robots in oil depot tank areas based on the Internet of Things, including:

[0006] Receive environmental data information collected by IoT sensor nodes in the oil depot tank area and real-time status information uploaded by multiple intelligent inspection robots, the environmental data information includes oil tank pressure data, temperature data and gas concentration data, and the real-time status information includes the remaining power of the robot, task execution progress and real-time position coordinates;

[0007] Construct a multi-objective collaborative scheduling model based on a deep neural network, input environmental data information and real-time status information into the multi-objective collaborative scheduling model, use the spatiotemporal attention mechanism to extract spatiotemporal fusion features, and establish a task priority evaluation sub-model and a path planning sub-model based on the spatiotemporal fusion features. The task priority evaluation sub-model calculates the task execution order based on the safety level of the oil tank area and the urgency of the task. The path planning sub-model calculates the collaborative obstacle avoidance path of multiple robots based on the improved ant colony algorithm. The task execution order and the collaborative obstacle avoidance path are input into the task allocation optimization module to generate the optimal task allocation strategy.

[0008] Based on the optimal task allocation strategy, inspection task instructions are issued to the robot, and the robot's task execution data and environmental data are collected in real time through the Internet of Things. When a robot task conflict or execution abnormality is detected, the emergency module is started. The emergency module performs situational awareness and risk assessment based on the task execution data and environmental data, and uses the robot's historical task completion rate and remaining power as constraints. A reinforcement learning algorithm is used to optimize and adjust the optimal task allocation strategy to generate an updated task allocation strategy.

[0009] In an optional embodiment,

[0010] Inputting environmental data information and real-time status information into the multi-objective collaborative scheduling model, adopting the spatiotemporal attention mechanism to extract spatiotemporal fusion features, establishing a task priority evaluation sub-model and a path planning sub-model based on the spatiotemporal fusion features, the task priority evaluation sub-model calculates the task execution order in combination with the safety level of the oil tank area and the urgency of the task, and the path planning sub-model calculates the collaborative obstacle avoidance path of multiple robots based on the improved ant colony algorithm, including the following steps:

[0011] For environmental data information and real-time status information, time windows are constructed to calculate the time attention weight respectively. Based on the IoT sensor node position and the robot's real-time position coordinates, a spatial correlation matrix is ​​constructed to calculate the spatial attention weight. The time attention weight and the spatial attention weight are fused to obtain the spatiotemporal fusion feature.

[0012] A task priority evaluation sub-model is constructed based on the spatiotemporal fusion feature, the task priority evaluation sub-model calculates the safety level of the oil tank farm based on the environmental data information, uses a long short-term memory network to predict the future state of the environmental data to obtain a risk trend index, calculates the task urgency in combination with the oil tank farm safety level and the real-time state information, calculates the feature similarity between tasks and the spatial distance between task target points to construct a task coupling matrix, inputs the risk trend index, task urgency and task coupling matrix into an adaptive weight network to obtain dynamic weights, calculates task priority scores based on the dynamic weights and outputs the task execution order;

[0013] A path planning sub-model is constructed based on the spatiotemporal fusion features and the task execution order. The path planning sub-model adopts an improved ant colony algorithm to perform multi-robot collaborative obstacle avoidance path planning, introduces environmental safety metrics into heuristic information, updates the pheromone matrix based on task priority, updates the dynamic taboo table according to the real-time position coordinates of the robots to record the robot's motion trajectory and occupied area, and generates a multi-robot collaborative obstacle avoidance path set.

[0014] In an optional embodiment,

[0015] The steps of calculating the feature similarity between tasks and the spatial distance between task target points to construct the task coupling matrix include:

[0016] Constructing a task feature vector, wherein the task feature vector includes environmental data information, a task type code, a time window parameter, and spatial location information, wherein the task type code includes type information of equipment inspection, pipeline inspection, and area inspection, the time window parameter includes the task start time and end time, and the spatial location information includes the coordinates of the task target point;

[0017] Based on the environmental data information, the environmental data similarity is calculated, and the cosine similarity method is used to calculate the comprehensive similarity of the oil tank pressure data, temperature data and gas concentration data between tasks; based on the task type code, a task type similarity matrix is ​​constructed, and the task type similarity is determined according to the degree of association between equipment inspection, pipeline inspection and area inspection; based on the time window parameter, the task time overlap is calculated, and the overlapping ratio of the task execution time is calculated according to the task start time and end time; based on the spatial position information, the actual spatial distance between the task target points is calculated, and the minimum spacing between adjacent oil tanks in the oil tank area is used as the characteristic distance of the oil tank area, and the actual spatial distance is divided by the characteristic distance of the oil tank area to obtain the standardized distance, and the standardized distance is converted into a distance attenuation coefficient through an exponential decay function;

[0018] The feature fusion method based on entropy weight method is used to calculate the weight of environmental data similarity, task type similarity, task time overlap and distance decay coefficient. The task similarity matrix is ​​obtained by weighted fusion of environmental data similarity, task type similarity, task time overlap and distance decay coefficient.

[0019] The task similarity matrix is ​​subjected to threshold sparse processing to obtain an initial coupling matrix, and the initial coupling matrix is ​​optimized using a Gaussian smoothing kernel to obtain a final task coupling matrix, wherein the kernel size of the Gaussian smoothing kernel is dynamically adjusted according to the number of tasks.

[0020] In an optional embodiment,

[0021] The path planning sub-model adopts an improved ant colony algorithm to perform multi-robot collaborative obstacle avoidance path planning, introduces environmental safety metrics into heuristic information, updates the pheromone matrix based on task priority, updates the dynamic taboo table based on the real-time position coordinates of the robots to record the robot's motion trajectory and occupied area, and generates a multi-robot collaborative obstacle avoidance path set, including the following steps:

[0022] An environmental safety metric is introduced into the heuristic information calculation, wherein the environmental safety metric is obtained by weighted combination of the oil tank pressure safety factor, the gas concentration safety factor and the temperature safety factor by weighted coefficients, wherein the oil tank pressure safety factor is calculated based on the deviation between the current pressure value and the safety pressure threshold, the gas concentration safety factor is calculated based on the ratio of the current concentration to the critical concentration, and the temperature safety factor is calculated based on the deviation between the current temperature and the normal temperature range, and the weighted coefficient is dynamically output by the task priority evaluation sub-model;

[0023] The pheromone matrix update rule is improved based on task priority. In local update, task priority and path length are used together as the calculation basis of pheromone increment. In global update, a task priority dynamic adjustment factor is introduced. The task priority dynamic adjustment factor changes exponentially with the ratio of the current task priority to the maximum priority.

[0024] Construct a dynamic taboo table in the time and space dimensions to handle multi-robot path conflicts. The dynamic taboo table includes time dimension taboo items and space dimension obstacle avoidance constraints. The time dimension taboo items record the occupied area of ​​each robot within the safe distance range at different times. The space dimension obstacle avoidance constraints are calculated based on the distance between any two robots at any time being greater than twice the safe distance.

[0025] Multi-robot path planning is performed based on the environmental safety metric, the updated pheromone matrix and the dynamic taboo table, the environmental safety metric and the task priority are used as input to calculate the state transition probability of the improved ant colony algorithm, and the path selection is performed according to the state transition probability during the path construction process. When it is detected that the next position violates the time dimension taboo item or the space dimension obstacle avoidance constraint in the dynamic taboo table, dynamic conflict resolution is performed; the dynamic conflict resolution includes: by calculating a set of potential conflicting spatiotemporal points, a temporary taboo area is constructed around the conflict point and a global taboo table is updated. When the distance between the robot and the temporary taboo area is less than a preset threshold, the environmental safety metric and the task priority are used as input to trigger local path re-planning. After the path construction is completed, global optimization is performed and the pheromone matrix and the dynamic taboo table are updated, and a set of collaborative obstacle avoidance paths for multiple robots is output.

[0026] In an optional embodiment,

[0027] The step of inputting the task execution order and the collaborative obstacle avoidance path into a task allocation optimization module to generate an optimal task allocation strategy comprises:

[0028] Converting the task execution order into a task priority matrix, wherein the task priority matrix includes the front-to-back dependency relationship and execution priority weights between tasks, and converting the collaborative obstacle avoidance path into a path time-space matrix, wherein the path time-space matrix includes the time-space resource occupancy information of the robot motion trajectory;

[0029] A multi-objective task allocation optimization model is constructed based on the task priority matrix and the path time-space matrix, wherein the multi-objective task allocation optimization model takes maximizing task priority satisfaction and minimizing path execution cost as optimization objectives; in the multi-objective task allocation optimization model, a task priority order constraint is established based on the task priority matrix; a path conflict constraint is established based on the path time-space matrix, wherein the path conflict constraint quantifies the collision risk by calculating the minimum safe distance between any two robots in the time-space dimension; a timing constraint is established based on the forward and backward dependencies between tasks, wherein the timing constraint is used to regulate the execution order of interdependent tasks; an energy constraint is established based on the remaining power of the robot and the task resource requirements, wherein the energy constraint is used to constrain the resource allocation of the task allocation strategy;

[0030] A mixed integer programming method is used to solve the multi-objective task allocation optimization model, and the task priority satisfaction, path execution cost and completion time are combined to construct a weighted objective function; the task priority order constraints, path conflict constraints, timing constraints and energy constraints are converted into penalty terms of the weighted objective function through the Lagrangian relaxation method to form an augmented Lagrangian function; the augmented Lagrangian function is iteratively solved using a column generation method, wherein the main problem is restricted to determine the current task allocation strategy, the pricing subproblem generates a new feasible allocation strategy and adds it to the feasible solution space, and the subgradient method is used to update the Lagrangian multiplier to adjust the penalty intensity of constraint violation to obtain the optimal task allocation strategy.

[0031] In an optional embodiment,

[0032] Based on the optimal task allocation strategy, the inspection task instruction is issued to the robot, and the task execution data and environmental data of the robot are collected in real time through the Internet of Things. When a robot task conflict or execution abnormality is detected, the emergency module is started. The emergency module performs situational awareness and risk assessment based on the task execution data and environmental data, and uses the robot's historical task completion rate and remaining power as constraints. The optimal task allocation strategy is optimized and adjusted using a reinforcement learning algorithm. The steps of generating an updated task allocation strategy include:

[0033] Collect robot layer data and environment layer data through the Internet of Things, where the robot layer data includes the robot's position coordinates, power status, task execution status and motion parameters, and the environment layer data includes gas concentration, temperature distribution, equipment vibration and pipeline pressure, and perform outlier processing, standardization processing and time alignment processing on the robot layer data and the environment layer data; build a situational awareness model based on the processed data, calculate the number of task conflicts by calculating the task time overlap and the minimum distance of the robot path, identify the degree of execution abnormality by task progress deviation and trajectory deviation, and calculate the comprehensive risk according to the weighted combination of environmental risk, task risk and equipment risk;

[0034] Construct a reinforcement learning network based on a dual deep Q network, and use the robot state vector, task state vector and environment state vector to form a state space, and use task reallocation, task suspension, task resumption and task termination as the action space. The robot state vector includes the position coordinates and the remaining power, the task state vector includes the task priority, deadline and execution progress, and the environment state vector includes the comprehensive risk, the execution abnormality degree and the number of task conflicts. Design a combined reward function of task completion reward, risk penalty, resource consumption penalty and time penalty, wherein the risk penalty is positively correlated with the execution abnormality degree. Train the dual deep Q network based on the combined reward function, use the experience replay mechanism to store state transition samples, select actions through the ε greedy strategy, and regularly update the target network parameters of the dual deep Q network.

[0035] When the comprehensive risk exceeds the risk threshold, the number of task conflicts exceeds the conflict threshold, or the remaining power of the robot is lower than the power threshold, the optimization of the optimal task allocation strategy is triggered, and the current robot state vector, task state vector and environment state vector are input into the trained dual deep Q network to obtain the optimized task allocation strategy. After the feasibility of the optimized task allocation strategy is verified, a scheduling instruction is issued to the robot.

[0036] In an optional embodiment,

[0037] Designing a combined reward function of task completion reward, risk penalty, resource consumption penalty and time penalty, wherein the risk penalty is positively correlated with the execution abnormality degree, training a dual deep Q network based on the combined reward function, using an experience replay mechanism to store state transition samples, selecting actions through an ε greedy strategy, and regularly updating the target network parameters of the dual deep Q network includes:

[0038] The task completion reward is calculated based on the product of task progress, task priority and quality score, the quality score is based on the task completion stability and timeliness assessment, the risk penalty includes execution abnormality penalty and safety distance breach penalty, the resource consumption penalty is calculated based on the power consumption rate, and the time penalty is calculated based on the task urgency and overtime duration; the dual deep Q network includes an evaluation network and a target network, both of which adopt a three-layer fully connected structure, the input layer receives the state vector of the state space, the hidden layer uses the ReLU activation function to extract features, and the output layer generates action value estimation;

[0039] Establish an experience replay buffer with a fixed capacity to store transition samples including the current state, executed actions, rewards obtained, and next state, and randomly sample from the experience replay buffer based on a preset batch size; execute an action selection strategy based on the action value estimation, select an action based on the ε greedy strategy at each decision moment, select the action with the maximum action value with a probability of 1-ε, and randomly select an action for exploration with a probability of ε, where the ε value decreases linearly with the training rounds;

[0040] The target Q value output by the target network is calculated using the temporal difference learning method and the Bellman equation. The error between the Q value predicted by the evaluation network and the target Q value is calculated based on the smooth L1 loss function. The evaluation network parameters are optimized through back propagation. The evaluation network parameters are soft-updated to the target network according to the preset update cycle. The training process is continuously iterated. In each round of training, the state vector of the state space is input into the evaluation network to obtain the action selection, the selected action is executed to obtain the new state and reward value, and the transferred samples are stored in the experience replay buffer. When the cumulative number of samples reaches the training batch requirement, the evaluation network parameters are updated through back propagation based on the loss function, and the updated evaluation network parameters are copied to the target network through the soft update mechanism according to the preset update cycle until the preset maximum number of iterations is reached.

[0041] According to a second aspect of the embodiments of the present invention,

[0042] Provides an IoT-based oil depot tank area inspection robot task allocation system, including:

[0043] The first unit is used to receive environmental data information collected by the IoT sensor nodes in the oil tank area and real-time status information uploaded by multiple intelligent inspection robots, wherein the environmental data information includes oil tank pressure data, temperature data and gas concentration data, and the real-time status information includes the remaining power of the robot, task execution progress and real-time position coordinates;

[0044] The second unit is used to build a multi-objective collaborative scheduling model based on a deep neural network, input environmental data information and real-time status information into the multi-objective collaborative scheduling model, use the spatiotemporal attention mechanism to extract spatiotemporal fusion features, and establish a task priority evaluation sub-model and a path planning sub-model based on the spatiotemporal fusion features. The task priority evaluation sub-model calculates the task execution order based on the safety level of the oil tank area and the urgency of the task. The path planning sub-model calculates the collaborative obstacle avoidance path of multiple robots based on the improved ant colony algorithm, and inputs the task execution order and the collaborative obstacle avoidance path into the task allocation optimization module to generate the optimal task allocation strategy;

[0045] The third unit is used to issue inspection task instructions to the robot based on the optimal task allocation strategy, collect the robot's task execution data and environmental data in real time through the Internet of Things, and start the emergency module when a robot task conflict or execution abnormality is detected. The emergency module performs situational awareness and risk assessment based on the task execution data and environmental data, and uses the robot's historical task completion rate and remaining power as constraints, and uses a reinforcement learning algorithm to optimize and adjust the optimal task allocation strategy to generate an updated task allocation strategy.

[0046] According to a third aspect of the embodiments of the present invention,

[0047] An electronic device is provided, comprising:

[0048] processor;

[0049] a memory for storing processor-executable instructions;

[0050] The processor is configured to call the instructions stored in the memory to execute the aforementioned method.

[0051] According to a fourth aspect of the embodiments of the present invention,

[0052] A computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the aforementioned method is implemented.

[0053] The present invention constructs a multi-objective collaborative scheduling model based on a deep neural network and extracts spatiotemporal fusion features in combination with the spatiotemporal attention mechanism, thereby achieving effective processing of complex environmental data and robot state information and improving the accuracy and efficiency of task allocation. The model can comprehensively consider multi-dimensional factors such as the safety level of the oil tank area and the urgency of the task, generate the optimal task execution sequence and collaborative obstacle avoidance path, and thus realize the intelligent allocation of robot inspection tasks.

[0054] The present invention introduces an emergency module that can monitor the robot's task execution and environmental changes in real time. When task conflicts or execution anomalies are detected, the emergency module will perform situational awareness and risk assessment based on the latest task execution data and environmental data, and use a reinforcement learning algorithm to dynamically optimize and adjust the task allocation strategy. This adaptive adjustment mechanism greatly improves the system's robustness and ability to respond to emergencies.

[0055] The present invention organically combines the Internet of Things technology with the intelligent inspection robot to achieve comprehensive perception and real-time monitoring of the oil depot tank area environment. By continuously collecting environmental data and robot status information, the system can continuously optimize the task allocation strategy and improve inspection efficiency and safety. At the same time, it takes into account constraints such as the robot's remaining power and historical task completion rate, ensuring the executability of task allocation and effectively extending the robot's service life. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] Figure 1 It is a flowchart of a method for allocating tasks of inspection robots in oil tank areas based on the Internet of Things according to an embodiment of the present invention;

[0057] Figure 2 The present invention is a schematic diagram of the structure of an oil depot tank area inspection robot task allocation system based on the Internet of Things according to an embodiment of the present invention. DETAILED DESCRIPTION

[0058] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0059] The technical solution of the present invention is described in detail with specific embodiments below. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments.

[0060] Figure 1 FIG. 1 is a flow chart of task allocation of inspection robots in oil tank areas based on the Internet of Things according to an embodiment of the present invention. Figure 1 As shown, the method includes:

[0061] S1. Receive environmental data information collected by IoT sensor nodes in the oil tank area and real-time status information uploaded by multiple intelligent inspection robots, the environmental data information includes oil tank pressure data, temperature data and gas concentration data, and the real-time status information includes the remaining power of the robot, task execution progress and real-time position coordinates;

[0062] S2. Construct a multi-objective collaborative scheduling model based on a deep neural network, input environmental data information and real-time status information into the multi-objective collaborative scheduling model, use the spatiotemporal attention mechanism to extract spatiotemporal fusion features, and establish a task priority evaluation sub-model and a path planning sub-model based on the spatiotemporal fusion features. The task priority evaluation sub-model calculates the task execution order based on the safety level of the oil tank area and the urgency of the task. The path planning sub-model calculates the collaborative obstacle avoidance path of multiple robots based on the improved ant colony algorithm. The task execution order and the collaborative obstacle avoidance path are input into the task allocation optimization module to generate the optimal task allocation strategy.

[0063] S3. Based on the optimal task allocation strategy, inspection task instructions are issued to the robot, and the robot's task execution data and environmental data are collected in real time through the Internet of Things. When a robot task conflict or execution abnormality is detected, the emergency module is started. The emergency module performs situational awareness and risk assessment based on the task execution data and environmental data, and uses the robot's historical task completion rate and remaining power as constraints. The optimal task allocation strategy is optimized and adjusted using a reinforcement learning algorithm to generate an updated task allocation strategy.

[0064] In an optional embodiment,

[0065] Inputting environmental data information and real-time status information into the multi-objective collaborative scheduling model, adopting the spatiotemporal attention mechanism to extract spatiotemporal fusion features, establishing a task priority evaluation sub-model and a path planning sub-model based on the spatiotemporal fusion features, the task priority evaluation sub-model calculates the task execution order in combination with the safety level of the oil tank area and the urgency of the task, and the path planning sub-model calculates the collaborative obstacle avoidance path of multiple robots based on the improved ant colony algorithm, including the following steps:

[0066] For environmental data information and real-time status information, time windows are constructed to calculate the time attention weight respectively. Based on the IoT sensor node position and the robot's real-time position coordinates, a spatial correlation matrix is ​​constructed to calculate the spatial attention weight. The time attention weight and the spatial attention weight are fused to obtain the spatiotemporal fusion feature.

[0067] A task priority evaluation sub-model is constructed based on the spatiotemporal fusion feature, the task priority evaluation sub-model calculates the safety level of the oil tank farm based on the environmental data information, uses a long short-term memory network to predict the future state of the environmental data to obtain a risk trend index, calculates the task urgency in combination with the oil tank farm safety level and the real-time state information, calculates the feature similarity between tasks and the spatial distance between task target points to construct a task coupling matrix, inputs the risk trend index, task urgency and task coupling matrix into an adaptive weight network to obtain dynamic weights, calculates task priority scores based on the dynamic weights and outputs the task execution order;

[0068] A path planning sub-model is constructed based on the spatiotemporal fusion features and the task execution order. The path planning sub-model adopts an improved ant colony algorithm to perform multi-robot collaborative obstacle avoidance path planning, introduces environmental safety metrics into heuristic information, updates the pheromone matrix based on task priority, updates the dynamic taboo table according to the real-time position coordinates of the robots to record the robot's motion trajectory and occupied area, and generates a multi-robot collaborative obstacle avoidance path set.

[0069] Exemplarily, first, a time window is constructed for the environmental data information and the real-time status information, and the time attention weight is calculated. Specifically, a sliding time window, such as 30 minutes, can be set to analyze the data within the time window. For the data at each time point, the time difference between it and the current moment is calculated, and the time difference is mapped to a weight value using a softmax function. For example, assuming that the current moment is 10:00, the weight of the data at 9:40 may be 0.2, and the weight of the data at 9:55 may be 0.5.

[0070] Next, a spatial correlation matrix is ​​constructed based on the IoT sensor node locations and the robot’s real-time location coordinates, and the spatial attention weight is calculated. The tank farm can be divided into grids, each corresponding to a sensor node. The Euclidean distance between each robot and each sensor node is calculated, and the distance is converted into a correlation score using a Gaussian kernel function. For example, a sensor node with a distance of 5 meters may have a score of 0.8, while a sensor node with a distance of 20 meters may have a score of 0.3.

[0071] The temporal attention weight and the spatial attention weight are fused to obtain the spatiotemporal fusion feature. A weighted summation method can be used, such as setting the temporal weight to 0.4 and the spatial weight to 0.6, and then adding the two to obtain the final fusion feature.

[0072] The task priority evaluation sub-model is constructed based on the spatiotemporal fusion features. First, the safety level of the oil tank area is calculated based on the environmental data information. Multiple safety indicators can be set, such as temperature, pressure, and combustible gas concentration, and thresholds and weights can be set for each indicator. The comprehensive safety score is obtained by weighted summation, and then the score is mapped to different safety levels, such as 1-5.

[0073] The long short-term memory network (LSTM) is used to predict the future state of environmental data and obtain the risk trend index. The environmental data of the past period of time (such as 2 hours) is used as input to predict the data change trend of the future period of time (such as 30 minutes). The risk trend index is calculated based on the comparison between the predicted results and the safety threshold. For example, if the temperature is predicted to exceed the safety threshold within 20 minutes, the risk trend index may be 0.8.

[0074] The urgency of the task is calculated by combining the safety level of the oil tank area and the real-time status information. Factors such as task type, execution time window, and current safety level can be considered. For example, for an area with a safety level of 4, the urgency of an inspection task may be 0.9, while the urgency of a maintenance task may be 0.7.

[0075] Calculate the feature similarity between tasks and the spatial distance between task target points to construct the task coupling matrix. Feature similarity can be obtained by calculating the cosine similarity of task attribute vectors, and spatial distance can be obtained by using the Euclidean distance. These two indicators are weighted and fused to obtain the degree of coupling between tasks.

[0076] The risk trend index, task urgency and task coupling matrix are input into the adaptive weight network to obtain dynamic weights. The network can be a simple fully connected neural network, with the number of input layer nodes being the sum of the dimensions of the above three indicators and the number of output layer nodes being the number of tasks. The weight value output by the network represents the importance of each task.

[0077] Calculate the task priority score based on the dynamic weight and output the task execution order. Multiply the dynamic weight of each task with other relevant factors (such as task type weight, execution difficulty, etc.) to get the final priority score. Sort the tasks from high to low according to the score to determine the execution order.

[0078] Next, a path planning sub-model is constructed based on the spatiotemporal fusion features and the task execution order. This sub-model uses the improved ant colony algorithm for multi-robot collaborative obstacle avoidance path planning. First, the environmental safety metric is introduced into the heuristic information. A safety factor can be assigned to each grid according to the safety level of the oil tank area and the distribution of obstacles. The higher the safety factor, the greater the probability that the ant will choose this path.

[0079] Update the pheromone matrix based on the task priority. For tasks with higher priority, the pheromone concentration around its target point can be appropriately increased to attract more ants to choose this path. For example, if the priority score of a task is 0.9, the pheromone concentration of the 3×3 grid around its target point can be increased by 50%.

[0080] Update the dynamic taboo table according to the real-time position coordinates of the robot, record the robot's movement trajectory and occupied area. Mark the grids that have been occupied by the robot as taboo areas to prevent other robots from entering. At the same time, record the historical trajectory of each robot for subsequent path optimization.

[0081] Finally, a set of collaborative obstacle avoidance paths for multiple robots is generated. Each ant represents a robot and chooses the next move location based on heuristic information, pheromone concentration, and dynamic taboo table. When all ants reach the target point, the total length and safety of the path are evaluated and the global optimal solution is updated. After repeated iterations, the final set of collaborative obstacle avoidance paths is obtained.

[0082] The multi-objective collaborative scheduling model of the present invention effectively integrates environmental data and real-time status information through a spatiotemporal attention mechanism, thereby improving the perception and understanding of complex dynamic environments; the task priority evaluation submodel comprehensively considers the safety level of the oil tank area, risk trends, task urgency and coupling relationships between tasks, and can dynamically adjust the task execution order to ensure that key tasks are processed in a timely manner, thereby improving the safety and efficiency of the system; the path planning submodel adopts an improved ant colony algorithm, introduces environmental safety metrics and task priorities into the heuristic information and pheromone update mechanism, and uses a dynamic taboo table to avoid conflicts, thereby realizing collaborative obstacle avoidance path planning for multiple robots and improving the safety and efficiency of the path.

[0083] In an optional embodiment,

[0084] The steps of calculating the feature similarity between tasks and the spatial distance between task target points to construct the task coupling matrix include:

[0085] Constructing a task feature vector, wherein the task feature vector includes environmental data information, a task type code, a time window parameter, and spatial location information, wherein the task type code includes type information of equipment inspection, pipeline inspection, and area inspection, the time window parameter includes the task start time and end time, and the spatial location information includes the coordinates of the task target point;

[0086] Based on the environmental data information, the environmental data similarity is calculated, and the cosine similarity method is used to calculate the comprehensive similarity of the oil tank pressure data, temperature data and gas concentration data between tasks; based on the task type code, a task type similarity matrix is ​​constructed, and the task type similarity is determined according to the degree of association between equipment inspection, pipeline inspection and area inspection; based on the time window parameter, the task time overlap is calculated, and the overlapping ratio of the task execution time is calculated according to the task start time and end time; based on the spatial position information, the actual spatial distance between the task target points is calculated, and the minimum spacing between adjacent oil tanks in the oil tank area is used as the characteristic distance of the oil tank area, and the actual spatial distance is divided by the characteristic distance of the oil tank area to obtain the standardized distance, and the standardized distance is converted into a distance attenuation coefficient through an exponential decay function;

[0087] The feature fusion method based on entropy weight method is used to calculate the weight of environmental data similarity, task type similarity, task time overlap and distance decay coefficient. The task similarity matrix is ​​obtained by weighted fusion of environmental data similarity, task type similarity, task time overlap and distance decay coefficient.

[0088] The task similarity matrix is ​​subjected to threshold sparse processing to obtain an initial coupling matrix, and the initial coupling matrix is ​​optimized using a Gaussian smoothing kernel to obtain a final task coupling matrix, wherein the kernel size of the Gaussian smoothing kernel is dynamically adjusted according to the number of tasks.

[0089] Exemplarily, a task feature vector is first constructed, which includes environmental data information, task type code, time window parameters and spatial location information. The environmental data information includes oil tank pressure data, temperature data and gas concentration data. The task type code includes type information of equipment inspection, pipeline inspection and regional inspection. The time window parameter includes the task start time and end time. The spatial location information includes the coordinates of the task target point.

[0090] Based on the constructed task feature vector, the similarity of environmental data is calculated. The cosine similarity method is used to calculate the comprehensive similarity of oil tank pressure data, temperature data, and gas concentration data between tasks. For example, for two tasks A and B, their environmental data are [1.5, 25, 0.03] and [1.6, 26, 0.04] respectively, and the calculated cosine similarity is 0.9998, indicating that the environmental data are highly similar.

[0091] Next, a task type similarity matrix is ​​constructed based on the task type coding. The task type similarity is determined based on the degree of association between equipment inspection, pipeline inspection, and regional inspection. For example, the similarity between equipment inspection and pipeline inspection can be set to 0.8, the similarity between equipment inspection and regional inspection can be set to 0.6, and the similarity between pipeline inspection and regional inspection can be set to 0.7.

[0092] Then, the task time overlap is calculated based on the time window parameters. The overlap ratio of task execution time is calculated based on the task start time and end time. For example, if the execution time of task A is 9:00-11:00 and the execution time of task B is 10:00-12:00, then the time overlap is (11:00-10:00) / (12:00-9:00)=1 / 3.

[0093] The actual spatial distance between the mission target points is calculated based on the spatial position information. The minimum spacing between adjacent oil tanks in the oil tank area is used as the characteristic distance of the oil tank area, for example, set to 10 meters. The actual spatial distance is divided by the characteristic distance of the oil tank area to obtain the standardized distance. For example, if the actual distance between two mission target points is 25 meters, the standardized distance is 25 / 10=2.5. The standardized distance is converted into a distance attenuation coefficient through an exponential decay function.

[0094] The feature fusion method based on entropy weight method is used to calculate the weight of environmental data similarity, task type similarity, task time overlap and distance decay coefficient. Assume that the calculated weights are 0.3, 0.25, 0.2 and 0.25 respectively. The task similarity matrix is ​​obtained by weighted fusion of environmental data similarity, task type similarity, task time overlap and distance decay coefficient.

[0095] The task similarity matrix is ​​subjected to threshold sparsification to obtain the initial coupling matrix. The threshold can be set to 0.5, and similarity values ​​less than 0.5 are set to 0, while those greater than or equal to 0.5 are retained as the original values. The initial coupling matrix is ​​optimized using a Gaussian smoothing kernel to obtain the final task coupling matrix. The kernel size of the Gaussian smoothing kernel is dynamically adjusted according to the number of tasks. For example, the kernel size can be set to the square root of the number of tasks rounded down. Assuming there are 100 tasks, the kernel size is 10. The initial coupling matrix is ​​convolved with a Gaussian smoothing kernel to obtain the final task coupling matrix.

[0096] The present invention constructs a task feature vector, comprehensively considers the task characteristics of multiple dimensions such as environmental data information, task type, time window and spatial position, and can comprehensively reflect the similarity and correlation between tasks; adopts a feature fusion method based on the entropy weight method to dynamically calculate the weights of different features, avoiding the subjectivity and inaccuracy that may be caused by artificially setting weights, and improving the objectivity and accuracy of task similarity calculation; through threshold sparsification processing and Gaussian smoothing kernel optimization, the density of the task coupling matrix is ​​effectively reduced, the coupling relationship between strongly correlated tasks is highlighted, and at the same time the local noise is smoothed, thereby improving the robustness and reliability of the task coupling matrix.

[0097] In an optional embodiment,

[0098] The path planning sub-model adopts an improved ant colony algorithm to perform multi-robot collaborative obstacle avoidance path planning, introduces environmental safety metrics into heuristic information, updates the pheromone matrix based on task priority, updates the dynamic taboo table based on the real-time position coordinates of the robots to record the robot's motion trajectory and occupied area, and generates a multi-robot collaborative obstacle avoidance path set, including the following steps:

[0099] An environmental safety metric is introduced into the heuristic information calculation, wherein the environmental safety metric is obtained by weighted combination of the oil tank pressure safety factor, the gas concentration safety factor and the temperature safety factor by weighted coefficients, wherein the oil tank pressure safety factor is calculated based on the deviation between the current pressure value and the safety pressure threshold, the gas concentration safety factor is calculated based on the ratio of the current concentration to the critical concentration, and the temperature safety factor is calculated based on the deviation between the current temperature and the normal temperature range, and the weighted coefficient is dynamically output by the task priority evaluation sub-model;

[0100] The pheromone matrix update rule is improved based on task priority. In local update, task priority and path length are used together as the calculation basis of pheromone increment. In global update, a task priority dynamic adjustment factor is introduced. The task priority dynamic adjustment factor changes exponentially with the ratio of the current task priority to the maximum priority.

[0101] Construct a dynamic taboo table in the time and space dimensions to handle multi-robot path conflicts. The dynamic taboo table includes time dimension taboo items and space dimension obstacle avoidance constraints. The time dimension taboo items record the occupied area of ​​each robot within the safe distance range at different times. The space dimension obstacle avoidance constraints are calculated based on the distance between any two robots at any time being greater than twice the safe distance.

[0102] Multi-robot path planning is performed based on the environmental safety metric, the updated pheromone matrix and the dynamic taboo table, the environmental safety metric and the task priority are used as input to calculate the state transition probability of the improved ant colony algorithm, and the path selection is performed according to the state transition probability during the path construction process. When it is detected that the next position violates the time dimension taboo item or the space dimension obstacle avoidance constraint in the dynamic taboo table, dynamic conflict resolution is performed; the dynamic conflict resolution includes: by calculating a set of potential conflicting spatiotemporal points, a temporary taboo area is constructed around the conflict point and a global taboo table is updated. When the distance between the robot and the temporary taboo area is less than a preset threshold, the environmental safety metric and the task priority are used as input to trigger local path re-planning. After the path construction is completed, global optimization is performed and the pheromone matrix and the dynamic taboo table are updated, and a set of collaborative obstacle avoidance paths for multiple robots is output.

[0103] Exemplarily, first, the environmental safety metric is introduced into the heuristic information calculation. The environmental safety metric is obtained by weighting the safety factor of the oil tank pressure, the safety factor of the gas concentration and the temperature safety factor by weighted coefficients. The oil tank pressure safety factor is calculated based on the deviation between the current pressure value and the safety pressure threshold. For example, if the current pressure is 0.8MPa and the safety pressure threshold is 1MPa, the pressure safety factor is 0.2. The gas concentration safety factor is calculated based on the ratio of the current concentration to the critical concentration. For example, if the current methane concentration is 2% and the critical concentration is 5%, the concentration safety factor is 0.6. The temperature safety factor is calculated based on the deviation between the current temperature and the normal temperature range. If the current temperature is 35°C and the normal range is 20-30°C, the temperature safety factor is 0.83. The weighted coefficient is dynamically output by the task priority evaluation submodel. For example, the weighted coefficients of pressure, concentration and temperature are 0.4, 0.3 and 0.3 respectively. The final environmental safety metric is 0.2×0.4+0.6×0.3+0.83×0.3=0.509.

[0104] Secondly, the pheromone matrix update rule is improved based on task priority. In local updates, the task priority and path length are used together as the basis for calculating the pheromone increment. For example, if the path length is 100m and the task priority is 0.8, the pheromone increment can be set to 0.8×(1 / 100)=0.008. In global updates, a dynamic adjustment factor of task priority is introduced, which changes exponentially with the ratio of the current task priority to the maximum priority.

[0105] Then, a dynamic taboo table of spatiotemporal dimensions is constructed to handle multi-robot path conflicts. The dynamic taboo table includes time dimension taboo items and space dimension obstacle avoidance constraints. The time dimension taboo items record the occupied area of ​​each robot within the safe distance range at different times. For example, at time t, the coordinates of robot A are (10, 20) and the safe distance is 2 meters. The taboo area is a circular area with a radius of 2 meters and a center of (10, 20). The space dimension obstacle avoidance constraint is calculated based on the distance between any two robots at any time being greater than twice the safe distance. For example, if the safe distances of robots A and B are both 2 meters, the distance between A and B should be greater than 4 meters.

[0106] Next, multi-robot path planning is performed based on the environmental safety metric, the updated pheromone matrix, and the dynamic taboo table. The state transition probability of the improved ant colony algorithm is calculated using the environmental safety metric and task priority as input. During the path construction process, path selection is performed based on the state transition probability. When it is detected that the next position violates the time dimension taboo item or the spatial dimension obstacle avoidance constraint in the dynamic taboo table, dynamic conflict resolution is performed.

[0107] The dynamic conflict resolution process is as follows: First, calculate the potential conflicting spatiotemporal point set. For example, the predicted positions of robots A and B at time t are (30, 40) and (32, 41), respectively, which are less than the safety distance of 4 meters. Then (31, 40.5) is a potential conflict point. Then, a temporary taboo area is constructed around the conflict point and the global taboo table is updated, such as a circular area with a radius of 2 meters and a center of (31, 40.5). When the distance between the robot and the temporary taboo area is less than the preset threshold (such as 1 meter), the environmental safety metric and task priority are used as input to trigger the local path replanning. After the path construction is completed, global optimization is performed and the pheromone matrix and dynamic taboo table are updated, and finally the collaborative obstacle avoidance path set of multiple robots is output.

[0108] By introducing environmental safety metrics as heuristic information, the present invention can consider the safety status of the oil tank area more comprehensively and improve the safety and reliability of path planning; constructing a dynamic taboo table in the time and space dimensions can effectively handle path conflicts between multiple robots and avoid collision risks; taking into account multiple factors such as environmental safety, task priority, path conflicts, etc., it can generate a safer, more efficient and coordinated multi-robot path planning scheme, which is suitable for multi-robot collaborative operations in complex oil tank area environments.

[0109] In an optional embodiment,

[0110] The step of inputting the task execution order and the collaborative obstacle avoidance path into a task allocation optimization module to generate an optimal task allocation strategy comprises:

[0111] Converting the task execution order into a task priority matrix, wherein the task priority matrix includes the front-to-back dependency relationship and execution priority weights between tasks, and converting the collaborative obstacle avoidance path into a path time-space matrix, wherein the path time-space matrix includes the time-space resource occupancy information of the robot motion trajectory;

[0112] A multi-objective task allocation optimization model is constructed based on the task priority matrix and the path time-space matrix, wherein the multi-objective task allocation optimization model takes maximizing task priority satisfaction and minimizing path execution cost as optimization objectives; in the multi-objective task allocation optimization model, a task priority order constraint is established based on the task priority matrix; a path conflict constraint is established based on the path time-space matrix, wherein the path conflict constraint quantifies the collision risk by calculating the minimum safe distance between any two robots in the time-space dimension; a timing constraint is established based on the forward and backward dependencies between tasks, wherein the timing constraint is used to regulate the execution order of interdependent tasks; an energy constraint is established based on the remaining power of the robot and the task resource requirements, wherein the energy constraint is used to constrain the resource allocation of the task allocation strategy;

[0113] A mixed integer programming method is used to solve the multi-objective task allocation optimization model, and the task priority satisfaction, path execution cost and completion time are combined to construct a weighted objective function; the task priority order constraints, path conflict constraints, timing constraints and energy constraints are converted into penalty terms of the weighted objective function through the Lagrangian relaxation method to form an augmented Lagrangian function; the augmented Lagrangian function is iteratively solved using a column generation method, wherein the main problem is restricted to determine the current task allocation strategy, the pricing subproblem generates a new feasible allocation strategy and adds it to the feasible solution space, and the subgradient method is used to update the Lagrangian multiplier to adjust the penalty intensity of constraint violation to obtain the optimal task allocation strategy.

[0114] Exemplarily, first, the task execution order is converted into a task priority matrix. The task priority matrix includes the dependencies between tasks and the execution priority weights.

[0115] Next, the collaborative obstacle avoidance path is converted into a path space-time matrix. The path space-time matrix contains the space-time resource occupancy information of the robot's motion trajectory. For example, for the collaborative obstacle avoidance path of three robots, a 3D matrix can be constructed, in which two dimensions represent plane coordinates and the third dimension represents time. A matrix element of 1 indicates that the space-time point is occupied by a robot, and a matrix element of 0 indicates that it is not occupied.

[0116] Then, a multi-objective task allocation optimization model is constructed based on the task priority matrix and the path time-space matrix. The optimization objectives of this model are to maximize the satisfaction of task priorities and minimize the cost of path execution. In the model, first, task priority order constraints are established based on the task priority matrix to ensure that high-priority tasks are executed first. Secondly, path conflict constraints are established based on the path time-space matrix. The collision risk is quantified by calculating the minimum safe distance between any two robots in the time-space dimension to ensure that there will be no collision between robots. Thirdly, timing constraints are established based on the front-to-back dependency relationship between tasks to regulate the execution order of interdependent tasks. Finally, energy constraints are established based on the remaining power of the robot and the task resource requirements to constrain the resource allocation of the task allocation strategy.

[0117] Next, the mixed integer programming method is used to solve the multi-objective task allocation optimization model. First, the task priority satisfaction, path execution cost and completion time are combined to construct a weighted objective function. For example, the weight coefficients w1, w2, and w3 can be set to represent the importance of these three objectives respectively. Then, the various constraints are converted into penalty terms of the weighted objective function through the Lagrangian relaxation method to form an augmented Lagrangian function.

[0118] Finally, the column generation method is used to iteratively solve the augmented Lagrangian function. In each iteration, the constraint master problem is first solved to determine the current task allocation strategy. Then the pricing subproblem is solved to generate a new feasible allocation strategy and add it to the feasible solution space. At the same time, the Lagrangian multiplier is updated using the subgradient method to adjust the penalty intensity for constraint violations. This process is repeated until convergence to obtain the optimal task allocation strategy.

[0119] For example, suppose there are 3 robots and 5 tasks. First, a 5x5 task priority matrix is ​​constructed according to the order of task execution, and a priority weight is assigned to each task. Then a 3D path space-time matrix is ​​constructed according to the collaborative obstacle avoidance path. An optimization model is constructed based on these two matrices, and the objective function weight coefficients w1=0.5, w2=0.3, and w3=0.2 are set. Through iterative solution, a task allocation strategy is finally obtained, for example, robot 1 performs tasks 1 and 4, robot 2 performs tasks 2 and 5, and robot 3 performs task 3.

[0120] The present invention effectively captures the dependencies between tasks and the spatiotemporal characteristics of robot motion by converting the task execution sequence and the collaborative obstacle avoidance path into a matrix form, thus providing a reliable mathematical basis for subsequent optimization. By constructing a multi-objective task allocation optimization model, it comprehensively considers multiple objectives such as task priority, path execution cost, and completion time, and introduces a variety of constraints to ensure that the generated task allocation strategy can not only meet the task requirements but also ensure the feasibility and safety of execution. The mixed integer programming and column generation method are used to solve the optimization model, which can efficiently handle large-scale task allocation problems, and continuously improve the quality of the solution through iterative optimization, ultimately obtaining a task allocation strategy close to the global optimal.

[0121] In an optional embodiment,

[0122] Based on the optimal task allocation strategy, the inspection task instruction is issued to the robot, and the task execution data and environmental data of the robot are collected in real time through the Internet of Things. When a robot task conflict or execution abnormality is detected, the emergency module is started. The emergency module performs situational awareness and risk assessment based on the task execution data and environmental data, and uses the robot's historical task completion rate and remaining power as constraints. The optimal task allocation strategy is optimized and adjusted using a reinforcement learning algorithm. The steps of generating an updated task allocation strategy include:

[0123] Collect robot layer data and environment layer data through the Internet of Things, where the robot layer data includes the robot's position coordinates, power status, task execution status and motion parameters, and the environment layer data includes gas concentration, temperature distribution, equipment vibration and pipeline pressure, and perform outlier processing, standardization processing and time alignment processing on the robot layer data and the environment layer data; build a situational awareness model based on the processed data, calculate the number of task conflicts by calculating the task time overlap and the minimum distance of the robot path, identify the degree of execution abnormality by task progress deviation and trajectory deviation, and calculate the comprehensive risk according to the weighted combination of environmental risk, task risk and equipment risk;

[0124] Construct a reinforcement learning network based on a dual deep Q network, and use the robot state vector, task state vector and environment state vector to form a state space, and use task reallocation, task suspension, task resumption and task termination as the action space. The robot state vector includes the position coordinates and the remaining power, the task state vector includes the task priority, deadline and execution progress, and the environment state vector includes the comprehensive risk, the execution abnormality degree and the number of task conflicts. Design a combined reward function of task completion reward, risk penalty, resource consumption penalty and time penalty, wherein the risk penalty is positively correlated with the execution abnormality degree. Train the dual deep Q network based on the combined reward function, use the experience replay mechanism to store state transition samples, select actions through the ε greedy strategy, and regularly update the target network parameters of the dual deep Q network.

[0125] When the comprehensive risk exceeds the risk threshold, the number of task conflicts exceeds the conflict threshold, or the remaining power of the robot is lower than the power threshold, the optimization of the optimal task allocation strategy is triggered, and the current robot state vector, task state vector and environment state vector are input into the trained dual deep Q network to obtain the optimized task allocation strategy. After the feasibility of the optimized task allocation strategy is verified, a scheduling instruction is issued to the robot.

[0126] Exemplarily, first, robot layer data and environment layer data are collected through the Internet of Things. Robot layer data includes the robot's position coordinates, power status, task execution status, and motion parameters. For example, the position coordinates of a robot are (100, 200), the power is 80%, it is executing task A, and the movement speed is 0.5m / s. Environmental layer data includes gas concentration, temperature distribution, equipment vibration, and pipeline pressure. For example, the methane concentration in a certain area is 500ppm, the temperature is 25℃, the equipment vibration frequency is 20Hz, and the pipeline pressure is 0.6MPa.

[0127] The collected data is preprocessed, including outlier processing, standardization, and time series alignment. The 3σ criterion is used for outlier processing, and data beyond the mean ±3 times the standard deviation is considered as an outlier and removed. The minimum-maximum standardization method is used for standardization, mapping the data to the [0, 1] interval. The linear interpolation method is used for time series alignment, unifying data of different frequencies to the same timestamp.

[0128] Next, a situational awareness model is built based on the processed data. Calculate the task time overlap, for example, the execution time overlap rate of task A and task B is 60%. Calculate the minimum distance of the robot path, for example, the minimum distance between robot 1 and robot 2 is 2 meters. Calculate the number of task conflicts based on the time overlap and the minimum path distance. Calculate the task progress deviation by comparing the actual task progress with the planned progress, for example, the actual progress of task C lags behind the planned progress by 15%. Calculate the trajectory deviation by comparing the actual trajectory with the planned trajectory, for example, the average deviation of a certain trajectory is 0.3 meters. Identify the degree of execution abnormality based on the task progress deviation and trajectory deviation.

[0129] When calculating the comprehensive risk, three aspects are considered: environmental risk, task risk, and equipment risk. Environmental risk is assessed based on parameters such as gas concentration and temperature, task risk is assessed based on factors such as task complexity and urgency, and equipment risk is assessed based on indicators such as equipment status and usage time. The three types of risks are weighted and combined to obtain a comprehensive risk. The weights can be adjusted according to actual conditions, for example, the environmental risk weight is 0.4, the task risk weight is 0.3, and the equipment risk weight is 0.3.

[0130] Then, a reinforcement learning network based on a dual deep Q network is constructed. The state space includes the robot state vector, the task state vector, and the environment state vector. The robot state vector includes the position coordinates and the remaining power, such as (100, 200, 80%). The task state vector includes the task priority, deadline and execution progress, such as (high priority, 2 hours later, 50%). The environment state vector includes the comprehensive risk, the degree of execution abnormality, and the number of task conflicts, such as (0.6, 0.2, 2). The action space includes four actions: task reallocation, task suspension, task resumption, and task termination.

[0131] Design a combined reward function, including task completion reward, risk penalty, resource consumption penalty, and time penalty. Task completion reward is proportional to task importance and completion quality. Risk penalty is positively correlated with the degree of execution abnormality. The higher the degree of abnormality, the greater the penalty. Resource consumption penalty is proportional to the robot's power consumption. Time penalty is proportional to task delay time.

[0132] The experience replay mechanism is used to store state transition samples, which include information such as the current state, executed actions, rewards, next state, and whether the training is over. An experience replay pool of 10,000 is used, and 256 samples are randomly sampled for learning each time. The ε greedy strategy is used for action selection, and the initial ε value is set to 1, which is gradually reduced to 0.1 as the training progresses to ensure the balance between exploration and utilization. The target network parameters are updated every 100 training steps to ensure the stability of learning.

[0133] Finally, set the trigger conditions to optimize the task allocation strategy. When the comprehensive risk exceeds 0.8, the number of task conflicts exceeds 3, or the remaining battery power of the robot is less than 20%, the optimization process is triggered. Input the current state into the trained dual-depth Q network to obtain the optimized task allocation strategy. For example, if the network output action is "task reallocation", the system will reallocate tasks to other available robots according to the current situation. Verify the feasibility of the optimized strategy and issue scheduling instructions to the relevant robots after ensuring that all constraints are met.

[0134] The present invention collects multi-dimensional data in real time through the Internet of Things technology, constructs a comprehensive situational awareness model, can promptly identify task conflicts and execution anomalies, and improves the execution efficiency and safety of inspection tasks; adopts a dual deep Q network as the core of the reinforcement learning algorithm, combined with technologies such as experience replay and target network, effectively improves the stability and convergence speed of learning, and realizes continuous optimization and adaptive adjustment of task allocation strategies; designs a combined reward function that comprehensively considers task completion, risk level, resource consumption and time delay, takes into account safety and resource utilization while ensuring task efficiency, and realizes balanced optimization of multiple objectives.

[0135] In an optional embodiment,

[0136] Designing a combined reward function of task completion reward, risk penalty, resource consumption penalty and time penalty, wherein the risk penalty is positively correlated with the execution abnormality degree, training a dual deep Q network based on the combined reward function, using an experience replay mechanism to store state transition samples, selecting actions through an ε greedy strategy, and regularly updating the target network parameters of the dual deep Q network includes:

[0137] The task completion reward is calculated based on the product of task progress, task priority and quality score, the quality score is based on the task completion stability and timeliness assessment, the risk penalty includes execution abnormality penalty and safety distance breach penalty, the resource consumption penalty is calculated based on the power consumption rate, and the time penalty is calculated based on the task urgency and overtime duration; the dual deep Q network includes an evaluation network and a target network, both of which adopt a three-layer fully connected structure, the input layer receives the state vector of the state space, the hidden layer uses the ReLU activation function to extract features, and the output layer generates action value estimation;

[0138] Establish an experience replay buffer with a fixed capacity to store transition samples including the current state, executed actions, rewards obtained, and next state, and randomly sample from the experience replay buffer based on a preset batch size; execute an action selection strategy based on the action value estimation, select an action based on the ε greedy strategy at each decision moment, select the action with the maximum action value with a probability of 1-ε, and randomly select an action for exploration with a probability of ε, where the ε value decreases linearly with the training rounds;

[0139] The target Q value output by the target network is calculated using the temporal difference learning method and the Bellman equation. The error between the Q value predicted by the evaluation network and the target Q value is calculated based on the smooth L1 loss function. The evaluation network parameters are optimized through back propagation. The evaluation network parameters are soft-updated to the target network according to the preset update cycle. The training process is continuously iterated. In each round of training, the state vector of the state space is input into the evaluation network to obtain the action selection, the selected action is executed to obtain the new state and reward value, and the transferred samples are stored in the experience replay buffer. When the cumulative number of samples reaches the training batch requirement, the evaluation network parameters are updated through back propagation based on the loss function, and the updated evaluation network parameters are copied to the target network through the soft update mechanism according to the preset update cycle until the preset maximum number of iterations is reached.

[0140] Exemplarily, first, a combined reward function is designed. The function contains four parts: task completion reward, risk penalty, resource consumption penalty and time penalty. The task completion reward is calculated based on the product of task progress, task priority and quality score. Among them, the task progress can be expressed as a decimal between 0 and 1 to represent the completion percentage, the task priority can be set to an integer from 1 to 10, and the quality score is based on the stability and timeliness of task completion, which can be expressed as a score from 0 to 100. For example, if the progress of a task is 0.8, the priority is 8, and the quality score is 90, then the task completion reward is 0.8×8×90=576. The risk penalty includes two parts: execution abnormality penalty and safety distance breach penalty. The execution abnormality penalty is positively correlated with the degree of execution abnormality and can be set to the square value of the degree of abnormality. The safety distance breach penalty can set a fixed penalty value based on the breach distance. The resource consumption penalty is calculated based on the power consumption rate and can be set as a linear function of the consumption rate. The time penalty is calculated based on the task urgency and the overtime duration, and can be set as a logarithmic function of the product of the two.

[0141] Next, a dual deep Q network is constructed. The network includes an evaluation network and a target network, both of which use a three-layer fully connected structure. The input layer receives the state vector of the state space, which can contain state information in multiple dimensions such as the current task progress, remaining resources, and execution time. The hidden layer uses the ReLU activation function to extract features, and 64 neurons can be set. The output layer generates action value estimates, and the number of neurons is equal to the size of the action space.

[0142] Then, a fixed-capacity experience replay buffer is established. The buffer is used to store transition samples including the current state, the action performed, the reward obtained, and the next state. The buffer capacity can be set to 10,000, and each sample contains 4 elements. When sampling, random sampling is performed from the buffer based on the preset batch size (such as 32).

[0143] In action selection, the ε greedy strategy is adopted. At each decision moment, the action with the maximum action value is selected with a probability of 1-ε, and the action is randomly selected for exploration with a probability of ε. The ε value decreases linearly with the training rounds, for example, it can decrease linearly from the initial value of 0.9 to the final value of 0.1.

[0144] Next, network training and parameter update are performed. The target Q value output by the target network is calculated using the temporal difference learning method and the Bellman equation. The error between the Q value predicted by the evaluation network and the target Q value is calculated based on the smoothed L1 loss function. The evaluation network parameters are optimized by back propagation. The evaluation network parameters are soft-updated to the target network according to the preset update cycle (e.g., every 100 steps), and the update weight can be set to 0.01.

[0145] Finally, the training process is iterated continuously. In each round of training, the state vector of the state space is input into the evaluation network to obtain the action selection. The selected action is executed to obtain the new state and reward value. The transfer sample is stored in the experience replay buffer. When the cumulative number of samples reaches the training batch requirement (such as 32), the evaluation network parameters are updated by back propagation based on the loss function. The updated evaluation network parameters are copied to the target network through the soft update mechanism according to the preset update cycle. The above process is repeated until the preset maximum number of iterations (such as 10,000 rounds) is reached.

[0146] By designing a combined reward function, the present invention comprehensively considers multiple aspects such as task completion, execution risk, resource consumption and time constraints, etc., which can more comprehensively evaluate and optimize the task execution process and improve the rationality of task allocation and execution efficiency.

[0147] Figure 2 FIG. 1 is a schematic diagram of a task allocation system for inspection robots in oil tank areas based on the Internet of Things according to an embodiment of the present invention. Figure 2 As shown, the system comprises:

[0148] The first unit is used to receive environmental data information collected by the IoT sensor nodes in the oil tank area and real-time status information uploaded by multiple intelligent inspection robots, wherein the environmental data information includes oil tank pressure data, temperature data and gas concentration data, and the real-time status information includes the remaining power of the robot, task execution progress and real-time position coordinates;

[0149] The second unit is used to build a multi-objective collaborative scheduling model based on a deep neural network, input environmental data information and real-time status information into the multi-objective collaborative scheduling model, use the spatiotemporal attention mechanism to extract spatiotemporal fusion features, and establish a task priority evaluation sub-model and a path planning sub-model based on the spatiotemporal fusion features. The task priority evaluation sub-model calculates the task execution order based on the safety level of the oil tank area and the urgency of the task. The path planning sub-model calculates the collaborative obstacle avoidance path of multiple robots based on the improved ant colony algorithm, and inputs the task execution order and the collaborative obstacle avoidance path into the task allocation optimization module to generate the optimal task allocation strategy;

[0150] The third unit is used to issue inspection task instructions to the robot based on the optimal task allocation strategy, collect the robot's task execution data and environmental data in real time through the Internet of Things, and start the emergency module when a robot task conflict or execution abnormality is detected. The emergency module performs situational awareness and risk assessment based on the task execution data and environmental data, and uses the robot's historical task completion rate and remaining power as constraints, and uses a reinforcement learning algorithm to optimize and adjust the optimal task allocation strategy to generate an updated task allocation strategy.

[0151] According to a third aspect of the embodiments of the present invention,

[0152] An electronic device is provided, comprising:

[0153] processor;

[0154] a memory for storing processor-executable instructions;

[0155] The processor is configured to call the instructions stored in the memory to execute the aforementioned method.

[0156] According to a fourth aspect of the embodiments of the present invention,

[0157] A computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the aforementioned method is implemented.

[0158] The present invention may be a method, an apparatus, a system and / or a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for executing various aspects of the present invention.

[0159] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for allocating tasks of inspection robots in oil depot tank areas based on the Internet of Things, characterized in that: include: Receive environmental data information collected by IoT sensor nodes in the oil tank area and real-time status information uploaded by multiple intelligent inspection robots, the environmental data information includes oil tank pressure data, temperature data and gas concentration data, and the real-time status information includes the remaining power of the robot, task execution progress and real-time position coordinates; Construct a multi-objective collaborative scheduling model based on a deep neural network, input environmental data information and real-time status information into the multi-objective collaborative scheduling model, use the spatiotemporal attention mechanism to extract spatiotemporal fusion features, and establish a task priority evaluation sub-model and a path planning sub-model based on the spatiotemporal fusion features. The task priority evaluation sub-model calculates the task execution order based on the safety level of the oil tank area and the urgency of the task. The path planning sub-model calculates the collaborative obstacle avoidance path of multiple robots based on the improved ant colony algorithm. The task execution order and the collaborative obstacle avoidance path are input into the task allocation optimization module to generate the optimal task allocation strategy. Based on the optimal task allocation strategy, inspection task instructions are issued to the robot, and task execution data and environmental data of the robot are collected in real time through the Internet of Things. When a robot task conflict or execution abnormality is detected, an emergency module is started. The emergency module performs situational awareness and risk assessment based on the task execution data and environmental data, and uses the robot's historical task completion rate and remaining power as constraints, and uses a reinforcement learning algorithm to optimize and adjust the optimal task allocation strategy to generate an updated task allocation strategy; Inputting environmental data information and real-time status information into the multi-objective collaborative scheduling model, adopting the spatiotemporal attention mechanism to extract spatiotemporal fusion features, establishing a task priority evaluation sub-model and a path planning sub-model based on the spatiotemporal fusion features, the task priority evaluation sub-model calculates the task execution order in combination with the safety level of the oil tank area and the urgency of the task, and the path planning sub-model calculates the collaborative obstacle avoidance path of multiple robots based on the improved ant colony algorithm, including the following steps: For environmental data information and real-time status information, time windows are constructed to calculate the time attention weight respectively. Based on the IoT sensor node position and the robot's real-time position coordinates, a spatial correlation matrix is ​​constructed to calculate the spatial attention weight. The time attention weight and the spatial attention weight are fused to obtain the spatiotemporal fusion feature. A task priority evaluation sub-model is constructed based on the spatiotemporal fusion feature, the task priority evaluation sub-model calculates the safety level of the oil tank farm based on the environmental data information, uses a long short-term memory network to predict the future state of the environmental data to obtain a risk trend index, calculates the task urgency in combination with the oil tank farm safety level and the real-time state information, calculates the feature similarity between tasks and the spatial distance between task target points to construct a task coupling matrix, inputs the risk trend index, task urgency and task coupling matrix into an adaptive weight network to obtain dynamic weights, calculates task priority scores based on the dynamic weights and outputs the task execution order; A path planning sub-model is constructed based on the spatiotemporal fusion features and the task execution order. The path planning sub-model adopts an improved ant colony algorithm to perform multi-robot collaborative obstacle avoidance path planning, introduces environmental safety metrics into heuristic information, updates the pheromone matrix based on task priority, updates the dynamic taboo table according to the real-time position coordinates of the robots to record the robot's motion trajectory and occupied area, and generates a multi-robot collaborative obstacle avoidance path set.

2. The method according to claim 1, characterized in that The steps of calculating the feature similarity between tasks and the spatial distance between task target points to construct the task coupling matrix include: Constructing a task feature vector, wherein the task feature vector includes environmental data information, a task type code, a time window parameter, and spatial location information, wherein the task type code includes type information of equipment inspection, pipeline inspection, and area inspection, the time window parameter includes the task start time and end time, and the spatial location information includes the coordinates of the task target point; Based on the environmental data information, the environmental data similarity is calculated, and the cosine similarity method is used to calculate the comprehensive similarity of the oil tank pressure data, temperature data and gas concentration data between tasks; based on the task type code, a task type similarity matrix is ​​constructed, and the task type similarity is determined according to the degree of association between equipment inspection, pipeline inspection and area inspection; based on the time window parameter, the task time overlap is calculated, and the overlapping ratio of the task execution time is calculated according to the task start time and end time; based on the spatial position information, the actual spatial distance between the task target points is calculated, and the minimum spacing between adjacent oil tanks in the oil tank area is used as the characteristic distance of the oil tank area, and the actual spatial distance is divided by the characteristic distance of the oil tank area to obtain the standardized distance, and the standardized distance is converted into a distance attenuation coefficient through an exponential decay function; The feature fusion method based on entropy weight method is used to calculate the weight of environmental data similarity, task type similarity, task time overlap and distance decay coefficient. The task similarity matrix is ​​obtained by weighted fusion of environmental data similarity, task type similarity, task time overlap and distance decay coefficient. The task similarity matrix is ​​subjected to threshold sparse processing to obtain an initial coupling matrix, and the initial coupling matrix is ​​optimized using a Gaussian smoothing kernel to obtain a final task coupling matrix, wherein the kernel size of the Gaussian smoothing kernel is dynamically adjusted according to the number of tasks.

3. The method according to claim 1, characterized in that The path planning sub-model adopts an improved ant colony algorithm to perform multi-robot collaborative obstacle avoidance path planning, introduces environmental safety metrics into heuristic information, updates the pheromone matrix based on task priority, updates the dynamic taboo table based on the real-time position coordinates of the robots to record the robot's motion trajectory and occupied area, and generates a multi-robot collaborative obstacle avoidance path set, including the following steps: An environmental safety metric is introduced into the heuristic information calculation, wherein the environmental safety metric is obtained by weighted combination of the oil tank pressure safety factor, the gas concentration safety factor and the temperature safety factor by weighted coefficients, wherein the oil tank pressure safety factor is calculated based on the deviation between the current pressure value and the safety pressure threshold, the gas concentration safety factor is calculated based on the ratio of the current concentration to the critical concentration, and the temperature safety factor is calculated based on the deviation between the current temperature and the normal temperature range, and the weighted coefficient is dynamically output by the task priority evaluation sub-model; The pheromone matrix update rule is improved based on task priority. In local update, task priority and path length are used together as the calculation basis of pheromone increment. In global update, a task priority dynamic adjustment factor is introduced. The task priority dynamic adjustment factor changes exponentially with the ratio of the current task priority to the maximum priority. Construct a dynamic taboo table in the time and space dimensions to handle multi-robot path conflicts. The dynamic taboo table includes time dimension taboo items and space dimension obstacle avoidance constraints. The time dimension taboo items record the occupied area of ​​each robot within the safe distance range at different times. The space dimension obstacle avoidance constraints are calculated based on the distance between any two robots at any time being greater than twice the safe distance. Multi-robot path planning is performed based on the environmental safety metric, the updated pheromone matrix and the dynamic taboo table, the environmental safety metric and the task priority are used as input to calculate the state transition probability of the improved ant colony algorithm, and the path selection is performed according to the state transition probability during the path construction process. When it is detected that the next position violates the time dimension taboo item or the space dimension obstacle avoidance constraint in the dynamic taboo table, dynamic conflict resolution is performed; the dynamic conflict resolution includes: by calculating a set of potential conflicting spatiotemporal points, a temporary taboo area is constructed around the conflict point and a global taboo table is updated. When the distance between the robot and the temporary taboo area is less than a preset threshold, the environmental safety metric and the task priority are used as input to trigger local path re-planning. After the path construction is completed, global optimization is performed and the pheromone matrix and the dynamic taboo table are updated, and a set of collaborative obstacle avoidance paths for multiple robots is output.

4. The method according to claim 1, characterized in that: The step of inputting the task execution order and the collaborative obstacle avoidance path into a task allocation optimization module to generate an optimal task allocation strategy comprises: Converting the task execution order into a task priority matrix, wherein the task priority matrix includes the front-to-back dependency relationship and execution priority weights between tasks, and converting the collaborative obstacle avoidance path into a path time-space matrix, wherein the path time-space matrix includes the time-space resource occupancy information of the robot motion trajectory; A multi-objective task allocation optimization model is constructed based on the task priority matrix and the path time-space matrix, wherein the multi-objective task allocation optimization model takes maximizing task priority satisfaction and minimizing path execution cost as optimization objectives; in the multi-objective task allocation optimization model, a task priority order constraint is established based on the task priority matrix; a path conflict constraint is established based on the path time-space matrix, wherein the path conflict constraint quantifies the collision risk by calculating the minimum safe distance between any two robots in the time-space dimension; a timing constraint is established based on the forward and backward dependencies between tasks, wherein the timing constraint is used to regulate the execution order of interdependent tasks; an energy constraint is established based on the remaining power of the robot and the task resource requirements, wherein the energy constraint is used to constrain the resource allocation of the task allocation strategy; A mixed integer programming method is used to solve the multi-objective task allocation optimization model, and the task priority satisfaction, path execution cost and completion time are combined to construct a weighted objective function; the task priority order constraints, path conflict constraints, timing constraints and energy constraints are converted into penalty terms of the weighted objective function through the Lagrangian relaxation method to form an augmented Lagrangian function; the augmented Lagrangian function is iteratively solved using a column generation method, wherein the main problem is restricted to determine the current task allocation strategy, the pricing subproblem generates a new feasible allocation strategy and adds it to the feasible solution space, and the subgradient method is used to update the Lagrangian multiplier to adjust the penalty intensity of constraint violation to obtain the optimal task allocation strategy.

5. The method according to claim 1, characterized in that Based on the optimal task allocation strategy, the inspection task instruction is issued to the robot, and the task execution data and environmental data of the robot are collected in real time through the Internet of Things. When a robot task conflict or execution abnormality is detected, the emergency module is started. The emergency module performs situational awareness and risk assessment based on the task execution data and environmental data, and uses the robot's historical task completion rate and remaining power as constraints. The optimal task allocation strategy is optimized and adjusted using a reinforcement learning algorithm. The steps of generating an updated task allocation strategy include: Collect robot layer data and environment layer data through the Internet of Things, where the robot layer data includes the robot's position coordinates, power status, task execution status and motion parameters, and the environment layer data includes gas concentration, temperature distribution, equipment vibration and pipeline pressure, and perform outlier processing, standardization processing and time alignment processing on the robot layer data and the environment layer data; build a situational awareness model based on the processed data, calculate the number of task conflicts by calculating the task time overlap and the minimum distance of the robot path, identify the degree of execution abnormality by task progress deviation and trajectory deviation, and calculate the comprehensive risk according to the weighted combination of environmental risk, task risk and equipment risk; Construct a reinforcement learning network based on a dual deep Q network, and use the robot state vector, task state vector and environment state vector to form a state space, and use task reallocation, task suspension, task resumption and task termination as the action space. The robot state vector includes the position coordinates and the remaining power, the task state vector includes the task priority, deadline and execution progress, and the environment state vector includes the comprehensive risk, the execution abnormality degree and the number of task conflicts. Design a combined reward function of task completion reward, risk penalty, resource consumption penalty and time penalty, wherein the risk penalty is positively correlated with the execution abnormality degree. Train the dual deep Q network based on the combined reward function, use the experience replay mechanism to store state transition samples, select actions through the ε greedy strategy, and regularly update the target network parameters of the dual deep Q network. When the comprehensive risk exceeds the risk threshold, the number of task conflicts exceeds the conflict threshold, or the remaining power of the robot is lower than the power threshold, the optimization of the optimal task allocation strategy is triggered, and the current robot state vector, task state vector and environment state vector are input into the trained dual deep Q network to obtain the optimized task allocation strategy. After the feasibility of the optimized task allocation strategy is verified, a scheduling instruction is issued to the robot.

6. The method according to claim 5, characterized in that Designing a combined reward function of task completion reward, risk penalty, resource consumption penalty and time penalty, wherein the risk penalty is positively correlated with the execution abnormality degree, training a dual deep Q network based on the combined reward function, using an experience replay mechanism to store state transition samples, performing action selection through an ε greedy strategy, and regularly updating the target network parameters of the dual deep Q network includes: The task completion reward is calculated based on the product of task progress, task priority and quality score, the quality score is based on the task completion stability and timeliness assessment, the risk penalty includes execution abnormality penalty and safety distance breach penalty, the resource consumption penalty is calculated based on the power consumption rate, and the time penalty is calculated based on the task urgency and overtime duration; the dual deep Q network includes an evaluation network and a target network, both of which adopt a three-layer fully connected structure, the input layer receives the state vector of the state space, the hidden layer uses the ReLU activation function to extract features, and the output layer generates action value estimation; Establish an experience replay buffer with a fixed capacity to store transition samples including the current state, executed actions, rewards obtained, and next state, and randomly sample from the experience replay buffer based on a preset batch size; execute an action selection strategy based on the action value estimation, select an action based on the ε greedy strategy at each decision moment, select the action with the maximum action value with a probability of 1-ε, and randomly select an action for exploration with a probability of ε, where the ε value decreases linearly with the training rounds; The target Q value output by the target network is calculated using the temporal difference learning method and the Bellman equation. The error between the Q value predicted by the evaluation network and the target Q value is calculated based on the smooth L1 loss function. The evaluation network parameters are optimized through back propagation. The evaluation network parameters are soft-updated to the target network according to the preset update cycle. The training process is continuously iterated. In each round of training, the state vector of the state space is input into the evaluation network to obtain the action selection, the selected action is executed to obtain the new state and reward value, and the transferred samples are stored in the experience replay buffer. When the cumulative number of samples reaches the training batch requirement, the evaluation network parameters are updated through back propagation based on the loss function, and the updated evaluation network parameters are copied to the target network through the soft update mechanism according to the preset update cycle until the preset maximum number of iterations is reached.

7. An oil depot tank area inspection robot task allocation system based on the Internet of Things, used to implement the method described in any one of claims 1 to 6, characterized in that: include: The first unit is used to receive environmental data information collected by the IoT sensor nodes in the oil depot tank area and real-time status information uploaded by multiple intelligent inspection robots, wherein the environmental data information includes oil tank pressure data, temperature data and gas concentration data, and the real-time status information includes the remaining power of the robot, task execution progress and real-time position coordinates; The second unit is used to build a multi-objective collaborative scheduling model based on a deep neural network, input environmental data information and real-time status information into the multi-objective collaborative scheduling model, use the spatiotemporal attention mechanism to extract spatiotemporal fusion features, and establish a task priority evaluation sub-model and a path planning sub-model based on the spatiotemporal fusion features. The task priority evaluation sub-model calculates the task execution order based on the safety level of the oil tank area and the urgency of the task. The path planning sub-model calculates the collaborative obstacle avoidance path of multiple robots based on the improved ant colony algorithm, and inputs the task execution order and the collaborative obstacle avoidance path into the task allocation optimization module to generate the optimal task allocation strategy; The third unit is used to issue inspection task instructions to the robot based on the optimal task allocation strategy, collect the robot's task execution data and environmental data in real time through the Internet of Things, and start the emergency module when a robot task conflict or execution abnormality is detected. The emergency module performs situational awareness and risk assessment based on the task execution data and environmental data, and uses the robot's historical task completion rate and remaining power as constraints, and uses a reinforcement learning algorithm to optimize and adjust the optimal task allocation strategy to generate an updated task allocation strategy.

8. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the method according to any one of claims 1 to 6.

9. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Robot inspection path planning method and system based on cloud platform

    CN117782097A

  • Deep reinforcement learning cooperative scheduling method and device for heterogeneous computing resources

    CN117909044A