Automatic closing robot control system for gas cylinder leakage in hazardous chemical accident site

By integrating high-precision sensors and deep learning networks into a hierarchical reinforcement learning framework, and optimizing the balance between exploration and utilization, the decision reliability problem of autonomous shut-off control of gas cylinder leaks at hazardous chemical accident sites was solved, achieving efficient and reliable valve shut-off operation.

CN121245845BActive Publication Date: 2026-04-24INST OF URBAN SAFETY & ENVIRONMENTAL SCI BEIJING ACAD OF SCI & TECH
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
INST OF URBAN SAFETY & ENVIRONMENTAL SCI BEIJING ACAD OF SCI & TECH
Filing Date
2025-11-17
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

In hazardous chemical accident sites where gas cylinders leak, existing robot control technologies struggle to balance exploration and utilization in unknown environments due to the difficulty of reinforcement learning algorithms, leading to reduced reliability of autonomous shutdown control decisions.

Method used

Employing a high-precision gas sensor array, multispectral vision sensor, and omnidirectional mobile robot platform, combined with deep convolutional neural networks and Bayesian inference networks, and through a hierarchical reinforcement learning framework and risk constraint strategy, the system optimizes the balance between exploration and utilization, generates a sequence of decision-making actions, and drives the robotic arm to perform valve closing operations.

Benefits of technology

In the context of unknown hazardous chemical accidents, it is essential to effectively balance exploration and utilization, reduce unnecessary risky behaviors, and improve the decision-making reliability and operational efficiency of autonomous shutdown control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121245845B_ABST
    Figure CN121245845B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of hazardous chemical emergency robot control, and particularly relates to a hazardous chemical accident site gas cylinder leakage automatic closing robot control system, which comprises an environment data acquisition module, a state reconstruction module, a dynamic risk assessment module, an adaptive exploration decision module and a control output module. The environment data acquisition module acquires multi-modal sensor data in real time through physical equipment devices and aligns the data by using a time stamp synchronization algorithm. The state reconstruction module reconstructs a three-dimensional environment map by using a deep convolutional neural network and a point cloud segmentation algorithm. The dynamic risk assessment module integrates a Bayesian inference network to predict a risk output dynamic risk field map. The adaptive exploration decision module uses a hierarchical reinforcement learning framework and a risk constraint strategy to adjust the exploration and utilization balance. The control output module converts a decision action sequence into a robot joint trajectory. The system coordination engine aggregates module outputs to coordinate operation through a federated learning framework. The system realizes global adaptive decision making through multi-module collaborative evolution, effectively balances exploration and utilization, and improves the reliability and efficiency of autonomous closing control.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of emergency robot control technology for hazardous chemicals, specifically to an automatic shut-off robot control system for gas cylinder leaks at hazardous chemical accident sites. Background Technology

[0002] In emergency response to hazardous chemical accidents, gas cylinder leaks can trigger a chain reaction of hazards, and existing manual shut-off operations face extremely high safety risks. Current robotic control technology employs autonomous or remotely operated robot systems that integrate gas sensors and vision recognition modules to detect leak sources and locate gas cylinder valves in real time. By processing environmental data through advanced control algorithms, the robot drives a robotic arm or specialized actuator to precisely execute the shut-off action, effectively isolating the spread of hazardous substances.

[0003] Existing robot control technologies suffer from the following technical challenges. Specifically, in emergency scenarios involving hazardous chemical accidents, the reinforcement learning algorithms used in robot control rely on environmental interactions to optimize decisions. However, the unknown accident scene presents complexities such as dynamic leaks and uncertain obstacle distribution, making it difficult for the algorithm to balance exploring new actions to acquire information with utilizing known safety strategies. Overexploration may cause the robot to execute unnecessary paths or test risky actions, such as attempting to detour around unknown areas instead of directly closing the valve when approaching a leaking gas cylinder, thereby delaying the response or exacerbating the leak. On the other hand, overutilization may prevent the algorithm from adapting to sudden changes such as signs of a secondary explosion, ultimately reducing the reliability of autonomous shutdown control decisions. Summary of the Invention

[0004] To address the shortcomings of existing technologies, this invention provides an automatic shut-off robot control system for gas cylinder leaks at hazardous chemical accident sites. This invention solves the technical problem of reduced reliability of autonomous shut-off control decisions caused by an imbalance between exploration and utilization in unknown hazardous chemical accident environments due to reinforcement learning algorithms.

[0005] To solve the above-mentioned technical problems, the specific contents of the present invention are as follows:

[0006] The automatic shut-off robot control system for gas cylinder leaks at hazardous chemical accident sites provided by this invention includes:

[0007] A control device and a physical device, wherein the control device establishes a communication connection with the physical device;

[0008] The physical device includes a high-precision gas sensor array, a multispectral vision sensor, an omnidirectional mobile robot platform, and a seven-degree-of-freedom robotic arm actuator.

[0009] The control device includes an environmental data acquisition module, a state reconstruction module, a dynamic risk assessment module, an adaptive exploration decision-making module, a control output module, and a system coordination engine;

[0010] The environmental data acquisition module collects gas concentration gradients, thermal imaging data, stereoscopic visual flow, and laser point cloud sequences at the hazardous chemical accident site in real time through the physical equipment, and uses a timestamp synchronization algorithm to align multimodal sensor data and outputs a synchronization data packet to the state reconstruction module.

[0011] The state reconstruction module receives the synchronization data packet, processes multimodal sensor data through a deep convolutional neural network and a point cloud segmentation algorithm, reconstructs a three-dimensional environment map, identifies the topology of the leak source, the pose of the gas cylinder valve and the trajectory of dynamic obstacles, generates an environmental state tensor with uncertainty measurement, and outputs the environmental state tensor to the dynamic risk assessment module and the adaptive exploration decision module.

[0012] The dynamic risk assessment module receives the environmental state tensor, integrates Bayesian inference networks and Monte Carlo simulation, predicts the probability of catastrophic consequences of potential action chains, and outputs a dynamic risk field map to the adaptive exploration decision module.

[0013] The adaptive exploration decision-making module receives the environmental state tensor and the dynamic risk field map. Through a hierarchical reinforcement learning framework and a curiosity-driven exploration mechanism, it uses a risk-constrained proximal strategy optimization algorithm to adjust the balance between exploration and utilization: a safety barrier function is introduced to constrain the exploration action space in high-risk areas, while controllable exploration is allowed in low-risk areas to optimize path planning. A decision action sequence is generated and output to the control output module. The control output module processes the data and drives the seven-degree-of-freedom robotic arm actuator to perform a valve closing operation.

[0014] Furthermore, in the aforementioned automatic shut-off robot control system for gas cylinder leaks at hazardous chemical accident sites, the control output module is configured to receive the decision action sequence, convert the decision action sequence into robot joint trajectory instructions using a model predictive control algorithm, drive the seven-degree-of-freedom robotic arm actuator to perform the valve shut-off operation, and feed back the execution status to the system coordination engine.

[0015] It also includes a system coordination engine;

[0016] The system coordination engine aggregates the output data of the environmental data acquisition module, state reconstruction module, dynamic risk assessment module, adaptive exploration decision-making module, and control output module through a federated learning framework, dynamically optimizes the weight allocation between modules, coordinates the operation of each module, and generates global adaptive decision-making capabilities, thereby improving the reliability of autonomous shutdown control decisions.

[0017] The environmental data acquisition module uses the IEEE 1588 protocol to synchronize the timestamps of multimodal sensor data.

[0018] The environmental data acquisition module is internally deployed with an adaptive sampling strategy optimizer, which dynamically adjusts the sensor sampling frequency according to the environmental change gradient: when the gas concentration gradient changes abruptly, the gas sensor sampling rate is increased, and the depth camera frame rate is increased during the high-speed movement phase of the robot.

[0019] The synchronized data stream is encapsulated into a standard data packet with precision calibration parameters and transmitted to the state reconstruction module.

[0020] Furthermore, in the aforementioned automatic shut-off robot control system for gas cylinder leaks at hazardous chemical accident sites, the state reconstruction module processes input data through point cloud registration algorithms and deep learning instance segmentation networks;

[0021] The state reconstruction module first performs voxelization noise reduction on the laser point cloud sequence, and then completes the stitching of multiple point cloud frames through the iterative nearest point algorithm to construct a dense three-dimensional environment map.

[0022] The instance segmentation network employs an improved Mask R-CNN architecture, specifically trained to identify gas cylinder valves and leakage features;

[0023] The state reconstruction module outputs an environmental state tensor including spatial coordinates, attitude quaternions, class probabilities, and covariance matrix, and runs an uncertainty evaluator to calculate the state estimation confidence in real time.

[0024] Furthermore, in the aforementioned automatic shut-off robot control system for gas cylinder leaks at hazardous chemical accident sites, when the uncertainty evaluator of the state reconstruction module is below a threshold, it sends a sensor reconfiguration command to the environmental data acquisition module, triggering an adaptive adjustment of the sampling frequency or resolution.

[0025] The environmental data acquisition module responds to the reconfiguration command, dynamically optimizes the data acquisition strategy, and forms a closed-loop feedback from state reconstruction to data acquisition.

[0026] Furthermore, in the aforementioned automatic shut-off robot control system for gas cylinder leaks at hazardous chemical accident sites, the dynamic risk assessment module constructs a risk reasoning engine based on a dynamic Bayesian network, with network nodes including environmental state variables, robot action variables, and consequence severity variables.

[0027] The dynamic risk assessment module uses the Markov chain Monte Carlo method to perform posterior probability sampling to simulate accident chains that may be caused by different action sequences;

[0028] The risk field map generator maps the probability calculation results to risk isosurfaces in three-dimensional space and marks high-risk restricted areas;

[0029] The dynamic risk assessment module integrates online learning capabilities, automatically updating the conditional probability distribution of the Bayesian network when the sensor detects new signs of danger.

[0030] Furthermore, in the aforementioned automatic shut-off robot control system for gas cylinder leaks at hazardous chemical accident sites, the adaptive exploration decision module implements a hierarchical reinforcement learning architecture, including a meta-controller and a bottom-level actuator two-layer decision unit;

[0031] The meta-controller employs a proximal policy optimization algorithm based on curiosity-driven exploration, and prunes the action space using safety constraints generated by the dynamic risk field graph.

[0032] The underlying actuator integrates a digital twin simulation environment to perform millisecond-level simulation verification before the action is executed;

[0033] The adaptive exploration decision module establishes an exploration utility evaluation model, which comprehensively considers information gain and risk cost. When the expected exploration utility is higher than the threshold, a controllable exploration strategy driven by a random forest decision tree is activated.

[0034] Furthermore, in the aforementioned automatic shut-off robot control system for gas cylinder leaks at hazardous chemical accident sites, the adaptive exploration decision-making module stores the new knowledge gained through exploration into a shared memory pool via an experience playback mechanism, and feeds it back to the state reconstruction module for online updating of deep learning model parameters;

[0035] The state reconstruction module utilizes exploration feedback data to optimize the instance segmentation network, improves the accuracy of environmental state estimation, and forms a reinforcement learning loop from exploration decision-making to state reconstruction.

[0036] Furthermore, in the aforementioned automatic shut-off robot control system for gas cylinder leaks at hazardous chemical accident sites, the control output module employs a constrained model predictive control algorithm to convert high-level decisions into robot joint trajectories.

[0037] The trajectory planner takes into account the dynamic constraints of the robotic arm and obstacle avoidance requirements to generate a smooth Cartesian space trajectory;

[0038] The control output module integrates an impedance controller to monitor the contact force between the end effector of the robotic arm and the valve in real time. When abnormal resistance is detected, it automatically triggers the compliant control mode.

[0039] All control commands are formally verified by a dynamic system verification tool and then sent to the seven-degree-of-freedom robotic arm actuator.

[0040] Furthermore, in the aforementioned automatic shut-off robot control system for gas cylinder leaks at hazardous chemical accident sites, the system coordination engine maintains the contribution weight matrix of each module and dynamically adjusts the weight allocation by analyzing historical decision-making effects.

[0041] When an output conflict between modules is detected, the evidence theory conflict resolution algorithm is activated.

[0042] The system coordination engine runs an adaptive learning rate adjustment strategy, which automatically adjusts the system response speed according to the complexity of the environment, and optimizes the overall decision-making efficiency while ensuring safety.

[0043] Furthermore, in the automatic shut-off robot control system for gas cylinder leaks at hazardous chemical accident sites described in this invention, the system coordination engine aggregates the outputs of each module through a federated averaging algorithm and feeds back the optimized weight parameters to the environmental data acquisition module, state reconstruction module, dynamic risk assessment module, adaptive exploration decision-making module, and control output module.

[0044] The environmental data acquisition module adjusts the data acquisition priority according to the weight parameters, the state reconstruction module optimizes the data processing flow, the dynamic risk assessment module updates the risk model, the adaptive exploration decision module adjusts the exploration strategy, and the control output module optimizes the control commands, forming a global adaptive mechanism of multi-module collaborative evolution.

[0045] Beneficial effects of this invention;

[0046] The automatic shut-off robot control system for gas cylinder leaks at hazardous chemical accident sites provided by this invention utilizes an environmental data acquisition module that synchronizes multimodal sensor data using the IEEE 1588 protocol and adaptively adjusts the sampling frequency. A state reconstruction module generates an environmental state tensor with uncertainty metrics using point cloud registration and instance segmentation networks. A dynamic risk assessment module integrates a Bayesian inference network and Monte Carlo simulation to output a dynamic risk field map. An adaptive exploration and decision-making module employs a hierarchical reinforcement learning framework and risk constraint strategies to balance exploration and utilization. A control output module converts decision-making action sequences into robot joint trajectories and integrates impedance control. A system coordination engine aggregates the outputs of each module and dynamically optimizes weight allocation through a federated learning framework, forming a global adaptive mechanism for multi-module collaborative evolution. This effectively balances exploration and utilization in unknown hazardous chemical accident environments, reduces unnecessary risky behaviors, and improves the decision reliability and operational efficiency of autonomous shut-off control. Attached Figure Description

[0047] To more clearly illustrate the technical solution of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on the accompanying drawings without creative effort.

[0048] Figure 1 This is a system architecture diagram of an automatic shut-off robot control system for gas cylinder leaks at hazardous chemical accident sites. Detailed Implementation

[0049] To make the technical solution of the present invention clearer, the present invention will be clearly and completely described below with reference to specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention. The present invention provided by various embodiments will be described in detail below with reference to the accompanying drawings. To better understand the purpose of the present invention, the present invention will be described in further detail below.

[0050] Please see Figure 1 The present invention provides an automatic shut-off robot control system for gas cylinder leaks at hazardous chemical accident sites, comprising:

[0051] A control device and a physical device, wherein the control device establishes a communication connection with the physical device;

[0052] The physical device includes a high-precision gas sensor array, a multispectral vision sensor, an omnidirectional mobile robot platform, and a seven-degree-of-freedom robotic arm actuator.

[0053] The control device includes an environmental data acquisition module, a state reconstruction module, a dynamic risk assessment module, an adaptive exploration decision-making module, a control output module, and a system coordination engine;

[0054] The environmental data acquisition module collects gas concentration gradients, thermal imaging data, stereoscopic visual flow, and laser point cloud sequences at the hazardous chemical accident site in real time through the physical equipment, and uses a timestamp synchronization algorithm to align multimodal sensor data and outputs a synchronization data packet to the state reconstruction module.

[0055] The state reconstruction module receives the synchronization data packet, processes multimodal sensor data through a deep convolutional neural network and a point cloud segmentation algorithm, reconstructs a three-dimensional environment map, identifies the topology of the leak source, the pose of the gas cylinder valve and the trajectory of dynamic obstacles, generates an environmental state tensor with uncertainty measurement, and outputs the environmental state tensor to the dynamic risk assessment module and the adaptive exploration decision module.

[0056] The dynamic risk assessment module receives the environmental state tensor, integrates Bayesian inference networks and Monte Carlo simulation, predicts the probability of catastrophic consequences of potential action chains, and outputs a dynamic risk field map to the adaptive exploration decision module.

[0057] The adaptive exploration decision-making module receives the environmental state tensor and the dynamic risk field map. Through a hierarchical reinforcement learning framework and a curiosity-driven exploration mechanism, it uses a risk-constrained proximal strategy optimization algorithm to adjust the balance between exploration and utilization: a safety barrier function is introduced to constrain the exploration action space in high-risk areas, while controllable exploration is allowed in low-risk areas to optimize path planning. A decision action sequence is generated and output to the control output module. The control output module processes the data and drives the seven-degree-of-freedom robotic arm actuator to perform a valve closing operation.

[0058] The automatic shut-off robot control system for gas cylinder leaks at hazardous chemical accident sites consists of a control unit and physical equipment. The control unit and physical equipment establish a real-time communication connection via industrial Ethernet. The physical equipment includes a high-precision gas sensor array, a multispectral vision sensor, an omnidirectional mobile robot platform, and a seven-degree-of-freedom robotic arm actuator. These hardware components work together to perceive the environment and perform operations. The control unit integrates an environmental data acquisition module, a state reconstruction module, a dynamic risk assessment module, an adaptive exploration and decision-making module, a control output module, and a system coordination engine. These modules interact deeply through a bidirectional data bus and an event-driven architecture.

[0059] After the environmental data acquisition module is activated, it monitors the gas concentration gradient distribution at the accident site in real time using a high-precision gas sensor array within the physical equipment. Simultaneously, it captures thermal imaging data and stereoscopic visual flow using a multispectral vision sensor, and generates a laser point cloud sequence in conjunction with a lidar system. This module employs a timestamp synchronization algorithm based on the IEEE 1588 protocol to perform hard synchronization of the multimodal sensor data, achieving data time alignment. The synchronized data is encapsulated into standard data packets, appended with accuracy calibration parameters, and then transmitted to the state reconstruction module. Internally, the module deploys an adaptive sampling strategy optimizer that dynamically adjusts the sensor sampling frequency based on environmental gradient changes; for example, it increases the sampling rate during sudden gas concentration changes and increases the visual frame rate during robot movement.

[0060] After receiving the synchronization data packet from the environmental data acquisition module, the state reconstruction module initiates a parallel processing pipeline. The point cloud processing thread uses an iterative nearest-point algorithm to perform voxelization denoising and multi-frame registration on the laser point cloud sequence, constructing a dense 3D environmental map. The image processing thread uses an instance segmentation network in a deep convolutional neural network to identify the topology of the leak source, the pose of the gas cylinder valve, and the trajectory of dynamic obstacles. The processed data is fused to generate an environmental state tensor with an uncertainty metric, which includes spatial coordinates, pose quaternions, class probabilities, and a covariance matrix. An uncertainty estimator runs internally within the module, calculating the state estimation confidence level in real time. When the confidence level falls below a threshold, a sensor reconfiguration command is sent to the environmental data acquisition module, forming a closed-loop feedback.

[0061] The dynamic risk assessment module subscribes to the environmental state tensor output by the state reconstruction module, constructing a risk inference engine based on a dynamic Bayesian network. Network nodes integrate environmental state variables, robot action variables, and consequence severity variables, and use the Markov chain Monte Carlo method for posterior probability sampling to simulate accident chains that may result from different action sequences. The risk field map generator maps the probability calculation results to a risk isosurface in three-dimensional space and marks high-risk no-go zones. This module has online learning capabilities; when sensors detect new dangerous signs, it automatically updates the conditional probability distribution of the Bayesian network and outputs a dynamic risk field map to the adaptive exploration decision-making module.

[0062] The adaptive exploration decision-making module simultaneously receives the environmental state tensor and the dynamic risk field graph, realizing a hierarchical reinforcement learning architecture. The meta-controller employs a proximal policy optimization algorithm based on curiosity-driven exploration, pruning the action space through safety constraints generated by the dynamic risk field graph. In high-risk areas, a safety barrier function is introduced to restrict exploration behavior, prioritizing the use of a digital twin simulation system to verify policy feasibility; in low-risk areas, controlled exploration based on random forest decision trees is allowed to optimize path planning. The underlying actuator integrates a simulation environment, performing millisecond-level simulation verification before action execution. New knowledge gained through exploration is stored in a shared memory pool through an experience replay mechanism and fed back to the state reconstruction module for online model updates.

[0063] The control output module receives the decision action sequence generated by the adaptive exploration decision module and uses a constrained model predictive control algorithm to convert the high-level decisions into robot joint trajectories. The trajectory planner considers the robot arm's dynamic constraints and obstacle avoidance requirements to generate a smooth Cartesian space trajectory. An integrated impedance controller monitors the contact force between the robot arm's end effector and the valve in real time, automatically triggering a compliant control mode when abnormal resistance is detected. All control commands are formally verified using a dynamic system verification tool before being sent to the seven-DOF robot arm actuator to drive it to complete the valve closing operation.

[0064] The system coordination engine aggregates the output data of each module through a federated learning framework and maintains a module contribution weight matrix. The engine analyzes historical decision-making effects to dynamically adjust weight allocation, and when output conflicts between modules are detected, it initiates an evidence-based conflict resolution algorithm. The coordination engine runs an adaptive learning rate adjustment strategy, adjusting the system response speed according to environmental complexity and feeding back optimization parameters to each module, forming a global adaptive mechanism for multi-module collaborative evolution, ultimately improving the reliability of autonomous shutdown control decisions.

[0065] Specifically, in the aforementioned automatic shut-off robot control system for gas cylinder leaks at hazardous chemical accident sites, the control output module is configured to receive the decision action sequence, use a model predictive control algorithm to convert the decision action sequence into robot joint trajectory instructions, drive the seven-degree-of-freedom robotic arm actuator to perform the valve shut-off operation, and feed back the execution status to the system coordination engine.

[0066] It also includes a system coordination engine;

[0067] The system coordination engine aggregates the output data of the environmental data acquisition module, state reconstruction module, dynamic risk assessment module, adaptive exploration decision-making module, and control output module through a federated learning framework, dynamically optimizes the weight allocation between modules, coordinates the operation of each module, and generates global adaptive decision-making capabilities, thereby improving the reliability of autonomous shutdown control decisions.

[0068] The environmental data acquisition module uses the IEEE 1588 protocol to synchronize the timestamps of multimodal sensor data.

[0069] The environmental data acquisition module is internally deployed with an adaptive sampling strategy optimizer, which dynamically adjusts the sensor sampling frequency according to the environmental change gradient: when the gas concentration gradient changes abruptly, the gas sensor sampling rate is increased, and the depth camera frame rate is increased during the high-speed movement phase of the robot.

[0070] The synchronized data stream is encapsulated into a standard data packet with precision calibration parameters and transmitted to the state reconstruction module.

[0071] After receiving the decision action sequence from the adaptive exploration decision module, the control output module uses a model predictive control algorithm for trajectory planning. This algorithm predicts system state changes over a future period by establishing a robot kinematic model and environmental constraints, and continuously optimizes joint trajectory commands. The trajectory planning process considers the dynamic characteristics of the seven-DOF robotic arm and obstacle avoidance requirements, generating a smooth and safe Cartesian space path. The robotic arm actuator performs valve closing operations according to commands, monitoring the contact force data between the end effector and the valve in real time. Execution status data, including position deviation, force feedback, and completion status, is fed back to the system coordination engine in real time.

[0072] The system coordination engine aggregates execution state data from the control output module using a federated learning framework, while also integrating output information from the environmental data acquisition module, state reconstruction module, dynamic risk assessment module, and adaptive exploration decision-making module. The engine maintains a dynamic weight matrix, adjusting the influence weights between modules based on their historical performance metrics and current environmental complexity. The weight allocation algorithm is based on gradient descent optimization, prioritizing the outputs of conservative strategy modules in high-risk scenarios and increasing the weights of exploration modules in stable environments. This coordination mechanism enables the system to develop global adaptive decision-making capabilities, significantly improving the reliability of shutdown operations.

[0073] The environmental data acquisition module employs the IEEE 1588 precision time protocol to synchronize timestamps of multimodal sensor data. Through a master-slave clock architecture and network latency compensation mechanism, the protocol ensures a unified time reference for gas concentration gradients, thermal imaging data, stereo vision streams, and laser point cloud sequences. An adaptive sampling strategy optimizer deployed within the module continuously analyzes environmental change gradients. When the gas concentration sensor detects a sudden change, it automatically increases the sampling rate to capture rapid changes; when the omnidirectional mobile robot platform enters a high-speed movement state, it increases the frame rate of the depth camera to ensure image continuity. The synchronized multi-source data streams are encapsulated into standard data packets with accuracy calibration parameters and transmitted to the state reconstruction module for further processing.

[0074] After receiving real-time data from the environmental data acquisition module, the system coordination engine updates the global model parameters using a federated averaging algorithm. The engine then feeds back the optimized weight parameters to the environmental data acquisition module, guiding it to adjust data acquisition priorities. For example, when the risk level increases, the engine increases the weight of gas sensor data, prompting the acquisition module to increase the frequency of gas monitoring. This closed-loop adjustment mechanism forms a reverse optimization path from decision execution to data acquisition, strengthening the system's adaptability to dynamic hazardous environments.

[0075] Specifically, in the aforementioned automatic shut-off robot control system for gas cylinder leaks at hazardous chemical accident sites, the state reconstruction module processes input data through point cloud registration algorithms and deep learning instance segmentation networks.

[0076] The state reconstruction module first performs voxelization noise reduction on the laser point cloud sequence, and then completes the stitching of multiple point cloud frames through the iterative nearest point algorithm to construct a dense three-dimensional environment map.

[0077] The instance segmentation network employs an improved Mask R-CNN architecture, specifically trained to identify gas cylinder valves and leakage features;

[0078] The state reconstruction module outputs an environmental state tensor including spatial coordinates, attitude quaternions, class probabilities, and covariance matrix, and runs an uncertainty evaluator to calculate the state estimation confidence in real time.

[0079] After receiving the synchronization data packet from the environmental data acquisition module, the state reconstruction module initiates a parallel processing pipeline to process multimodal sensor data. The synchronization data packet includes a laser point cloud sequence and a stereo vision stream. The module first performs voxelization denoising on the laser point cloud sequence, dividing the point cloud space into voxel grids and applying statistical filtering to remove outliers, reducing noise interference in subsequent steps. The denoised point cloud data enters the point cloud registration stage, using an iterative nearest-point algorithm to align the point clouds frame by frame. By minimizing the point-to-point distance error, multiple frames of point clouds are stitched together to construct a dense 3D environmental map. The point cloud registration process simultaneously estimates sensor pose changes, providing a spatial consistency basis for the environmental map.

[0080] Simultaneously with point cloud processing, the image processing thread employs a deep learning instance segmentation network to analyze the stereo visual flow. This instance segmentation network is based on an improved Mask R-CNN architecture, with its backbone network using a residual structure to enhance feature extraction capabilities and specifically trained for gas cylinder valve and leakage features. After receiving image input, the network generates candidate regions through a region proposal network, then extracts features using a ROI alignment layer, ultimately outputting pixel-level segmentation masks and category labels. The segmentation results are aligned with the point cloud data using a spatial transformation matrix, achieving precise mapping between visual information and a 3D map.

[0081] Point cloud registration results and instance segmentation outputs are fused at the feature layer to generate an environment state tensor. The fusion process employs a multimodal attention mechanism to weightedly integrate the spatial accuracy of the point cloud and the semantic information of the segmentation. The environment state tensor includes spatial coordinates, pose quaternions, class probabilities, and a covariance matrix. The spatial coordinates represent the 3D position of the target object, the pose quaternions describe the valve orientation, the class probabilities reflect the recognition confidence, and the covariance matrix quantifies the estimation uncertainty. Kalman filtering is applied during tensor construction to smooth the temporal data and improve the stability of the state estimation.

[0082] The state reconstruction module internally runs an uncertainty estimator that calculates the state estimation confidence level in real time. Based on a Bayesian inference framework, the estimator analyzes factors such as point cloud registration error, segmentation network confidence, and sensor measurement noise, estimating the posterior probability distribution through Monte Carlo sampling. The confidence level calculation outputs a scalar value to measure the reliability of the environmental state tensor. When the confidence level falls below a preset threshold, the uncertainty estimator sends a sensor reconfiguration command to the environmental data acquisition module, triggering an adjustment to the data acquisition strategy, thus forming a closed-loop feedback loop from state reconstruction to data acquisition.

[0083] Throughout the processing flow, point cloud denoising and registration provide geometric constraints for instance segmentation, which in turn optimizes point cloud classification using the segmentation results. Multimodal fusion ensures the integrity and accuracy of the environmental state tensor. Uncertainty assessment dynamically monitors processing quality, enabling the system to adapt to environmental changes and improve the robustness of state reconstruction.

[0084] Specifically, in the aforementioned automatic shut-off robot control system for gas cylinder leaks at hazardous chemical accident sites, when the uncertainty evaluator of the state reconstruction module is lower than a threshold, it sends a sensor reconfiguration command to the environmental data acquisition module, triggering an adaptive adjustment of the sampling frequency or resolution.

[0085] The environmental data acquisition module responds to the reconfiguration command, dynamically optimizes the data acquisition strategy, and forms a closed-loop feedback from state reconstruction to data acquisition.

[0086] The uncertainty estimator in the state reconstruction module calculates the estimated confidence level of the environmental state tensor in real time. The confidence level calculation is based on a joint analysis of point cloud registration error, instance segmentation network output probability, and sensor measurement noise. The estimator employs a Bayesian inference framework and quantifies the uncertainty of the state estimation using a Monte Carlo sampling method to generate a scalar confidence value. When the confidence value falls below a preset threshold, the uncertainty estimator sends a sensor reconfiguration command to the environmental data acquisition module, which includes adjustments to the target sampling frequency or resolution parameters.

[0087] After receiving the reconfiguration command, the environmental data acquisition module parses the command content and initiates an adaptive adjustment process. The module's internal adaptive sampling strategy optimizer dynamically reconstructs the data acquisition strategy based on the command parameters. For example, it might increase the gas sensor sampling rate to capture rapid concentration changes or increase the depth camera resolution to enhance image details. During the adjustment process, the optimizer simultaneously calibrates the sensor accuracy calibration parameters to ensure data quality consistency.

[0088] The sensor reconfiguration command response mechanism is based on an event-driven architecture. The environmental data acquisition module monitors the command queue in real time and prioritizes high-priority reconfiguration requests. The adjusted sensor parameters take effect immediately, the multimodal sensor data acquisition process is updated accordingly, and the acquired data stream is repackaged into a standard data packet with accuracy calibration parameters and transmitted to the state reconstruction module.

[0089] After receiving a new data packet, the state reconstruction module restarts the processing pipeline and regenerates the environmental state tensor using the updated sensor data. The uncertainty evaluator simultaneously verifies the confidence level of the new tensor. If the confidence level increases, the current acquisition strategy is maintained; otherwise, it triggers a new round of reconfiguration instructions. This cyclical process forms a closed-loop feedback loop from state reconstruction to data acquisition, continuously improving the accuracy of state estimation through iterative optimization of data acquisition parameters.

[0090] The closed-loop feedback mechanism enables the system to adapt to dynamic environmental changes, such as in scenarios involving escalating gas leaks or moving obstacles, by rapidly adjusting sensor parameters to ensure data validity. The tight coupling of data flow and control flow in the feedback loop enhances the system's robustness to unknown hazardous environments, providing reliable input for subsequent risk assessment and decision-making modules.

[0091] Specifically, the automatic shut-off robot control system for gas cylinder leakage at hazardous chemical accident sites includes a risk inference engine that constructs a dynamic Bayesian network in the dynamic risk assessment module. The network nodes include environmental state variables, robot action variables, and consequence severity variables.

[0092] The dynamic risk assessment module uses the Markov chain Monte Carlo method to perform posterior probability sampling to simulate accident chains that may be caused by different action sequences;

[0093] The risk field map generator maps the probability calculation results to risk isosurfaces in three-dimensional space and marks high-risk restricted areas;

[0094] The dynamic risk assessment module integrates online learning capabilities, automatically updating the conditional probability distribution of the Bayesian network when the sensor detects new signs of danger.

[0095] The dynamic risk assessment module receives the environmental state tensor from the state reconstruction module as input data for the dynamic Bayesian network risk inference engine. The environmental state tensor includes spatial coordinates, pose quaternions, class probabilities, and a covariance matrix. This data is used to initialize the network nodes of the risk inference engine. Network node definitions include environmental state variables, robot action variables, and consequence severity variables. The environmental state variables describe the real-time conditions at the accident site, the robot action variables represent the possible sequences of robotic arm operations, and the consequence severity variables quantify the potential hazard levels of different action chains.

[0096] The dynamic risk assessment module uses the Markov chain Monte Carlo method for posterior probability sampling to simulate accident chains that may result from different action sequences. The sampling process is based on a Bayesian inference framework, utilizing prior probability distributions and real-time environmental state data to iteratively generate a large number of possible scenario paths. Each scenario path evaluates the intermediate steps from the current state to the completion of the valve closure operation, calculating the probability of catastrophic consequences such as escalated leakage, explosion, or obstacle collision. The probability sampling results form a posterior distribution reflecting the risk level of each action sequence.

[0097] The risk field map generator maps the posterior probability calculation results to risk isosurfaces in 3D space. The mapping process uses a spatial interpolation algorithm to convert the probability values ​​into a continuous risk gradient field, where high-risk areas correspond to spatial locations where the probability of the consequence exceeds a safety threshold. The risk field map is overlaid on the 3D environment map as a visual layer, marking high-risk restricted areas, whose boundaries are automatically generated based on the probability isosurfaces. The risk field map is output to the adaptive exploration decision module to guide the robot's action planning.

[0098] The dynamic risk assessment module integrates online learning capabilities, automatically updating the conditional probability distribution of the Bayesian network when sensors detect new hazards. The online learning mechanism is based on incremental Bayesian update rules; new sensor data, such as sudden changes in gas concentration or thermal imaging anomalies, triggers adjustments to the probability distribution. The conditional probability table update process considers the weighted balance between historical and real-time data, retaining recent key information through a sliding window mechanism. The updated Bayesian network optimizes risk prediction accuracy in real time, adapting to dynamic changes at the accident site.

[0099] The entire risk assessment process forms a closed-loop learning system. The risk field map output is fed back to the sensor data acquisition stage. When new changes are identified in high-risk restricted areas, the environmental state tensor is recalculated. This closed-loop mechanism enables the risk model to continuously evolve, improving its adaptability to unknown hazardous chemical environments.

[0100] Specifically, the automatic shut-off robot control system for gas cylinder leakage at hazardous chemical accident sites includes an adaptive exploration decision-making module that implements a hierarchical reinforcement learning architecture, comprising two decision-making units: a meta-controller and a bottom-level actuator.

[0101] The meta-controller employs a proximal policy optimization algorithm based on curiosity-driven exploration, and prunes the action space using safety constraints generated by the dynamic risk field graph.

[0102] The underlying actuator integrates a digital twin simulation environment to perform millisecond-level simulation verification before the action is executed;

[0103] The adaptive exploration decision module establishes an exploration utility evaluation model, which comprehensively considers information gain and risk cost. When the expected exploration utility is higher than the threshold, a controllable exploration strategy driven by a random forest decision tree is activated.

[0104] The adaptive exploration decision-making module employs a hierarchical reinforcement learning architecture, comprising two decision-making units: a meta-controller and a low-level executor. The meta-controller is responsible for high-level policy planning, while the low-level executor is responsible for executing and verifying specific actions. The meta-controller makes decisions based on a curiosity-driven exploration mechanism and a proximal policy optimization algorithm. Curiosity-driven exploration encourages exploration of the unknown environment by calculating the information gain of new states, while the proximal policy optimization algorithm ensures stability during policy updates. The meta-controller receives a dynamic risk field map from the dynamic risk assessment module and uses the safety constraints in the field map to prune the action space, eliminating high-risk action options to ensure the safety of the exploration process.

[0105] The underlying actuator integrates a digital twin simulation environment, which is a virtual copy of the physical system. Before executing any actual action, the underlying actuator inputs candidate action sequences generated by the meta-controller into the digital twin simulation environment for millisecond-level simulation verification. The simulation verification process predicts the consequences of the actions, assesses their feasibility and potential risks, and only actions that pass verification are submitted to the control output module for execution. The digital twin simulation environment uses real-time sensor data to update model parameters, achieving simulation accuracy consistent with the real environment.

[0106] The adaptive exploration decision-making module constructs an exploration utility evaluation model that comprehensively considers information gain and risk cost. Information gain measures the increase in environmental knowledge resulting from exploration actions, while risk cost assesses the potential harm of the actions. The exploration utility evaluation model uses a random forest decision tree algorithm to process information gain and risk cost factors, generating an expected exploration utility value. When the expected exploration utility value exceeds a preset threshold, the module activates a controlled exploration strategy driven by the random forest decision tree. The controlled exploration strategy allows the system to conduct purposeful exploration in low-risk areas to optimize path planning and learning efficiency.

[0107] The exploration utility evaluation model is updated in real time, adjusting the parameters of the random forest decision tree based on historical exploration data and environmental feedback. The model evaluation results are fed back to the meta-controller to optimize the weight allocation for curiosity-driven exploration. The meta-controller adjusts the exploration parameters of the proximal policy optimization algorithm based on the evaluation results, forming a closed-loop learning mechanism from evaluation to decision-making. The underlying actuator uses the evaluation model's output to optimize the verification rules for the digital twin simulation environment, improving the accuracy of simulation verification.

[0108] Through a hierarchical architecture and evaluation model, the adaptive exploration decision-making module dynamically balances exploration and utilization, improving the reliability of decisions in unknown hazardous environments. The collaborative work of the meta-controller and the underlying actuators ensures that exploration operations are both safe and efficient, while the exploration utility evaluation model provides data-driven decision support, enabling the system to adapt to complex scenario changes.

[0109] Specifically, in the aforementioned automatic shut-off robot control system for gas cylinder leaks at hazardous chemical accident sites, the adaptive exploration decision module stores the new knowledge gained through exploration into a shared memory pool through an experience playback mechanism, and feeds it back to the state reconstruction module for online updating of deep learning model parameters.

[0110] The state reconstruction module utilizes exploration feedback data to optimize the instance segmentation network, improves the accuracy of environmental state estimation, and forms a reinforcement learning loop from exploration decision-making to state reconstruction.

[0111] Specifically, the automatic shut-off robot control system for gas cylinder leaks at hazardous chemical accident sites uses a constrained model predictive control algorithm in its control output module to convert high-level decisions into robot joint trajectories.

[0112] The trajectory planner takes into account the dynamic constraints of the robotic arm and obstacle avoidance requirements to generate a smooth Cartesian space trajectory;

[0113] The control output module integrates an impedance controller to monitor the contact force between the end effector of the robotic arm and the valve in real time. When abnormal resistance is detected, it automatically triggers the compliant control mode.

[0114] All control commands are formally verified by a dynamic system verification tool and then sent to the seven-degree-of-freedom robotic arm actuator.

[0115] After receiving the decision action sequence from the adaptive exploration decision module, the control output module uses a constrained model predictive control algorithm for trajectory planning. This algorithm predicts system state changes over a future period by establishing a robot kinematic model and environmental constraints, and continuously optimizes joint trajectory commands. The trajectory planning process considers the dynamic characteristics of the seven-DOF robotic arm and obstacle avoidance requirements, generating a smooth and safe Cartesian space path. The robotic arm actuator performs a valve closing operation according to the commands, monitoring the contact force data between the end effector and the valve in real time. Execution status data, including position deviation, force feedback, and completion status, is fed back to the system coordination engine in real time.

[0116] The system coordination engine aggregates execution state data from the control output module using a federated learning framework, while also integrating output information from the environmental data acquisition module, state reconstruction module, dynamic risk assessment module, and adaptive exploration decision-making module. The engine maintains a dynamic weight matrix, adjusting the influence weights between modules based on their historical performance metrics and current environmental complexity. The weight allocation algorithm is based on gradient descent optimization, prioritizing the outputs of conservative strategy modules in high-risk scenarios and increasing the weights of exploration modules in stable environments. This coordination mechanism enables the system to develop global adaptive decision-making capabilities, significantly improving the reliability of shutdown operations.

[0117] The environmental data acquisition module employs the IEEE 1588 precision time protocol to synchronize timestamps of multimodal sensor data. Through a master-slave clock architecture and network latency compensation mechanism, the protocol ensures a unified time reference for gas concentration gradients, thermal imaging data, stereo vision streams, and laser point cloud sequences. An adaptive sampling strategy optimizer deployed within the module continuously analyzes environmental change gradients. When the gas concentration sensor detects a sudden change, it automatically increases the sampling rate to capture rapid changes; when the omnidirectional mobile robot platform enters a high-speed movement state, it increases the frame rate of the depth camera to ensure image continuity. The synchronized multi-source data streams are encapsulated into standard data packets with accuracy calibration parameters and transmitted to the state reconstruction module for further processing.

[0118] After receiving real-time data from the environmental data acquisition module, the system coordination engine updates the global model parameters using a federated averaging algorithm. The engine then feeds back the optimized weight parameters to the environmental data acquisition module, guiding it to adjust data acquisition priorities. For example, when the risk level increases, the engine increases the weight of gas sensor data, prompting the acquisition module to increase the frequency of gas monitoring. This closed-loop adjustment mechanism forms a reverse optimization path from decision execution to data acquisition, strengthening the system's adaptability to dynamic hazardous environments.

[0119] Specifically, the automatic shut-off robot control system for gas cylinder leaks at hazardous chemical accident sites includes a system coordination engine that maintains the contribution weight matrix of each module and dynamically adjusts the weight allocation by analyzing historical decision-making effects.

[0120] When an output conflict between modules is detected, the evidence theory conflict resolution algorithm is activated.

[0121] The system coordination engine runs an adaptive learning rate adjustment strategy, which automatically adjusts the system response speed according to the complexity of the environment, and optimizes the overall decision-making efficiency while ensuring safety.

[0122] The system coordination engine maintains a contribution weight matrix, which records the real-time contribution values ​​of the environmental data acquisition module, state reconstruction module, dynamic risk assessment module, adaptive exploration decision-making module, and control output module. The weight matrix is ​​continuously updated by collecting output data and behavioral records from each module, such as the accuracy of module decisions, response time, and task completion status. The engine periodically analyzes historical decision performance, evaluating module performance metrics in past tasks, including decision success rate, risk avoidance effectiveness, and operational efficiency. Based on the analysis results, the engine uses a gradient descent optimization algorithm to dynamically adjust the weight allocation, increasing the weight ratio of high-performing modules and reducing the impact factors of poorly performing modules.

[0123] When the system coordination engine detects output conflicts between modules—for example, inconsistencies in data interpretation between the environmental data acquisition module and the state reconstruction module, or conflicting action recommendations between the dynamic risk assessment module and the adaptive exploration decision-making module—the engine activates the evidence theory conflict resolution algorithm. This algorithm collects confidence-level evidence from each module's output, fuses different evidence sources using Dempster's combination rule, and calculates a joint trust function. The conflict resolution process prioritizes the outputs of modules with high confidence and strong evidence consistency to eliminate decision-making contradictions. The evidence theory conflict resolution algorithm also considers historical collaboration records between modules, assigning higher trust weights to modules with a long history of good collaboration.

[0124] The system coordination engine operates an adaptive learning rate adjustment strategy, which monitors environmental complexity metrics in real time, including gas leakage rate, obstacle density, and sensor data fluctuations. The engine dynamically adjusts the learning rate parameters of the federated learning framework based on environmental complexity, reducing the learning rate to maintain decision stability in high-complexity scenarios and increasing it to accelerate model convergence in low-complexity scenarios. The learning rate adjustment strategy is synchronized with the weight matrix update, forming a dual-loop optimization mechanism. The engine monitors system response speed metrics, such as decision latency and execution efficiency, and adjusts the learning rate parameters accordingly to match system response with environmental requirements.

[0125] An adaptive learning rate adjustment strategy is integrated into the weight allocation process. The engine uses changes in the learning rate to infer the adaptability of each module and further optimizes the weight matrix. The module contribution weights and learning rate parameters jointly influence the global decision output of the system coordination engine, creating a synergistic optimization effect. This integrated mechanism enables the system to maintain efficient decision-making amidst dynamic changes at hazardous chemical accident sites, while ensuring operational safety. Ultimately, the system coordination engine improves the overall decision-making efficiency of autonomous shutdown control by continuously optimizing weight allocation and conflict resolution.

[0126] Specifically, the automatic shut-off robot control system for gas cylinder leakage at hazardous chemical accident sites described in this invention has a system coordination engine that aggregates the outputs of each module through a federated averaging algorithm and feeds back the optimized weight parameters to the environmental data acquisition module, state reconstruction module, dynamic risk assessment module, adaptive exploration decision-making module, and control output module.

[0127] The environmental data acquisition module adjusts the data acquisition priority according to the weight parameters, the state reconstruction module optimizes the data processing flow, the dynamic risk assessment module updates the risk model, the adaptive exploration decision module adjusts the exploration strategy, and the control output module optimizes the control commands, forming a global adaptive mechanism of multi-module collaborative evolution.

[0128] The system coordination engine aggregates the output data from the environmental data acquisition module, state reconstruction module, dynamic risk assessment module, adaptive exploratory decision-making module, and control output module using a federated averaging algorithm. The federated averaging algorithm employs a distributed learning framework, using the output data of each module as local model parameters and calculating the global model parameters through weighted averaging. The aggregation process considers the real-time performance metrics and historical contributions of each module to generate a comprehensive global state assessment.

[0129] The system coordination engine dynamically optimizes weight parameters based on the aggregation results. These weight parameters reflect the relative importance of each module in the decision-making process. The optimization process uses gradient descent to minimize system decision-making errors and improve overall performance. The optimized weight parameters are fed back to each module in real time via communication links. The feedback mechanism adopts an event-driven model to ensure timely parameter updates.

[0130] After receiving the weight parameters, the environmental data acquisition module adjusts the data acquisition priority based on the parameter values. Higher weights correspond to higher-priority data sources, and the module reallocates sensor resources to focus on key data streams. The state reconstruction module uses the weight parameters to optimize the data processing flow, adjusting the computational resource allocation for point cloud registration and instance segmentation algorithms to improve processing efficiency. The dynamic risk assessment module updates the risk model based on the weight parameters, corrects the conditional probability distribution of the Bayesian network, and enhances the accuracy of risk prediction. The adaptive exploration decision-making module adjusts the exploration strategy based on the weight parameters, balancing exploration and utilization, and optimizing action sequence generation. The control output module uses the weight parameters to optimize control commands, refine joint trajectory planning, and improve control accuracy.

[0131] The adjustments made by each module form a closed-loop feedback loop, with the system coordination engine continuously monitoring module performance and iteratively optimizing weight parameters. This multi-module collaborative evolution mechanism enables the system to adapt to environmental changes. Through parameter sharing and coordinated optimization, it develops global adaptive decision-making capabilities, improving the reliability and efficiency of autonomous shutdown control at hazardous chemical accident sites.

[0132] This invention addresses the balance between exploration and utilization by constructing a hierarchical reinforcement learning architecture and a multi-module collaborative mechanism. The system coordination engine aggregates data from the environmental data acquisition module, state reconstruction module, dynamic risk assessment module, adaptive exploration decision-making module, and control output module through a federated learning framework, dynamically optimizing the weight allocation among modules. The adaptive exploration decision-making module employs a risk-constrained proximal policy optimization algorithm, tailoring the action space based on safety constraints generated from the dynamic risk field graph: in high-risk areas, a safety barrier function is introduced to restrict exploration behavior, prioritizing the use of a digital twin simulation system to verify policy feasibility; in low-risk areas, controlled exploration based on random forest decision trees is allowed to optimize path planning. The state reconstruction module calculates the confidence level of the environmental state estimate in real time through an uncertainty evaluator. When the confidence level falls below a threshold, a sensor reconfiguration command is triggered, forming a closed-loop feedback from state to data acquisition. The dynamic risk assessment module integrates online learning capabilities, automatically updating the conditional probability distribution of the Bayesian network when new hazards are detected. This multi-level coordination mechanism enables the system to acquire environmental information through controlled exploration while utilizing known safety policies to mitigate risks, thereby maintaining decision reliability in unknown hazardous environments.

Claims

1. An automatic shut-off robot control system for gas cylinder leaks at hazardous chemical accident sites, characterized in that: include: A control device and a physical device, wherein the control device establishes a communication connection with the physical device; The physical device includes a high-precision gas sensor array, a multispectral vision sensor, an omnidirectional mobile robot platform, and a seven-degree-of-freedom robotic arm actuator. The control device includes an environmental data acquisition module, a state reconstruction module, a dynamic risk assessment module, an adaptive exploration decision-making module, a control output module, and a system coordination engine; The environmental data acquisition module collects gas concentration gradients, thermal imaging data, stereoscopic visual flow, and laser point cloud sequences at the hazardous chemical accident site in real time through the physical equipment, and uses a timestamp synchronization algorithm to align multimodal sensor data and outputs a synchronization data packet to the state reconstruction module. The state reconstruction module receives the synchronization data packet, processes multimodal sensor data through a deep convolutional neural network and a point cloud segmentation algorithm, reconstructs a three-dimensional environment map, identifies the topology of the leak source, the pose of the gas cylinder valve and the trajectory of dynamic obstacles, generates an environmental state tensor with uncertainty measurement, and outputs the environmental state tensor to the dynamic risk assessment module and the adaptive exploration decision module. The dynamic risk assessment module receives the environmental state tensor, integrates Bayesian inference networks and Monte Carlo simulation, predicts the probability of catastrophic consequences of potential action chains, and outputs a dynamic risk field map to the adaptive exploration decision module. The adaptive exploration decision module receives the environmental state tensor and the dynamic risk field map. Through a hierarchical reinforcement learning framework and a curiosity-driven exploration mechanism, it uses a risk-constrained proximal strategy optimization algorithm to adjust the balance between exploration and utilization: a safety barrier function is introduced to constrain the exploration action space in high-risk areas, and controllable exploration is allowed in low-risk areas to optimize path planning. The module generates a decision action sequence and outputs the decision action sequence to the control output module. The control output module processes the data and drives the seven-degree-of-freedom robotic arm actuator to perform a valve closing operation. The adaptive exploration decision module implements a hierarchical reinforcement learning architecture, including two decision units: a meta-controller and a bottom-level executor. The meta-controller employs a proximal policy optimization algorithm based on curiosity-driven exploration, and prunes the action space using safety constraints generated by the dynamic risk field graph. The underlying actuator integrates a digital twin simulation environment to perform millisecond-level simulation verification before the action is executed; The adaptive exploration decision module establishes an exploration utility evaluation model, which comprehensively considers information gain and risk cost. When the expected exploration utility is higher than the threshold, a controllable exploration strategy driven by a random forest decision tree is activated. The control output module employs a constrained model predictive control algorithm to convert high-level decisions into robot joint trajectories. The trajectory planner takes into account the dynamic constraints of the robotic arm and obstacle avoidance requirements to generate a smooth Cartesian space trajectory; The control output module integrates an impedance controller to monitor the contact force between the end effector of the robotic arm and the valve in real time. When abnormal resistance is detected, the compliant control mode is automatically triggered. All control commands are formally verified by a dynamic system verification tool and then sent to the seven-degree-of-freedom robotic arm actuator.

2. The automatic shut-off robot control system for gas cylinder leaks at hazardous chemical accident sites according to claim 1, characterized in that, The control output module is configured to receive the decision action sequence, use a model predictive control algorithm to convert the decision action sequence into robot joint trajectory instructions, drive the seven-degree-of-freedom robotic arm actuator to perform the valve closing operation, and feed back the execution status to the system coordination engine. It also includes a system coordination engine; The system coordination engine is configured to aggregate the output data of the environmental data acquisition module, state reconstruction module, dynamic risk assessment module, adaptive exploration decision module and control output module through a federated learning framework, dynamically optimize the weight allocation between modules, coordinate the operation of each module, and generate global adaptive decision-making capabilities, thereby improving the reliability of autonomous shutdown control decisions. The environmental data acquisition module uses the IEEE 1588 protocol to synchronize the timestamps of multimodal sensor data. The environmental data acquisition module is internally deployed with an adaptive sampling strategy optimizer, which dynamically adjusts the sensor sampling frequency according to the environmental change gradient: when the gas concentration gradient changes abruptly, the gas sensor sampling rate is increased, and the depth camera frame rate is increased during the high-speed movement phase of the robot. The synchronized data stream is encapsulated into a standard data packet with precision calibration parameters and transmitted to the state reconstruction module.

3. The automatic shut-off robot control system for gas cylinder leaks at hazardous chemical accident sites according to claim 1, characterized in that, The state reconstruction module processes input data using a point cloud registration algorithm and a deep learning instance segmentation network. The state reconstruction module first performs voxelization noise reduction on the laser point cloud sequence, and then completes the stitching of multiple frame point clouds through the iterative nearest point algorithm to construct a dense three-dimensional environment map. The instance segmentation network employs an improved MaskR-CNN architecture, specifically trained to identify gas cylinder valves and leakage features; The state reconstruction module outputs an environmental state tensor including spatial coordinates, attitude quaternions, class probabilities, and covariance matrix, and runs an uncertainty evaluator to calculate the state estimation confidence in real time.

4. The automatic shut-off robot control system for gas cylinder leaks at hazardous chemical accident sites according to claim 3, characterized in that, When the confidence level of the state reconstruction module is lower than the threshold, the uncertainty evaluator sends a sensor reconfiguration command to the environmental data acquisition module, triggering an adaptive adjustment of the sampling frequency or resolution. The environmental data acquisition module responds to the reconfiguration command, dynamically optimizes the data acquisition strategy, and forms a closed-loop feedback from state reconstruction to data acquisition.

5. The automatic shut-off robot control system for gas cylinder leaks at hazardous chemical accident sites according to claim 1, characterized in that, The dynamic risk assessment module constructs a risk inference engine using a dynamic Bayesian network. The network nodes include environmental state variables, robot action variables, and consequence severity variables. The dynamic risk assessment module uses the Markov chain Monte Carlo method to perform posterior probability sampling to simulate accident chains that may be caused by different action sequences; The risk field map generator maps the probability calculation results to risk isosurfaces in three-dimensional space and marks high-risk restricted areas; The dynamic risk assessment module integrates online learning capabilities, automatically updating the conditional probability distribution of the Bayesian network when the sensor detects new signs of danger.

6. The automatic shut-off robot control system for gas cylinder leaks at hazardous chemical accident sites according to claim 1, characterized in that, The adaptive exploration decision module stores the new knowledge gained through exploration into a shared memory pool through an experience replay mechanism, and feeds it back to the state reconstruction module for online updating of deep learning model parameters; The state reconstruction module utilizes exploration feedback data to optimize the instance segmentation network, improves the accuracy of environmental state estimation, and forms a reinforcement learning loop from exploration decision-making to state reconstruction.

7. The automatic shut-off robot control system for gas cylinder leaks at hazardous chemical accident sites according to claim 1, characterized in that, The system coordination engine maintains the contribution weight matrix of each module and dynamically adjusts the weight allocation by analyzing the effects of historical decisions. When an output conflict between modules is detected, the evidence theory conflict resolution algorithm is activated. The system coordination engine runs an adaptive learning rate adjustment strategy, which automatically adjusts the system response speed according to the complexity of the environment, and optimizes the overall decision-making efficiency while ensuring safety.

8. The automatic shut-off robot control system for gas cylinder leaks at hazardous chemical accident sites according to claim 1, characterized in that, The system coordination engine aggregates the outputs of each module through a federated averaging algorithm and feeds back the optimized weight parameters to the environmental data acquisition module, state reconstruction module, dynamic risk assessment module, adaptive exploration decision-making module, and control output module. The environmental data acquisition module adjusts the data acquisition priority according to the weight parameters, the state reconstruction module optimizes the data processing flow, the dynamic risk assessment module updates the risk model, the adaptive exploration decision-making module adjusts the exploration strategy, and the control output module optimizes the control commands, forming a global adaptive mechanism of multi-module collaborative evolution.

Citation Information

Patent Citations

  • Driving situation prediction and adaptive strategy generation system based on cloud multi-mode large model

    CN119408566A

  • Robot dynamic risk assessment and decision-making system and method based on multi-modal perception

    CN120680531A

  • Double-arm robot autonomous control system and method based on remote operation and visual features

    CN120816484A