Limited space facility safety risk patrol method and system based on intelligent system
By using an embodied intelligent robot dog vehicle to establish a 3D environment model and semantic topology map in a confined space, the optimal inspection path is planned, and multimodal perception data collection and feature fusion analysis are performed. This solves the problems of poor environmental adaptability and insufficient safety assurance in existing technologies, and achieves efficient and accurate safety risk inspection.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANGHAI RESEARCH INSTITUTE OF BUILDING SCIENCES CO LTD
- Filing Date
- 2026-03-24
- Publication Date
- 2026-04-21
AI Technical Summary
Existing technologies for safety inspection of facilities in confined spaces suffer from poor environmental adaptability, limited detection methods, low levels of intelligence, and insufficient safety assurance. In particular, they are difficult to achieve comprehensive detection and real-time intelligent diagnosis in complex environments such as dampness and narrow spaces.
By employing an embodied intelligent robot dog vehicle combined with synchronous positioning and mapping sensors, a 3D environment model and semantic topology map are established. Multimodal perception data is acquired and feature fusion analysis is performed to plan the optimal inspection path and maintain a stable posture in complex environments to collect multimodal perception data.
It enables efficient and accurate safety risk inspections in complex and confined spaces, avoiding direct exposure of personnel to potential dangers, improving detection accuracy and safety, and reducing the time wasted on blind exploration.
Smart Images

Figure CN121900428A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent inspection technology, and in particular to a method and system for inspecting safety risks in confined space facilities based on embodied intelligence. Background Technology
[0002] Currently, safety inspections of confined space facilities mainly employ the following methods: Manual inspection: Staff enter the confined space for visual inspection, which suffers from high operational risks, low efficiency, and poor detection accuracy. Fixed sensor monitoring: Fixed sensors are deployed at key locations, but the coverage is limited and cannot comprehensively detect structural defects. Traditional robot inspection: Wheeled or tracked robots are used, but their mobility in complex environments is poor, and their functions are limited. Problems with existing technologies: Poor environmental adaptability: Traditional equipment struggles to adapt to complex confined space environments such as dampness, water accumulation, and narrow spaces; Limited detection methods: Lacking multimodal perception capabilities, it cannot achieve comprehensive detection such as geometric measurement, defect detection, and gas analysis; Low level of intelligence: Data acquisition and analysis are separated, preventing real-time intelligent diagnosis and decision-making; Insufficient safety guarantees: There is a lack of effective anti-collision and anti-flooding mechanisms, and equipment cannot be safely recovered in case of failure. Summary of the Invention
[0003] To address the aforementioned shortcomings in existing technologies, this invention provides a confined space facility safety risk inspection method based on embodied intelligence, which solves the problems of low inspection safety and efficiency in confined space environments.
[0004] To achieve the aforementioned objectives, the technical solution adopted by this invention is: a method for inspecting safety risks in confined space facilities based on embodied intelligence, comprising: S1: Based on synchronous positioning and mapping, sensor-collected confined space environment data are used to build a three-dimensional environment model and semantic topology map; S2: Based on the 3D environment model and semantic topology map, combined with the preset risk perception path assessment strategy, the optimal inspection path is obtained through analysis. S3: Control the embodied intelligent robot dog vehicle to move along the optimal inspection path, acquire terrain data and water depth data in real time during the movement, execute the corresponding motion control strategy according to the terrain data and water depth data to maintain the stable attitude of the embodied intelligent robot dog vehicle, and acquire multimodal perception data of the confined space facility under the stable attitude. S4: Input the multimodal sensing data into a pre-trained multimodal large model for feature fusion analysis to obtain the safety risk inspection results of confined space facilities and complete the safety risk inspection of confined space facilities.
[0005] The beneficial effects of this invention are as follows: This invention provides a method for inspecting the safety risks of confined space facilities based on embodied intelligence. By controlling an embodied intelligent robot dog vehicle to acquire terrain and water depth data in real time during movement, and executing corresponding motion control strategies accordingly, the stable attitude of the vehicle is effectively maintained. Multimodal perception data of the confined space facility is acquired under a stable attitude, and this data is input into a pre-trained multimodal large model for feature fusion analysis, avoiding the limitations of a single data source, thus obtaining more accurate safety risk inspection results for confined space facilities. Based on the confined space environment data collected by synchronous positioning and mapping sensors, a three-dimensional environment model and semantic topology map are established, and analyzed in conjunction with a preset risk perception path evaluation strategy to obtain the optimal inspection path. Controlling the embodied intelligent robot dog vehicle to move along this optimal inspection path effectively avoids the time loss of blind exploration. Using an embodied intelligent robot dog vehicle to perform inspection tasks along the planned optimal inspection path instead of manual entry into the confined space environment, and maintaining a stable attitude for data collection under complex terrain and water depth conditions, fundamentally avoids direct exposure of personnel to potential safety risks in confined spaces.
[0006] Further, S2 includes: Based on the three-dimensional environment model, the standard deviation of height, standard deviation of slope, and obstacle density of the ground in the confined space are extracted. Combined with the ground friction coefficient, the terrain complexity assessment model is used to calculate the terrain complexity value. Based on the semantic topology map, the terrain complexity value is used as the path cost weight, and the graph search algorithm is used to plan the minimum risk cost to obtain inspection path data containing the spatial coordinate sequence of monitoring points. By using a pre-defined risk perception path assessment strategy, the inspection path data containing the spatial coordinate sequence of monitoring points is analyzed to obtain the optimal inspection path.
[0007] Furthermore, the expression for the terrain complexity value is: ; in, This represents the terrain complexity value. to All are preset weighting coefficients. For high standard deviation, For the standard deviation of slope, For obstacle density, is the coefficient of friction of the ground.
[0008] Further, S3 includes: Control the embodied intelligent robot dog vehicle to move along the optimal inspection path, and obtain the current step frequency, stride length and foot lift height of the embodied intelligent robot dog vehicle in real time as the reference gait parameters; During the movement, the multimodal terrain complexity assessment model outputs real-time terrain data and water depth data. Based on the terrain data and water depth data, the gait adaptive adjustment algorithm is used to dynamically correct the baseline gait parameters to obtain the adjusted gait parameters adapted to the current terrain, which serve as the corresponding motion control strategy. Using a motion control strategy, the movement of the embodied intelligent robot dog vehicle is controlled to traverse the spatial coordinate sequence of monitoring points in order to maintain the stable posture of the embodied intelligent robot dog vehicle. Multimodal sensor arrays are used to acquire multimodal perception data of the confined space facility.
[0009] Furthermore, the expression for the adjusted gait parameters is: ; ; ; in, , and These are the adjusted cadence, stride length, and foot lift height. , and All are baseline parameters. , and All are adjustment coefficients. This represents the terrain complexity value.
[0010] Furthermore, in step S3, the movement of the embodied intelligent robot dog vehicle is controlled using a motion control strategy to traverse the spatial coordinate sequence of monitoring points, and the water / land mode switching step is also included: Real-time acquisition of water depth data in confined spaces, as well as pitch angle data and foot force data of the embodied intelligent robot dog vehicle; Based on water depth data, pitch angle data, and foot force data, the environmental state of the embodied intelligent robot dog vehicle is determined, and the judgment result is obtained. When the judgment result is that the vehicle is in the water, the intelligent robot dog vehicle is controlled to deploy its buoyancy device and switch to underwater propulsion mode; when the judgment result is that the vehicle is on land, the intelligent robot dog vehicle is controlled to retract its buoyancy device and switch to land walking mode, and the intelligent robot dog vehicle is controlled to move and traverse the spatial coordinate sequence of monitoring points.
[0011] Furthermore, the expression for the judgment result is: and It was determined to be in a submerged state; and This indicates a logged-in status. in, For real-time water depth data, The threshold for the center height of the wheel hub. For pitch angle data, The pitch angle threshold, For data on plantar force, This is the plantar force threshold.
[0012] Further, S4 includes: Obtain observation data from various sensors in multimodal sensing data and their corresponding observation noise covariance, and construct a factor graph fusion objective function that includes kinematic prediction terms and sensor observation terms; Based on a pre-trained multimodal large model, the optimal state estimate and unified environmental characteristics of the embodied intelligent robot dog vehicle at each time step are obtained by minimizing the factor graph fusion objective function, thus completing the safety risk inspection of the confined space facility.
[0013] Furthermore, the objective function for factor graph fusion is: ; in, Let t be the vehicle state at time t. For control input, Q is the process noise covariance. Let be the observation data of the i-th type of sensor at time t. For the observation model of the i-th type of sensor, This represents the objective function for factor graph fusion. This represents the vehicle motion control function. Represents the observation noise covariance. This represents the kinematic prediction, where Q is the process noise covariance. This indicates the vehicle state at time t-1.
[0014] This invention provides a confined space facility safety risk inspection system based on embodied intelligence, comprising: The environmental perception and modeling unit is used to build a three-dimensional environmental model and a semantic topology map based on the limited spatial environmental data collected by the sensor based on synchronous positioning and mapping. The path planning unit is used to analyze the optimal inspection path based on the three-dimensional environment model and semantic topology map, combined with a preset risk perception path assessment strategy. The embodied inspection execution subsystem includes an embodied intelligent robot dog vehicle, a multimodal sensor group, and a motion controller. The motion controller controls the embodied intelligent robot dog vehicle to move along the optimal inspection path, acquires terrain data and water depth data in real time during movement, and executes corresponding motion control strategies based on the terrain data and water depth data to maintain the stable attitude of the embodied intelligent robot dog vehicle. The multimodal sensor group is used to acquire multimodal perception data of the confined space facility under the stable attitude. The intelligent analysis unit is used to input the multimodal sensing data into a pre-trained multimodal large model for feature fusion analysis to obtain the safety risk inspection results of confined space facilities. Attached Figure Description
[0015] This specification will be further described by way of exemplary embodiments, which will be described in detail with reference to the accompanying drawings. These embodiments are not limiting; in these embodiments, the same reference numerals denote the same structures, wherein: Figure 1 This is an exemplary flowchart illustrating a confined space facility safety risk inspection method based on embodied intelligence, according to some embodiments of this specification. Figure 2 This is a schematic diagram of a confined space facility safety risk inspection system based on embodied intelligence, as shown in some embodiments of this specification. Detailed Implementation
[0016] The specific embodiments of the present invention are described below to enable those skilled in the art to understand the present invention. However, it should be understood that the present invention is not limited to the scope of the specific embodiments. For those skilled in the art, various changes are obvious as long as they are within the spirit and scope of the present invention as defined and determined by the appended claims. All inventions utilizing the concept of the present invention are protected.
[0017] Example 1 Figure 1 This is an exemplary flowchart illustrating a method for inspecting safety risks in confined space facilities based on embodied intelligence, according to some embodiments of this specification. Figure 1 As shown, the process includes the following steps. In some embodiments, the process may be executed by a processor.
[0018] S1: Based on synchronous positioning and mapping, sensor-collected confined space environment data are used to build a 3D environment model and semantic topology map.
[0019] Confined space environment data refers to the set of raw observation data that reflects the characteristics of the physical environment inside a confined space. For example, confined space environment data may include distance point cloud data collected by lidar, environmental texture image data collected by visual cameras, and vehicle attitude data collected by inertial measurement units.
[0020] In some embodiments, the processor can acquire the confined space environment data from a group of simultaneous localization and mapping (SLAM) sensors via a communication interface.
[0021] A three-dimensional environment model refers to a three-dimensional spatial representation obtained by digitally reconstructing a confined space, used to describe the geometric structure and boundary information of the confined space. For example, a three-dimensional environment model can be a raster map, octree map, or TSDF (truncated symbolic distance field) model based on point cloud data.
[0022] In some embodiments, the processor can construct the three-dimensional environment model based on the confined space environment data by running a SLAM algorithm (such as LIO-SAM or Fast-LIO).
[0023] A semantic topology map is a graph structure abstracted from a three-dimensional environment model, containing spatial connectivity and functional attributes. For example, a semantic topology map can be composed of nodes and edges, where nodes represent key locations within a confined space (such as intersections, pipe entrances, and dead ends), and edges represent walkable paths between nodes, along with semantic information such as path width and length.
[0024] In some embodiments, the processor can identify key topological nodes by performing distance field analysis and skeleton extraction on a 3D environment model, thereby generating the semantic topological map.
[0025] Specifically, this step includes comprehensive environmental perception using a multimodal sensor array: acquiring information such as facility structural dimensions and deformation through geometric measurement methods; measuring water levels using depth sensors to detect water accumulation; detecting leaks by combining infrared thermal imagers and visual sensors to identify leak points; analyzing harmful gas concentrations using a comprehensive gas detector; detecting structural defects such as cracks and corrosion using image sensors combined with AI (artificial intelligence) algorithms; and measuring deformation through comparison of multiple inspection data to analyze structural deformation trends. In some embodiments, the topology map construction process includes converting SLAM point cloud data into a semantic topology map to identify topological structures such as passages, intersections, and dead ends.
[0026] In some embodiments, the expression for the topology map node generation algorithm is: ; in, This is the candidate set of nodes for the topological map. For the i-th spatial point, This is the set of collision-free edges obtained after performing Delaunay triangulation on the node set, which will be used for subsequent graph search. The Laplace operator for the distance field has maxima at corners, doorways, and forks in the path, corresponding to key waypoints. Store the distance to the nearest obstacle for each grid cell. For curvature threshold, To perform Delaunay triangulation on the node set.
[0027] S2: Based on the 3D environment model and semantic topology map, combined with the preset risk perception path assessment strategy, the optimal inspection path is obtained through analysis.
[0028] A risk-aware path assessment strategy refers to a computational rule or algorithmic logic used to quantify the travel cost and safety risks of a path. For example, a risk-aware path assessment strategy could be a weighted cost function that comprehensively considers factors such as the path's geometric length, terrain complexity risk, gas concentration risk, and water depth risk.
[0029] In some embodiments, the risk-aware path evaluation strategy is pre-stored in memory, and the processor invokes the strategy when planning a path to calculate the comprehensive cost of different path candidate segments.
[0030] The optimal inspection path refers to the sequence of paths with the lowest overall cost, calculated based on a risk perception path assessment strategy, while meeting coverage requirements. For example, the optimal inspection path is a collision-free trajectory that starts from the starting point, traverses all target monitoring areas, and returns to the endpoint, and has the lowest cumulative risk value among all feasible trajectories.
[0031] In some embodiments, the processor may search for the optimal inspection path based on a semantic topology map, using a heuristic search algorithm (such as A or DLite) in conjunction with the risk perception path evaluation strategy.
[0032] In some embodiments, the processor can extract the standard deviation of height, standard deviation of slope, and obstacle density of the confined space ground based on a three-dimensional environment model, and combine them with the ground friction coefficient to calculate the terrain complexity value using a multimodal terrain complexity assessment model; based on a semantic topological map, the terrain complexity value is used as a path cost weight, and a graph search algorithm is used to plan the minimum risk cost to obtain inspection path data containing the spatial coordinate sequence of monitoring points; the inspection path data containing the spatial coordinate sequence of monitoring points is analyzed using a preset risk perception path assessment strategy to obtain the optimal inspection path.
[0033] Inspection path data refers to a set of discretized coordinate points that describe the specific direction of the optimal inspection path. For example, inspection path data may include a series of sequentially arranged three-dimensional spatial coordinate points (x, y, z) or pose points containing attitude information (x, y, z, roll, pitch, yaw).
[0034] In some embodiments, the processor generates the inspection path data by interpolating and smoothing the optimal inspection path found.
[0035] In some embodiments, the expression for the terrain complexity value is: ; in, This represents the terrain complexity value. to All are preset weighting coefficients. For high standard deviation, For the standard deviation of slope, For obstacle density, is the coefficient of friction of the ground.
[0036] In some embodiments, the gait adaptive adjustment algorithm based on deep reinforcement learning can automatically adjust the robot dog's gait parameters according to the slipperiness of the ground, the slope, and the height of obstacles, and is multimodal; the environmental feature encoder encodes SLAM point cloud data, visual images, and tactile information into a unified environmental feature vector to achieve millisecond-level response to environmental changes.
[0037] In some embodiments, the expression for the risk perception path evaluation function is: ; in, The total cost of the path. This is the terrain weighting coefficient. Let be the terrain complexity at path location s. For gas weighting coefficients, For the permissible concentration threshold of the gas, This represents the actual gas concentration measurement at path location s. This is the water depth weighting coefficient. For the maximum permissible water depth, This represents the actual water depth measurement at location s along the path. For the risk of gas concentration exceeding the standard multiple / normalized, The relative water depth risk / normalized penalty at path location s is used. The cost of sections exceeding the fuselage height tends to infinity to prohibit passage. Recommended weight values: λ1=0.4, λ2=0.4, λ3=0.2, which can be fine-tuned online through reinforcement learning.
[0038] In some embodiments, the processor can support dynamic replanning, completing local path replanning within 100ms when a new obstacle or dangerous area is detected; and perform inspection path optimization based on a coverage optimization genetic algorithm, whose fitness function takes into account low repetition, low risk and low corner targets.
[0039] In some embodiments, the fitness function is expressed as: ; in, This is the path fit value. These are non-overlapping weighting coefficients. This refers to the path overlap rate. This is the risk cost weighting coefficient. The total cost of the risk perception path. For the steering smoothness weighting coefficient, This represents the percentage of the total turning angle along the path.
[0040] In some embodiments, the shock resistance stability margin index is as follows: ; in, For stability margin, the larger the value, the stronger the anti-interference ability. The zero moment point (ZMP) is represented by its two-dimensional coordinates within the plane of the supporting polygon, indicating the point of application of the resultant ground reaction force. The point on the nearest edge of the supporting polygon that is closest to the zero moment point (ZMP) is also called the "support boundary point". The outward normal vector of this supporting edge (pointing outward from the polygon) is used to measure the direction of the risk of the zero moment point (ZMP) going out of bounds. This is the minimum value for stability margin. The formula quantifies the "safety margin" between the zero moment point (ZMP) and the boundary of the support area in real time, ensuring that the robot dog can maintain an internal safety distance of ≥15mm after being subjected to a lateral impact of 30kg·m / s, thereby recovering stable standing within 2s.
[0041] S3: Control the embodied intelligent robot dog vehicle to move along the optimal inspection path, acquire terrain data and water depth data in real time during the movement, execute the corresponding motion control strategy according to the terrain data and water depth data to maintain the stable attitude of the embodied intelligent robot dog vehicle, and acquire multimodal perception data of the confined space facility under the stable attitude.
[0042] Topographic data refers to a set of parameters that describe the geometric features and physical properties of the ground in a confined space. For example, topographic data may include the standard deviation of ground height, the standard deviation of slope, obstacle density, and the coefficient of ground friction.
[0043] In some embodiments, the processor can perform real-time analysis and calculation on SLAM point cloud data and visual image data using a multimodal terrain complexity assessment model to obtain the terrain data at the current location.
[0044] Water depth data refers to the numerical value of the liquid surface depth in a water-filled area within a confined space. For example, water depth data is a scalar value measured in meters or centimeters, used to characterize the risk of flooding of a vehicle due to water accumulation.
[0045] In some embodiments, the processor can acquire the water depth data by reading the analog signals from the level sensor or pressure sensor in the gas / water depth integrated probe and processing them through an analog-to-digital converter (ADC).
[0046] Motion control strategy refers to a set of control instructions or algorithmic logic used to guide an embodied intelligent robot dog vehicle to perform specific actions. For example, motion control strategy may include gait parameter adjustment strategies (such as adjusting stride frequency, stride length, and foot lift height) and mode switching strategies (such as switching from land walking mode to underwater propulsion mode).
[0047] In some embodiments, the processor dynamically generates the corresponding motion control strategy based on real-time terrain data and water depth data, using a gait adaptive adjustment algorithm or a land-water switching logic.
[0048] Multimodal sensing data refers to a heterogeneous collection of data about the status of confined space facilities, acquired by different types of sensors at the same time or within a synchronous time window. For example, multimodal sensing data may include visible light images, infrared thermal images, high-precision 3D point clouds, gas concentration values, and water depth values.
[0049] In some embodiments, the processor triggers synchronous sampling of the multimodal sensor group via an FPGA synchronization board and acquires the multimodal sensing data after spatial alignment.
[0050] In some embodiments, the processor can control the embodied intelligent robot dog vehicle to move along the optimal inspection path, and acquire the current stride frequency, stride length, and foot lift height of the embodied intelligent robot dog vehicle in real time as reference gait parameters; during the movement, real-time terrain data and water depth data are output based on a multimodal terrain complexity assessment model, and the reference gait parameters are dynamically corrected using a gait adaptive adjustment algorithm based on the terrain data and water depth data to obtain adjusted gait parameters adapted to the current terrain, which serve as the corresponding motion control strategy; using the motion control strategy, the embodied intelligent robot dog vehicle is controlled to move, traversing the spatial coordinate sequence of monitoring points to maintain the stable posture of the embodied intelligent robot dog vehicle, and multimodal perception data of the confined space facility is acquired using a multimodal sensor group.
[0051] In some embodiments, the expression for the adjusted gait parameters is: ; ; ; in, , and These are the adjusted cadence, stride length, and foot lift height. , and All are baseline parameters. , and All are adjustment coefficients. This represents the terrain complexity value.
[0052] In some embodiments, a motion control strategy is used to control the movement of the embodied intelligent robot dog vehicle, traversing the spatial coordinate sequence of monitoring points. The method also includes a water-land mode switching step: real-time acquisition of water depth data within the confined space, as well as pitch angle data and foot force data of the embodied intelligent robot dog vehicle; based on the water depth data, pitch angle data, and foot force data, determining the environmental state of the embodied intelligent robot dog vehicle and obtaining a determination result; when the determination result indicates a water entry state, controlling the embodied intelligent robot dog vehicle to deploy its buoyancy device and switch to underwater propulsion mode; when the determination result indicates a land landing state, controlling the embodied intelligent robot dog vehicle to retract its buoyancy device and switch to land walking mode, controlling the movement of the embodied intelligent robot dog vehicle, and traversing the spatial coordinate sequence of monitoring points.
[0053] In some embodiments, the expression for the judgment result is: and It was determined to be in a submerged state; and This indicates a logged-in status. in, For real-time water depth data, The threshold for the center height of the wheel hub. For pitch angle data, The pitch angle threshold, For data on plantar force, This is the plantar force threshold.
[0054] S4: Input the multimodal sensing data into a pre-trained multimodal large model for feature fusion analysis to obtain the safety risk inspection results of confined space facilities and complete the safety risk inspection of confined space facilities.
[0055] A multimodal large model refers to a neural network model built on a deep learning architecture that can simultaneously process and understand input data from multiple modalities. For example, a multimodal large model may include a convolutional neural network branch for extracting image features, a PointNet branch for extracting point cloud features, and an attention mechanism module for cross-modal feature fusion.
[0056] In some embodiments, the multimodal large model is pre-trained with a large number of confined space disease samples and deployed in the main control computing unit of the embodied intelligent robot dog vehicle.
[0057] The results of a safety risk inspection of confined space facilities refer to the final assessment information output by the system after comprehensive detection and analysis of the confined space facilities. For example, the results of a safety risk inspection of confined space facilities may include the type of defects (such as cracks, leaks, and deformation), the specific location coordinates of the defects, the severity level of the defects, the predicted trend of structural deformation, and an overall safety status assessment report.
[0058] In some embodiments, the processor inputs multimodal sensing data into a multimodal large model for inference, parses the output layer data of the model, and thus obtains the safety risk inspection results of the confined space facility.
[0059] In some embodiments, the processor can acquire various sensor observation data and their corresponding observation noise covariance from the multimodal sensing data, construct a factor graph fusion objective function that includes kinematic prediction terms and sensor observation terms; based on a pre-trained multimodal large model, by minimizing the factor graph fusion objective function, solve for the optimal state estimate and unified environmental characteristics of the embodied intelligent robot dog vehicle at each time step, obtain the safety risk inspection results of the confined space facility, and complete the safety risk inspection of the confined space facility.
[0060] In some embodiments, the objective function for factor graph fusion is: ; in, Let t be the vehicle state at time t. For control input, Q is the process noise covariance. Let be the observation data of the i-th type of sensor at time t. For the observation model of the i-th type of sensor, This represents the objective function for factor graph fusion. This represents the vehicle motion control function. Represents the observation noise covariance. This represents the kinematic prediction, where Q is the process noise covariance. This indicates the vehicle state at time t-1.
[0061] In some embodiments, the expression for the microsecond-level hardware-triggered synchronization error is: ; in, To ensure the maximum deviation at all sensor sampling times, the time drift of each modal data at a 30Hz frame rate is much smaller than the motion displacement of one pixel or one laser point (typically 0.01mm). To find the maximum value function, This is the hardware trigger timestamp for sensor i (generated by the TTL edge output uniformly by the FPGA). This is the hardware trigger timestamp for sensor j.
[0062] In some embodiments, the expression for the spatial calibration extrinsic self-calibration error function is: ; in, Let be the camera projection model (including radial-tangential distortion), R be the rotation matrix from the LiDAR coordinate system to the camera coordinate system, and be the external parameters to be optimized. The objective function for reprojection error in multi-sensor spatial calibration is: Let N be the actual pixel coordinates of the k-th feature point on the camera's 2D image plane, where k is the feature point's index and N is the total number of 3D-2D matching feature points participating in the calibration calculation. This is the intrinsic parameter matrix of the camera. This represents the k-th 3D point in the laser point cloud (distortion removed). The camera-laser extrinsics to be optimized. The target value of 0.5px corresponds to the reprojection error, ensuring pixel-level alignment of the infrared, visible light, and laser three-modal sensors to prevent ghosting during subsequent fusion.
[0063] In some embodiments, the expression for the objective function of the fusion of Kalman filtering and factor graph optimization is: ; in, For optimal state estimation, To minimize the parameters, For the observation value of the i-th type of sensor, For the i-th state node, This refers to the state at the previous moment. Let i be the observation model for the i-th type of sensor (laser, vision, IMU, gas, water depth). For "subtraction" on the manifold, the Lie group structure of robot state xt∈SE(3) is guaranteed. For kinematic prediction, Q represents process noise, achieving optimal estimation of multimodal data at a unified coordinate and time, and Σi represents the corresponding observation noise covariance, which is estimated in real time adaptively.
[0064] In some embodiments, the expression for online noise covariance estimation is: ; in, Let be the online noise covariance matrix of the i-th type of sensor. The length of the sliding window. The original observation value of the i-th type of sensor at time k is... For the predicted observation value of the i-th type of sensor, the sliding window length M = 30 (approximately 1 second).
[0065] In some embodiments, the processor can update Σi in real time and substitute it into 3-C, so that the fusion weights change dynamically with the environment (e.g., when fog causes increased laser scattering, the laser factor weight is automatically reduced and the infrared and visual weights are increased).
[0066] In some embodiments, topology map construction involves converting SLAM point cloud data into a semantic topology map to identify topological structures such as passages, intersections, and dead ends; risk-aware path assessment utilizes multi-objective path optimization that comprehensively considers factors such as path length, terrain complexity, and potential risks; dynamic replanning capability enables local path replanning within 100ms when new obstacles or hazardous areas are detected; and coverage optimization algorithm uses a genetic algorithm-based inspection path optimization to ensure a detection coverage rate of ≥99% and duplicate paths of ≤5%.
[0067] In some embodiments, the real-time constraint function for dynamic replanning is as follows: Tre-plan≤100ms, Coverage≥99%, Overlap≤5%; The incremental D*Lite algorithm is used to repair paths locally only within a 2m influence zone of newly added obstacles, with an average CPU time of 42ms (STM32H7+CMSIS-DSP). Tre-plan is the replanning time, Coverage is the inspection area ratio relative to the complete grid map, and Overlap is the ratio of repeated path lengths, used to measure energy redundancy.
[0068] In some embodiments, the Model Predictive Obstacle Avoidance (MPC) constraint formula is as follows: ; in, Let this be the system state vector at the next time step (k+1). Here is the state transition matrix. Let k be the system state vector at the current time (k). To control the input matrix, The control input vector at the current time (k) For the robot's turning speed, For maximum angular velocity constraints, For the robot's forward / backward speed, For maximum linear velocity constraints, For free space sets, It is a two-dimensional Euclidean plane. Let be the center coordinates of the i-th obstacle. Let be the original radius of the i-th obstacle. For the robot's equivalent safety radius, state =[x,y,θ] ,control =[v,ω] For linear velocity and angular velocity; ω max =2rad / s,v max =1m / s. The expanded obstacle disk has an expansion radius equal to the equivalent fuselage radius of 0.25m plus laser uncertainty of 0.05m to ensure a safe containment. The prediction time domain is N=15 (0.5s), with a rolling optimization step size of 50ms, matching the real-time constraint step size of the dynamic replanning.
[0069] In some embodiments, the expression for the overriding optimized genetic operator is: ; in, This is the path fit value. This refers to the path overlap rate. The total cost of the risk perception path. This represents the percentage of total turning angles along the path. Chromosome encoding: sequential sequence of topological nodes → variable-length crossover (OX) + 2-opt local flip. Weights α=0.5, β=0.3, γ=0.2, balancing the three objectives of "low repetition, low risk, and low turning angles"; population size 50, 30 iterations, single evolution time 38ms, satisfying online reprogramming requirements.
[0070] In some embodiments, the multimodal feature extraction network is designed with a specialized multi-branch CNN architecture to process visual images, infrared thermal images, and 3D point cloud data respectively; a cross-modal attention mechanism is implemented to achieve adaptive weighted fusion of features from different modalities, improving the accuracy of disease identification to ≥95%; a few-shot learning technique is adopted to solve the problem of scarce disease samples, requiring only 50 samples to identify new disease types; and edge computing optimization is achieved with a model compression rate of ≥90% and an inference time of ≤50ms on an embedded GPU, meeting the requirements for real-time detection.
[0071] In some embodiments, the overall loss function of a multimodal large model includes: ; ; ; ; ; ; ; ; ; ; ; ; ; in, For the total loss, It is a multi-class cross-entropy used for pixel-level disease classification (cracks, seepage, peeling, gas corrosion, etc.). Using Smooth-L1 loss, the regression bounding box center offset and width / height are calculated, with a positioning accuracy ≤3cm. Binary cross-entropy is used to supervise the consistency between the infrared heatmap and the visual mask, thereby improving edge accuracy. and As weight, Total number of pixels The true label for pixel i belonging to category c. For the unnormalized value of pixel i of class c, Let C be the unnormalized value of pixel i for category k, and C be the number of disease categories. This represents the total number of positive sample bounding boxes within the batch. For loss function, For deviation, The normalized prediction bias of the j-th box in the x-axis is... The x-axis of the center coordinate of the j-th positive sample bounding box is the model predicted value. The x-axis value of the center coordinate of the bounding box of the j-th positive sample is the true value. Let be the width annotation value of the bounding box of the j-th positive sample. This is the normalized vertical deviation. The model predicted value is the center coordinate ordinate of the bounding box of the j-th positive sample. The true value of the y-axis coordinate of the center of the bounding box of the j-th positive sample is given. Let j be the height annotation value of the bounding box of the j-th positive sample. This represents the normalized width deviation. The model predicted value for the width of the bounding box of the j-th positive sample. The height of the bounding box of the j-th positive sample is the model predicted value. For binary cross-entropy loss, For Dice coefficient loss, Let i be the true label of the i-th pixel. Let be the predicted probability of the i-th pixel. For numerically stable terms, This is an activation function used to predict probabilities. The query vector (N×dk) generated for the visible light branch. This is the transpose of the key matrix. Let be the dimension of the key vector. A value matrix, For optimal meta-parameters, For all feasible parameters θ In the process, we search for the parameter that minimizes the objective function. For task T i The loss function on the support set. θ For meta-parameters, Step size (learning rate) For parameters θ gradient descent, To support set loss on parameters θ gradient, For task distribution, The quantized INT8 precision weights, For the cropping operation, For the FP32 precision weights to be quantized, This is the quantization scaling factor (Scale). It is the weight tensor or matrix of the original FP32 precision.
[0072] BCE is used to optimize pixel-level probability, and Dice is used to optimize region overlap. Empirical weight values are as follows: =2, =1, obtained through grid search on the validation set. KIR, VIR: Key and value vectors of the infrared branch (M×d) k ,M×d v The attention matrix has an N×M dimension, achieving pixel-level alignment between visual questions and infrared responses, with simultaneous weight enhancement for crack and thermal anomaly regions. k =256, Scaling prevents softmax saturation; computational cost O(NMd) kThe number of multiply-accumulate operations (MAC) is approximately 45M, completed in 38ms on Jetson Xavier. Supported set Ti: only 50 512×512 images (including annotations) per disease type. The inner loop learning rate α=0.01, allowing for task-adaptive parameters to be obtained with a single gradient update; the outer loop calculates the expectation of all tasks, achieving "learning to learn". New disease deployment time is reduced from 3 days to 15 minutes, while the recognition accuracy remains ≥92% (compared to 78% with traditional fine-tuning). The weight quantization ratio s is calculated per channel, and activation uses layer-by-layer dynamic quantization; the model size is compressed from 92MB to 11MB. int8 inference is 3.4 times faster on embedded GPUs (graphics processors), with a 42% reduction in power consumption, meeting the 50ms real-time constraint.
[0073] In some embodiments, the processor can achieve accurate measurement and trend prediction of structural deformation based on multi-temporal data comparison and analysis, adopt a phase correlation algorithm to achieve deformation detection accuracy at the level of 0.1mm, use multi-period point cloud fine registration based on ICP algorithm with registration error ≤0.5mm, use LSTM (Long Short-Term Memory) neural network model to predict the deformation development trend in the next 3-6 months, establish a deformation rate threshold model, and automatically alarm when the deformation rate exceeds the set threshold.
[0074] In some embodiments, the expression for sub-pixel level deformation measurement based on the phase correlation method is: ; in, This refers to the subpixel displacement vector of the same texture in two images; the patent constrains the peak search range to ≤0.1px to prevent false matching. For candidate displacement values u Within the range, find the cross-correlation function. ρ Take the displacement value of the maximum value. The normalized cross-correlation peak function, calculated using the frequency domain phase spectrum, achieves a positioning resolution of 0.01px. The reference image was acquired before deformation. For the image acquired after deformation, in the coordinates of the reference image x The texture at that location has been displaced in the target image. u Converted to physical size: 0.1mm pixel size at a 1m viewing distance. Theoretical deformation resolution is 0.01 mm, and actual measured (root mean square error) RMSE = 0.05 mm.
[0075] In some embodiments, the expression for the multi-phase point cloud fine registration algorithm based on ICP variants is: ; in, For ICP registration error, For rotation matrix, This refers to the i-th 3D coordinate point in the second phase point cloud (point cloud to be registered). For time, For the first phase of point cloud (baseline point cloud) and P i scan2 The corresponding matching point (nearest neighbor). These are the regularization weight coefficients. It is the identity matrix. The first term represents the ICP registration error. The second term is Lie algebra regularization to prevent large turns from getting trapped in local minima; λ=0.01. The 0.5mm threshold corresponds to 1 / 20 of the crack width on the tunnel surface, meeting the requirements of general industry safety specifications.
[0076] In some embodiments, the expression for LSTM-based temporal deformation prediction is: ; in, The network outputs a sequence of deformation increments for the next k=12 steps (i.e., the next 3 months), with dimensions 3×k=3×12. The root mean square error, It is a two-layer LSTM with 128 hidden units. This is a sequence of 3D deformations over the past N periods (60 sampling points in total), with each period lasting 30 days and dimensions of 3×N=3×60. Dropout=0.2; the parameter set θ contains approximately 0.18M trainable weights. The input gate, forget gate, and output gate work together to capture long-term dependencies, avoiding the gradient vanishing problem of traditional RNNs (Recurrent Neural Networks).
[0077] In some embodiments, the processor can employ a sequence-to-sequence (Seq2Seq) architecture, which can obtain multiple predictions in a single forward pass, reducing accumulated errors. RMSE ≤ 0.5 mm, and within a 90% confidence interval, the root mean square error of the three-dimensional Euclidean error between the predicted value and the subsequent true value does not exceed 0.5 mm.
[0078] In some embodiments, the processor can employ Huber loss (δ=1mm) to be robust to outliers; Adam's initial lr=1e-3, cosine annealing. The training set contains ≥800 historical tunnel sequences, with data augmentation: Gaussian noise σ=0.02mm, randomly discarding 10% of frames to improve the model's generalization ability.
[0079] The fault self-diagnosis network is a fault diagnosis model based on Bayesian networks, with a fault detection accuracy of ≥98%.
[0080] The tiered emergency response strategy involves activating different levels of emergency measures (continue mission, suspend mission, emergency return, send out distress signal) based on the severity of the malfunction.
[0081] Autonomous return navigation is a method of returning autonomously based on pre-built maps and dead reckoning when communication is interrupted.
[0082] The emergency rescue mechanism is to automatically release rescue signals when equipment is trapped, including audible and visual alarms, radio beacons, etc.
[0083] In some embodiments, the expression for the Bayesian network-based fault diagnosis model is: ; Where F∈{motor-fault,comm-loss,water-inrush,roll-over} is the set of fault modes (4 classes), O={ωIMU,Imotor,hwater,Mstab} is the observable evidence vector (IMU angular velocity, motor current, water depth, stability margin), and the node CPT (conditional probability table) is 1.2×10 6 The data was trained using real-world operational data; when the posterior probability is ≥0.98, the corresponding emergency strategy is triggered, with a false negative rate of <0.5%.
[0084] In some embodiments, the expression for the threshold of the tiered emergency response strategy is: ; in, This is the stability margin. The angular velocity vector measured by the IMU (Inertial Measurement Unit). - For communication link interruption time, Level-1 is when the zero moment point (ZMP) is less than 15mm from the support boundary, immediately reducing the walking speed by 50% and increasing the support polygon (Trot→Walk). Level-2 is when the angular velocity exceeds the limit, triggering single-leg impedance control, restoring attitude within 0.5s; if not restored within 2s, it is upgraded. Level-3 is when communication with the ground station is lost for 5s, initiating autonomous return mode: backtracking along the topology map, with a maximum backtracking distance of 500m.
[0085] In some embodiments, the expression for autonomous return navigation under no-communication conditions is: ; in, To estimate the position for dead-reckoning, For the actual location, For positioning drift, To control the drift rate, a wheel-mounted odometer, IMU, and magnetometer are integrated, and EKF-SLAM (Extended Kalman Filter Simultaneous Localization and Mapping) is used to maintain the local map; the wheel diameter is 0.15m, and the error model is automatically corrected every 100m. A drift rate of 1cm / s ensures that the error at the end of the 50s backtracking path is <0.5m, which is sufficient to re-enter the communication area.
[0086] In some embodiments, the expression for the emergency rescue mechanism under trapped conditions is: ; in, The trapped trigger flag, - The duration of no movement, For the foot / track ground reaction force sensor readings, Given the current water depth, If the robot's knee joint height and three conditions are met simultaneously, a "stuck + floating" accident is determined, and immediately: all floats are deployed and the water mode is switched; an audible and visual alarm is triggered (120dB, 1kHz, lasting 30s); a 433MHz radio beacon broadcasts GPS (Global Positioning System) + UWB (Ultra-Wideband) coordinates every 3s for 6 hours to facilitate search and rescue.
[0087] Example 2 Figure 2 This is a schematic diagram of a confined space facility safety risk inspection system based on embodied intelligence, as shown in some embodiments of this specification.
[0088] In some embodiments, a confined space facility safety risk inspection system based on embodied intelligence includes: an environmental perception and modeling unit, used to build a three-dimensional environmental model and a semantic topology map based on the confined space environmental data collected by sensors using synchronous positioning and mapping; The path planning unit is used to analyze the optimal inspection path based on the three-dimensional environment model and semantic topology map, combined with a preset risk perception path assessment strategy. The embodied inspection execution subsystem includes an embodied intelligent robot dog vehicle, a multimodal sensor group, and a motion controller. The motion controller controls the embodied intelligent robot dog vehicle to move along the optimal inspection path, acquires terrain data and water depth data in real time during movement, and executes corresponding motion control strategies based on the terrain data and water depth data to maintain the stable attitude of the embodied intelligent robot dog vehicle. The multimodal sensor group is used to acquire multimodal perception data of the confined space facility under the stable attitude. The intelligent analysis unit is used to input the multimodal sensing data into a pre-trained multimodal large model for feature fusion analysis to obtain the safety risk inspection results of confined space facilities.
[0089] In some embodiments, the system includes: an embodied intelligent robot dog vehicle: using a quadruped robot dog as a mobile platform, possessing excellent obstacle-crossing ability and environmental adaptability; a multi-sensor integrated module: SLAM (Simultaneous Localization and Mapping) sensor, high-precision measuring instrument (added if necessary), infrared thermal imager, robot dog's own image sensor, gas comprehensive detector, water depth sensor, intelligent control and data analysis module, remote communication and monitoring module, and safety protection module (anti-collision, anti-flood, abnormal recovery).
[0090] In some embodiments, the deployable buoyancy device employs a float structure made of shape memory alloy material, which automatically deploys upon entering the water, providing buoyancy ≥150% of the vehicle's weight. The underwater propulsion system integrates a brushless motor-driven propeller, achieving an underwater speed ≥0.3 m / s and an endurance ≥45 minutes. The land-water transition control algorithm utilizes attitude stabilization control based on a hydrodynamic model to achieve a smooth transition from land-based walking to water-based floating. Float deployment criteria: ; in, For total buoyancy, The density of water, The volume of water displaced, This refers to gravitational acceleration. Archimedes' buoyancy is determined by the volume of water displaced by the shape memory alloy float after it enters the water. The formula on the right side: 1.5 times the total weight of the robot (safety factor 1.5) ensures the robot dog can maintain a floating posture on the water surface even under maximum load, and provides a 50% redundancy to cope with wave impacts.
[0091] In some embodiments, the underwater propulsion dynamics formula for the robot is: ; in, For the mass of the robot body, For the robot's linear acceleration underwater, The density of water, The linear velocity of the robot underwater. The thrust generated by the brushless motor and ducted propeller is rated at 30N, with a duty cycle linearly adjustable from 0 to 1. The value is 0.45 (measured value of fuselage + float assembly). With a frontal area of 0.12 m², The value is km, and k = 0.32 (an empirical coefficient for the ellipsoid's major-to-diameter ratio of 2:1), reflecting the resistance of water inertia to acceleration. A constant upward force is applied, which, after offsetting 100% of its own weight, leaves a net buoyancy of 0.5 mg, ensuring that the antenna and sensor remain above the water surface at all times. drag This is a square-damped model. F addedThis is to add mass force.
[0092] In some embodiments, the switching logic from water to land and from land to water is as follows: and It was determined to be in a submerged state; and This indicates a logged-in status. in, Given the current water depth, The pitch angle of the aircraft. For foot contact force, 0.08m: Hub center height, to prevent misinterpreting splashing water as water ingress. A pitch angle of 12° indicates that the aircraft has begun its dive into the water. 20N: Threshold of the six-dimensional force sensor on the sole of the foot, confirming that the leg has made contact with a solid ground.
[0093] In some embodiments, the Sigmoid mixed weight transition formula is as follows: ; in, To ultimately control the output force / torque, For land-based control output, For waterborne control output, It is a sigmoid activation function. Given the current water depth, To control the weights for land use and ensure a smooth transition to avoid mode jumps, the range is [0,1]. 0.06m: Water depth at the center of the transition zone (depth at which the float begins to draft). The transition zone width parameter is 0.012m, which corresponds to a sigmoid slope range of 5%–95% of approximately 50mm, ensuring that the mode switch is completed within 0.3s.
Claims
1. A method for inspecting safety risks in confined space facilities based on embodied intelligence, characterized in that, Includes the following steps: S1: Based on synchronous positioning and mapping, sensor-collected confined space environment data are used to build a three-dimensional environment model and semantic topology map; S2: Based on the 3D environment model and semantic topology map, combined with the preset risk perception path assessment strategy, the optimal inspection path is obtained through analysis. S3: Control the embodied intelligent robot dog vehicle to move along the optimal inspection path, acquire terrain data and water depth data in real time during the movement, execute the corresponding motion control strategy according to the terrain data and water depth data to maintain the stable attitude of the embodied intelligent robot dog vehicle, and acquire multimodal perception data of the confined space facility under the stable attitude. S4: Input the multimodal sensing data into a pre-trained multimodal large model for feature fusion analysis to obtain the safety risk inspection results of confined space facilities and complete the safety risk inspection of confined space facilities.
2. The confined space facility safety risk inspection method based on embodied intelligence according to claim 1, characterized in that, S2 includes: Based on the three-dimensional environment model, the standard deviation of height, standard deviation of slope, and obstacle density of the ground in the confined space are extracted. Combined with the ground friction coefficient, the terrain complexity assessment model is used to calculate the terrain complexity value. Based on the semantic topology map, the terrain complexity value is used as the path cost weight, and the graph search algorithm is used to plan the minimum risk cost to obtain inspection path data containing the spatial coordinate sequence of monitoring points. By using a pre-defined risk perception path assessment strategy, the inspection path data containing the spatial coordinate sequence of monitoring points is analyzed to obtain the optimal inspection path.
3. The confined space facility safety risk inspection method based on embodied intelligence according to claim 2, characterized in that, The expression for the terrain complexity value is: ; in, This represents the terrain complexity value. to All are preset weighting coefficients. For high standard deviation, For the standard deviation of slope, For obstacle density, is the coefficient of friction of the ground.
4. The confined space facility safety risk inspection method based on embodied intelligence according to claim 3, characterized in that, S3 includes: Control the embodied intelligent robot dog vehicle to move along the optimal inspection path, and obtain the current step frequency, stride length and foot lift height of the embodied intelligent robot dog vehicle in real time as the reference gait parameters; During the movement, the multimodal terrain complexity assessment model outputs real-time terrain data and water depth data. Based on the terrain data and water depth data, the gait adaptive adjustment algorithm is used to dynamically correct the baseline gait parameters to obtain the adjusted gait parameters adapted to the current terrain, which serve as the corresponding motion control strategy. Using a motion control strategy, the movement of the embodied intelligent robot dog vehicle is controlled to traverse the spatial coordinate sequence of monitoring points in order to maintain the stable posture of the embodied intelligent robot dog vehicle. Multimodal sensor arrays are used to acquire multimodal perception data of the confined space facility.
5. The confined space facility safety risk inspection method based on embodied intelligence according to claim 4, characterized in that, The expression for the adjusted gait parameters is: ; ; ; in, To adjust the step frequency, To adjust the stride, To adjust the height of the hindfoot lift, , and All are baseline parameters. , and All are adjustment coefficients. This represents the terrain complexity value.
6. The confined space facility safety risk inspection method based on embodied intelligence according to claim 4, characterized in that, In step S3, a motion control strategy is used to control the movement of the embodied intelligent robot dog vehicle, traversing the spatial coordinate sequence of monitoring points, and also includes a water / land mode switching step: Real-time acquisition of water depth data in confined spaces, as well as pitch angle data and foot force data of the embodied intelligent robot dog vehicle; Based on water depth data, pitch angle data, and foot force data, the environmental state of the embodied intelligent robot dog vehicle is determined, and the judgment result is obtained. When the judgment result is that the vehicle is in the water, the intelligent robot dog vehicle is controlled to deploy its buoyancy device and switch to underwater propulsion mode; when the judgment result is that the vehicle is on land, the intelligent robot dog vehicle is controlled to retract its buoyancy device and switch to land walking mode, and the intelligent robot dog vehicle is controlled to move and traverse the spatial coordinate sequence of monitoring points.
7. The confined space facility safety risk inspection method based on embodied intelligence according to claim 6, characterized in that, The expression for the judgment result is: and It was determined to be in a submerged state; and This indicates a logged-in status. in, For real-time water depth data, The threshold for the center height of the wheel hub. For pitch angle data, The pitch angle threshold, For data on plantar force, This is the plantar force threshold.
8. The method for inspecting safety risks of confined space facilities based on embodied intelligence according to claim 1, characterized in that, S4 includes: Obtain observation data from various sensors in multimodal sensing data and their corresponding observation noise covariance, and construct a factor graph fusion objective function that includes kinematic prediction terms and sensor observation terms; Based on a pre-trained multimodal large model, the optimal state estimate and unified environmental characteristics of the embodied intelligent robot dog vehicle at each moment are obtained by minimizing the factor graph fusion objective function, thus completing the safety risk inspection of the confined space facility.
9. The method for inspecting safety risks of confined space facilities based on embodied intelligence according to claim 8, characterized in that, The objective function for factor graph fusion is: ; in, Let t be the vehicle state at time t. For control input, Q is the process noise covariance. Let be the observation data of the i-th type of sensor at time t. For the observation model of the i-th type of sensor, This represents the objective function for factor graph fusion. This represents the vehicle motion control function. Represents the observation noise covariance. Indicates kinematic prediction, This indicates the vehicle state at time t-1.
10. A confined space facility safety risk inspection system based on embodied intelligence, used to execute the confined space facility safety risk inspection method based on embodied intelligence as described in any one of claims 1 to 9, characterized in that, include: The environmental perception and modeling unit is used to build a three-dimensional environmental model and a semantic topology map based on the limited spatial environmental data collected by the sensor based on synchronous positioning and mapping. The path planning unit is used to analyze the optimal inspection path based on the three-dimensional environment model and semantic topology map, combined with a preset risk perception path assessment strategy. The embodied inspection execution subsystem includes an embodied intelligent robot dog vehicle, a multimodal sensor group, and a motion controller. The motion controller is used to control the embodied intelligent robot dog vehicle to move along the optimal inspection path, and to acquire terrain data and water depth data in real time during the movement. Based on the terrain data and water depth data, the controller executes the corresponding motion control strategy to maintain the stable posture of the embodied intelligent robot dog vehicle. The multimodal sensor array is used to acquire multimodal sensing data of the confined space facility under the stable attitude. The intelligent analysis unit is used to input the multimodal sensing data into a pre-trained multimodal large model for feature fusion analysis to obtain the safety risk inspection results of confined space facilities.
Citation Information
Patent Citations
Underwater hexapod robot gait generation and conversion method based on CPG-Hopf network coupling algorithm
CN113985874A
Control method of operation type intelligent foot robot applied to electric power inspection
CN119472641A
Dynamic stable overturn-preventing device for intelligent inspection robot dog
CN119690116A
Quadruped robot inspection method based on multi-modal sensing fusion
CN121433214A
Hexapod robot with amphibious function
CN121469202A