Multi-axis coordination and multi-sensor fusion subway tunnel roof inspection method and system

CN122574801BActive Publication Date: 2026-09-08SHANDONG JIANZHU UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202611054636.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-07-16
Publication Date
2026-09-08
Estimated Expiration
2046-07-16

AI Technical Summary

Technical Problem

[0008]针对现有隧道巡检技术缺乏多维空间自适应避障能力、单一视觉易受环境干扰以及运行状态闭环管理缺失等问题,本发明提供一种多轴协同与多传感器融合的地铁隧道顶板巡检方法,以实现从底层硬件自检到云端交互的全流程闭环管控,更通过多模态数据融合与逆运动学路径规划,实现复杂管线遮挡下对微小缺陷的高精度贴面探测与端侧实时预警

Benefits of technology

[0084] This invention overcomes the limitations of complex pipeline obstruction, achieving blind-spot-free surface detection. For redundant degree-of-freedom systems, it innovatively employs an improved RRT* multi-objective optimization path planning algorithm based on target bias and adaptive step size. Compared to traditional detection equipment that uses hard-coded fixed trajectories, this method can autonomously extract the boundaries of obstacles such as pipelines at the tunnel top, and calculate a collision-free multi-axis linkage trajectory in a virtual configuration space, guiding the detection device to deftly probe into unobstructed entry gaps behind pipelines. This method effectively avoids interference from rigid contact wires and complex pipelines, completely eliminating the visual blind spots of traditional equipment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122574801B_ABST
    Figure CN122574801B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of rail transit route inspection, and discloses a subway tunnel roof inspection method and system based on multi-axis cooperation and multi-sensor fusion, which comprises the following steps: acquiring environmental point cloud data of the tunnel roof, extracting a three-dimensional bounding box of an obstacle, and guiding a detection device to probe into an unobstructed cut gap; performing spatial registration on the collected 2D texture image and 3D depth point cloud, generating fused RGB-D multi-modal tensor data, and inputting the data into a lightweight defect recognition network to output the classification and three-dimensional geometric features of the roof defect. The system comprises an initialization and self-checking module, an active safety and abnormal fuse module, a dynamic addressing and positioning module, a spatial perception and obstacle avoidance planning module, a multi-source data acquisition and fusion module, and an evaluation and edge cloud cooperation module. The present application breaks through the limitation of complex pipeline shielding, effectively suppresses water stain reflection noise, and realizes close-up surface high-precision detection and real-time early warning from the side.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of subway tunnel roof inspection technology, and more specifically, to an intelligent inspection method and system for subway tunnel roof based on multi-axis collaborative path planning and multi-source heterogeneous data fusion. Background Technology

[0002] Structural health monitoring of subway tunnels is a crucial aspect of ensuring safe train operation. The tunnel roof is constantly subjected to geological stress and train vibrations, making it prone to micro-cracks, segment misalignment, or spalling. Furthermore, this area is directly overhead, so any collapse could easily lead to major accidents such as power outages in the overhead contact system. Therefore, high-frequency, high-precision inspections of the tunnel roof are of irreplaceable importance.

[0003] Currently, various track-mounted automated inspection devices have been introduced in the industry to replace traditional manual inspections. However, existing inspection methods and underlying algorithms still reveal the following significant technical shortcomings when faced with the complex real-world conditions of subway tunnels:

[0004] First, existing motion control logic is rigid and lacks the ability to dynamically avoid obstacles and adaptively detect in complex three-dimensional spaces. Most existing tunnel inspection systems use hard-coded control logic of "fixed trajectory and fixed-point shooting". Due to the dense network of rigid contact wires, lighting cables, communication cables, and fire pipes on the roof of subway tunnels, traditional inspection equipment cannot autonomously detect and replan its path when encountering these obstacles. Furthermore, existing control systems lack the ability to perform multi-axis collaborative inverse kinematics calculations when facing the narrow, concealed spaces behind pipelines, making it impossible to guide the detection end to perform flexible penetration and close-range "surface" exploration, resulting in numerous and potentially fatal blind spots on the tunnel roof.

[0005] Second, existing visual inspection algorithms are modal-dependent, have weak anti-interference capabilities, and cannot acquire depth features of defects. Subway tunnels are typically low-light, high-dust environments, often accompanied by localized water seepage and reflections from tunnel segments. Existing image recognition methods generally rely solely on two-dimensional visible light cameras to acquire images and use conventional thresholding or basic deep learning networks to extract cracks. Under such complex environmental noise interference, water stain boundaries and cable shadows are easily misjudged as structural cracks by existing algorithms, leading to a very high false alarm rate. Furthermore, a single two-dimensional vision system cannot acquire three-dimensional depth information of segment misalignment and spalling, preventing the system from quantitatively assessing the severity of defects from a geometric and mechanical perspective.

[0006] Third, the existing data processing architecture suffers from computing power bottlenecks, making it difficult to achieve real-time early warning at the edge. Traditional inspection data streams often employ a processing mode of "blind front-end collection, full data transmission, and offline cloud analysis." Since most areas of the tunnel roof are in a normal state (defect-free), this indiscriminate full data transmission method not only puts immense pressure on the limited communication bandwidth within the tunnel but also results in massive amounts of redundant data consuming enormous storage and computing resources. This high-latency offline processing mode makes it impossible to achieve real-time edge-side response and alarm interception for high-risk defects while inspection operations are underway.

[0007] In summary, existing subway tunnel inspection technologies have significant technical bottlenecks in multi-degree-of-freedom collaborative obstacle avoidance control, multi-sensor anti-interference fusion, and real-time computing architecture. There is an urgent need for an intelligent inspection method that can actively sense obstacle gaps, fuse multi-modal sensor data, and output high-precision detection results in real time at the end. Summary of the Invention

[0008] To address the shortcomings of existing tunnel inspection technologies, such as the lack of multi-dimensional spatial adaptive obstacle avoidance, susceptibility to environmental interference due to single-vision imaging, and the absence of closed-loop management of operational status, this invention provides a multi-axis collaborative and multi-sensor fusion method for inspecting the roof slab of subway tunnels. This method achieves closed-loop control throughout the entire process, from underlying hardware self-inspection to cloud interaction. Furthermore, through multi-modal data fusion and inverse kinematics path planning, it enables high-precision surface detection and real-time end-side early warning of minute defects even under complex pipeline obstructions. The invention also provides a multi-axis collaborative and multi-sensor fusion subway tunnel roof slab inspection system that implements this method.

[0009] This invention relates to a multi-axis collaborative and multi-sensor fusion method for inspecting the roof slab of a subway tunnel. The method involves mounting a detection device on the end of a robotic arm, which includes a camera and a laser profilometer. The robotic arm is carried by an inspection vehicle that travels on a track for inspection. The inspection vehicle is equipped with obstacle detection radar and an environmental perception camera. The method includes the following steps:

[0010] S1: The inspection vehicle cruises along the subway tunnel track, simultaneously activating the obstacle detection radar and camera on the inspection vehicle to actively perceive the environmental safety status and trigger the abnormal circuit breaker mechanism when an intruding obstacle is detected.

[0011] S2: Under normal patrol conditions, the Kalman filter algorithm is used to fuse the absolute odometer of the patrol vehicle with the visual features of the tunnel for joint positioning, update the global coordinates, and park the patrol vehicle after matching the target segment.

[0012] S3: Obtain environmental point cloud data of the tunnel top through a laser profilometer, extract the three-dimensional bounding box of the obstacle, and calculate the collision-free multi-axis linkage trajectory in three-dimensional space based on the multi-objective optimization path planning algorithm with target offset and adaptive step size, so as to guide the detection device to probe into the obstacle-free cutting gap.

[0013] S4: The camera and laser profilometer in the detection device work synchronously to acquire 2D texture images and 3D depth point clouds respectively, and spatially register the 2D texture images and 3D depth point clouds to generate fused RGB-D multimodal tensor data.

[0014] S5: Input the RGB-D multimodal tensor data into the lightweight defect recognition network, output the classification and three-dimensional geometric features of the top plate defects, and upload them to the cloud scheduling server based on the end-cloud collaboration mechanism.

[0015] Each step is further defined as follows.

[0016] The abnormal circuit breaker mechanism in step S1 is:

[0017] Real-time acquisition of obstacle point cloud distance d measured by the obstacle detection radar on the inspection vehicle t and the current speed v of the inspection vehicle t Determine if the collision threshold condition is met: D safe To preset a safe distance, a max The maximum braking deceleration is set; if this is met, all moving parts are locked (all electric components in the inspection vehicle and robotic arm are de-energized, and the moving mechanism is locked to prevent secondary collisions), and the camera on the inspection vehicle captures on-site images as an over-authorization alarm data packet and uploads them to the cloud dispatch server.

[0018] The specific process of step S2 is as follows:

[0019] Define state variables , where p k v represents the current absolute mileage. k Let be the speed of the inspection vehicle in the kth sampling period; the state transition equation is described as:

[0020] ;

[0021] Where A is the state transition matrix; B represents the state variables of the previous time step; B is the control input matrix; µ k For the inspection vehicle servo drive command; W k Represented as system process noise, in Kalman filtering, it represents the inherent uncertainty of the kinematic model caused by slippage or vibration when the inspection vehicle is traveling on the track;

[0022] Simultaneously, environmental perception cameras are used to capture the joints of tunnel segments, and the standard spacing between the joints is extracted as an observation variable. The mileage coordinates p are then updated in real time using Kalman gain. k When p kThe target tunnel segment coordinates are matched to those sent by the cloud scheduling server. When a suspected crack is detected in the initial screening, the inspection vehicle parks at the detection origin.

[0023] The specific process of step S3 is as follows:

[0024] After the inspection vehicle is parked, the camera and laser profilometer are activated to scan the top of the tunnel; the background point cloud of the tunnel's curved surface is removed using region growing or RANSAC algorithms, and the set of 3D bounding boxes of obstacles is extracted. obs The improved Rapid Expanding Random Tree (RRT*) algorithm is used for inverse kinematic path planning. The specific process is as follows:

[0025] (1) Target bias sampling strategy:

[0026] Generate random nodes q in the configuration space rand When, a bias probability P is introduced. bias Let the mapping attitude of the target detection slot in the configuration space be q. goal The sampling function is defined as:

[0027] ;

[0028] Where ξ is a uniformly random number in the interval (0,1). Represents a free-configuration space without collisions;

[0029] (2) Collision detection and adaptive step size expansion:

[0030] Let the distance in the current tree be q. rand The nearest node is q near , in q near To q rand Expand to generate a new node q new At that time, an adaptive step size λ based on the obstacle distance is adopted:

[0031] ;

[0032] The step size λ is determined by the distance to the nearest obstacle:

[0033] ;

[0034] In the formula, λ max To set the maximum safe step size, d obs (q near The step size is λ, calculated from the end effector cylinder of the robotic arm at the current node to the nearest pipeline. k is a smoothing coefficient. As the robotic arm moves away from obstacles (such as overhead contact lines), the step size tends towards λ. max This enables rapid, wide-range movement; when the robotic arm probes into the gap in the pipeline (d obs (qnear When the step size is extremely small, the step size λ decreases exponentially, enabling smooth micro-interlacing;

[0035] Randomly sample joint angle nodes q in configuration space rand To balance obstacle avoidance safety and robotic arm motion smoothness, a composite cost function C(q) is defined to select the optimal expansion node:

[0036] ;

[0037] In the formula, L is the joint angle spatial distance (i.e., path length) from the robot arm's initial posture to its current posture; dist is the boundary box O of the robot arm's link envelope cylinders to the obstacle. obs,i The shortest physical Euclidean distance; E(q) is the energy consumption penalty term, used to suppress the ineffective reciprocating motion of the large inertia joints of the robotic arm; α, β, γ are preset weight coefficients; This indicates the initial joint angle configuration of the robotic arm; This indicates the current evaluation of the robotic arm joint angle configuration; M represents the total number of extracted obstacle 3D bounding boxes (representing the total number of obstacles extracted by the current spatial clustering algorithm).

[0038] By minimizing the aforementioned cost function, a set of servo commands for each joint axis of the robotic arm that enable interference-free motion of each link in three-dimensional space is calculated, guiding the detection device to precisely penetrate into the gap and reach the effective focal length range from the top plate.

[0039] The step S4, which generates the fused RGB-D multimodal tensor data, specifically includes:

[0040] By using the pre-calibrated intrinsic parameter matrix of the camera (including focal length and optical center offset) and the extrinsic parameter rotation and translation matrix of the lidar on the inspection vehicle to the camera, the 3D depth point cloud in the world coordinate system is reprojected onto the 2D image pixel plane; through interpolation alignment operations, the depth value is assigned to the corresponding RGB pixel point to generate four-dimensional tensor data.

[0041] The specific process is as follows:

[0042] After the camera and laser profilometer reach the designated pose, they work synchronously. The camera acquires a 2D texture image (pixel coordinate system (u,v)), and the laser profilometer acquires a 3D depth point cloud, i.e., a world coordinate system (u,v). , , );

[0043] By using a joint calibration matrix, the 3D point cloud is reprojected onto the 2D image plane, and the transformation relationship is as follows:

[0044] ;

[0045] In the formula, K is the intrinsic parameter matrix of the camera, R and T are the rotation matrix and translation vector of the laser radar on the inspection vehicle to the camera, respectively, and is the depth value in the camera coordinate system. Through interpolation and alignment, four-dimensional tensor data (RGB-D) containing RGB and Depth channels is generated.

[0046] The lightweight defect recognition network in step S5 is a deep feature-guided convolutional neural network, and its processing includes:

[0047] The depth channel matrix is ​​extracted from the RGB-D multimodal tensor data and sequentially processed by global average pooling, convolutional dimensionality reduction and expansion, and activation function to generate a spatial attention weight matrix. The spatial attention weight matrix is ​​then multiplied element-wise with the RGB texture feature map extracted by the network backbone to physically suppress environmental noise interference caused by water stain reflection. The defect classification, 2D bounding box size, and 3D depth value are output synchronously through a multi-task output head.

[0048] The specific process of the aforementioned lightweight defect identification network includes:

[0049] (1) Four-channel depthwise separable convolution:

[0050] The RGB-D four-dimensional tensor generated in step S5 is used as the network input. To control the number of parameters, depthwise separable convolutions are used in the feature extraction backbone; let the input feature map size be D. F ×D F ×M, where M is the number of input channels, initially set to 4, and the kernel size is D. k ×D k Its standard computational cost is:

[0051] ;

[0052] N represents the number of output channels of the convolutional layer;

[0053] (2) Attention mechanism guided by deep features:

[0054] The input deep feature channel matrix F depth The extracted spatial attention weight matrix W is generated after dimensionality reduction and expansion via global average pooling and convolutional layers, and then activated by the sigmoid activation function σ. attn :

[0055] ;

[0056] Where W1 and W2 are network learning parameters, and δ is the ReLU activation function. This represents a global average pooling operation; subsequently, the attention weights are compared with the RGB texture feature map F.RGB Feature recalibration is performed by element-wise multiplication:

[0057] ;

[0058] This represents the output feature map after spatial attention weight recalibration;

[0059] To simultaneously predict the 2D contour and 3D depth of defects, the network's loss function is Loss. total Optimized to:

[0060] ;

[0061] Among them, L cls The cross-entropy classification loss (used to distinguish between cracks, seepage, and spalling) is L. box L is the regression loss for the target bounding box (used to calculate the length and width of the defect in the image). depth This is the depth regression loss (mean squared error, used to quantify crack depth or misalignment height). For depth loss weights.

[0062] If it is deduced that the crack width exceeds 0.3mm or there is a risk of block falling, it is judged as a high-risk defect; the high-risk data message is added to the highest priority queue and transmitted to the cloud dispatch server in real time through the network, so as to realize efficient anomaly early warning from the inspection vehicle to the cloud dispatch server.

[0063] The process of uploading to the cloud scheduling server based on the end-to-cloud collaboration mechanism in step S5 includes hierarchical storage and breakpoint resume mechanism:

[0064] The collected raw data stream is written to the memory ring buffer on the edge side in real time; based on the output of the lightweight defect identification network, the normal segment data that is determined to be defect-free is overwritten and cleaned at the memory level; the data that is determined to be suspected or high-risk defects is compressed and encapsulated with metadata containing absolute coordinates and timestamps before being persisted locally to disk; the current network bandwidth is monitored, and when the bandwidth is sufficient, the defect data that has been written to disk is asynchronously transmitted to the cloud scheduling server, and the interruption is resumed based on the local log after the network interruption is restored;

[0065] The specific process includes:

[0066] (1) End-side annular buffer and data cleaning:

[0067] During the patrol and detection process, the status stream, high-definition video stream and point cloud data of the drive motors of the inspection vehicle and the robotic arm are first written into the memory ring buffer of the main control computer. After the defect assessment in step S5 is performed, the normal segment data that is determined to be defect-free is directly overwritten and cleaned at the memory level and is not stored.

[0068] (2) Local solidification and metadata encapsulation of high-value data:

[0069] For RGB-D fusion tensors identified as potentially defective or highly defective, they are retrieved from the circular buffer and stored in the main control computer; before local storage, the data is lightweight compressed according to a preset protocol and tagged with an image containing absolute odometry coordinates p. k Metadata tags for current timestamp, ambient temperature and humidity, and main control computer fault codes are used to generate standardized JSON or XML defect description files;

[0070] (3) Asynchronous pass-through with adaptive bandwidth and breakpoint resumption:

[0071] When sufficient network bandwidth is detected in the current tunnel section, high-priority defect feature maps and alarm logs are asynchronously transmitted to the cloud. When the train travels to a signal blind spot and the network connection is interrupted, the transmission queue is automatically suspended. After the network handshake is restored, the interrupted transmission is resumed based on the local log database.

[0072] (4) Cloud-based digital twin mapping and dimensionality reduction storage:

[0073] After receiving the defect description file with absolute coordinates and point cloud features, the cloud scheduling server no longer saves the original video stream. Instead, it directly maps and updates the three-dimensional geometric features of the defect to the BIM (Building Information Modeling) or digital twin system of the entire subway tunnel, realizing full life-cycle management of tunnel defects with data lightweighting as the core.

[0074] The system for implementing the above-mentioned multi-axis collaborative and multi-sensor fusion method for subway tunnel roof inspection includes:

[0075] The initialization and self-test module is used to control the robotic arm to find and reset zero and perform cloud communication self-tests.

[0076] The active safety and abnormal circuit breaker module is used to trigger emergency braking and unauthorized alarm based on the perception threshold during cruise.

[0077] The dynamic addressing and positioning module is used to fuse multi-source features to update global coordinates and perform precise parking;

[0078] The spatial perception and obstacle avoidance planning module is used to extract three-dimensional obstacle boundaries and use an improved path planning algorithm to solve multi-axis cooperative obstacle avoidance trajectories.

[0079] A multi-source data acquisition and fusion module is used to synchronously acquire and spatially register RGB-D multimodal tensor data;

[0080] The evaluation and edge-cloud collaboration module is used to quantify network output defects based on deep feature guidance and to perform hierarchical transmission.

[0081] The present invention also provides an electronic device, including at least one processor; and a memory communicatively connected to the processor; the memory storing a computer program executed by the processor, the computer program being executed by the processor to enable the processor to execute the multi-objective optimization path planning algorithm based on target bias and adaptive step size of the present invention.

[0082] The present invention also provides a computer-readable storage medium storing computer instructions for causing a computer to execute the multi-objective optimization path planning algorithm based on target bias and adaptive step size of the present invention.

[0083] Compared with the prior art, the present invention has the following significant advantages:

[0084] This invention overcomes the limitations of complex pipeline obstruction, achieving blind-spot-free surface detection. For redundant degree-of-freedom systems, it innovatively employs an improved RRT* multi-objective optimization path planning algorithm based on target bias and adaptive step size. Compared to traditional detection equipment that uses hard-coded fixed trajectories, this method can autonomously extract the boundaries of obstacles such as pipelines at the tunnel top, and calculate a collision-free multi-axis linkage trajectory in a virtual configuration space, guiding the detection device to deftly probe into unobstructed entry gaps behind pipelines. This method effectively avoids interference from rigid contact wires and complex pipelines, completely eliminating the visual blind spots of traditional equipment.

[0085] By integrating multimodal depth perception, this invention effectively suppresses environmental noise and reduces false alarm rates. It overcomes the limitations of a single visible light camera by spatially registering 2D textures with 3D depth point clouds to generate RGB-D multimodal tensor data, and constructs a lightweight convolutional neural network guided by depth features. This network utilizes the numerical steps generated on the depth map by real structural defects to generate spatial attention weights, effectively suppressing environmental noise interference caused by low illumination, water stain reflections, etc., from a physical mechanism perspective. Simultaneously, the multi-task output head enables high-precision quantitative evaluation of three-dimensional geometric features such as crack depth and misalignment, significantly reducing the false alarm rate for defects.

[0086] This invention addresses communication and computing bottlenecks by constructing an edge-cloud collaborative and hierarchical storage mechanism. Recognizing the pain points of limited communication bandwidth and onboard storage space in subway tunnels, it designs a hierarchical storage and breakpoint resume mechanism based on edge pre-screening. Raw data is cleaned within a circular memory buffer of the onboard edge computing platform; defect-free and normal data is directly overwritten and discarded, while only high-risk or suspected defective data is compressed, solidified, and asynchronously transmitted. This mechanism saves a significant amount of ineffective storage overhead, greatly alleviates network pressure, and, combined with breakpoint resume, ensures "zero loss" of high-value defective data packets, enabling real-time edge response and early warning.

[0087] By introducing physical circuit breaking and multi-source positioning mechanisms, this invention comprehensively ensures the safety of unmanned operations. It constructs an active safety anomaly circuit breaking mechanism using obstacle detection radar and environmental perception cameras. When collision threshold conditions are met, it directly bypasses the conventional planning to issue an emergency braking command and issue an overreach alarm, ensuring the absolute safety of the equipment. Simultaneously, this invention employs a Kalman filter algorithm to fuse the absolute encoder of the inspection vehicle's drive shaft with visual features for joint positioning. This overcomes the cumulative error caused by slippage from a single odometer, achieving precise "point-to-point" parking and deployment of target segments, improving the accuracy and automation level of the inspection closed loop. Attached Figure Description

[0088] Figure 1 This is a flowchart of the subway tunnel roof inspection method in this invention;

[0089] Figure 2 This is a block diagram of the active safety and state machine logic in this invention;

[0090] Figure 3 This is a flowchart of the improved RRT* multi-axis cooperative obstacle avoidance algorithm in this invention;

[0091] Figure 4 This is a diagram of the lightweight neural network architecture guided by deep features in this invention;

[0092] Figure 5 This is a data flow diagram of mid-cloud collaborative hierarchical storage and transmission in this invention;

[0093] Figure 6 This is a virtual modular structure diagram of the inspection system in this invention. Detailed Implementation

[0094] This invention relates to track-based inspection, comprising an inspection vehicle, a robotic arm, and a detection device. The robotic arm is mounted on the inspection vehicle, and the detection device is located at the end of the robotic arm. An absolute encoder and a braking mechanism are mounted on the drive shaft of the inspection vehicle, which also houses an obstacle detection lidar and an environmental perception camera. The robotic arm, capable of using a six-axis multi-degree-of-freedom design, drives the detection device to the inspection area. The detection device includes a high-definition camera and a laser profilometer, both existing technologies. The method is executed by a main control computer on the inspection vehicle.

[0095] The multi-axis collaborative and multi-sensor fusion method for inspecting the roof of a subway tunnel, as described in this invention, is as follows: Figure 1 As shown, it can be divided into six steps, S1-S6, which are explained in detail below.

[0096] Step S1: Initialization and cloud-based closed-loop self-check.

[0097] First, perform hardware topology scanning and node wake-up, and then perform zero-finding and reset operations on the robotic arm and detection device in sequence.

[0098] The system extracts the initial readings of the absolute encoders of each drive shaft of the inspection vehicle and performs a self-test using the baseline conditions (baseline temperature) of the sensors (specifically, these sensors include: the servo motor's built-in temperature sensor (to prevent overheating of each drive shaft motor); the energy storage battery pack temperature sensor (to ensure battery charging and discharging safety); and the main control computer motherboard temperature sensor (to ensure the stability of edge computing power). The self-test results are packaged into a heartbeat data packet and fed back to the cloud dispatch server in real time via a 5G communication module at a preset frequency (e.g., 1Hz). If the self-test fails, the drive wheels of the inspection vehicle are locked, and a self-test log containing fault codes is reported to the cloud dispatch server, requesting manual intervention. If the self-test passes, the vehicle enters a patrol standby state.

[0099] Step S2: Active safety cruise and abnormal circuit breaker mechanism.

[0100] The inspection vehicle travels along the subway tunnel track, using the obstacle detection lidar and cameras on the vehicle to acquire real-time 3D point cloud and video stream of the track and clearance space ahead; if it detects an intruding obstacle, foreign object on the track, or abnormal equipment condition, it immediately triggers the abnormal circuit breaker mechanism (emergency braking command) and uploads the captured on-site environmental images and radar alarm data to the cloud dispatch server in real time, entering the safety suspension mode.

[0101] After passing the self-inspection, the inspection vehicle's obstacle detection radar emits detection waves in real time at a preset sampling rate during its patrol. For example... Figure 2 As shown, the radar safe distance threshold D is set. safe Let d be the distance of the nearest obstacle point cloud acquired by the radar at the current moment. t The current speed of the inspection vehicle is vt The maximum braking deceleration is a max The dynamic braking solution logic formula running inside the main control computer is as follows:

[0102] ;

[0103] If d is satisfied t Less than this value (i.e.: If the abnormal circuit breaker is triggered, the inspection vehicle will brake suddenly. At the same time, the camera on the inspection vehicle will capture the current frame image and upload it to the cloud dispatch server along with the radar point cloud as an unauthorized alarm data packet.

[0104] Step S3: Dynamic addressing and precise target parking.

[0105] Based on the Kalman filter fusion algorithm of the absolute odometer of the inspection vehicle (encoder installed on the drive shaft of the inspection vehicle) and the visual features of the tunnel segment joints, the global coordinates of the inspection vehicle are updated in real time. When the inspection vehicle cruises to the preset key inspection segment area, or when the camera initially identifies the location of the roof plate suspected of having abnormal defects, the inspection vehicle is controlled to decelerate smoothly, brake precisely and park at the target inspection origin.

[0106] To address the cumulative error caused by slippage from a single encoder on smooth rails during normal patrol operation of the inspection vehicle, this invention employs a Kalman filter algorithm to fuse the inspection vehicle's drive shaft encoder with visual features for joint positioning. State variables are defined. , where p k v represents the current absolute mileage. k The speed of the inspection vehicle in the kth sampling period (the value is derived from the aforementioned current speed v). t Discretized sampled values); the state transition equation is described as:

[0107] ;

[0108] in, Let A represent the state variables of the previous time step; A is the state transition matrix, B is the control input matrix, and µ represents the variable. k For the servo drive commands of the inspection vehicle, W k Represented as system process noise, in Kalman filtering, it represents the inherent uncertainty of the kinematic model caused by slippage or vibration when the inspection vehicle is traveling on the track.

[0109] Simultaneously, the environmental perception camera on the inspection vehicle captures the segment joints, extracts the standard spacing between the joints as the observation variable, and updates the mileage coordinates p in real time through Kalman gain. k When p kMatching the target segment coordinates issued by the cloud scheduling server, when a suspected crack is detected in the initial screening, the inspection vehicle smoothly decelerates and precisely parks at the detection origin.

[0110] Step S4: Local spatial perception and multi-axis cooperative obstacle avoidance trajectory planning.

[0111] Three-dimensional depth information of the tunnel top environment is collected, and three-dimensional bounding boxes of static obstacles such as catenary and pipelines are extracted using spatial clustering algorithms. The unobstructed cut-in gap between pipelines is calculated. A redundant kinematic model of "lifting mechanism - horizontal displacement mechanism - robotic arm" is established. Based on the cut-in gap coordinates, a multi-objective optimization path planning algorithm (RRT algorithm) is used to solve a set of collision-free multi-axis linkage trajectories in three-dimensional space to guide the detection device to bypass obstacles and approach the roof surface of suspected defects.

[0112] The multi-objective optimization path planning algorithm based on target bias and adaptive step size is as follows: Establish the inverse kinematics configuration space of the robotic arm; when sampling nodes in the configuration space, introduce a target bias probability so that randomly sampled nodes directly take the mapped posture of the unobstructed entry gap with a preset probability; during node expansion, calculate the Euclidean distance between the current node and the three-dimensional bounding box of the nearest obstacle, and dynamically adjust the expansion step size based on the Euclidean distance. The expansion step size is positively correlated with the Euclidean distance, so as to adaptively reduce the step size for smooth insertion when approaching the pipeline gap.

[0113] Solving collision-free multi-axis linkage trajectories includes: defining a composite cost function to evaluate extended nodes. The composite cost function includes: a joint angle spatial distance penalty term from the robot arm's initial posture to its current posture, a shortest Euclidean distance penalty term from the envelope of each link to the obstacle bounding box, and an energy consumption penalty term to suppress ineffective motion of large inertia joints.

[0114] The specific process is as follows.

[0115] After the inspection vehicle is parked, the 3D sensors (high-definition camera and laser profilometer) are activated to scan the tunnel ceiling. Region growing or RANSAC algorithms are used to remove the background point cloud on the tunnel's curved surface, extracting a set of 3D bounding boxes for obstacles such as overhead contact lines and cables. obs .

[0116] For a six-axis robotic arm with redundant degrees of freedom, if the traditional RRT* algorithm is used to randomly scatter points in the whole space, the computational load is large and the convergence is extremely slow in the "narrow and pipeline-filled" environment at the top of the tunnel. Therefore, this invention uses an improved fast expanding random tree (RRT*) algorithm for inverse kinematic path planning.

[0117] In multi-axis cooperative obstacle avoidance trajectory planning, this invention addresses the challenges of dense pipelines on the top of subway tunnels, narrow detection gaps, and a large inverse kinematics space for multi-degree-of-freedom robotic arms. It employs an improved RRT* algorithm based on target bias and adaptive step size. The principle and process are as follows: Figure 3 As shown, the specific process is as follows.

[0118] (1) Target bias sampling strategy:

[0119] Generate random nodes q in the configuration space rand In this case, instead of using a globally uniform distribution, a bias probability P is introduced. bias Let the mapping attitude of the target detection slot in the configuration space be q. goal The sampling function is defined as:

[0120] ;

[0121] Where ξ is a uniformly random number in the interval (0,1). This represents a collision-free, free configuration space. This strategy gives the random tree a strong goal orientation in the early stages of growth, significantly reducing the expansion of ineffective nodes in unobstructed, open areas.

[0122] (2) Collision detection and adaptive step size expansion:

[0123] Let the distance in the current tree be q. rand The nearest node is q near , in q near To q rand Expand to generate a new node q new In this invention, a fixed step size is abandoned in favor of an adaptive step size λ based on the distance to the obstacle.

[0124] ;

[0125] The step size λ is determined by the distance to the nearest obstacle:

[0126] ;

[0127] In the formula, λ max To set the maximum safe step size, d obs (q near The step size is λ, calculated from the end effector cylinder of the robotic arm at the current node to the nearest pipeline. k is a smoothing coefficient. As the robotic arm moves away from obstacles (such as overhead contact lines), the step size tends towards λ. max This enables rapid, wide-range movement; when the robotic arm probes into the gap in the pipeline (d obs (q nearWhen the step size is extremely small, the step size λ decreases exponentially, enabling smooth micro-interlacing and effectively avoiding "patterning" collisions caused by fixed large step sizes.

[0128] Randomly sample joint angle nodes q in configuration space rand To balance obstacle avoidance safety and robotic arm motion smoothness, this invention defines a composite cost function C(q) to select the optimal expansion node:

[0129] ;

[0130] In the formula, L is the joint angle spatial distance (i.e., path length) from the initial posture to the current posture. This indicates the initial joint angle configuration of the robotic arm; This indicates the current evaluation of the robot arm's joint angle configuration; dist represents the boundary box O from the enveloping cylinders of each link of the robot arm to the obstacle. obs,i The shortest physical Euclidean distance; E(q) is the energy consumption penalty term, used to suppress the ineffective reciprocating motion of the large inertia joints of the robotic arm; α, β, γ are preset weight coefficients.

[0131] By minimizing the aforementioned cost function, a set of interference-free robotic arm axis follow-up servo commands are calculated, guiding the end-effector detection device to precisely insert into the gap and reach within the effective focal length (e.g., 150mm to 200mm) from the top plate.

[0132] Step S5: Multiple Sources Heterogeneous data fusion and acquisition.

[0133] After the high-definition camera and laser profilometer reach the designated pose, they work synchronously. The camera acquires 2D texture images (pixel coordinate system (u,v)), and the laser profilometer acquires 3D depth point clouds (world coordinate system (u,v)). , , )).

[0134] By using a joint calibration matrix, the 3D point cloud is reprojected onto the 2D image plane, and the transformation relationship is as follows:

[0135] ;

[0136] In the formula, K is the intrinsic parameter matrix of the camera (including focal length and optical center offset), R and T are the rotation matrix and translation vector of the lidar on the inspection vehicle to the camera, respectively, and is the depth value in the camera coordinate system. Through interpolation and alignment, four-dimensional tensor data (RGB-D) containing RGB and depth channels is generated. This step effectively compensates for the lack of spatial geometric features in the reflection of water stains in a single visible light image.

[0137] Step S6: End-side defect assessment and feedback;

[0138] The generated RGB-D four-dimensional tensor is input into a lightweight convolutional neural network deployed on the main control computer. Considering the computing power limitations at the edge, this invention uses a lightweight backbone network (improved MobileNet) to extract features. Traditional MobileNet only accepts RGB three-channel input. Because this invention incorporates a laser profilometer (to obtain depth) in step S5, and considering the computing power limitations of the main control computer, as well as the severe interference of tunnel low light and water stain reflections on single 2D vision, the improvement lies in designing a lightweight network structure that can process 4-channel (RGB-D) features and utilizes depth information as "attention weights".

[0139] This invention constructs an improved lightweight backbone network based on deep feature guidance, the principle and process of which are as follows: Figure 4 As shown, the specific structure and algorithm model are as follows.

[0140] (1) Four-channel depthwise separable convolution:

[0141] The RGB-D four-dimensional tensor generated in step S5 is used as the network input. To control the number of parameters, depthwise separable convolutions are used in the feature extraction backbone; let the input feature map size be D. F ×D F ×M, where M is the number of input channels, initially set to 4, and the kernel size is D. k ×D k Its standard computational cost is:

[0142] ;

[0143] N represents the number of output channels of the convolutional layer.

[0144] Compared to standard convolution, it significantly reduces the power consumption and inference latency of industrial control computers, and the computational load is reduced to only a fraction of the original amount. .

[0145] (2) Attention mechanism guided by deep features:

[0146] To address the issue of water stain reflections being easily misidentified as cracks, the network incorporates a self-developed deep feature attention branch after each inverted residual block. The input deep feature channel matrix F... depth The extracted spatial attention weight matrix W is generated after dimensionality reduction and expansion via global average pooling and convolutional layers, and then activated by the sigmoid activation function σ. attn :

[0147] ;

[0148] Where W1 and W2 are network learning parameters, and δ is the ReLU activation function. This indicates a global average pooling operation.

[0149] Subsequently, this attention weight is compared with the RGB texture feature map F RGB Feature recalibration is performed by element-wise multiplication:

[0150] ;

[0151] This represents the output feature map after spatial attention weight recalibration.

[0152] Through the optimized algorithm model described above, real structural cracks or debris will produce a significant numerical step in the depth map. Using the above formula, the network utilizes depth abrupt change information to generate high weights, enhancing the response of real physical defects in the feature map; while water stains, although appearing as cracks in RGB, are smooth in the depth map (with extremely low weights), and are thus directly suppressed by the physical mechanism during network forward propagation, fundamentally reducing the false alarm rate in complex environments.

[0153] To simultaneously predict the 2D contour and 3D depth of defects, the network's loss function is Loss. total Optimized to:

[0154] ;

[0155] Among them, L cls The cross-entropy classification loss (used to distinguish between cracks, seepage, and spalling) is L. box L is the regression loss for the target bounding box (used to calculate the length and width of the defect in the image). depth This is the depth regression loss (mean squared error, used to quantify crack depth or misalignment height). For depth loss weights.

[0156] The main control computer outputs defect results with physical dimension markings. If the crack width is determined to exceed 0.3mm or there is a risk of chipping, it is classified as a high-risk defect. The high-risk data message is added to the highest priority queue and transmitted in real time to the cloud dispatch server via the 5G network, thereby achieving efficient anomaly early warning from the inspection vehicle to the cloud dispatch server.

[0157] In summary, the core working principle of this invention lies in constructing a three-level closed-loop link of "macro-adaptive scheduling - micro-flexible perception - edge-cloud collaborative decision-making".

[0158] In actual subway nighttime inspection operations, the overall collaborative logic is as follows.

[0159] During the inspection of the entire tunnel, the inspection vehicle is no longer a blindly moving vehicle. Instead, it achieves precise "point-to-point" delivery of target segments in the tunnel, which is tens of kilometers long, through multimodal Kalman filtering (step S3). At the same time, the radar and vision on the inspection vehicle construct a physical circuit breaker mechanism for abnormal working conditions (step S2), ensuring the absolute safety of unmanned operation.

[0160] When facing the complex overhead contact network of a tunnel, the limitations of the traditional fixed perspective are broken. Instead of forcibly driving the robotic arm to extend blindly, it first uses 3D clustering to find spatial gaps, and then uses an improved RRT* algorithm (step S4) with target bias and adaptive step size to calculate an "optimal" trajectory to avoid the pipeline in the virtual configuration space, directing the physical robotic arm to approach the surface of the tunnel segment like a human eye.

[0161] For real-time data analysis and decision-making, the detection device simultaneously collects texture and depth data for spatiotemporal registration (step S5). Finally, by introducing a lightweight convolutional network with a deep feature attention mechanism (step S6), noise such as water stains and reflections is physically suppressed at the feature map level under the limitation of edge computing power, and only the confirmed high-risk cracks or chips are transmitted to the cloud in real time.

[0162] Addressing the technical challenges of 5G / private network signal attenuation, limited hard drive storage space, and the massive data load generated by detection modules within subway tunnels, this invention employs a hierarchical storage and breakpoint resume mechanism based on edge pre-screening in the information storage and processing stage. The principle and process are as follows: Figure 5 As shown.

[0163] (1) End-side annular buffer and data cleaning:

[0164] During the patrol and detection process, the status streams of the inspection vehicle's drive motor and the motors of each joint of the robotic arm, as well as the high-definition video stream and point cloud data, are first written into the main control computer's memory ring buffer. After performing the defect assessment in step S6, normal segment data that is determined to be "defect-free" are directly overwritten and cleaned at the memory level without being written to disk, thus saving excessive invalid storage overhead.

[0165] (2) Local solidification and metadata encapsulation of high-value data:

[0166] For RGB-D fusion tensors identified as "suspected defects" or "high-risk defects," they are retrieved from the circular buffer and stored in the main control computer. Before storage, the data is lightweight compressed according to a preset protocol and tagged with an image containing absolute odometry coordinates p. k Metadata tags for current timestamp, ambient temperature and humidity, and main control computer fault codes are used to generate standardized JSON or XML defect description files.

[0167] (3) Asynchronous pass-through with adaptive bandwidth and breakpoint resumption:

[0168] The main control computer runs a network heartbeat monitoring daemon. When sufficient network bandwidth is detected in the current tunnel section, high-priority defect feature maps and alarm logs are asynchronously transmitted to the cloud. When a train travels to a signal dead zone, causing a network connection interruption, the transmission queue is automatically suspended. After the network handshake is restored, interrupted transmission is resumed based on the local log database, ensuring "zero loss" of high-value defect data packets during the end-to-cloud interaction process.

[0169] (4) Cloud-based digital twin mapping and dimensionality reduction storage:

[0170] After receiving the defect description file with absolute coordinates and point cloud features, the cloud scheduling server no longer saves the original massive video stream. Instead, it directly maps and updates the three-dimensional geometric features of the defect to the BIM (Building Information Modeling) or digital twin system of the entire subway tunnel, realizing full life-cycle management of tunnel defects with data lightweighting as the core.

[0171] Corresponding to the above-mentioned subway tunnel roof inspection method, this invention also provides a subway tunnel roof inspection system based on multi-axis collaboration and multi-sensor fusion. For example... Figure 6 As shown, it mainly includes the following virtual functional modules.

[0172] (1) Initialization and self-test module: used to control the zero-finding and reset of each actuator, and to perform real-time self-test communication with the cloud, and to perform the functions of the aforementioned step S1;

[0173] (2) Active safety and abnormal circuit breaker module: used to acquire forward-looking radar and visual environment perception data in real time during cruise, trigger emergency braking based on physical collision threshold and issue over-authorization alarm, and perform the functions of the aforementioned step S2.

[0174] (3) Dynamic addressing and precise positioning module: used to fuse absolute odometer and tunnel visual features, update the robot's global coordinates, and control the chassis to precisely park after matching the target segment, and perform the functions of the aforementioned step S3;

[0175] (4) Spatial perception and obstacle avoidance planning module: used to extract the three-dimensional obstacle bounding box at the top of the tunnel, and use a multi-objective optimization path planning algorithm based on target offset and adaptive step size (such as the improved RRT* algorithm) to solve the multi-axis cooperative obstacle avoidance trajectory and perform the functions of the aforementioned step S4;

[0176] (5) Multi-source heterogeneous data acquisition and fusion module: used to control the camera and laser profilometer to work synchronously, spatially register 2D texture image and 3D depth point cloud to generate RGB-D tensor data, and perform the functions of the aforementioned step S5;

[0177] (6) Evaluation and hierarchical end-to-cloud collaborative transmission module: used to use a lightweight neural network guided by deep features to infer the RGB-D data, output defect quantification index, and realize end-to-cloud collaboration based on hierarchical caching and network interruption resumption mechanism, and perform the functions of the aforementioned step S6.

[0178] The specific working logic and mathematical calculation process of each virtual module in the above system embodiment have been described in detail in the corresponding method embodiments (steps S1-S6) above, and will not be repeated here.

Claims

1. A method for inspecting the roof slab of a subway tunnel using multi-axis collaboration and multi-sensor fusion, comprising mounting a detection device on the end of a robotic arm, the detection device including a camera and a laser profilometer, and using an inspection vehicle carrying the robotic arm to travel on a track for inspection, the inspection vehicle being equipped with obstacle detection radar and an environmental perception camera; characterized in that, Includes the following steps: S1: The inspection vehicle cruises along the subway tunnel track, simultaneously activating the obstacle detection radar and camera on the inspection vehicle to actively perceive the environmental safety status and trigger the abnormal circuit breaker mechanism when an intruding obstacle is detected. S2: Under normal patrol conditions, the Kalman filter algorithm is used to fuse the absolute odometer of the patrol vehicle with the visual features of the tunnel for joint positioning, update the global coordinates, and park the patrol vehicle after matching the target segment. S3: Obtain environmental point cloud data of the tunnel top through a laser profilometer, extract the three-dimensional bounding box of the obstacle, and calculate the collision-free multi-axis linkage trajectory in three-dimensional space based on the multi-objective optimization path planning algorithm with target offset and adaptive step size, so as to guide the detection device to probe into the obstacle-free cutting gap. S4: The camera and laser profilometer in the detection device work synchronously to acquire 2D texture images and 3D depth point clouds respectively, and spatially register the 2D texture images and 3D depth point clouds to generate fused RGB-D multimodal tensor data. S5: Input the RGB-D multimodal tensor data into the lightweight defect recognition network, output the classification and three-dimensional geometric features of the top plate defects, and upload them to the cloud scheduling server based on the end-cloud collaboration mechanism; The specific process of step S3 is as follows: After the inspection vehicle is parked, the camera and laser profilometer are activated to scan the top of the tunnel; the background point cloud of the tunnel's curved surface is removed using region growing or RANSAC algorithms, and the set of 3D bounding boxes of obstacles is extracted. obs An improved fast expanding random tree algorithm is used for inverse kinematic path planning. The specific process is as follows: (1) Target bias sampling strategy: Generate random nodes q in the configuration space rand When, a bias probability P is introduced. bias Let the mapping attitude of the target detection slot in the configuration space be q. goal The sampling function is defined as: ; Where ξ is a uniformly random number in the interval (0,1). Represents a free-configuration space without collisions; (2) Collision detection and adaptive step size expansion: Let the distance in the current tree be q. rand The nearest node is q near , in q near To q rand Expand to generate a new node q new At that time, an adaptive step size λ based on the obstacle distance is adopted: ; The step size λ is determined by the distance to the nearest obstacle: ; In the formula, λ max To set the maximum safe step size, d obs (q near ) represents the Euclidean distance from the end effector cylinder of the robotic arm to the nearest pipeline calculated at the current node, and k is the smoothing coefficient; as the robotic arm moves away from the obstacle, the step size tends to λ. max This enables rapid movement; when the robotic arm probes into the gaps in the pipeline, the step size λ decreases exponentially, achieving smooth micro-distance insertion. Randomly sample joint angle nodes q in configuration space rand Define a composite cost function C(q) to select the optimal expansion node: ; In the formula, L is the joint angle spatial distance from the robot arm's initial posture to its current posture; dist is the distance from the robot arm's link envelope cylinder to the obstacle boundary box O. obs,i The shortest physical Euclidean distance; E(q) is the energy consumption penalty term, used to suppress the ineffective reciprocating motion of the large inertia joints of the robotic arm; α, β, γ are preset weight coefficients; This indicates the initial joint angle configuration of the robotic arm; This indicates the current evaluation of the robotic arm joint angle configuration; M represents the total number of extracted obstacle 3D bounding boxes. By minimizing the aforementioned cost function, a set of servo commands for each joint axis of the robotic arm that enable interference-free motion of each link in three-dimensional space is calculated, guiding the detection device to precisely penetrate into the gap and reach the effective focal length range from the top plate.

2. The method for inspecting the roof of a subway tunnel using multi-axis collaboration and multi-sensor fusion as described in claim 1, characterized in that, The abnormal circuit breaker mechanism in step S1 is: Real-time acquisition of obstacle point cloud distance d measured by the obstacle detection radar on the inspection vehicle t and the current speed v of the inspection vehicle t Determine if the collision threshold condition is met: D safe To preset a safe distance, a max The maximum braking deceleration is set; if this is met, all moving parts are locked, and the camera on the inspection vehicle captures images of the scene as an over-authorization alarm data packet, which is then uploaded to the cloud dispatch server.

3. The method for inspecting the roof of a subway tunnel using multi-axis collaboration and multi-sensor fusion as described in claim 1, characterized in that, The specific process of step S2 is as follows: Define state variables , where p k v represents the current absolute mileage. k Let be the speed of the inspection vehicle in the kth sampling period; the state transition equation is described as: ; Where A is the state transition matrix; B represents the state variables of the previous time step; B is the control input matrix; µ k For the servo drive commands of the inspection vehicle; W k Represented as system process noise, in Kalman filtering, it represents the inherent uncertainty of the kinematic model caused by slippage or vibration when the inspection vehicle is traveling on the track; Simultaneously, environmental perception cameras are used to capture the joints of tunnel segments, and the standard spacing between the joints is extracted as an observation variable. The mileage coordinates p are then updated in real time using Kalman gain. k When p k The target tunnel segment coordinates are matched to those sent by the cloud scheduling server. When a suspected crack is detected in the initial screening, the inspection vehicle parks at the detection origin.

4. The method for inspecting the roof of a subway tunnel using multi-axis collaboration and multi-sensor fusion as described in claim 1, characterized in that, The step S4, which generates the fused RGB-D multimodal tensor data, specifically includes: By using the pre-calibrated intrinsic parameter matrix of the camera and the extrinsic parameter rotation and translation matrix of the lidar on the inspection vehicle to the camera, the 3D depth point cloud in the world coordinate system is reprojected onto the 2D image pixel plane; through interpolation alignment operations, the depth value is assigned to the corresponding RGB pixel point to generate four-dimensional tensor data.

5. The method for inspecting the roof of a subway tunnel using multi-axis collaboration and multi-sensor fusion as described in claim 4, characterized in that, The specific process for generating four-dimensional tensor data is as follows: After the camera and laser profilometer reach the designated pose, they work synchronously. The camera acquires 2D texture images, and the laser profilometer acquires 3D depth point clouds, i.e., world coordinates. , , ); By using a joint calibration matrix, the 3D point cloud is reprojected onto the 2D image plane, and the transformation relationship is as follows: ; In the formula, K is the intrinsic parameter matrix of the camera, R and T are the rotation matrix and translation vector of the laser radar on the inspection vehicle to the camera, respectively, and is the depth value in the camera coordinate system. Through interpolation and alignment, four-dimensional tensor data containing RGB and Depth channels are generated.

6. The method for inspecting the roof of a subway tunnel using multi-axis collaboration and multi-sensor fusion as described in claim 1, characterized in that, The lightweight defect recognition network in step S5 is a deep feature-guided convolutional neural network, and its processing includes: The depth channel matrix is ​​extracted from the RGB-D multimodal tensor data and processed sequentially by global average pooling, convolutional dimensionality reduction and expansion, and activation function to generate a spatial attention weight matrix. The spatial attention weight matrix is ​​then multiplied element-wise with the RGB texture feature map extracted from the network backbone to physically suppress environmental noise interference caused by water stain reflection. The defect classification, 2D bounding box size, and 3D depth value are output synchronously through the multi-task output head.

7. The method for inspecting the roof of a subway tunnel using multi-axis collaboration and multi-sensor fusion as described in claim 6, characterized in that, The specific process of the lightweight defect identification network includes: (1) Four-channel depthwise separable convolution: The RGB-D four-dimensional tensor generated in step S5 is used as the network input. To control the number of parameters, depthwise separable convolution is used in the feature extraction backbone. Let the size of the input feature map be D. F ×D F ×M, where M is the number of input channels, initially set to 4, and the kernel size is D. k ×D k Its standard computational cost is: ; N represents the number of output channels of the convolutional layer; (2) Attention mechanism guided by deep features: The input deep feature channel matrix F depth The extracted spatial attention weight matrix W is generated after dimensionality reduction and expansion via global average pooling and convolutional layers, and then activated by the sigmoid activation function σ. attn : ; Where W1 and W2 are network learning parameters, and δ is the ReLU activation function. This represents a global average pooling operation; subsequently, the attention weights are compared with the RGB texture feature map F. RGB Feature recalibration is performed by element-wise multiplication: ; This represents the output feature map after spatial attention weight recalibration; To simultaneously predict the 2D contour and 3D depth of defects, the network's loss function is Loss. total Optimized to: ; Among them, L cls For cross-entropy classification loss, L box L is the regression loss for the target bounding box. depth For deep regression loss, Weights for depth loss; If it is deduced that the crack width exceeds 0.3mm or there is a risk of block falling, it is judged as a high-risk defect; the high-risk data message is added to the highest priority queue and transmitted to the cloud dispatch server in real time through the network, so as to realize efficient anomaly early warning from the inspection vehicle to the cloud dispatch server.

8. The method for inspecting the roof of a subway tunnel using multi-axis collaboration and multi-sensor fusion as described in claim 1, characterized in that, The process of uploading to the cloud scheduling server based on the end-to-cloud collaboration mechanism in step S5 includes hierarchical storage and breakpoint resume mechanism: The collected raw data stream is written to the edge-side memory circular buffer in real time; Based on the output of the lightweight defect identification network, the normal segment data that is determined to be defect-free is overwritten and cleaned at the memory level. Data identified as suspected or high-risk defects is compressed and encapsulated with metadata including absolute coordinates and timestamps before being persisted locally to disk; the current network bandwidth is monitored, and when bandwidth is sufficient, the defect data to disk is asynchronously transmitted to the cloud scheduling server, and the interrupted transmission is resumed based on the local logs after the network interruption is restored. The specific process includes: (1) End-side annular buffer and data cleaning: During the patrol and detection process, the status stream, high-definition video stream and point cloud data of the drive motors of the inspection vehicle and the robotic arm are first written into the memory ring buffer of the main control computer. After the defect assessment in step S5 is performed, the normal segment data that is determined to be defect-free is directly overwritten and cleaned at the memory level and is not stored. (2) Local solidification and metadata encapsulation of high-value data: For RGB-D fusion tensors identified as potentially defective or highly defective, they are retrieved from the circular buffer and stored in the main control computer; before local storage, the data is lightweight compressed according to a preset protocol and tagged with an image containing absolute odometry coordinates p. k Metadata tags for current timestamp, ambient temperature and humidity, and main control computer fault codes are used to generate standardized JSON or XML defect description files; (3) Asynchronous pass-through with adaptive bandwidth and breakpoint resumption: When sufficient network bandwidth is detected in the current tunnel section, high-priority defect feature maps and alarm logs are asynchronously transmitted to the cloud. When the train travels to a signal blind spot and the network connection is interrupted, the transmission queue is automatically suspended. After the network handshake is restored, the interrupted transmission is resumed based on the local log database. (4) Cloud-based digital twin mapping and dimensionality reduction storage: After receiving the defect description file with absolute coordinates and point cloud features, the cloud scheduling server no longer saves the original video stream. Instead, it directly maps and updates the three-dimensional geometric features of the defect to the BIM or digital twin system of the entire subway tunnel, realizing full life-cycle management of tunnel defects with data lightweighting as the core.

9. A multi-axis collaborative and multi-sensor fusion subway tunnel roof inspection system, used to implement the multi-axis collaborative and multi-sensor fusion subway tunnel roof inspection method according to any one of claims 1-8, characterized in that, include: The initialization and self-test module is used to control the robotic arm to find and reset zero and perform cloud communication self-tests. The active safety and abnormal circuit breaker module is used to trigger emergency braking and unauthorized alarm based on the perception threshold during cruise. The dynamic addressing and positioning module is used to fuse multi-source features to update global coordinates and perform precise parking; The spatial perception and obstacle avoidance planning module is used to extract three-dimensional obstacle boundaries and use an improved path planning algorithm to solve multi-axis cooperative obstacle avoidance trajectories. A multi-source data acquisition and fusion module is used to synchronously acquire and spatially register RGB-D multimodal tensor data; The evaluation and edge-cloud collaboration module is used to quantify network output defects based on deep feature guidance and to perform hierarchical transmission.

Citation Information

Patent Citations

  • Subway tunnel unmanned aerial vehicle inspection system

    CN120669735A

  • Inspection robot detection system based on image recognition

    CN121061853A