A multimodal collaborative control system and method for an intelligent inspection robot in coal mines
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-30
- Publication Date
- 2026-08-14
AI Technical Summary
现有技术中多模态数据融合方法多采用简单的加权叠加或特征拼接方式,未能充分挖掘不同模态数据之间的互补性和关联性,导致融合后的场景地图存在特征缺失、边界模糊等问题,无法准确重建巷道三维结构模型,最终使得规划出的避障轨迹精度不足,在狭窄巷道或复杂障碍区域容易出现碰撞风险
[0046]本发明提供的一种煤矿井下智能巡检机器人多模态协同控制方法,包括:通过巡检机器人上的RGB-D相机和激光雷达对煤矿井下巷道进行扫描,得到环境图深数据和点云数据,并对所述环境图深数据和点云数据进行时空标定,得到标定后的异构数据;基于所述标定后的异构数据进行特征提取,得到场景特征集合,并对所述场景特征集合进行多源信息对齐,得到统一坐标系下的场景表征;对所述统一坐标系下的场景表征进行多模态数据融合,得到融合场景地图,并根据所述融合场景地图进行通道环境重建,得到巷道结构模型;基于所述巷道结构模型对巡检机器人的路径进行规划,得到避障轨迹,并根据所述避障轨迹调整巡检机器人的运动状态,解决了现有技术中无法准确重建巷道三维结构模型,使得规划出的避障轨迹精度不足,在狭窄巷道或复杂障碍区域容易出现碰撞风险的问题,实现基于高精度巷道结构模型规划的避障轨迹更加贴合实际环境,能够准确识别支护、管线、积水等各类障碍物的空间位置和形态,有效降低了机器人在狭窄巷道和复杂障碍区域的碰撞风险,提升了巡检作业的安全性和可靠性。
Smart Images

Figure CN121946489B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of robotics, and in particular to a multimodal collaborative control system and method for an intelligent underground inspection robot in coal mines. Background Technology
[0002] Coal mine inspection robots, as crucial equipment for ensuring safe production in mines, have been widely applied in the intelligent construction of coal mines. In existing technologies, inspection robots typically employ single sensors or simple multi-sensor combinations for environmental perception. For example, they may rely solely on LiDAR to acquire distance information and construct a two-dimensional grid map, or use RGB cameras alone for visual recognition. These solutions can achieve basic autonomous navigation in highly structured surface environments, but in the complex working conditions of underground coal mines, the limited information dimensionality of single-modal data makes it difficult to comprehensively depict the spatial characteristics of the tunnels. Some technologies attempt to fuse RGB-D camera and LiDAR data, but the lack of a precise spatiotemporal calibration mechanism for heterogeneous sensor data leads to discrepancies in timestamps and coordinate systems between different modal data. This prevents accurate alignment of depth information, color texture, and point cloud data, thus affecting the accuracy of environmental reconstruction and the reliability of path planning.
[0003] In the unique environment of underground coal mines, the tunnel structures are complex and varied, containing various obstacles such as support equipment, pipelines, and accumulated water. Furthermore, poor lighting conditions and high dust concentrations place higher demands on the environmental perception capabilities of robots. Existing multimodal data fusion methods often employ simple weighted superposition or feature stitching, failing to fully exploit the complementarity and correlation between different modalities. This results in fused scene maps with missing features and blurred boundaries, making it impossible to accurately reconstruct the 3D structural model of the tunnel. Ultimately, this leads to insufficient accuracy in the planned obstacle avoidance trajectories, increasing the risk of collisions in narrow tunnels or areas with complex obstacles. Summary of the Invention
[0004] The purpose of this invention is to at least partially solve one of the technical problems existing in the prior art.
[0005] To achieve the above objectives, the present invention provides a multimodal collaborative control system for an intelligent underground inspection robot in coal mines, comprising:
[0006] The scanning module is used to scan the underground roadways of the coal mine using the RGB-D camera and lidar on the inspection robot to obtain environmental depth data and point cloud data, and to perform spatiotemporal calibration on the environmental depth data and point cloud data to obtain calibrated heterogeneous data.
[0007] The alignment module is used to extract features based on the calibrated heterogeneous data to obtain a scene feature set, and to align the scene feature set with multi-source information to obtain a scene representation under a unified coordinate system.
[0008] The fusion module is used to perform multimodal data fusion on the scene representation under the unified coordinate system to obtain a fused scene map, and to reconstruct the channel environment based on the fused scene map to obtain a tunnel structure model.
[0009] The planning module is used to plan the path of the inspection robot based on the alleyway structure model, obtain the obstacle avoidance trajectory, and adjust the motion state of the inspection robot according to the obstacle avoidance trajectory.
[0010] This invention also provides a multimodal collaborative control method for an intelligent inspection robot in underground coal mines, comprising the following steps:
[0011] The inspection robot scans the underground roadways of the coal mine using an RGB-D camera and a lidar to obtain environmental depth data and point cloud data. The environmental depth data and point cloud data are then spatiotemporally calibrated to obtain calibrated heterogeneous data.
[0012] Feature extraction is performed based on the calibrated heterogeneous data to obtain a scene feature set, and the scene feature set is aligned with multi-source information to obtain a scene representation under a unified coordinate system.
[0013] Multimodal data fusion is performed on the scene representation under the unified coordinate system to obtain a fused scene map, and the channel environment is reconstructed based on the fused scene map to obtain a tunnel structure model;
[0014] The path of the inspection robot is planned based on the tunnel structure model to obtain the obstacle avoidance trajectory, and the motion state of the inspection robot is adjusted according to the obstacle avoidance trajectory.
[0015] Furthermore, the inspection robot uses an RGB-D camera and LiDAR to scan the underground roadways of the coal mine, obtaining environmental depth data and point cloud data, including:
[0016] The inspection robot continuously captures multiple frames of images from underground coal mine roadways using an RGB-D camera, obtaining original color image sequences and original depth image sequences. The original color image sequences are then subjected to grayscale histogram equalization to obtain enhanced color image sequences.
[0017] Pixel-level registration is performed between the enhanced color image sequence and the original depth image sequence to obtain image-depth registration data. Neighborhood pixel interpolation is then performed to fill the invalid depth regions in the image-depth registration data to obtain environmental image-depth data.
[0018] The laser radar is used to perform rotational scanning and ranging of the underground roadway in the coal mine to obtain the original laser echo signal. The original laser echo signal is then subjected to intensity threshold filtering to remove dust scattering interference, resulting in point cloud data.
[0019] Furthermore, the spatiotemporal calibration of the environmental map-depth data and point cloud data to obtain calibrated heterogeneous data includes:
[0020] The image frames in the environmental depth data are timestamped to obtain timestamped image data. At the same time, each point cloud cluster in the point cloud data is timestamped to obtain timestamped point cloud data.
[0021] Based on timestamped image data and timestamped point cloud data, the data is aligned according to the chronological order to obtain time-synchronized image depth point cloud data;
[0022] Image feature points and point cloud feature points are extracted from the time-synchronized image-depth-point cloud data to obtain image feature point sets and point cloud feature point sets;
[0023] Calculate the Euclidean distance between the image feature point set and the point cloud feature point set, find the closest feature point pair, and calculate the coordinate transformation parameters based on the feature point pair to obtain the transformation parameters from the point cloud coordinate system to the image coordinate system. Then, perform coordinate transformation on the time-synchronized image-depth point cloud data based on the transformation parameters to obtain calibrated heterogeneous data.
[0024] Furthermore, the feature extraction based on the calibrated heterogeneous data to obtain a scene feature set includes:
[0025] Image edge detection is performed on the environmental depth data in the calibrated heterogeneous data to obtain an edge contour image. Connectivity analysis is then performed on the edge contour image to connect adjacent edge pixels into continuous line segments, resulting in a set of line segments including the tunnel wall contour line and the equipment edge line.
[0026] Based on the set of line segments, the point cloud data in the calibrated heterogeneous data is filtered to obtain the local point cloud corresponding to the line segment, and the local point cloud is fitted with a plane to obtain planar features including the tunnel wall plane and the equipment installation plane.
[0027] The planar features are subjected to geometric parameter calculation to obtain planar geometric parameters. These planar geometric parameters are then classified and grouped according to their area and normal vector direction to obtain a set of scene features.
[0028] Furthermore, the step of aligning the scene feature set with multi-source information to obtain a scene representation in a unified coordinate system includes:
[0029] The endpoint coordinates of the line segment set in the scene feature set are extracted to obtain the line segment endpoint coordinate sequence, and the boundary point coordinates of the planar features in the scene feature set are extracted to obtain the planar boundary point coordinate sequence.
[0030] Spatial position matching is performed based on the coordinate sequence of the line segment endpoints and the coordinate sequence of the plane boundary points to obtain a set of feature matching point pairs. Based on the set of feature matching point pairs, the rotation matrix and translation vector from the image coordinate system to the point cloud coordinate system are calculated to obtain the coordinate transformation parameters.
[0031] Based on the coordinate transformation parameters, the coordinates of each line segment in the line segment set are transformed to obtain the transformed line segment set. The transformed line segment set is then associated and combined with the planar features according to their spatial positional relationship to obtain a scene representation under a unified coordinate system.
[0032] Furthermore, the step of performing multimodal data fusion on the scene representation under the unified coordinate system to obtain a fused scene map includes:
[0033] The transformed line segment set in the scene representation under the unified coordinate system is subjected to pixel rasterization to obtain a rasterized line segment map, and the planar features in the scene representation are superimposed with texture information to obtain textured planar data.
[0034] Spatial position calibration is performed on the textured plane data based on the rasterized line segment map to obtain a calibrated textured plane. Then, the calibrated textured plane and the rasterized line segment map are overlaid and synthesized to obtain a fused scene map.
[0035] Furthermore, the process of reconstructing the passageway environment based on the fused scene map to obtain the alleyway structure model includes:
[0036] The topological nodes of the rasterized line segment graph in the fused scene map are extracted to obtain the coordinates of the alleyway nodes, and the connectivity of the alleyway node coordinates is sorted out to obtain the node connectivity table.
[0037] Based on the node connectivity table, the textured planar data in the fused scene map is subjected to three-dimensional stretching to obtain the alleyway three-dimensional surface block, and the adjacent surface of the alleyway three-dimensional surface block is calibrated to obtain the alleyway three-dimensional shell.
[0038] By combining the node connectivity list with the three-dimensional shell of the tunnel, a topological relationship binding is performed to obtain the tunnel topological model. Then, dimension annotations and spatial constraint information are added to the tunnel topological model to obtain the tunnel structure model.
[0039] Furthermore, the step of planning the path of the inspection robot based on the alleyway structure model to obtain an obstacle avoidance trajectory, and adjusting the motion state of the inspection robot according to the obstacle avoidance trajectory, includes:
[0040] The inner wall surface of the three-dimensional shell of the tunnel in the tunnel structure model is extracted to obtain the inner wall boundary surface of the tunnel. Based on the inner wall boundary surface of the tunnel, a safe distance offset and contraction is performed in the direction of the tunnel center to obtain the passable boundary.
[0041] The alleyway structure model is divided into grids based on the passable boundary line to obtain a passable grid map. Then, a grid-by-grid search is performed in the passable grid map from the current position of the inspection robot to the target inspection position to obtain an obstacle avoidance trajectory point sequence. The obstacle avoidance trajectory point sequence is used as the obstacle avoidance trajectory.
[0042] Vector difference calculation is performed on adjacent trajectory points in the obstacle avoidance trajectory point sequence to obtain the travel direction angle and travel distance of each trajectory segment. Based on the travel direction angle and travel distance, the motion state parameters of the robot are generated, and the motion state of the inspection robot is controlled based on the motion state parameters.
[0043] Furthermore, the step of dividing the alleyway structure model into grids based on the traversable boundary lines to obtain a traversable grid map includes:
[0044] The passable boundary is horizontally projected along the bottom surface of the alley to obtain a two-dimensional passable area outline. The grid side length is determined based on the dimension annotation information in the alley structure model. The area enclosed by the two-dimensional passable area outline is divided into equidistant grids according to the grid side length to obtain a grid cell array. The grid cell array includes the row and column number and center point coordinates of each grid cell.
[0045] Based on the three-dimensional shell of the tunnel structure model, the installation location of the underground equipment is extracted to obtain the obstacle occupancy area. The position inclusion of each grid cell in the grid cell array with the two-dimensional passage area outline and the obstacle occupancy area is determined. Grid cells located inside the two-dimensional passage area outline and not overlapping with the obstacle occupancy area are marked as passable grid cells, and the remaining grid cells are marked as obstacle grid cells, thus obtaining a passage grid map.
[0046] This invention provides a multimodal collaborative control method for an intelligent underground coal mine inspection robot, comprising: scanning underground coal mine roadways using an RGB-D camera and a lidar on the inspection robot to obtain environmental depth data and point cloud data; performing spatiotemporal calibration on the environmental depth data and point cloud data to obtain calibrated heterogeneous data; extracting features based on the calibrated heterogeneous data to obtain a scene feature set; aligning the scene feature set with multi-source information to obtain a scene representation in a unified coordinate system; fusing multimodal data on the scene representation in the unified coordinate system to obtain a fused scene map; and reconstructing the tunnel environment based on the fused scene map. The process involves obtaining a tunnel structure model, planning the path of the inspection robot based on this model to obtain an obstacle avoidance trajectory, and adjusting the robot's motion state according to the trajectory. This solves the problem in existing technologies where the three-dimensional tunnel structure model cannot be accurately reconstructed, resulting in insufficient accuracy of the planned obstacle avoidance trajectory and a risk of collision in narrow tunnels or complex obstacle areas. The new method achieves obstacle avoidance trajectories planned based on a high-precision tunnel structure model that better fit the actual environment, accurately identifying the spatial location and shape of various obstacles such as supports, pipelines, and water accumulation. This effectively reduces the risk of collisions for the robot in narrow tunnels and complex obstacle areas, improving the safety and reliability of inspection operations. Attached Figure Description
[0047] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0048] Figure 1 This is a schematic diagram of a multimodal collaborative control system for an intelligent underground inspection robot in a coal mine, according to one embodiment of the present invention.
[0049] Figure 2 This is a schematic diagram of a multimodal collaborative control method for an intelligent underground inspection robot in a coal mine according to an embodiment of the present invention;
[0050] The objectives, features, and advantages of this invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0051] The embodiments of the present invention are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention. The step numbers in the following embodiments are set only for ease of explanation, and there is no limitation on the order between the steps. The execution order of each step in the embodiments can be adaptively adjusted according to the understanding of those skilled in the art.
[0052] The following describes in detail, with reference to the accompanying drawings, a multimodal collaborative control device for an intelligent underground inspection robot in a coal mine according to an embodiment of the present invention. First, the multimodal collaborative control device for an intelligent underground inspection robot in a coal mine according to an embodiment of the present invention will be described with reference to the accompanying drawings.
[0053] Figure 2 This invention provides a multimodal collaborative control method for an intelligent underground inspection robot in a coal mine, comprising the following steps:
[0054] Step S1: The underground roadway of the coal mine is scanned by the RGB-D camera and lidar on the inspection robot to obtain environmental depth data and point cloud data. The environmental depth data and point cloud data are then spatiotemporally calibrated to obtain calibrated heterogeneous data.
[0055] Specifically, after the inspection robot starts, the RGB-D camera installed at the front of the robot body begins to work. It emits coded light spots into the tunnel space through an infrared structured light projection module. The receiving module captures the reflected light spot pattern and calculates the depth value. At the same time, the RGB sensor collects color images of the corresponding locations. The two are combined to form environmental depth data containing pixel coordinates, color values, and depth values. Meanwhile, the LiDAR emits laser pulses in a rotating scanning manner, calculates the distance based on the time difference of the beam return, and acquires tens of thousands of spatial coordinate points in each scanning cycle. These coordinate points constitute point cloud data describing the tunnel contour. Because the two sensors work on different principles, the RGB-D camera typically has a sampling frequency of 30Hz while the LiDAR is 10Hz. The data generation times are not synchronized, and their independent coordinate origins and axis definitions also differ. Therefore, spatiotemporal calibration processing is required. In practice, a calibration board is first placed in the tunnel. The surface of the calibration board has a feature point array. The robot scans the calibration board from multiple angles. The RGB-D camera identifies the pixel positions of the feature points and calculates the depth, while the LiDAR records the three-dimensional coordinates of the feature points. By establishing a mapping relationship between corresponding feature points in two sets of data, the rotation matrix and translation vector from the RGB-D camera coordinate system to the LiDAR coordinate system are calculated. This transformation parameter is the spatial calibration result. For time calibration, hardware trigger signals are used to make the two sensors collect data at the same time, or timestamp interpolation methods are used to align data from different times to a unified time reference. After completing these operations, the heterogeneous data achieves accurate spatiotemporal correspondence.
[0056] Step S2: Based on the calibrated heterogeneous data, feature extraction is performed to obtain a scene feature set, and the scene feature set is aligned with multi-source information to obtain a scene representation under a unified coordinate system.
[0057] Specifically, after obtaining the calibrated heterogeneous data, features are first extracted from the environmental depth data. Operationally, the SIFT algorithm is used to detect key points in the image. These key points often appear at locations with drastic grayscale changes, such as the edges of tunnel supports and pipeline corners. Each key point generates a 128-dimensional descriptive vector recording the gradient distribution of surrounding pixels. The point cloud data is processed differently. A normal vector estimation method is used to calculate the curvature characteristics of each point. Points with high curvature are usually located at tunnel bends or on obstacle surfaces. These points, along with their normal vectors and curvature values, are combined to form point cloud features. Although the two types of features are extracted, their representations differ significantly. Image features use pixel coordinates and grayscale descriptors, while point cloud features use three-dimensional spatial coordinates with geometric attributes, making them incompatible. Next, multi-source information alignment is performed. The key is to find corresponding relationships. For example, a vertical support in a tunnel appears as a straight line segment with a sudden change in grayscale in an RGB-D image. The corresponding point cloud data will contain a series of points with high curvature arranged in a columnar shape. Matching these two sets of feature points establishes a connection. The matching process uses the rotation matrix and translation vector obtained from the previous spatiotemporal calibration to transform the pixel coordinates of image feature points into three-dimensional coordinates by combining them with depth values. This transformation parameter is then projected onto the LiDAR coordinate system, so that both types of features are expressed in the LiDAR coordinate system. In the resulting scene feature set, the image features and point cloud features of the same object point to the same spatial location; this is the scene representation under a unified coordinate system.
[0058] Step S3: Perform multimodal data fusion on the scene representation under the unified coordinate system to obtain a fused scene map, and reconstruct the channel environment based on the fused scene map to obtain a tunnel structure model.
[0059] Specifically, after obtaining the scene representation, the multimodal data fusion step is mainly achieved through a weighted fusion strategy. In practice, different weights are assigned to image features and point cloud features, and these weights are dynamically adjusted based on the current environmental conditions. For example, if the lighting in a section of the alley is weak, and the image captured by the RGB-D camera is blurry, the weight of the point cloud features is increased to 0.7, while the weight of the image features is decreased to 0.3. Conversely, in areas with high dust concentration, the laser penetration ability decreases, so the weight of the image features is increased. During the fusion calculation, the image feature vector and the point cloud feature vector at the same spatial location are added according to their weights. Assuming the image feature at a certain pillar location is a 128-dimensional vector A, and the point cloud feature is a 33-dimensional vector B (including xyz coordinates, normal vector, curvature, etc.), B is first expanded to 128 dimensions for easier computation, and then the fused feature is calculated using a formula like 0.5A + 0.5B. This process is repeated for all spatial locations, and the resulting dataset is the fused scene map. Each grid point in this map contains both color and texture information as well as precise distance data. To reconstruct the alleyway environment using a fused scene map, the process involves first setting the spatial resolution and dividing the alleyway space into a 10cm × 10cm × 10cm voxel grid. Feature points from the fused scene map are then iterated through, filling the corresponding voxels with data. If multiple feature points exist within a voxel, the average value is taken. After filling, a moving cube algorithm scans the voxel grid. This algorithm checks every eight adjacent voxels forming a cube and generates triangular facets on the cube's surface based on whether the voxels are occupied. For example, voxel configurations at the junction of the alleyway roof and sidewalls will form angled triangular meshes. Connecting all the triangular facets yields the alleyway structure model. Texture information from the fused scene map can also be applied to the model's surface. This results in a reconstructed model with accurate geometry, where even rust marks on the support surfaces are clearly visible.
[0060] Step S4: Plan the path of the inspection robot based on the alleyway structure model to obtain the obstacle avoidance trajectory, and adjust the motion state of the inspection robot according to the obstacle avoidance trajectory.
[0061] Specifically, when planning the path using the alleyway structure model, the robot's current coordinates (5.25 meters, 2.15 meters) are read first. These are then converted to grid numbers. The horizontal coordinate (5.25 meters) divided by the side length (50 centimeters) equals 10.5, rounded down to the 10th column. The vertical coordinate (2.15 meters) divided by 50 centimeters equals 4.3, rounded down to the 4th row. The current position is in grid number 4-10. The target inspection location is at 40 meters, checking the electrical control cabinet. Coordinates 40 meters and 2 meters correspond to grid number 80-4. The search begins from the starting point 4-10, checking the surrounding 8 neighboring grids. The passable grid map is checked to see which grids are passable. Assuming a pipeline occupies grid 3-11 (marked 0), it's an obstacle grid; the remaining 7 grids (marked 1) are passable. The cost of these 7 grids is calculated using the formula g + h, where g is the actual distance traveled (e.g., moving 0.5 meters to 5-10), and h is the estimated distance to the destination (approximately 37.25 meters as a straight line). The total cost f equals 37.75 meters. After calculating the values of all neighboring cells, they are sorted by their f values. The cell with the smallest f value is then used to expand the search area. This process is repeated to search outwards in cycles.
[0062] Finding the target number 80-4 indicates a path has been found. Tracing back from the endpoint 80-4, we see it originated from 79-4, which in turn originated from 78-4, and so on back to the starting point 4-10. The coordinates of the grid center points along the path are extracted: the center of the starting point 4-10 is at 2 meters and 2.25 meters; the next possible center is at 2 meters and 2.75 meters; and the center of the endpoint 80-4 is at 2 meters and 40 meters. These coordinates are strung together to form a sequence of obstacle avoidance trajectory points. The motion parameters between adjacent points are calculated. Taking the first point at 2 meters and 2.25 meters and the second point at 2 meters and 2.75 meters, with a horizontal coordinate difference of 0 and a vertical coordinate difference of 0.5 meters, the direction angle is calculated using atan2(0.5, 0), where 90 degrees represents true north and the distance is 0.5 meters, resulting in a set of parameters. The second point is 2 meters and 2.75 meters, and the third point is assumed to be 2.5 meters and 3.25 meters, with a difference of 0.5 meters between the horizontal and vertical coordinates. atan2(0.5, 0.5) is used to calculate the 45-degree northeast direction, with a distance of 0.707 meters, which is another set of parameters. After calculating all adjacent points, dozens of sets of travel direction angle and travel distance data are obtained and sent to the control module.
[0063] The control module reads the first set of parameters and moves 0.5 meters at a 90-degree angle, causing the left and right wheel motors to rotate at the same speed of 120 revolutions per minute. With a wheel diameter of 20 centimeters, each wheel rotates 0.8 times, covering approximately 50 centimeters. After the encoder counts to 0.5 meters, it reads the second set of parameters. A 45-degree turn to the right is required to change the direction angle to 45 degrees. The right wheel slows to 60 revolutions per minute, while the left wheel maintains 120 revolutions per minute, achieving differential steering and rotating the vehicle 45 degrees. After the steering is complete, both wheels move forward at the same speed for 0.707 meters. Repeating the execution of subsequent parameters, the robot follows the obstacle avoidance trajectory from the starting point to the end point, completing the inspection.
[0064] In a specific embodiment, the underground roadways of the coal mine are scanned using an RGB-D camera and a LiDAR on the inspection robot to obtain environmental depth data and point cloud data, including:
[0065] The inspection robot continuously captures multiple frames of images from underground coal mine roadways using an RGB-D camera, obtaining original color image sequences and original depth image sequences. The original color image sequences are then subjected to grayscale histogram equalization to obtain enhanced color image sequences.
[0066] Pixel-level registration is performed between the enhanced color image sequence and the original depth image sequence to obtain image-depth registration data. Neighborhood pixel interpolation is then performed to fill the invalid depth regions in the image-depth registration data to obtain environmental image-depth data.
[0067] The laser radar is used to perform rotational scanning and ranging of the underground roadway in the coal mine to obtain the original laser echo signal. The original laser echo signal is then subjected to intensity threshold filtering to remove dust scattering interference, resulting in point cloud data.
[0068] Specifically, after the robot enters the tunnel, the RGB-D camera starts working at a frequency of 30Hz, collecting 30 frames of data per second. The RGB sensor captures color images to form the original color image sequence, while the infrared emitter projects coded structured light. The receiver calculates the distance based on the light spot deformation to generate the original depth image sequence. Underground lighting conditions are very poor, resulting in numerous dark areas in the captured color images. Histogram analysis reveals that pixel grayscale values are concentrated in the low-brightness range of 0 to 80, with almost no distribution in the high-brightness range. Therefore, grayscale histogram equalization is needed. Operationally, the grayscale distribution of each frame in the original color image sequence is first analyzed to establish a cumulative distribution function. This function is then used to perform a mapping transformation, redistributing the pixels concentrated in low brightness across the entire grayscale range of 0 to 255. A pixel with a grayscale value of 40 might be mapped to 120, thus revealing details in the dark areas. After processing all frames, an enhanced color image sequence is obtained.
[0069] After obtaining the enhanced color image sequence and the original depth image sequence, pixel-level registration is required. This is because the RGB and infrared sensors are positioned differently within the camera, causing pixel coordinates to misalign when photographing the same object. The registration process utilizes the camera's factory-calibrated intrinsic parameter matrix. This matrix records parameters such as the focal length and principal point position of the two sensors; for example, the RGB sensor's focal length is 525 pixels, and the infrared sensor's focal length is 580 pixels. Specifically, during registration, a depth image frame is extracted from the original depth image sequence, and the depth value of a specific pixel is read. This pixel is then converted into 3D spatial coordinates using the infrared sensor's focal length and principal point parameters. Finally, the 3D coordinates are projected back onto the color image plane using the RGB sensor's focal length and principal point parameters to locate the corresponding color pixel position. This process aligns the depth value with the pixel in the color image. All pixels are processed in this way to generate image-depth registration data. However, some areas in the registration data may have invalid depth values because those areas have surfaces that are too smooth or are beyond the measurement range, preventing the infrared structured light from reflecting back.
[0070] The handling of invalid depth regions employs a neighborhood pixel interpolation filling method. The depth-map registration data is scanned to identify invalid pixels with a depth value of 0. Then, a 3x3 neighborhood window is taken around each pixel, and the depth values of the valid pixels within the window are counted. For example, if an invalid point is located at the edge of a roadway support, and 5 out of its 8 neighboring pixels are valid with depths of 2.3 meters, 2.35 meters, 2.28 meters, 2.32 meters, and 2.31 meters respectively, the average of these 5 values (approximately 2.31 meters) is used to fill the invalid point's location. If the neighborhood window contains fewer than 3 valid pixels, it indicates a large invalid area, so the window is expanded to 5x5, and the search continues until enough valid neighbors are found. This interpolation filling is performed on every frame of the entire sequence, ultimately yielding complete environmental depth-map data.
[0071] The lidar uses a mechanical rotating scanning method. A motor drives the laser and detector to rotate horizontally at 10Hz for one revolution, emitting a pulse every 0.25 degrees during the rotation, for a total of 1440 pulses per revolution. After each emission, a timer starts counting, waiting for the echo signal to return. The time difference between the return time and the emission time is calculated. This time difference is multiplied by the speed of light and divided by 2 to calculate the distance. The speed of light is assumed to be 300,000 kilometers per second. The received raw laser echo signal contains a significant amount of clutter caused by dust scattering, and the intensity of this clutter signal is noticeably weaker. The solution is to set an intensity threshold. The measured dust scattering echo intensity in the mine is generally below 15%, while the normal echo intensity on the tunnel wall and equipment surface is above 40%. Therefore, the threshold is set at 20%. The intensity parameters of the original laser echo signal are scanned. Those below 20% are discarded, while those above 20% are retained. The distance and angle values corresponding to these retained echo signals are combined to form three-dimensional coordinate points. The distance multiplied by the cosine of the angle gives the horizontal coordinate, and the distance multiplied by the sine of the angle gives the vertical coordinate. The radar installation height is the vertical coordinate. All points are summed up to obtain the point cloud data after removing dust interference.
[0072] In a specific embodiment, the spatiotemporal calibration of the environmental map depth data and point cloud data to obtain calibrated heterogeneous data includes:
[0073] The image frames in the environmental depth data are timestamped to obtain timestamped image data. At the same time, each point cloud cluster in the point cloud data is timestamped to obtain timestamped point cloud data.
[0074] Based on timestamped image data and timestamped point cloud data, the data is aligned according to the chronological order to obtain time-synchronized image depth point cloud data;
[0075] Image feature points and point cloud feature points are extracted from the time-synchronized image-depth-point cloud data to obtain image feature point sets and point cloud feature point sets;
[0076] Calculate the Euclidean distance between the image feature point set and the point cloud feature point set, find the closest feature point pair, and calculate the coordinate transformation parameters based on the feature point pair to obtain the transformation parameters from the point cloud coordinate system to the image coordinate system. Then, perform coordinate transformation on the time-synchronized image-depth point cloud data based on the transformation parameters to obtain calibrated heterogeneous data.
[0077] Specifically, during the environmental depth data acquisition process, the RGB-D camera's processor records the system clock value at the instant each image frame is generated. This value is accurate to the millisecond level; for example, the generation time of a certain image frame is recorded as 14:32:05.127. Processing the entire sequence yields timestamped image data. The LiDAR operates slightly differently. Each rotation collects all ranging points and packages them into a point cloud cluster. When this cluster is generated, a timestamp of 14:32:05.234 is recorded. Continuous scanning generates multiple point cloud clusters, each with its own timestamp, forming timestamped point cloud data. After obtaining these two sets of data, they are sorted and aligned chronologically. However, because the RGB-D camera operates at 30Hz, producing one frame every 33 milliseconds, while the LiDAR rotates at 10Hz, producing one point cloud cluster every 100 milliseconds, this frequency mismatch makes it difficult for the timestamps to perfectly overlap. During the alignment operation, the system finds the pair with the smallest time difference. For example, if the timestamp of an image frame is 14:32:05.127 and the timestamp of the nearest point cloud cluster is 14:32:05.134, which is only 7 milliseconds apart, the two data sets are paired together. After all the data is paired in this way, the time-synchronized image depth point cloud data is obtained.
[0078] Next, features are extracted from the time-synchronized image depth point cloud data. For the image, the Harris corner detection algorithm scans each frame. The algorithm calculates the grayscale change rate of the area surrounding each pixel; points with large changes in both horizontal and vertical dimensions are corners. Corners are frequently detected at locations like support corners and pipeline joints in tunnels. All corner coordinates are collected to form an image feature point set. Point cloud data is processed differently. Each point cloud cluster is traversed, and the normal vector and curvature of each point are calculated. Points with curvature exceeding a set threshold of 0.1 are considered to have significant features, and these points constitute the point cloud feature point set. Once the two feature point sets are obtained, matching begins. Each point in the image feature point set is traversed, and its 3D coordinates are compared with all points in the point cloud feature point set. Euclidean distance is calculated to determine which point cloud feature point is closest. For example, if the image feature point coordinates are 2.5 meters, 1.2 meters, and 0.8 meters, and a point cloud feature point has coordinates of 2.48 meters, 1.23 meters, and 0.82 meters, the calculated distance is only 0.05 meters, and these two points are paired. After finding matching points for all image feature points, the coordinate transformation parameters are calculated using these feature point pairs. This is done by establishing a system of equations to solve for the rotation matrix and translation vector. The rotation matrix adjusts the axial orientation of the point cloud coordinate system, and the translation vector adjusts the origin position. The resulting parameters are the transformation parameters from the point cloud coordinate system to the image coordinate system. Finally, these parameters are used to process the time-synchronized image-depth point cloud data. The coordinates of each point in the point cloud data are rotated using the rotation matrix and then combined with the translation vector. The transformed point cloud data and image data are now expressed in the same coordinate system; this is the calibrated heterogeneous data.
[0079] In a specific embodiment, the step of extracting features based on the calibrated heterogeneous data to obtain a scene feature set includes:
[0080] Image edge detection is performed on the environmental depth data in the calibrated heterogeneous data to obtain an edge contour image. Connectivity analysis is then performed on the edge contour image to connect adjacent edge pixels into continuous line segments, resulting in a set of line segments including the tunnel wall contour line and the equipment edge line.
[0081] Based on the set of line segments, the point cloud data in the calibrated heterogeneous data is filtered to obtain the local point cloud corresponding to the line segment, and the local point cloud is fitted with a plane to obtain planar features including the tunnel wall plane and the equipment installation plane.
[0082] The planar features are subjected to geometric parameter calculation to obtain planar geometric parameters. These planar geometric parameters are then classified and grouped according to their area and normal vector direction to obtain a set of scene features.
[0083] Specifically, once the calibrated heterogeneous data is obtained, the environmental depth data is processed for edge detection. The Canny operator is used to scan the image; this operator calculates the gradient value of each pixel. Areas with large gradients and drastic grayscale changes indicate edge locations. The detection process consists of two steps: first, a Gaussian filter is used to smooth the image and remove noise; then, the gradient magnitude and direction are calculated. A dual threshold is set: pixels with gradient values greater than 150 are marked as strong edges, those between 50 and 150 are marked as weak edges, and those below 50 are discarded. This extracts the edge contour image. The edge pixels in the contour image are relatively scattered, so connected component analysis is needed to connect them. During the operation, starting from a certain edge pixel, its eight neighboring pixels in each direction are checked to see which are also edge points. Once found, the search continues outward until no new edge pixels are found. This circle of connected edge pixels forms a continuous line segment. For example, the edge of the left wall of the tunnel is represented as a vertical edge band in the image. Connected component analysis will string the pixels in this edge band together into a line segment and record it as the tunnel wall outline. The same method is used to process the edge of the equipment surface to obtain the equipment edge line. All line segments are collected to form a line segment set.
[0084] Once the set of line segments is obtained, the point cloud data is filtered. A specific line segment is extracted from the set, and its pixel coordinates in the image are known. By combining this with the depth values of these pixel positions in the environmental depth data, the corresponding 3D spatial range of the line segment can be calculated. Then, the point cloud data is traversed to see which points' coordinates fall within this 3D spatial range. For example, if the spatial range corresponding to the tunnel wall outline is 0 to 0.5 meters horizontally, 2 to 6 meters vertically, and 0 to 3 meters vertically, points in the point cloud with coordinates of 0.3 meters, 4.2 meters, and 1.5 meters fall within this range and are selected. All line segments are processed in this way, and the point clouds from each line segment form a local point cloud. After obtaining the local point cloud, a plane fitting is performed. The method is to randomly select 3 points from the local point cloud to determine an initial plane, calculate the distance from other points to this plane, and consider points with a distance less than 2 centimeters as being on the plane. The number of points that meet this condition is counted. Repeat this process, selecting three different points to try different planes, and finally select the plane containing the most points as the fitting result. If the local point cloud of the tunnel wall is fitted to a vertical plane, it is recorded as the tunnel wall plane. If the local point cloud of the equipment surface is fitted to a horizontal or inclined plane, it is recorded as the equipment installation plane. These fitted planes are collectively referred to as plane features.
[0085] After the planar features are obtained, geometric parameters need to be calculated. Each plane is described by the equation ax + by + cz = d, where a, b, and c form the normal vector representing the plane's orientation, and d is the distance from the origin to the plane. The plane area also needs to be calculated by counting the number of point cloud points on the fitted plane. The area occupied by each point is estimated to be approximately 4 square centimeters based on the point cloud density. Multiplying the number of points by 4 gives the plane area. For example, a tunnel wall containing 5000 points has an area of approximately 20 square meters. The normal vector and area data combined constitute the planar geometric parameters. Next, the planar geometric parameters are categorized and organized. First, they are sorted by area size: areas greater than 10 square meters are grouped into large planes, typically tunnel walls; areas between 1 and 10 square meters are grouped into medium planes, possibly equipment casings; and areas less than 1 square meter are grouped into small planes, such as pipeline supports. Within each group, they are further subdivided by the direction of the normal vector: vertically upward normals represent horizontal planes, vertically forward normals represent longitudinal walls, and vertically left and right normals represent transverse walls. All planar features grouped and summarized together yield the scene feature set.
[0086] In a specific embodiment, aligning the scene feature set with multi-source information to obtain a scene representation in a unified coordinate system includes:
[0087] The endpoint coordinates of the line segment set in the scene feature set are extracted to obtain the line segment endpoint coordinate sequence, and the boundary point coordinates of the planar features in the scene feature set are extracted to obtain the planar boundary point coordinate sequence.
[0088] Spatial position matching is performed based on the coordinate sequence of the line segment endpoints and the coordinate sequence of the plane boundary points to obtain a set of feature matching point pairs. Based on the set of feature matching point pairs, the rotation matrix and translation vector from the image coordinate system to the point cloud coordinate system are calculated to obtain the coordinate transformation parameters.
[0089] Based on the coordinate transformation parameters, the coordinates of each line segment in the line segment set are transformed to obtain the transformed line segment set. The transformed line segment set is then associated and combined with the planar features according to their spatial positional relationship to obtain a scene representation under a unified coordinate system.
[0090] Specifically, the line segment set in the scene feature set is processed. Each line segment has a start and end point. The line segment set is scanned to record the coordinates of all endpoints. For example, the start coordinates of a certain tunnel wall outline are 2.1 meters, 3.5 meters, and 0.8 meters in the image coordinate system, and the end coordinates are 2.1 meters, 3.5 meters, and 2.6 meters. Arranging all endpoint coordinates in order yields the line segment endpoint coordinate sequence. The processing method for planar features is slightly more complex. Each planar feature is fitted from a point cloud. The coordinates of these point cloud points at the edge of the plane are the boundary points. The point cloud in the planar features is traversed to find the points located at the boundary. The judgment method is to see if the point cloud density suddenly decreases within a certain range around a point. A decrease in density indicates that the boundary has been reached. Assuming that the tunnel wall plane contains 5000 points, there may be 200 points at the boundary. The coordinates of these 200 points are extracted to form the planar boundary point coordinate sequence. These coordinates are all expressed in the point cloud coordinate system.
[0091] Once the line segment endpoint coordinate sequence and the plane boundary point coordinate sequence are obtained, matching begins. Since the line segments are derived from image data using an image coordinate system, and the plane is derived from point cloud data using a point cloud coordinate system, the different coordinate systems result in different coordinate values for the same physical location. During the matching operation, points with similar geometric positions are paired. Each endpoint in the line segment endpoint coordinate sequence is traversed, and the image coordinates of that endpoint are roughly converted to point cloud coordinates using the preliminary transformation parameters obtained from the previous spatiotemporal calibration. Then, the nearest boundary point is searched in the plane boundary point coordinate sequence. For example, if the converted coordinates of a line segment endpoint are 2.15 meters, 3.48 meters, and 0.82 meters, and a point in the plane boundary point sequence has coordinates of 2.12 meters, 3.51 meters, and 0.79 meters, the calculated distance is only 5 centimeters. These two points are paired and recorded in the feature matching point pair set. After finding matching points for all endpoints, the feature matching point pair set may contain dozens or even hundreds of pairs of points.
[0092] These matching point pairs are used to optimize the coordinate transformation parameters. Each pair of matching points provides a constraint, meaning that a point in the image coordinate system should coincide with its corresponding point in the point cloud coordinate system after rotation and translation. A system of equations is established, listing all the constraints. The unknowns in the equations are the 9 elements of the rotation matrix and the 3 elements of the translation vector. The least squares method is used to solve this system of equations to calculate the optimal rotation matrix and translation vector. This set of parameters is the precise coordinate transformation parameter. After obtaining the coordinate transformation parameter, the set of line segments is processed. The start and end coordinates of each line segment are extracted. The coordinates in the image coordinate system are transformed into the point cloud coordinate system by multiplying the rotation matrix by the coordinate values and adding the translation vector. For example, the original coordinates of the starting point of a device edge line are 1.8 meters, 2.3 meters, and 1.2 meters. After the rotation matrix transformation, they become 1.75 meters, 2.28 meters, and 1.18 meters. The translation vectors are 0.05 meters, 0.03 meters, and 0.02 meters. After adding these, the final coordinates are 1.8 meters, 2.31 meters, and 1.2 meters. After processing all line segments in this way, we get the transformed set of line segments.
[0093] The transformed set of line segments and planar features are now represented in the point cloud coordinate system. Next, they are associated according to their spatial relationships. During operation, each line segment is checked to see which planar feature it is closest to, and the distance from points on the line segment to each plane is calculated. The plane with the smallest distance is associated with this line segment. For example, if a point on the left side wall outline of the tunnel is only 3 cm from the tunnel wall plane, but more than 50 cm from other planes, this outline is associated with the tunnel wall plane. After all line segments and planes are associated, the line segment features and planar features of the same object are combined. This combined data is the scene representation in a unified coordinate system, containing both edge information described by line segments and shape information described by planes, all expressed using point cloud coordinates.
[0094] In a specific embodiment, the step of performing multimodal data fusion on the scene representation under the unified coordinate system to obtain a fused scene map includes:
[0095] The transformed line segment set in the scene representation under the unified coordinate system is subjected to pixel rasterization to obtain a rasterized line segment map, and the planar features in the scene representation are superimposed with texture information to obtain textured planar data.
[0096] Spatial position calibration is performed on the textured plane data based on the rasterized line segment map to obtain a calibrated textured plane. Then, the calibrated textured plane and the rasterized line segment map are overlaid and synthesized to obtain a fused scene map.
[0097] Specifically, after obtaining the scene representation in a unified coordinate system, the first step is to process the transformed set of line segments by rasterization. During this process, a three-dimensional spatial grid is created with a resolution of 10 centimeters, meaning the entire tunnel space is divided into small squares at 10-centimeter intervals. A specific line segment is extracted from the set. For example, if the starting coordinates of this tunnel wall outline are 2.1 meters, 3.5 meters, and 0.8 meters, and the ending coordinates are 2.1 meters, 3.5 meters, and 2.6 meters, a sampling point is taken every 10 centimeters along the line segment from the starting point to the ending point. The grid square in which each sampling point falls is calculated, and the corresponding grid is marked as occupied. After processing all line segments in this way, the grid squares traversed by the line segments will form a three-dimensional raster map. This is the rasterized line segment map, which clearly shows the spatial distribution of linear features such as tunnel wall edges and equipment outlines.
[0098] Texture information needs to be added to the planar features. Each planar feature corresponds to a region fitted from the previous point cloud, and this region also has a corresponding image region in the environmental depth data. During operation, the coordinates of the point cloud points contained in the planar feature are found, and the pixel positions of these points in the RGB-D image are looked up. The color values of those pixels are read. For example, the coordinates of a point cloud point on the tunnel wall plane are 2.3 meters, 4.1 meters, and 1.5 meters. It corresponds to the pixel in the 320th row and 240th column in the image. The RGB value of this pixel is gray (R120, G115, B110). The color values are assigned to the point cloud points. After all points on the plane are assigned values, textured planar data is obtained. In this data, each point has both three-dimensional coordinates and color information, which can preserve details such as the concrete texture of the tunnel wall and the rust spots on the equipment surface.
[0099] Although both textured planar data and rasterized line segment maps exist in the point cloud coordinate system, there may be slight discrepancies, requiring calibration using the rasterized line segment map. Specifically, feature points are located at the boundaries of the textured planar data. These points should coincide with the positions of line segments in the rasterized line segment map. For example, if the coordinates of a point on the planar boundary are 2.15 meters, 3.52 meters, and 0.85 meters, the coordinates of the nearest occupied grid center in the rasterized line segment map are 2.1 meters, 3.5 meters, and 0.8 meters, a difference of 5 centimeters. The deviations between all boundary points and their corresponding grids are calculated. These deviations are then used to fine-tune the position of the planar data, perhaps by translating it by 2 centimeters or rotating it by 0.5 degrees. The adjusted textured planar data is the calibrated textured planar, whose boundary positions perfectly align with the line segments in the rasterized line segment map.
[0100] After calibration, layer overlay begins, using the rasterized line segment map as the underlying framework and the calibrated texture plane as the surface overlay. During the process, each grid cell is traversed. First, it checks if the cell in the rasterized line segment map is occupied by a line segment. If so, edge attributes are marked within the cell. Then, it checks if any points in the calibrated texture plane data fall within this cell. If so, the color information of these points is filled into the cell. For example, if a grid cell is both traversed by the alleyway wall outline and within the alleyway wall plane's coverage area, this cell is both marked as an edge position and filled with gray texture in the blended scene map. Once all grid cells are processed, the entire blended scene map is generated. The map displays both the structural framework outlined by line segments and the texture details laid out on the plane, fully presenting the geometry and surface features of the alleyway environment.
[0101] In a specific embodiment, the step of reconstructing the passage environment based on the fused scene map to obtain the alleyway structure model includes:
[0102] The topological nodes of the rasterized line segment graph in the fused scene map are extracted to obtain the coordinates of the alleyway nodes, and the connectivity of the alleyway node coordinates is sorted out to obtain the node connectivity table.
[0103] Based on the node connectivity table, the textured planar data in the fused scene map is subjected to three-dimensional stretching to obtain the alleyway three-dimensional surface block, and the adjacent surface of the alleyway three-dimensional surface block is calibrated to obtain the alleyway three-dimensional shell.
[0104] By combining the node connectivity list with the three-dimensional shell of the tunnel, a topological relationship binding is performed to obtain the tunnel topological model. Then, dimension annotations and spatial constraint information are added to the tunnel topological model to obtain the tunnel structure model.
[0105] Specifically, the rasterized line segment graph from the integrated scene map is used to extract topological nodes. This involves scanning grid squares occupied by line segments, focusing on finding intersections and endpoints. For example, a grid square at the intersection of the tunnel roof and the edge of the left side wall, with coordinates of 2.1 meters, 3.5 meters, and 2.8 meters, is a node. The starting position of a line segment in the tunnel is also considered a node. Extracting the coordinates of all these key locations—potentially dozens of nodes for the entire tunnel—constitutes the tunnel node coordinates. After node extraction, the connectivity relationships between them are analyzed. The rasterized line segment graph is traversed to check which nodes are connected by line segments. For instance, if a vertical line segment exists between nodes with coordinates of 2.1 meters, 3.5 meters, and 0.8 meters and nodes with coordinates of 2.1 meters, 3.5 meters, and 2.8 meters, this connection is recorded in a table. All node pair connectivity is recorded to form a table—the node connectivity table—clearly indicating which nodes are directly connected to which other nodes.
[0106] After obtaining the node connectivity table, the textured planar data in the scene map is processed and fused. This planar data is currently a two-dimensional point cloud and needs to be stretched into three-dimensional blocks. The stretching direction is determined first, based on the plane's normal vector. Vertical planes are stretched 20 cm in the thickness direction, and horizontal planes are stretched vertically. During stretching, the coordinate range of adjacent nodes in the connectivity table is referenced. For example, if the vertical coordinates of the two endpoint nodes of a tunnel wall plane are 0.8 meters and 2.8 meters respectively, the stretching is performed within this height range; it stops when the range is exceeded. After stretching, each planar data becomes a three-dimensional block. The tunnel wall plane, after stretching, becomes a cuboid block, and the equipment installation plane, after stretching, might become a thin plate-like block. These blocks are collectively called tunnel three-dimensional blocks. However, the edges of adjacent three-dimensional blocks may not align perfectly; there may be gaps of a few centimeters at the junction of the tunnel top plate and side wall blocks. Adjacent surface alignment calibration is needed to eliminate these gaps. During calibration, the contact boundary between two adjacent three-dimensional blocks is checked, the deviation of the boundary position is calculated, and then the vertex coordinates of one of the blocks are finely adjusted to ensure a tight fit between the two blocks. Once all adjacent facets have been calibrated, the entire tunnel space is enclosed by these three-dimensional facets, forming a closed shell structure. This is the three-dimensional shell of the tunnel.
[0107] The 3D shell of the tunnel and the node connectivity table are combined to perform topological relationship binding. During the operation, each 3D facet in the shell is numbered, and then the node connectivity table is consulted to find which nodes are located on the boundaries of that facet. For example, if facet 3 is the left side wall, and the nodes at coordinates 2.1m, 3.5m, 0.8m and 2.1m, 3.5m, 2.8m in the node connectivity table are on its boundary, then the binding with these two nodes is recorded in the facet data. After all faces are bound to nodes, the entire structure has both geometric shape and topological connections; this data constitutes the tunnel topology model. After the topology model is built, dimensional information needs to be added, measuring parameters such as tunnel width and height. For example, the lateral distance from the left side wall nodes (2.1m, 3.5m, 0.8m) to the corresponding right side wall node is 4.2m, which is marked as a tunnel width of 4.2m; the vertical distance from the floor to the roof (3m) is marked as a net height of 3m. Spatial constraint information also needs to be added, indicating certain areas where passage is prohibited, such as a no-entry constraint within 50cm below pipelines and a deceleration constraint in waterlogged areas. All dimensions and constraint information are added to the tunnel topology model, and the final product is a complete tunnel structure model. This model contains accurate geometric data, topological connections, dimension annotations, and spatial constraint rules.
[0108] In a specific embodiment, the step of planning the path of the inspection robot based on the alleyway structure model to obtain an obstacle avoidance trajectory, and adjusting the motion state of the inspection robot according to the obstacle avoidance trajectory, includes:
[0109] The inner wall surface of the three-dimensional shell of the tunnel in the tunnel structure model is extracted to obtain the inner wall boundary surface of the tunnel. Based on the inner wall boundary surface of the tunnel, a safe distance offset and contraction is performed in the direction of the tunnel center to obtain the passable boundary.
[0110] The alleyway structure model is divided into grids based on the passable boundary line to obtain a passable grid map. Then, a grid-by-grid search is performed in the passable grid map from the current position of the inspection robot to the target inspection position to obtain an obstacle avoidance trajectory point sequence. The obstacle avoidance trajectory point sequence is used as the obstacle avoidance trajectory.
[0111] Vector difference calculation is performed on adjacent trajectory points in the obstacle avoidance trajectory point sequence to obtain the travel direction angle and travel distance of each trajectory segment. Based on the travel direction angle and travel distance, the motion state parameters of the robot are generated, and the motion state of the inspection robot is controlled based on the motion state parameters.
[0112] Specifically, the 3D shell of the tunnel structure model includes facets such as the top plate, sidewalls, and bottom plate. When extracting the inner wall surface, the surface facing inwards from the tunnel must be identified. Operationally, each facet is traversed, and its normal vector is checked. The surface whose normal vector points to the center of the tunnel is the inner wall surface. For example, the left wall facet has two surfaces, one with a normal vector pointing inwards and the other pointing to the center of the tunnel; the surface pointing to the center is selected. After extracting the inner wall surfaces of all facets, they are combined to form the inner wall boundary surface of the tunnel. This boundary surface delineates the actual spatial extent of the tunnel. However, the robot cannot move too close to the wall; a safe distance must be maintained. The solution is to shrink the inner wall boundary surface of the tunnel towards the center. Specifically, starting from each point on the boundary surface, the robot moves inwards 40 centimeters along the normal vector. This 40 centimeters is determined by the robot's width plus a safety margin; for example, a robot 35 centimeters wide with a 5-centimeter buffer. After shrinking, all points are reconnected to form a surface. The area enclosed by this surface is the passable boundary, and the robot can only move within this boundary.
[0113] After determining the passable boundaries, the alleyway structure model is divided into grids, with the alleyway space cut into small grids of 50 cm squares. Each grid's center point is checked to see if it falls within the passable boundaries. Grids within the boundaries are marked as passable (1), while grids outside the boundaries or occupied by obstacles are marked as impassable (0). All grids are processed to form a two-dimensional array, which is the passable grid map. Assuming the alleyway is 50 meters long and 4 meters wide, after dividing into 50 cm grids, the map has 100 rows and 8 columns, totaling 800 grids. Path searching starts from the robot's current grid, which might be at row 10, column 4, while the target inspection location is at row 80, column 4. The search process uses the A* algorithm to expand grid by grid. First, it checks which of the current 8 neighboring grids are passable. For passable grids, a cost is calculated, equal to the actual distance traveled from the starting point to that grid plus the estimated distance from that grid to the destination. For example, if you expand to the 11th row, 4th column cell, the distance from the starting point to this cell is 50 centimeters, and the straight-line distance from here to the ending point is approximately 34.5 meters. The cost is 0.5 plus 34.5, which equals 35 meters. Sort all the expanded cells by cost, and continue expanding with the cell that has the lowest cost each time. If you encounter an obstacle cell, go around it and search the adjacent cells.
[0114] After finding the destination grid, the entire path is traced back, and the coordinates of the center points of the grids passed through are recorded to form an obstacle avoidance trajectory point sequence, which may contain dozens of trajectory points. Once the trajectory point sequence is obtained, the motion parameters between adjacent points are calculated. Taking the coordinates of the first point in the sequence as 5 meters and 2 meters, and the coordinates of the second point as 5.4 meters and 2 meters, the difference between the x-coordinates of the two points is 0.4 meters, and the difference between the y-coordinates is 0 meters. Using the arctangent function, the travel direction angle is calculated to be 0 degrees, representing due east, and the distance between the two points is 0.4 meters, i.e., the travel distance is 0.4 meters. This process is repeated for all adjacent point pairs. For some sections, the direction angle may be 90 degrees, representing due north, and for others, 45 degrees, representing northeast. The travel distance varies from 0.5 meters to 0.7 meters depending on the grid distribution. These direction angle and distance data form the motion state parameters and are sent to the robot control system. The control system reads the first set of parameters and knows that it needs to move 0.4 meters in the 0-degree direction, so it controls the left and right wheels to move forward at the same speed. After the encoder counts 0.4 meters traveled, it reads the second set of parameters. If the second steering angle becomes 45 degrees, the control system will reduce the speed of the right wheel while keeping the left wheel spinning, and use differential steering to turn the robot body 45 degrees before continuing to move forward the corresponding distance. By executing all motion state parameters in sequence according to this logic, the robot can travel from the starting point to the target position along the route planned by the obstacle avoidance trajectory point sequence to complete the inspection task.
[0115] In a specific embodiment, the step of dividing the alleyway structure model into grids based on the traversable boundary lines to obtain a traversable grid map includes:
[0116] The passable boundary is horizontally projected along the bottom surface of the alley to obtain a two-dimensional passable area outline. The grid side length is determined based on the dimension annotation information in the alley structure model. The area enclosed by the two-dimensional passable area outline is divided into equidistant grids according to the grid side length to obtain a grid cell array. The grid cell array includes the row and column number and center point coordinates of each grid cell.
[0117] Based on the three-dimensional shell of the tunnel structure model, the installation location of the underground equipment is extracted to obtain the obstacle occupancy area. The position inclusion of each grid cell in the grid cell array with the two-dimensional passage area outline and the obstacle occupancy area is determined. Grid cells located inside the two-dimensional passage area outline and not overlapping with the obstacle occupancy area are marked as passable grid cells, and the remaining grid cells are marked as obstacle grid cells, thus obtaining a passage grid map.
[0118] Specifically, the passable boundary is a three-dimensional closed surface, which needs to be converted into a two-dimensional plane for path planning. Operationally, the passable boundary is projected onto the bottom surface of the tunnel. The projection method is to take the coordinates of each point on the passable boundary, retain its x and y coordinates, and set the y coordinate to the height of the bottom plate, for example, 0 meters. In this way, the three-dimensional boundary points are compressed onto the bottom plane. After all the boundary points are projected, these two-dimensional coordinate points are connected in sequence to form a closed curve. The area enclosed by this curve is the outline of the two-dimensional passable area. Next, the grid size is determined by checking the dimension annotation information from the tunnel structure model. The model records a tunnel width of 4.2 meters. Considering the robot's turning radius and computational efficiency, the grid side length is set to 50 centimeters. Too small a grid would result in a large computational load, and too large a grid would not be accurate enough. After determining the grid side length, the outline of the two-dimensional passable area is divided into a grid, and the boundary range of the outline is found. Horizontally, it extends from 0 meters on the left to 4.2 meters on the right, and vertically, it extends from 0 meters at the start to 50 meters at the end. Starting from the origin, draw horizontal and vertical lines at 50-centimeter intervals. Draw 9 horizontal lines and 101 vertical lines, intersecting to divide the area into 800 small squares (8 x 100). Number each square, starting from the bottom left corner, with the first row and first column numbered 1-1, then 1-2, 1-3, and so on, increasing from the second row upwards to 2-1, 2-2, and so on, until the 100th row and 8th column, numbered 100-8. Simultaneously, calculate the center point coordinates of each grid cell. For example, for grid cell 3-5, row number 3 indicates a vertical position between 1 and 1.5 meters, with a center of 1.25 meters; column number 5 indicates a horizontal position between 2 and 2.5 meters, with a center of 2.25 meters. The center point coordinates are recorded as 2.25 meters and 1.25 meters. After processing all grid cells, each cell has a row and column number and center point coordinates; these data form a grid cell array.
[0119] After the grid cell array is built, it's necessary to mark which cells are traversable and which are not. First, locate the equipment installation positions within the 3D shell of the tunnel. Scan the shell data to see which 3D faces are marked with equipment attributes. For example, a face might be recorded as a ventilation duct located on the right side of the tunnel, with coordinates ranging from 3.5 to 4 meters horizontally and 20 to 25 meters vertically. Projecting this coordinate range onto the base plane forms a rectangular area, which represents the obstacle occupancy area for that duct. There may be multiple pieces of equipment underground, such as fans, electrical control cabinets, and water pipes; extract all of these and summarize the occupancy areas for all equipment. Then, iterate through each grid cell in the grid cell array to determine its position. Take a specific grid cell, for example, cell number 50-6 with center coordinates of 2.75 meters and 24.75 meters. First, check if this point is inside the 2D traversable area outline. This is done by drawing a ray outward from this point and counting the number of times the ray crosses the outline boundary. An odd number of crossings indicates the point is inside the outline, and an even number indicates it is outside. Assuming the ray crosses the outline once for the center point of cell number 50-6, it is determined to be inside the outline. Next, check if it overlaps with the obstacle area. See if the center point coordinates of 2.75 meters and 24.75 meters fall within any obstacle area. As mentioned earlier, the ventilation duct's horizontal area is 3.5 to 4 meters; 2.75 meters is not within this range. The vertical area is 20 to 25 meters, including 24.75 meters, but the horizontal area does not meet the requirement, so it's not considered an overlap. If the grid cell is both within the outline and does not overlap with an obstacle, assign a passable attribute value of 1 to this cell, marking it as a passable grid. Conversely, the center point of grid number 50-7 is 3.25 meters and 24.75 meters. The horizontal 3.25 meters falls near the edge of the duct's horizontal range of 3.5 to 4 meters. For a stricter determination, we need to consider the entire range of the grid cell. This grid horizontally covers 3 to 3.5 meters. Although the center point of 3.25 meters is not in the duct area, the grid boundary of 3.5 meters just touches the boundary of the duct's occupied area of 3.5 meters, which is considered an overlap. Assign a value of 0 to this cell, marking it as an obstacle grid. After all 800 grid cells have been checked, those that can be walked are marked with 1 and those that cannot be walked are marked with 0. The entire array forms a passable grid map. In the map, the positions of 1 form a passage and the positions of 0 are walls and obstacles.
[0120] The above describes a multimodal collaborative control method for an intelligent underground inspection robot in a coal mine, as described in an embodiment of the present invention. The following describes a multimodal collaborative control system for an intelligent underground inspection robot in a coal mine, as described in an embodiment of the present invention. Please refer to [link / reference]. Figure 1 One embodiment of the multimodal collaborative control system for an intelligent underground inspection robot in coal mines according to the present invention includes:
[0121] The scanning module 1 is used to scan the underground roadway of the coal mine using the RGB-D camera and lidar on the inspection robot to obtain environmental depth data and point cloud data, and to perform spatiotemporal calibration on the environmental depth data and point cloud data to obtain calibrated heterogeneous data.
[0122] Alignment module 2 is used to extract features based on the calibrated heterogeneous data to obtain a scene feature set, and to align the scene feature set with multi-source information to obtain a scene representation under a unified coordinate system.
[0123] The fusion module 3 is used to perform multimodal data fusion on the scene representation under the unified coordinate system to obtain a fused scene map, and to reconstruct the channel environment based on the fused scene map to obtain a tunnel structure model.
[0124] Planning module 4 is used to plan the path of the inspection robot based on the alleyway structure model, obtain the obstacle avoidance trajectory, and adjust the motion state of the inspection robot according to the obstacle avoidance trajectory.
[0125] In this embodiment, the specific implementation of each module in the above system embodiment is described in the above method embodiment, and will not be repeated here.
Claims
1. A multimodal collaborative control method for an intelligent underground inspection robot in coal mines, characterized in that, Includes the following steps: The inspection robot scans the underground roadways of the coal mine using an RGB-D camera and a lidar to obtain environmental depth data and point cloud data. The environmental depth data and point cloud data are then spatiotemporally calibrated to obtain calibrated heterogeneous data. Feature extraction is performed based on the calibrated heterogeneous data to obtain a scene feature set, and the scene feature set is aligned with multi-source information to obtain a scene representation under a unified coordinate system. Multimodal data fusion is performed on the scene representation under the unified coordinate system to obtain a fused scene map, and the channel environment is reconstructed based on the fused scene map to obtain a tunnel structure model; The path of the inspection robot is planned based on the alleyway structure model to obtain the obstacle avoidance trajectory, and the motion state of the inspection robot is adjusted according to the obstacle avoidance trajectory. The feature extraction based on the calibrated heterogeneous data yields a scene feature set, including: Image edge detection is performed on the environmental depth data in the calibrated heterogeneous data to obtain an edge contour image. Connectivity analysis is then performed on the edge contour image to connect adjacent edge pixels into continuous line segments, resulting in a set of line segments including the tunnel wall contour line and the equipment edge line. Based on the set of line segments, the point cloud data in the calibrated heterogeneous data is filtered to obtain the local point cloud corresponding to the line segment, and the local point cloud is fitted with a plane to obtain planar features including the tunnel wall plane and the equipment installation plane. The planar features are subjected to geometric parameter calculation to obtain planar geometric parameters. The planar geometric parameters are then classified and organized according to their area and normal vector direction to obtain a set of scene features. The step of aligning the scene feature set with multi-source information to obtain a scene representation in a unified coordinate system includes: The endpoint coordinates of the line segment set in the scene feature set are extracted to obtain the line segment endpoint coordinate sequence, and the boundary point coordinates of the planar features in the scene feature set are extracted to obtain the planar boundary point coordinate sequence. Spatial position matching is performed based on the coordinate sequence of the line segment endpoints and the coordinate sequence of the plane boundary points to obtain a set of feature matching point pairs. Based on the set of feature matching point pairs, the rotation matrix and translation vector from the image coordinate system to the point cloud coordinate system are calculated to obtain the coordinate transformation parameters. Based on the coordinate transformation parameters, the coordinates of each line segment in the line segment set are transformed to obtain the transformed line segment set. The transformed line segment set is then associated and combined with the planar features according to their spatial positional relationship to obtain a scene representation under a unified coordinate system.
2. The multimodal collaborative control method for an intelligent underground inspection robot in coal mines according to claim 1, characterized in that, The inspection robot scans underground mine roadways using an RGB-D camera and LiDAR to obtain environmental depth data and point cloud data, including: The inspection robot continuously captures multiple frames of images from underground coal mine roadways using an RGB-D camera, obtaining original color image sequences and original depth image sequences. The original color image sequences are then subjected to grayscale histogram equalization to obtain enhanced color image sequences. Pixel-level registration is performed between the enhanced color image sequence and the original depth image sequence to obtain image-depth registration data. Neighborhood pixel interpolation is then performed to fill the invalid depth regions in the image-depth registration data to obtain environmental image-depth data. The laser radar is used to perform rotational scanning and ranging of the underground roadway in the coal mine to obtain the original laser echo signal. The original laser echo signal is then subjected to intensity threshold filtering to remove dust scattering interference, resulting in point cloud data.
3. The multimodal collaborative control method for an intelligent underground inspection robot in coal mines according to claim 1, characterized in that, The process of performing spatiotemporal calibration on the environmental map depth data and point cloud data to obtain calibrated heterogeneous data includes: The image frames in the environmental depth data are timestamped to obtain timestamped image data. At the same time, each point cloud cluster in the point cloud data is timestamped to obtain timestamped point cloud data. Based on timestamped image data and timestamped point cloud data, the data is aligned according to the chronological order to obtain time-synchronized image depth point cloud data; Image feature points and point cloud feature points are extracted from the time-synchronized image-depth-point cloud data to obtain image feature point sets and point cloud feature point sets; Calculate the Euclidean distance between the image feature point set and the point cloud feature point set, find the closest feature point pair, and calculate the coordinate transformation parameters based on the feature point pair to obtain the transformation parameters from the point cloud coordinate system to the image coordinate system. Then, perform coordinate transformation on the time-synchronized image-depth point cloud data based on the transformation parameters to obtain calibrated heterogeneous data.
4. The multimodal collaborative control method for an intelligent underground inspection robot in coal mines according to claim 3, characterized in that, The process of fusing multimodal data on the scene representation under the unified coordinate system to obtain a fused scene map includes: The transformed line segment set in the scene representation under the unified coordinate system is subjected to pixel rasterization to obtain a rasterized line segment map, and the planar features in the scene representation are superimposed with texture information to obtain textured planar data. Spatial position calibration is performed on the textured plane data based on the rasterized line segment map to obtain a calibrated textured plane. Then, the calibrated textured plane and the rasterized line segment map are overlaid and synthesized to obtain a fused scene map.
5. The multimodal collaborative control method for an intelligent underground inspection robot in coal mines according to claim 4, characterized in that, The process of reconstructing the passage environment based on the fused scene map to obtain the alleyway structure model includes: The topological nodes of the rasterized line segment graph in the fused scene map are extracted to obtain the coordinates of the alleyway nodes, and the connectivity of the alleyway node coordinates is sorted out to obtain the node connectivity table. Based on the node connectivity table, the textured planar data in the fused scene map is subjected to three-dimensional stretching to obtain the alleyway three-dimensional surface block, and the adjacent surface of the alleyway three-dimensional surface block is calibrated to obtain the alleyway three-dimensional shell. By combining the node connectivity list with the three-dimensional shell of the tunnel, a topological relationship binding is performed to obtain the tunnel topological model. Then, dimension annotations and spatial constraint information are added to the tunnel topological model to obtain the tunnel structure model.
6. The multimodal collaborative control method for an intelligent underground inspection robot in coal mines according to claim 5, characterized in that, The process of planning the path of the inspection robot based on the alleyway structure model to obtain an obstacle avoidance trajectory, and adjusting the motion state of the inspection robot according to the obstacle avoidance trajectory, includes: The inner wall surface of the three-dimensional shell of the tunnel in the tunnel structure model is extracted to obtain the inner wall boundary surface of the tunnel. Based on the inner wall boundary surface of the tunnel, a safe distance offset and contraction is performed in the direction of the tunnel center to obtain the passable boundary. The alleyway structure model is divided into grids based on the passable boundary line to obtain a passable grid map. Then, a grid-by-grid search is performed in the passable grid map from the current position of the inspection robot to the target inspection position to obtain an obstacle avoidance trajectory point sequence. The obstacle avoidance trajectory point sequence is used as the obstacle avoidance trajectory. Vector difference calculation is performed on adjacent trajectory points in the obstacle avoidance trajectory point sequence to obtain the travel direction angle and travel distance of each trajectory segment. Based on the travel direction angle and travel distance, the motion state parameters of the robot are generated, and the motion state of the inspection robot is controlled based on the motion state parameters.
7. The multimodal collaborative control method for an intelligent underground inspection robot in coal mines according to claim 6, characterized in that, The process of dividing the alleyway structure model into grids based on the traversable boundary lines to obtain a traversable grid map includes: The passable boundary is horizontally projected along the bottom surface of the alley to obtain a two-dimensional passable area outline. The grid side length is determined based on the dimension annotation information in the alley structure model. The area enclosed by the two-dimensional passable area outline is divided into equidistant grids according to the grid side length to obtain a grid cell array. The grid cell array includes the row and column number and center point coordinates of each grid cell. Based on the three-dimensional shell of the tunnel structure model, the installation location of the underground equipment is extracted to obtain the obstacle occupancy area. The position inclusion of each grid cell in the grid cell array with the two-dimensional passage area outline and the obstacle occupancy area is determined. Grid cells located inside the two-dimensional passage area outline and not overlapping with the obstacle occupancy area are marked as passable grid cells, and the remaining grid cells are marked as obstacle grid cells, thus obtaining a passage grid map.
8. A multimodal collaborative control system for an intelligent underground inspection robot in coal mines, characterized in that, The method for implementing the multimodal collaborative control of an intelligent underground inspection robot in coal mines according to any one of claims 1-7 includes: The scanning module is used to scan the underground roadways of the coal mine using the RGB-D camera and lidar on the inspection robot to obtain environmental depth data and point cloud data, and to perform spatiotemporal calibration on the environmental depth data and point cloud data to obtain calibrated heterogeneous data. The alignment module is used to extract features based on the calibrated heterogeneous data to obtain a scene feature set, and to align the scene feature set with multi-source information to obtain a scene representation under a unified coordinate system. The fusion module is used to perform multimodal data fusion on the scene representation under the unified coordinate system to obtain a fused scene map, and to reconstruct the channel environment based on the fused scene map to obtain a tunnel structure model. The planning module is used to plan the path of the inspection robot based on the alleyway structure model, obtain the obstacle avoidance trajectory, and adjust the motion state of the inspection robot according to the obstacle avoidance trajectory.
Citation Information
Patent Citations
Roadway space contact type collision obstacle avoidance method under coal mine underground scene degradation condition
CN118034277A
Intelligent obstacle detection and avoidance method for power transmission line inspection unmanned aerial vehicle
CN121386844A