Drilling environment sensing system
By using multimodal sensing units and data fusion technology, a high-precision 3D point cloud map is generated, which solves the problem of inaccurate drilling positioning of rock drilling rigs in complex environments and improves construction quality and efficiency.
Patent Information
- Application Number
- CN202511835203.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-08
- Publication Date
- 2026-03-06
AI Technical Summary
In complex underground construction environments, the existing sensors on rock drilling rigs are insufficient, leading to inaccurate borehole positioning, borehole bottom deviation, and affecting construction quality and efficiency.
By employing a multimodal sensing unit, combined with LiDAR, HDR camera and thermal imaging camera, a high-precision 3D point cloud map is generated through multimodal data fusion and real-time dynamic mapping, enabling precise positioning and status feedback of the drilling environment.
It improves the accuracy of borehole positioning, reduces construction errors caused by changes in rock mass structure or equipment posture deviations, and provides reliable technical support for intelligent rock drilling operations.
Smart Images

Figure CN121611433A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of engineering surveying and intelligent construction, and in particular to a borehole environment sensing system. Background Technology
[0002] During drilling in peripheral holes, rock drilling rigs exhibit significant limitations in environmental perception, leading to substantial borehole bottom deviations and poor control over- and under-drilling. Traditional borehole positioning and environmental perception methods often rely on a single type of sensor, such as a laser sensor or a vision camera. Laser sensors are susceptible to interference from dust and water mist in underground construction environments, resulting in limited measurement distance or decreased accuracy. Vision cameras, on the other hand, struggle to acquire clear and effective image information in low-light, high-dust, and complex underground engineering structures (such as those with numerous anchor bolts and steel arches), often encountering problems like difficulty in feature extraction and matching errors.
[0003] This single-sensor-based sensing method struggles to achieve high-precision, real-time modeling of the construction area and accurate borehole location in complex underground construction environments characterized by low light, high dust levels. Insufficient environmental perception prevents the drilling rig from accurately grasping key environmental information such as the rock mass structure, existing support structure distribution, and the location of adjacent boreholes. This leads to inaccurate borehole trajectory control and significant deviations between the actual and designed borehole bottom positions. This not only affects the blasting effect of surrounding holes and causes irregular excavation profiles but also triggers serious over-excavation or under-excavation problems, increasing support material consumption, prolonging the construction period, and reducing the stability and economic efficiency of the engineering structure. To address these issues, this application proposes a borehole environmental perception system. Summary of the Invention
[0004] The purpose of this invention is to provide a drilling environment sensing system to solve the problem of weak sensing capabilities of current rock drilling rigs during drilling.
[0005] To achieve the above objectives, the present invention provides the following technical solution: A borehole environment sensing system, the system comprising: The multimodal sensing unit acquires information about the construction environment and the real-time status parameters of the drill arm based on various types of sensors; The fusion computing unit constructs a three-dimensional point cloud map of the construction environment based on information about the construction environment, and calculates the borehole outline and the three-dimensional structural surface state of the rock mass at the working face based on the three-dimensional point cloud map. The task mapping unit calculates the mapping position of the drill bit in the 3D point cloud map based on the pre-input drilling design drawing and the attitude of the drill arm, and dynamically maps each drilling point to the 3D point cloud map.
[0006] Furthermore, the multimodal sensing unit includes a lidar, an HDR camera, and a thermal imaging camera.
[0007] Furthermore, the multimodal sensing unit also includes: Dust concentration sensor is used to monitor the dust concentration in the drilling operation area in real time.
[0008] Furthermore, the power of the lidar is dynamically adjusted based on the dust concentration; when the dust concentration increases, the transmission power is automatically increased.
[0009] Furthermore, the multimodal sensing unit also includes: A light sensor monitors the light intensity inside the cave in real time and synchronizes the data to the exposure parameter control system of the HDR camera.
[0010] Furthermore, the method for constructing a 3D point cloud map using the fusion computing unit includes the following steps: Step A: Denoise and register the point cloud data from the LiDAR, and perform pixel-level alignment and fusion on the images from the HDR camera and thermal imaging camera. Step B: A multimodal feature fusion model based on the Transformer fusion network with cross-attention mechanism is constructed to weight and align point clouds and images in spatial and temporal dimensions. Step C: Reconstruct the construction scene using the fused high-dimensional features to generate a dense three-dimensional point cloud map containing geometric structure and temperature distribution.
[0011] Furthermore, a Transformer fusion network based on the cross-attention mechanism is used to extract the cavern outline, rock mass structure coordinates, and environmental disturbance parameters. The specific process is as follows: For four types of input data—LiDAR point cloud, enhanced visual image, infrared thermal imaging, and environmental parameters—feature extraction branches are designed to retain the core information of each modality. The attention weight of each pixel in the feature map is calculated by a self-attention mechanism. Environmental parameter features are introduced to construct a cross-attention matrix. Through cross-attention calculation, the features extracted from LiDAR, HDR camera and thermal imaging camera data are fused into a multimodal fusion feature map. This feature map contains the spatial geometry, visual texture and thermal structure information of the cavern. Set up an independent output head to decode the target parameters from the fused feature map, ensuring the extraction accuracy and real-time performance of each parameter.
[0012] Furthermore, the extracted visual features and thermal imaging features are mapped to the lidar coordinate system; Based on the timestamp of LiDAR acquisition, visual features, infrared features, and attitude data at the same moment are matched through a unified system clock. The aligned three types of features are input into the Transformer fusion network, and attention weights are dynamically allocated based on environmental parameters. After processing by the Transformer encoder, the integrated features are output. Based on sparse laser point clouds and integrated features, blank areas are filled in to generate high-density point clouds, and temperature and structural information are assigned to each point. By integrating the device's dynamic attitude data into the point cloud, the cumulative error caused by device movement / tilt can be eliminated; Global map stitching and optimization stitches together multiple dense point clouds into a complete cavern map, eliminating accumulated errors.
[0013] In summary, the present invention has the following advantages compared with the prior art: The borehole environment perception system disclosed in this invention achieves high-precision guidance and status feedback for borehole operations in complex underground construction environments through multimodal data fusion and real-time dynamic mapping. This system not only improves the accuracy of borehole positioning but also effectively reduces construction errors caused by changes in rock structure or equipment posture deviations, providing reliable technical support for intelligent rock drilling operations. Attached Figure Description
[0014] Figure 1 This is a framework diagram of the borehole environment sensing system disclosed in an embodiment of the present invention.
[0015] Figure 2 This is a flowchart of the three-dimensional point cloud map generated in the borehole environment perception system disclosed in an embodiment of the present invention.
[0016] Figure 3 This is a flowchart illustrating the extraction of cavern outline and rock mass structure coordinates from a borehole environment perception system disclosed in an embodiment of the present invention.
[0017] Figure 4 This is a flowchart of the construction scene reconstruction in the borehole environment perception system disclosed in an embodiment of the present invention. Detailed Implementation
[0018] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.
[0019] Figure 1As shown, an embodiment of the present invention provides a borehole environment perception system. The perception system includes a multimodal perception unit, a fusion computing unit, and a task mapping unit. The multimodal perception unit acquires information about the construction environment and real-time status parameters of the drill arm based on various types of sensors. The fusion computing unit constructs a three-dimensional point cloud map of the construction environment based on the information of the construction environment and calculates the borehole chamber outline and the three-dimensional structural surface state of the rock mass at the working face based on the three-dimensional point cloud map. The task mapping unit calculates the mapping position of the drill bit in the three-dimensional point cloud map based on the pre-input borehole design drawing and the posture of the drill arm, and dynamically maps each borehole point to the three-dimensional point cloud map.
[0020] Specifically, in this embodiment, during the drilling process of the rock drilling rig in the peripheral holes, the multimodal sensing unit is deployed on the rock drilling rig. During drilling, the multimodal sensing unit acquires information about the construction environment through various sensors, such as 3D point cloud maps, laser scattering maps, and infrared images. The multimodal sensing unit transmits the acquired construction environment information and the real-time status parameters of the drill arm to the fusion computing unit. The fusion computing unit uses the construction environment information to construct a high-precision 3D point cloud map, calculates the borehole contour and the 3D structural surface state of the rock mass at the working face based on the 3D point cloud map, and identifies the posture of the drill arm according to the real-time status parameters. The task mapping unit calculates the mapping position of the drill bit in the 3D point cloud map based on the pre-input borehole design drawing and the posture of the drill arm, and dynamically maps each borehole point to the 3D point cloud map.
[0021] The borehole environment perception system disclosed in this invention achieves high-precision guidance and status feedback for borehole operations in complex underground construction environments through multimodal data fusion and real-time dynamic mapping. This system not only improves the accuracy of borehole positioning but also effectively reduces construction errors caused by changes in rock structure or equipment posture deviations, providing reliable technical support for intelligent rock drilling operations.
[0022] Specifically, in this embodiment, the multimodal perception unit includes a lidar, a depth camera, and a thermal imaging camera. The lidar is used to acquire point cloud data of the construction area and construct the basic geometric framework of the environment; the HDR camera is used to acquire color images; and the thermal imaging camera senses the contours of objects through differences in thermal radiation and assists in identifying seepage areas in the rock mass.
[0023] Specifically, the multimodal sensing unit is mounted on a fixed gimbal on the front support of the rock drilling rig, ensuring that the field of view covers the drilling arm's working range and the face in front. The fusion computing unit is an embedded high-performance computing module with the ability to process multi-source sensor data in real time.
[0024] like Figure 2 As shown, the fusion computing unit generates a 3D point cloud map through the following steps: Step A: Denoise and register the point cloud data from the LiDAR, and perform pixel-level alignment and fusion on the images from the HDR camera and thermal imaging camera. Step B: A multimodal feature fusion model based on the Transformer fusion network with cross-attention mechanism is constructed to weight and align point clouds and images in spatial and temporal dimensions. Step C: Reconstruct the construction scene using the fused high-dimensional features to generate a dense three-dimensional point cloud map containing geometric structure and temperature distribution.
[0025] Specifically, in step A, the lidar point cloud data is denoised using statistical filtering and voxel grid downsampling methods, and the precise registration of multiple frame point clouds is achieved based on the ICP algorithm.
[0026] Preferably, in this embodiment, the multimodal sensing unit further includes a dust concentration sensor for real-time monitoring of dust concentration in the drilling area and synchronizing the data to the fusion computing unit for environmental condition assessment. This sensor is integrated into the front end of the pan-tilt unit and shares power and communication interfaces with other sensing devices, ensuring system compactness and stability.
[0027] When processing lidar point cloud data, a correlation model between lidar distance error and dust concentration is established, as shown in formula (1): ;Formula (1) in, This is the distance correction amount. The scattering coefficient (calibrated experimentally) is the scattering coefficient. This represents the real-time dust concentration. This refers to the initial measurement distance of the lidar; for example, when the dust concentration c = 15 mg / m³, the initial measurement distance... When =10m, =0.002×15×10=0.3m, the original distance needs to be corrected to 10-0.3=9.7m. After correcting the point cloud data, a statistical filtering algorithm is used to remove isolated noise points (in this embodiment, a threshold for the number of neighboring points is set: when a point has less than 5 neighboring points, it is determined to be noise). Then, voxel grid resampling (in this embodiment, the voxel size is set to 5mm×5mm×5mm) is used to reduce the amount of data, and finally outputs three-dimensional point cloud data of the cavern with uniform density and an error ≤±2mm.
[0028] Preferably, the power of the lidar is dynamically adjusted based on dust concentration. When the dust concentration increases, the transmission power is automatically increased to enhance the echo signal strength, thus enabling the acquisition of effective point cloud data even in high-dust environments. Simultaneously, the point cloud resolution is reduced to increase the sampling frequency. In this embodiment, multiple different thresholds are set for dust concentration. When the dust concentration increases, the lidar power is increased to a preset value. For example, dust concentration is divided into three levels: when the dust concentration is below 10 mg / m³, the lidar power maintains the baseline value; when the concentration is between 10 and 25 mg / m³, the power is increased by 20%; and when it exceeds 25 mg / m³, it is increased by 40%. Simultaneously, the point cloud resolution is adjusted from 5 mm to 10 mm, thereby ensuring the real-time performance and stability of data acquisition. This dynamic adjustment mechanism effectively extends the continuous working capability of the equipment in harsh environments and ensures the availability and accuracy consistency of point cloud data in key areas. For RGB images captured by HDR cameras, Block Adaptive Histogram Equalization (CLAHE) is used. In this embodiment, the image is divided into 8×8 pixel sub-blocks, a histogram is calculated for each sub-block, and a contrast threshold is limited to avoid noise amplification caused by traditional equalization and improve the clarity of rock texture in dark areas (such as corners of caves). In this embodiment, the contrast threshold is dynamically adjusted according to the real-time light intensity. For example, the threshold is set to 2.0 when the light intensity is <10 lux, and the threshold is set to 4.0 when the light intensity is >500 lux. When the light intensity is between 10 and 500 lux, the threshold is linearly interpolated between 2.0 and 4.0.
[0029] In this embodiment, the multimodal sensing unit further includes a light sensor, which monitors the indoor light intensity in real time and synchronizes the data to the exposure parameter control system of the HDR camera to achieve automatic gain adjustment. When the light sensor detects that the brightness of a local area is less than 10 lux, it triggers the HDR camera to extend the exposure time and start the fill light to assist imaging; when the light intensity is higher than 500 lux, it shortens the exposure to avoid overexposure.
[0030] For the thermal imaging data of the drilling face acquired by the thermal imaging camera, the infrared thermal gradient features (thermal gradients exist at rock fissures due to differences in heat dissipation) are superimposed onto the corresponding pixel positions of the RGB image using a grayscale mapping algorithm. If the texture of a certain area in the RGB image is blurred (e.g., obscured by dust), the structural details of that area (e.g., the edge of a fissure) are supplemented by the infrared thermal gradient features, ultimately outputting an enhanced visual feature map of "visible texture + infrared structure," providing detailed support for the identification of rock mass structure surfaces. Preferably, in this embodiment, the pixel coordinates of the lidar, HDR camera, and thermal imaging camera are also jointly calibrated and compensated by real-time acquisition of the pitch angle (θ) and roll angle (φ) of the drilling equipment using an tilt sensor, eliminating perception deviations caused by changes in equipment attitude.
[0031] Specifically, the calibration compensation method includes the following steps: Step A1, obtaining the translation vector (T) and initial rotation matrix (T) of the lidar, HDR camera, and thermal imaging camera relative to the tilt sensor mounting reference. The external parameter matrices (T, T, ...) of the lidar, HDR camera, and thermal imaging camera relative to the tilt sensor are calculated using the hand-eye calibration method. Step A2: Obtain the pitch angle in real time. (Rotation around the Y-axis), roll angle (Rotate around the X-axis), and construct the rotation matrix R, as shown in formulas (2), (3) and (4): ;Formula (2) ;Formula (3) ;Formula (4) Formula 2 is the rotation matrix about the X-axis (roll angle φ). Formula 3 is the rotation matrix about the Y-axis (pitch angle θ). Formula 4 is the combined attitude matrix R. Step A3: Correct the lidar coordinate system and read the original coordinates of a single point on the lidar ( ), convert to homogeneous coordinates ( ,1), multiply the original homogeneous coordinates by the attitude matrix R to obtain the corrected coordinates ( , , The translation vector T(Tx,Ty,Tz) calibrated by the extrinsic parameters is superimposed to obtain the final world coordinate system coordinates. ,in, , , (Using coordinates in the world coordinate system), repeat the above steps for all point clouds of the LiDAR to complete the overall coordinate system correction; correct the camera (including HDR and infrared cameras) coordinate system and obtain the camera intrinsic parameter matrix. The camera intrinsic parameter matrix Including focal length f and principal point coordinates Based on the camera imaging model (as shown in Formula 5), the pixel (u,v) is converted into the original 3D coordinates in the camera's original coordinate system. ): Formula (5) where, The distance to the corresponding point of this pixel as measured by the lidar; The pose matrix R is left-multiplied by the original 3D coordinates to correct the original 3D coordinates, resulting in the corrected coordinates. The same method used for LiDAR coordinate correction is employed to superimpose the translation vector T to obtain the true 3D coordinates of the structural surface feature points. In this embodiment, in step B, the LiDAR point cloud, HDR camera RGB image, and thermal imaging data are aligned in the spatiotemporal dimension according to the fusion strategy. Simultaneously, a Transformer fusion network based on a cross-attention mechanism is used to extract the cavern outline, rock mass structural surface coordinates, and environmental interference parameters. The specific process is as follows: B1. Multimodal feature extraction: For four types of input data, namely lidar point cloud, visual image, infrared thermal imaging and environmental parameters, feature extraction branches are designed to retain the core information of each modality.
[0032] Step B1-1, LiDAR Point Cloud Features: Based on VoxelNet, the point cloud data is output as a 3D feature map. Specifically, the point cloud data is divided into voxels of 0.1m × 0.1m × 0.1m. For each voxel, 8-12 dimensional local features such as "point cloud quantity, average coordinates, variance, maximum / minimum distance, and normal vector" are extracted. This transforms the sparse point cloud into a dense "voxel-feature" representation, highlighting key geometric information such as cavern wall edges and local rock protrusions. A 3-layer 3D convolutional network (Conv3D) is then used to further refine the feature map. Voxel features are convolutionally processed. The Conv3D convolution kernel size is set to 3×3×3 with a stride of 1×1×1. Padding is used to keep the feature map size constant. The first convolution layer increases the voxel features from 12 dimensions to 64 dimensions, capturing basic geometric features (such as planes and edges). The second convolution layer increases the feature dimension to 128 dimensions, enhancing local correlation features (such as curvature changes of adjacent voxels). The third convolution layer outputs a 256-dimensional 3D feature map, which includes the global geometric trends of the cave space (such as the arcuate features of the cave outline). High-level feature maps are extracted using the PointPillars extraction module; In this embodiment, firstly, pillars are constructed. The 3D space of the cave (e.g., 10m×10m×10m) is divided into uniform 2D pillar grids along the Z-axis (perpendicular to the ground). The pillar size is set to 0.1m×0.1m×H (H is the height of the cave, e.g., 5-10m). Each pillar corresponds to a 0.1m×0.1m grid on the ground, containing all point clouds (or voxel features output by VoxelNet) of that grid in the Z-axis direction. Next, feature encoding is performed within the cylinders. For the point cloud (or voxel features) within each cylinder, key features in the Z-axis direction are extracted and compressed into a fixed-dimensional feature vector. These key features include height features (average height, maximum / minimum height, and height variance of the point cloud within the cylinder, used to capture the undulations of the cavern walls and rock protrusions / concavities), density features (number of point clouds and point cloud density within the cylinder, used to reflect the reliability of the point cloud in that area, such as low density in high-dust areas), and geometric features (average normal vector and curvature of the point cloud within the cylinder, used to associate with local geometric features of VoxelNet and enhance structural surface recognition). Max-pooling is used to compress the multi-dimensional features in the Z-axis direction into 256 dimensions, and each cylinder ultimately outputs a 256-dimensional feature vector. Finally, 2D convolutional feature enhancement is performed. A lightweight 2D convolutional network is used to enhance the features of the pseudo-image. A network structure of 4 layers of 2D convolution (Conv2D) + 2 layers of max pooling (MaxPool2D) is adopted, with a convolution kernel size of 3×3 and a stride of 1×1, and a pooling kernel of 2×2 and a stride of 2×2. The 256-dimensional feature vector is gradually increased from "100×100×256" to "256×256×512", and the final output is a high-dimensional 2D feature map. Step B1-2, visual feature extraction: ResNet-50 + FPN (Feature Pyramid Network) is used to extract features from the preprocessed enhanced visual image. The multi-scale output of FPN (128×128, 256×256, 512×512) can cover the feature requirements of "rock fissures (small scale) - cave outline (large scale)", and the final output is a visual texture and structural feature map.
[0033] Step B1-3, Infrared thermal imaging feature branch: Using a lightweight CNN (MobileNetV3), extract the thermal gradient features of the face (significant differences in thermal gradient at rock fissures) and output a thermal structure feature map (size 256×256); Step B1-4, Environmental Parameter Feature Branch: The quantified dust concentration, light intensity, and equipment tilt angle are converted into 1×10 feature vectors (each parameter corresponds to 2 dimensions: real-time value + hierarchical weight), which serve as the "global interference weight signal" of the fusion network.
[0034] Step B2: Construct a two-stage Transformer encoder. Specifically, self-attention processing is performed on the lidar geometric features, visual structural features, and infrared thermal features respectively: The attention weight of each pixel in the feature map is calculated through the self-attention mechanism. For example, in this embodiment, the pixel weight of the cave edge in the lidar feature map is increased (highlighting the outline); in the visual feature map, the pixel weight of the rock fissure texture is increased (highlighting the structural surface), and finally, a single-modal feature with "key feature enhancement" is output. Environmental parameter features are introduced to construct a cross-attention matrix. For example, in this embodiment, when the dust concentration is high (>15mg / m³), the cross-attention matrix... The system automatically increases the weight of geometric features from the lidar (due to the strong dust interference resistance of lidar) and decreases the weight of visual features. When the illumination is extremely low (<10 lux), it increases the weight of infrared thermal features and decreases the weight of visual features. Through cross-attention calculation, the features extracted from lidar, HDR camera, and thermal imaging camera data are fused into a 256×256×512 multimodal fusion feature map. This feature map simultaneously contains the spatial geometry, visual texture, and thermal structure information of the cavern. Three independent output heads (cavern outline head, rock mass structure surface head, and environmental interference head) are set to decode target parameters from the fusion feature map, ensuring the extraction accuracy and real-time performance of each parameter. Step B2-1: Extract feature points of the cavern edge from the fused feature map (localized by edge detection algorithm), regress and output the 3D bounding box of the cavern (x / y / z axis coordinate range), and then combine it with the corrected lidar point cloud to reconstruct the 3D contour model of the cavern (triangular mesh density ≥1000 faces / m²). Step B2-2: First, the semantic segmentation network (U-Net structure) is used to segment three types of regions from the fused feature map: "rock matrix", "fracture", and "joint surface". Then, key point detection (such as fracture endpoints and joint surface intersections) is performed on the "fracture" and "joint surface" regions, and the three-dimensional coordinates (x / y / z) and key parameters (fracture width and joint surface dip angle) of the structural surface are output. Steps B2-3 map the "interference features" (such as point cloud noise caused by dust and visual blur caused by low light) in the fused feature map into quantified environmental interference parameters, and correct them by combining the original data of the environmental sensor. Finally, the output is "dust concentration (accuracy ±0.5mg / m³), light intensity (accuracy ±1lux), equipment vibration level (0-5 level)". When environmental interference parameters exceed a set threshold (e.g., dust concentration > 15 mg / m³), the system triggers adaptive adjustments to the sensors (e.g., increasing lidar power and sampling frequency) and simultaneously sends an "environmental interference warning" to the drilling system. In this embodiment, the specific implementation path for construction scene reconstruction in step C is as follows: Step C1: Map the extracted visual and thermal imaging features to the LiDAR coordinate system. Using the camera intrinsic (K) and extrinsic (R0, T0) matrix, map the pixel coordinates of the 2D feature map output by ResNet-50+FPN to 3D feature points in the LiDAR coordinate system. Map the temperature value (T) of each pixel in the infrared temperature feature map to the LiDAR coordinate system using the infrared sensor extrinsic matrix and bind it to the corresponding 3D point. Step C2: Using the LiDAR acquisition timestamp as a reference (set as...). The system uses a unified clock (error ≤ 1ms) to match visual features, infrared features, and attitude data at the same time. Step C3: Input the aligned three types of features into the Transformer fusion network, dynamically allocate attention weights based on environmental parameters, process them through the Transformer encoder (6 layers of self-attention + feedforward network), and output a 512-dimensional "geometry-texture-temperature" integrated feature. This feature simultaneously contains the spatial geometric information of the laser, the visual texture structure information, and the infrared temperature distribution information. Step C4: Based on the sparse laser point cloud and the integrated feature, fill in the blank areas to generate a high-density point cloud, and assign temperature and structural information to each point. Specifically, this step includes the following: Step C4-1: Retain the original sparse point cloud of the lidar (density of about 50 points / m²), and remove abnormal points determined by environmental parameters (such as isolated points caused by dust, invalid points with a distance > 50m); for the remaining valid points, bind the integrated features output in step C3 to form a "sparse feature point cloud". Step C4-2: Using the "K-Nearest Neighbors (KNN) + Feature Interpolation" algorithm, fill in the blank areas of the sparse point cloud: Divide the space into 0.05m×0.05m×0.05m grids (ensuring a density of ≥250 points / m² after filling); For each blank grid, search for its K=8 nearest neighbor sparse feature points, and extract the integrated features and coordinates of these points; Calculate the coordinates (X,Y,Z) of the filled points of the blank grid based on feature similarity weighted interpolation (the more similar the features, the higher the weight), ensuring that the filled points fit the geometric contour of the cavern (error ≤±3mm); Simultaneously interpolate the temperature value (T): Based on the temperature features of neighboring points and distance weighting, ensure the temperature distribution is continuous (avoiding abrupt changes).
[0035] Step C4-3: For the dense point cloud after point filling, add "structural labels" (such as "fracture point", "wall point", "joint surface point") to each point based on visual texture features (such as fracture edge heat map); for fracture points, further optimize the coordinate accuracy (combined with the normal vector information of laser geometric features) to ensure that the positioning accuracy of the structural surface is ≤ ±1mm; Step C5: Integrate the device's dynamic attitude data into the point cloud to eliminate accumulated errors caused by device movement / tilt, ensuring global accuracy. Figure 1 Consistency. Specifically, during the acquisition of each dense point cloud frame, the θ and φ values of the tilt sensor are recorded synchronously, and the world coordinate system coordinates of all points in that frame are corrected in real time using the attitude matrix R, as shown in formula (6): Formula (6) in which, The translation vector of the cavern origin is obtained through calibration. If the equipment moves (e.g., the drill arm extends or retracts), the origin coordinates are corrected using odometer data to ensure that the point cloud is aligned with the actual position of the cavern. When the attitude data deviation between two adjacent point cloud frames is greater than 0.5°, inter-frame registration is performed using the ICP (Iterative Closest Point) algorithm to align the overlapping areas of adjacent frames (registration error ≤ ±2mm). The registration process is constrained by attitude data to avoid registration drift (e.g., inter-frame misalignment caused by equipment tilt).
[0036] Step C6: Global map stitching and optimization. This step stitches together multiple dense point cloud frames into a complete cavern map, eliminating accumulated errors and ensuring global accuracy.
[0037] Specifically, in this embodiment, the "ICP registration + graph optimization (g2o)" scheme is adopted, as shown below: Step C6-1: Inter-frame registration. The ICP algorithm is used to align overlapping areas of adjacent frame point clouds (overlap rate ≥ 30%) to obtain the inter-frame transformation matrix. The transformation matrices of all frames are constructed into a graph model (nodes represent frames, edges represent inter-frame transformations), minimizing global reprojection error (target ≤ ±2mm) and eliminating accumulated errors. After stitching, the map covers the entire cavern construction area (e.g., 10m × 10m × 10m) without obvious stitching gaps. Step C6-2: Duplicate points are removed (one point is retained for distances < 0.01m). The temperature distribution is Gaussian smoothed (σ = 0.5) to avoid local abrupt changes. Each point cloud ultimately contains four core data categories: geometric coordinates (X, Y, Z), temperature value (T), attitude information (θ, φ), and structural labels. The coordinates of the cavern's pre-set calibration blocks are measured using a total station to verify the map's geometric accuracy (error ≤ ±5mm). The temperature accuracy is verified using an infrared thermometer (error ≤ ±0.5℃).
[0038] In this embodiment, the task mapping unit obtains the attitude of the drill arm based on the built-in IMU sensor of the drill arm and obtains the relative position of the drill bit relative to the multimodal sensing unit based on the parameters of the drill arm, such as the rotation angle and length of the drill arm. The world coordinate system is the unified coordinate system for reconstructing the map, which is consistent with the world coordinate system of the lidar and the camera. The relative position of the drill bit relative to the multimodal sensing unit is converted into the coordinates of the drill bit in the world coordinate system and mapped onto the reconstructed map. At the same time, the task mapping unit maps the drilling information of the drilling design drawing onto the reconstructed map and compares it with the collected drill arm information to determine whether the drilling is deviated.
[0039] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention. The singular forms “a,” “the,” and “the” used in this invention and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any or all possible combinations of one or more of the associated listed items.
[0040] It should be understood that although the terms first, second, third, etc., may be used in this invention to describe various information, this information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, first information may also be referred to as second information without departing from the scope of this invention, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to a determination."
[0041] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A borehole environment perception system, characterized by, The system comprises: A multi-modal perception unit acquires information of a construction environment and real-time state parameters of a drilling boom based on multiple different types of sensors; A fusion computing unit constructs a three-dimensional point cloud map of the construction environment based on the information of the construction environment and calculates a chamber profile of a drill hole and a three-dimensional structure surface state of a rock mass of a working face based on the three-dimensional point cloud map; A task mapping unit dynamically maps each drilling point to the three-dimensional point cloud map based on a pre-input drilling design drawing and a mapping position of a drill bit in the three-dimensional point cloud map calculated according to the attitude of the drilling boom.
2. The borehole environment perception system of claim 1, wherein, The multi-modal perception unit comprises a laser radar, an HDR camera and a thermal imaging camera.
3. The borehole environment perception system of claim 2, wherein, The multi-modal perception unit further comprises: A dust concentration sensor for monitoring the dust concentration of a drilling operation area in real time.
4. The borehole environment perception system of claim 3, wherein, The power of the laser radar is also dynamically adjusted based on the dust concentration, and the emission power is automatically increased when the dust concentration increases.
5. The borehole environment perception system of claim 2, wherein, The multi-modal perception unit further comprises: An illumination sensor for monitoring the illumination intensity inside the chamber in real time and synchronizing the data to the exposure parameter control system of the HDR camera.
6. The borehole environment perception system of claim 2, wherein, The method for constructing the three-dimensional point cloud map by the fusion computing unit comprises the following steps: Step A: denoising and registration of point cloud data from the laser radar, and pixel-level alignment and fusion of images from the HDR camera and the thermal imaging camera; Step B: a multi-modal feature fusion model constructed by a Transformer fusion network based on a cross-attention mechanism for weighted alignment of the point cloud and the image in the spatial and temporal dimensions; Step C: reconstruction of the construction scene by using the fused high-dimensional features to generate a dense three-dimensional point cloud map containing geometric structures and temperature distribution.
7. The borehole environment perception system of claim 6, wherein, The Transformer fusion network based on the cross-attention mechanism extracts the chamber profile, the rock mass structure surface coordinates and the environmental interference parameters, and the specific process is as follows: For the four types of input data of the laser radar point cloud, the enhanced visual image, the infrared thermal image and the environmental parameters, feature extraction branches are respectively designed to retain the core information of each modality; The attention weights of each pixel in the feature map are calculated through the self-attention mechanism, the environmental parameter features are introduced to construct a cross-attention matrix, and the features extracted from the laser radar, the HDR camera and the thermal imaging camera are fused into a multi-modal fusion feature map through cross-attention calculation, which contains the spatial geometry, visual texture and thermal structure information of the chamber; An independent output head is set to decode the target parameters from the fusion feature map to ensure the extraction accuracy and real-time performance of each parameter.
8. The drill hole environment perception system according to claim 7, characterized in that: the extracted visual features and thermal imaging features are mapped into the laser radar coordinate system; the visual features, the infrared features and the attitude data at the same time are matched through the system unified clock based on the laser radar acquisition timestamp; the aligned three types of features are input into the Transformer fusion network, the attention weights are dynamically assigned based on the environmental parameters, and the integrated features are output after processing by the Transformer encoder. Based on sparse laser point cloud and integrated features, fill in the blank area, generate high-density point cloud, and assign temperature and structure information to each point; Integrate device dynamic posture data into point cloud to eliminate accumulated errors caused by device movement / tilt; Global map stitching and optimization, stitching multiple frames of dense point cloud into a complete cave map, eliminating accumulated errors.