Vision and point cloud data fusion modeling coal face mining auxiliary method

By using a multi-sensor fusion device and a cross-modal data fusion algorithm, the problem of perception error in the spatial dimensions and connection relationships between equipment on the coal mining face was solved, realizing high-precision three-dimensional digital twin modeling and equipment identification, supporting remote control decision-making, and meeting the needs of intelligent unmanned mining.

CN120997631APending Publication Date: 2025-11-21CHINA COAL RES INST
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202511061630.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-30
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

Existing video stitching methods lack 3D point cloud data fusion on the coal mining face, resulting in large perception errors in the spatial dimensions and connection relationships between equipment, affecting the accuracy of the remote control system. Furthermore, the point cloud data lacks semantic annotation, making it difficult to identify equipment types and functional components.

Method used

By deploying a multi-sensor fusion device, the spatiotemporal synchronous acquisition of visual images and point cloud data is achieved. A cross-modal fusion algorithm is used to fuse image semantic information with point cloud geometric information to generate an attribute point cloud model containing geometry, color, and semantics. Furthermore, global pose optimization is used to eliminate accumulated errors and construct a high-precision global three-dimensional digital twin model, which can annotate the spatial dimensions and connection relationships of the device in real time.

Benefits of technology

It achieves high-precision, real-time, and visual modeling of the working space and equipment status in coal mining faces, supports remote control and decision-making, improves the accuracy of equipment identification and motion parameter calculation, and meets the needs of intelligent unmanned mining.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120997631A_ABST
    Figure CN120997631A_ABST
Patent Text Reader

Abstract

The invention provides a coal face mining auxiliary method based on visual and point cloud data fusion modeling, and the method comprises the steps: collecting the visual image data of a coal face, collecting the three-dimensional point cloud data through a point cloud sensor, and obtaining the motion state information through an inertial sensor; based on internal and external parameters and IMU data of joint calibration, mapping point cloud data to an image plane, establishing a geometric mapping relationship between three-dimensional points and pixels, generating an attribute point cloud model containing geometry, color and semantics, and generating a high-precision global three-dimensional digital twinning model; and identifying and tracking equipment instances in the digital twinborn model, and resolving motion parameters of key components based on a URDF kinematics model and a principal component analysis algorithm to support remote control decisions. According to the invention, high-precision real-time sensing and three-dimensional visualization of the working space size, the equipment connection relation and the movement position of the coal face can be realized, and the remote control precision and the response efficiency of unmanned mining are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of multi-modal data fusion perception and digital twin modeling of coal mining face, and particularly relates to a coal mining face mining auxiliary method based on visual and point cloud data fusion modeling. BACKGROUND

[0002] With the continuous promotion of intelligent transformation of the coal industry, less people and unmanned mining of the coal mining face has become a key direction to improve production efficiency and safety guarantee capability. In this technical system, real-time perception of the working environment, accurate measurement of the spatial size, and dynamic modeling of the device connection relationship constitute the core support of intelligent mining. In related technologies, through the cooperative operation of video stitching, point cloud acquisition, and inertial measurement unit (IMU), two-dimensional image display and local three-dimensional modeling of the coal mining face are initially realized. Specifically, this technology covers the whole process from data acquisition, feature extraction to spatial registration and visualization display, including key links such as high-definition image acquisition, laser radar point cloud modeling, and motion state monitoring, providing basic information support for remote control.

[0003] However, in the existing video stitching method, two-dimensional images are directly used for working space display, and three-dimensional point cloud data are not fused to obtain accurate spatial size and connection relationship between devices, which may result in large perception errors (such as ±5 cm) of key parameters such as the length of the support plate extension and the displacement of the push rod, or insufficient modeling accuracy (such as positioning error up to ±15 cm) of the device motion trajectory, thereby affecting the spatial decision-making ability and operation precision of the remote control system. In addition, although point cloud data can provide geometric information, they lack semantic annotation and are difficult to accurately identify device types and functional components, limiting their application value in complex dynamic scenarios. Therefore, it is urgent to build a multi-modal perception system that fuses visual and point cloud data to realize high-precision, real-time, and visual modeling of the working space and device state of the coal mining face. SUMMARY

[0004] The present application aims to at least partially solve one of the technical problems in the related art.

[0005] To this end, the first object of the present application is to propose a coal mining face mining auxiliary method based on visual and point cloud data fusion modeling.

[0006] The second object of the present application is to propose a coal mining face mining auxiliary device based on visual and point cloud data fusion modeling.

[0007] To achieve the above-mentioned objects, the present application proposes, in one aspect, a coal mining face mining auxiliary method based on visual and point cloud data fusion modeling, comprising:

[0008] To achieve the above purpose, the first aspect of the embodiment of the present application provides a coal mining face mining auxiliary method based on visual and point cloud data fusion modeling, comprising: S1, deploying a multi-sensor fusion device, collecting visual image data of the coal mining face through a high-definition industrial camera, collecting three-dimensional point cloud data through a point cloud sensor, and acquiring motion state information through an inertial sensor to realize the spatio-temporal synchronous collection of visual images, point cloud data and motion state; S2, based on the internal and external parameters and IMU data of joint calibration, mapping the point cloud data to the image plane, establishing the geometric mapping relationship between three-dimensional points and pixels, and fusing the image semantic information and point cloud geometric information through a cross-modal fusion algorithm to generate an attribute point cloud model containing geometry, color and semantics; S3, constructing a global pose graph and introducing structure prior constraints and closed loop constraints, using a nonlinear optimization algorithm to globally optimize the local three-dimensional data subsets collected by each sensor, eliminating cumulative errors, and generating a high-precision global three-dimensional digital twin model; S4, identifying and tracking the equipment instances in the digital twin model, calculating the motion parameters of the key components based on the URDF kinematic model and principal component analysis algorithm, and labeling the spatial dimensions and connection relationships between the equipment in real time to support remote control decision-making.

[0009] In an embodiment of the present application, the multi-sensor fusion device is deployed, the visual image data of the coal mining face is collected through a high-definition industrial camera, the three-dimensional point cloud data is collected through a point cloud sensor, and the motion state information is acquired through an inertial sensor to realize the spatio-temporal synchronous collection of visual images, point cloud data and motion state, further comprising: S11, integrating the high-definition industrial camera, the point cloud sensor and the inertial sensor through a rigid support for integrated packaging, and centering and calibrating the coordinate system of the industrial camera and the point cloud sensor to ensure that they have fixed relative poses and overlapping fields of view; S12, applying accurate time stamps to each frame of image and each group of point cloud data through the built-in micro data processor to realize microsecond-level time synchronization of visual images and point cloud data.

[0010] In an embodiment of the present application, based on the internal and external parameters and IMU data of joint calibration, the point cloud data is mapped to the image plane, the geometric mapping relationship between three-dimensional points and pixels is established, and the image semantic information and point cloud geometric information are fused through a cross-modal fusion algorithm to generate an attribute point cloud model containing geometry, color and semantics, further comprising: S21, projecting the point cloud data to the two-dimensional image plane through the camera internal and external parameter matrices, and judging the effectiveness of the projected point cloud coordinates to retain the mapping relationship within the image effective resolution range; S22, weighting and fusing the semantic features of the two-dimensional image and the geometric features of the three-dimensional point cloud through an attention mechanism to generate multi-modal fusion features, and inputting the features into an instance prediction head to output device category labels, confidence and three-dimensional instance masks.

[0011] In an embodiment of the present application, the global pose graph is constructed, the structure prior constraint and the closed loop constraint are introduced, a nonlinear optimization algorithm is used to globally optimize the local three-dimensional data subsets collected by each sensor, the accumulated error is eliminated, and a high-precision global three-dimensional digital twin model is generated, further comprising: S31, introducing structure prior constraints, including local smoothness / rigidity constraints and sequential road point constraints, wherein the sequential road point constraints use the ordered hydraulic support recognized by semantic recognition as a natural road mark, combined with the distance constraint of the known design distance of the support and the sequence integrity constraint based on the known total number of supports; S32, constructing a closed loop constraint, when using a mobile continuous scanning mode, a constraint is established between data nodes collected at the same physical location but at different times through feature matching to eliminate the accumulated drift of one-way scanning.

[0012] In an embodiment of the present application, further comprising: S5, GPU accelerated rendering buffer processing of the two-dimensional pixel coordinates input by the user on the visualization interface, instantaneous analysis into the unique instance identifier of the corresponding key component in the three-dimensional scene, and querying the spatial size and motion parameters calculated in real time according to the identifier, and superimposing and rendering the query results in text and graphic forms to the visualization interface.

[0013] To achieve the above purpose, the second aspect embodiment of the present application provides a coal mining face mining auxiliary device based on visual and point cloud data fusion modeling, comprising: a multi-source data synchronous acquisition module, used for deploying a multi-sensor fusion device, acquiring visual image data of a coal mining face through a high-definition industrial camera, acquiring three-dimensional point cloud data through a point cloud sensor, and acquiring motion state information through an inertial sensor, realizing the spatio-temporal synchronous acquisition of visual image, point cloud data and motion state; a cross-modal data fusion module, used for mapping the point cloud data to the image plane based on the internal and external parameters of joint calibration and IMU data, establishing a geometric mapping relationship between three-dimensional points and pixels, and fusing image semantic information and point cloud geometric information through a cross-modal fusion algorithm to generate an attribute point cloud model containing geometry, color and semantics; a global pose optimization module, used for constructing a global pose graph and introducing structure prior constraints and closed loop constraints, using a nonlinear optimization algorithm to globally optimize the local three-dimensional data subsets collected by each sensor, eliminating accumulated error, and generating a high-precision global three-dimensional digital twin model; a device motion parameter solving and labeling module, used for identifying and tracking device instances in the digital twin model, solving the motion parameters of key components based on a URDF kinematics model and a principal component analysis algorithm, and labeling the spatial size and connection relationship between devices in real time to support remote control decision-making.

[0014] The coal mining face mining auxiliary method and device of the visual and point cloud data fusion modeling embodiment of the present application realize high-precision real-time perception and three-dimensional visualization of the working space size, equipment connection relationship and motion position of the coal mining face, and effectively support the remote control requirements of intelligent unmanned mining.

[0015] Additional aspects and advantages of the present application will be set forth in part in the following description, will become apparent from the following description, or will be learned by practice of the present application. BRIEF DESCRIPTION OF DRAWINGS

[0016] The above and / or additional aspects and advantages of the present application will become apparent and be readily understood from the following description, taken in conjunction with the accompanying drawings, in which:

[0017] Figure 1 is a flow chart of the coal mining face mining auxiliary method of visual and point cloud data fusion modeling according to the embodiment of the present application;

[0018] Figure 2 is a hardware architecture diagram according to the embodiment of the present application;

[0019] Figure 3 is a visual-point cloud fusion sensor main body structure schematic diagram according to the embodiment of the present application;

[0020] Figure 4 is a structure diagram of the coal mining face mining auxiliary device of visual and point cloud data fusion modeling according to the embodiment of the present application. DETAILED DESCRIPTION

[0021] It should be noted that the embodiments in the present application and the features in the embodiments can be combined with each other without conflict. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0022] In order to enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor should belong to the protection scope of the present application.

[0023] The coal mining face mining auxiliary method and device of visual and point cloud data fusion modeling according to the embodiments of the present application will be described below with reference to the accompanying drawings.

[0024] Embodiment 1

[0025] Figure 1is a flowchart of a coal mining face mining auxiliary method based on visual and point cloud data fusion modeling according to an embodiment of the present application, as shown, comprising: Figure 1

[0026] S1, deploy a multi-sensor fusion device, collect visual image data of the coal mining face through a high-definition industrial camera, collect three-dimensional point cloud data through a point cloud sensor, and obtain motion state information through an inertial sensor, to realize time-space synchronous collection of visual images, point cloud data, and motion states.

[0027] Specifically, deploying a multi-sensor fusion device and realizing time-space synchronous collection of visual images, point cloud data, and motion states are key prerequisites for the present application to realize high-precision three-dimensional modeling and dynamic perception of the coal mining face. The technical implementation principle of this step is based on the collaborative work and time-space synchronization mechanism of multi-source heterogeneous sensors.

[0028] In some implementations, the multi-sensor fusion device is composed of a high-definition industrial camera, a laser radar (LiDAR), and an inertial measurement unit (IMU), which are integrated and packaged through a rigid support to ensure their fixed relative poses in space. The high-definition industrial camera uses an industrial-grade CMOS image sensor, has a resolution of 4K (3840x2160), a frame rate of up to 30 fps, and supports HDR mode to adapt to complex lighting conditions underground. The point cloud sensor is a multi-line laser radar with a resolution of 128 lines or higher, a ranging accuracy of ±2 cm, a scanning frequency of 10 Hz, and a point cloud density of up to 100,000 points per second. The IMU provides angular velocity and acceleration information with a sampling frequency of 200 Hz, an angular random walk (ARW) of less than 0.1° / √hr, and a linear acceleration error of less than 0.05 m / s 2 .

[0029] To realize time-space synchronization, the data output interfaces of the three types of sensors are connected to an embedded micro data processor (such as an embedded system based on ARM architecture), which runs a time synchronization protocol (such as PTP precise time protocol) to assign uniform timestamps to each frame of image and point cloud data, with a time synchronization accuracy of sub-millisecond (<1 ms). At the same time, through joint calibration (including internal and external parameters and space-time alignment of IMU and visual / point cloud sensors), a unified coordinate system conversion relationship is established to ensure the consistency of different modal data in the spatial dimension.

[0030] This step is deployed on the hydraulic support or mobile inspection device of the coal mining face in actual application, and synchronously transmits the data to the data storage module in real time through a network base station (such as 5G or industrial Ethernet). Its technical value lies in providing a high-precision, high-consistency raw data basis for subsequent three-dimensional modeling, equipment identification, and size measurement, which is a key supporting link for realizing unmanned and intelligent control of the coal mining face.​

[0031] Further, S1 comprises:

[0032] S11, integrating the high-definition industrial camera, the point cloud sensor, and the inertial sensor through a rigid support, and centering and calibrating the coordinate system of the industrial camera and the point cloud sensor to ensure that they have a fixed relative pose and overlapping fields of view.

[0033] Specifically, this step involves integrating the high-definition industrial camera, the point cloud sensor (such as a laser radar), and the inertial measurement unit (IMU) through a rigid support, and centering and calibrating the coordinate system of the industrial camera and the point cloud sensor to ensure that they have a fixed relative pose and overlapping fields of view. This step is the basis for multi-modal data fusion modeling and plays a key role in subsequent three-dimensional reconstruction, device recognition, and spatial dimension measurement.

[0034] In some implementations, the integrated packaging uses a rigid support made of carbon fiber composite material to ensure that the sensors can maintain a stable physical installation relationship under complex working conditions (such as vibration, impact, temperature and humidity changes). The installation positions of the industrial camera and the point cloud sensor need to meet certain geometric constraints, usually aligning the camera lens optical axis with the center axis of the laser radar scanning plane to maximize the overlapping area of their fields of view in space, generally requiring an overlapping angle of no less than 60° to ensure the accuracy of subsequent cross-modal data mapping.

[0035] Further, the centering and coordinate system calibration include camera intrinsic calibration and extrinsic calibration. Intrinsic calibration uses a checkerboard calibration board to obtain parameters such as camera focal length (fx, fy), principal point coordinates (cx, cy), distortion coefficients (k1, k2, p1, p2) through Zhang's Method, with a calibration accuracy typically reaching sub-pixel level (<0.1 pixel). Extrinsic calibration uses joint calibration methods such as multi-view acquisition based on calibration spheres or calibration targets to calculate the rotation matrix (R) and translation vector (T) between the camera coordinate system and the laser radar coordinate system, establishing a rigid transformation relationship, with a calibration error controlled within millimeters (<5 mm) to meet the needs of high-precision spatial modeling.

[0036] Optionally, the joint calibration of the IMU with the camera / laser radar involves synchronously collecting motion data and image / point cloud data, using a timestamp synchronization mechanism (such as PTP protocol or hardware triggering) to achieve time alignment, and simultaneously fusing the attitude data of the IMU with the extrinsic matrix of the camera / laser radar to improve the pose estimation accuracy in dynamic scenarios. This step ensures the strict consistency of multiple sensors in space and time dimensions, providing a reliable data foundation for subsequent point cloud and image fusion, device motion calculation, and dimension measurement.

[0037] S12, by the built-in micro data processor for each frame of image and each group of point cloud data to impose precise time stamp, realize the visual image and point cloud data time synchronization of microsecond level.

[0038] Specifically, by the built-in micro data processor for each frame of image and each group of point cloud data to impose precise time stamp, realize the visual image and point cloud data time synchronization of microsecond level. This step is the basis of data alignment and fusion in the whole visual and point cloud fusion modeling system, its technical implementation depends on the high-precision time synchronization mechanism and the coordination control of multi-sensor data acquisition.

[0039] The micro data processor usually adopts embedded system with hardware time stamp function, such as real-time processing unit based on ARM Cortex-A series or FPGA architecture. High-definition industrial camera and laser radar (LiDAR) access the processor through their respective trigger interface (such as GPIO or PTP protocol). After receiving the external synchronization signal (such as from the main control PLC), the processor immediately generates time stamp for image frame and point cloud data packet respectively. The precision of time stamp usually reaches microsecond (μs) level, which meets the time synchronization requirements of ISO 15897 standard for industrial vision and three-dimensional perception system. In some implementations, the system can also use IEEE 1588 precision time protocol (PTP) to further improve the stability and precision of time synchronization.

[0040] At the parameter index level, the image acquisition frequency is usually set to 30-60 frames / s, the point cloud data acquisition frequency is 10-30 frames / s, and the time stamp error needs to be controlled within ±1 μs to ensure that the image and point cloud data are strictly aligned in time dimension in the subsequent cross-modal data mapping and fusion process. In addition, the time stamp needs to contain sensor ID, acquisition time stamp, frame number and other metadata, which is convenient for subsequent data retrieval and processing.

[0041] At the application scenario level, this time synchronization mechanism is widely used in multi-sensor collaborative perception system of coal mining face, especially in the distributed synchronous acquisition and mobile continuous scanning modes, to ensure that the data from different positions or different times are fused and modeled under the unified time reference, thereby improving the accuracy of spatial dimension measurement and equipment state identification.

[0042] The technical effect of this step is that through microsecond-level time synchronization, the misalignment problem of visual image and point cloud data in time dimension is effectively solved, which provides reliable time consistency guarantee for subsequent cross-modal data mapping, three-dimensional modeling and dynamic target tracking, and is the key technical support for realizing high-precision, real-time and intelligent perception of coal mining face.

[0043] S2, based on the jointly calibrated intrinsic and extrinsic parameters and the IMU data, mapping the point cloud data to the image plane, establishing a geometric mapping relationship between the three-dimensional points and the pixels, and fusing the image semantic information and the point cloud geometric information through a cross-modal fusion algorithm to generate an attribute point cloud model containing geometry, color and semantics.

[0044] Specifically, this step aims to map the point cloud data to the image plane based on the jointly calibrated intrinsic and extrinsic parameters and the IMU data, establish a geometric mapping relationship between the three-dimensional points and the pixels, and fuse the image semantic information and the point cloud geometric information through a cross-modal fusion algorithm to generate an attribute point cloud model containing geometry, color and semantics. This process is a key link to realize multi-modal data collaborative perception and intelligent recognition.

[0045] In terms of technical implementation, first, the joint calibration of multiple sensors needs to be completed. Among them, the camera intrinsic parameter calibration adopts a checkerboard calibration board, and the focal length (fx, fy), principal point coordinates (cx, cy) and distortion coefficients (k1, k2, p1, p2, k3) of the camera are obtained through Zhang's Calibration Method to correct lens distortion and establish an accurate pinhole imaging model. The extrinsic parameter calibration calculates the rotation matrix (R ∈ R 3 ×3) and translation vector (T ∈ R 3 ) between the camera and the lidar through a calibration target or known spatial points to establish a rigid transformation relationship. The joint calibration of the IMU and the camera / lidar acquires the spatial transformation relationship between the IMU and the visual point cloud sensor through the synchronous acquisition of the IMU's acceleration and angular velocity data and the image / point cloud timestamps, and adopts a motion trajectory-based calibration method (such as joint calibration based on visual-inertial odometry VIO) to ensure the consistency of the motion state and the spatial data.

[0046] In the cross-modal mapping process, the point cloud data is first transformed to the camera coordinate system through the extrinsic parameter matrix, and then projected and transformed using the camera intrinsic parameter matrix to map the three-dimensional points (x, y, z) to two-dimensional pixel coordinates (u, v). The effectiveness of the depth value needs to be considered in the projection process, and usually the image resolution range is set to 1920x1080, and only the points with z values between 0.5m and 20m are retained to ensure the mapping accuracy. The mapped point cloud and the image pixels establish a one-to-one correspondence, forming a sparse point-pixel mapping table.

[0047] Further, the cross-modal fusion algorithm assigns the RGB color information and semantic segmentation results (such as semantic labels extracted by MaskR-CNN or PointNet++) in the image to the corresponding three-dimensional points. The semantic labels include device categories (such as hydraulic supports, coal mining machines), component states (such as shield plate expansion / withdrawal), etc. The attribute point cloud model generated after fusion not only has high-precision geometric structure, but also contains rich visual and semantic information, providing a data basis for subsequent instance segmentation, motion association, and size labeling.

[0048] In application scenarios, this step is widely used in remote monitoring and intelligent decision-making systems of coal mining faces, especially in device state recognition, spatial relationship modeling, and unmanned operation control. Through the attribute point cloud model after fusion, the system can realize high-precision three-dimensional reconstruction and semantic understanding of the working face equipment, improving the comprehensiveness and accuracy of environmental perception.

[0049] The technical effect of this step is reflected in: on the one hand, through accurate geometric mapping and cross-modal fusion, the visualization quality and semantic expression ability of the point cloud model are significantly improved; on the other hand, it provides structured and semantic data support for subsequent device recognition, pose calculation, and size measurement, which is a core technical link to realize intelligent unmanned mining of coal mining faces.

[0050] Further, S2 includes:

[0051] S21, project the point cloud data to the two-dimensional image plane through the camera intrinsic and extrinsic matrices, and perform validity judgment on the projected point cloud coordinates, and retain the mapping relationship within the effective resolution range of the image.

[0052] Specifically, in the method steps of the present application, the point cloud data is projected to the two-dimensional image plane through the camera intrinsic and extrinsic matrices, and the validity of the projected point cloud coordinates is judged, which is a key link to realize the cross-modal fusion of vision and point cloud data. Based on the camera imaging geometric model and the sensor joint calibration results, this step establishes the mapping relationship between three-dimensional space points and two-dimensional image pixels, thereby providing a basic support for subsequent semantic fusion, instance segmentation, and size labeling.

[0053] In terms of technical implementation, first, the intrinsic matrix K of the camera needs to be obtained, which contains the focal length f x f y , principal point coordinates c x c yand distortion coefficients k1, k2, p1, p2, etc., are usually calibrated by a checkerboard calibration board, which meets the camera calibration standards in OpenCV or ROS. The extrinsic matrix [R|t] represents the rotation and translation relationship between the LiDAR coordinate system and the camera coordinate system, which is obtained through joint calibration to ensure that the two have a fixed and accurate relative pose in space. The point cloud data is represented in the LiDAR coordinate system as P_{\text{lidar}} = (x, y, z), which is converted to the camera coordinate system by the coordinate transformation formula P_{\text{camera}} = R·P_{\text{lidar}} + t.

[0054] Subsequently, the three-dimensional points are projected onto the two-dimensional image plane using the pinhole camera model, with the formula p = K·P_{\text{camera}}, where p = (u, v) is the image pixel coordinate. After projection, the validity of the point cloud coordinates needs to be judged, i.e., checking whether u and v fall within the effective resolution range of the image (such as 1920×1080 or 4K resolution), and excluding invalid mappings caused by distortion or occlusion. This process is usually implemented in a GPU-accelerated point cloud processing framework, such as using PCL (Point Cloud Library) or Open3D for batch processing.

[0055] This step has a dual role in the system: on the one hand, it gives the point cloud image semantic information, improving the visualization and interpretability of the three-dimensional model; on the other hand, it provides a pixel-point cloud correspondence for device instance segmentation and motion parameter solving, which is the basis for realizing digital twin modeling and remote control of the coal mining face.

[0056] S22, the semantic features of the two-dimensional image are weighted and fused with the geometric features of the three-dimensional point cloud through an attention mechanism to generate multi-modal fusion features, which are input into an instance prediction head to output device class labels, confidence and three-dimensional instance masks.

[0057] Specifically, in one embodiment of the present application, the key operation is to weight and fuse the semantic features of the two-dimensional image with the geometric features of the three-dimensional point cloud through an attention mechanism to generate multi-modal fusion features, which are input into an instance prediction head to output device class labels, confidence and three-dimensional instance masks. The core of this step is to achieve efficient fusion of cross-modal features and accurate instance recognition, thereby providing a structured and semantic three-dimensional data basis for subsequent spatial analysis and dimension solving.

[0058] At the technical implementation level, the fusion process first extracts the local geometric features and global structural features of the point cloud data through a three-dimensional sparse convolutional network (such as MinkowskiNet), while simultaneously extracting the semantic features in the image using a two-dimensional convolutional neural network (such as ResNet-50 or FPN structure). To achieve cross-modal alignment, the system uses pre-calibrated camera internal and external parameters to project the three-dimensional point cloud onto the two-dimensional image plane, establishing a sparse mapping relationship between points and pixels. On this basis, an attention mechanism (such as a multi-head self-attention or cross-attention module) is used to dynamically weight the fusion of point cloud geometric features and image semantic features. Specifically, the feature vector of each three-dimensional point is calculated with the semantic feature vector of the corresponding image region to generate attention weights, thereby achieving adaptive feature fusion.

[0059] In terms of parameter indicators, image semantic features are usually 256-dimensional or 512-dimensional feature vectors, while point cloud geometric features are high-dimensional features including normal vectors, curvature, local density, etc. In the attention mechanism, the number of attention heads can be set to 8 or 16, the feature dimension is 256, and the multi-modal feature dimension after fusion is 512. The instance prediction head usually adopts the FCN or Mask R-CNN structure, outputting device class labels (such as hydraulic support, coal mining machine, scraper conveyor, etc.), confidence (0-1 interval), and three-dimensional instance mask (point set belonging to the same instance in the point cloud).

[0060] In actual application scenarios, this step runs on the global three-dimensional model of the coal mining face, supporting real-time identification and tracking of fixed equipment and moving parts (such as support plates and push rods). By fusing image semantics and point cloud geometry information, the system can effectively deal with problems such as light changes, occlusions, and geometric similarities in coal mining scenarios, significantly improving the robustness and accuracy of identification.

[0061] The technical effect of this step is to achieve high-precision instance segmentation and semantic recognition of coal mining face equipment, providing key data support for subsequent motion calculation, size labeling, and remote control. Through multi-modal feature fusion, the system not only improves the recognition accuracy (which can reach more than 95%), but also enhances the perception of device spatial position and state, which is an important technical link for realizing intelligent unmanned mining of the coal mining face.

[0062] S3, construct a global pose graph and introduce structure prior constraints and loop closure constraints, use a nonlinear optimization algorithm to globally optimize the local three-dimensional data subsets collected by each sensor, eliminate cumulative errors, and generate a high-precision global three-dimensional digital twin model.

[0063] Specifically, in the method of the present application, a global pose graph is constructed and structural prior constraints and loop closure constraints are introduced, a nonlinear optimization algorithm is used to globally optimize the local three-dimensional data subsets collected by each sensor, to eliminate cumulative errors and generate a high-precision global three-dimensional digital twin model, which is one of the core steps to realize precise modeling and dynamic perception of spatial information of the coal mining face.

[0064] In some implementations, this step first regards the local three-dimensional point cloud data collected by each sensor as a node in a graph, and the edges between nodes are established through relative pose transformation to form a global pose graph. In the graph, the nodes represent the pose of the sensor at a certain time, and the edges represent the relative transformation relationship between adjacent nodes. To improve modeling accuracy, the system introduces structural prior constraints and loop closure constraints. The structural prior constraints include local smoothness constraints and serialized road point constraints. The former suppresses local deformation caused by sensor noise or motion jitter by limiting the pose change amplitude between adjacent nodes; the latter establishes global structural constraints by using known hydraulic support design distances (such as 1.5m between adjacent supports) and total number of supports to ensure overall geometric consistency of the model.

[0065] Further, in the mobile scanning mode, the system identifies the same physical location collected by the device at different time points through feature matching, thereby establishing loop closure constraints and effectively eliminating the drift problem caused by cumulative error in single-pass scanning. In terms of parameter settings, the weight coefficient of the structural prior constraint is usually set to 0.1-0.5, and the weight of the loop closure constraint is 1.0-2.0, to balance the local and global optimization effects.

[0066] A nonlinear optimization algorithm (such as Gauss-Newton or Levenberg-Marquardt) is used to jointly optimize the poses of all nodes in the graph to solve the optimal solution that satisfies all constraints. During optimization, the system uses a factor graph optimization framework (such as g2o or Ceres Solver) to model the pose error function as a least squares problem and iteratively converges to the globally optimal pose configuration.

[0067] This step is applicable to two data acquisition modes in practical applications: distributed static acquisition (such as installing sensors on hydraulic supports) and mobile continuous scanning (such as handheld or inspection robots). Through global optimization, the system can accurately register all local point cloud data to a unified global coordinate system to generate a three-dimensional digital twin model with no internal deformation and high geometric fidelity, providing a reliable spatial reference for subsequent equipment identification, motion tracking and dimension measurement. The technical value lies in significantly improving the modeling accuracy and supporting remote control and intelligent unmanned mining of the coal mining face.

[0068] Further, S3 comprises:

[0069] S31, introducing structural priors, including local smoothness / rigidity constraints and sequential road marker constraints, where the sequential road marker constraints utilize the ordered hydraulic supports recognized by semantic identification as natural road markers, combined with distance constraints of known design spacing of supports and sequence integrity constraints based on known total number of supports.

[0070] Specifically, in one embodiment of the present application, introducing structural priors is a key step to achieve high-precision three-dimensional modeling and spatial registration. This step effectively improves the geometric consistency and topological integrity of point cloud data in the splicing process by fusing local smoothness / rigidity constraints and sequential road marker constraints.

[0071] In some implementations, the local smoothness / rigidity constraints are based on the rigid characteristics of the equipment structure in the coal mining face, for example, the hydraulic support should maintain the geometric continuity of its surface when it does not deform. This constraint limits the relative pose change amplitude between adjacent nodes by introducing a smoothness penalty term in the pose graph optimization, usually using a second-order Taylor expansion nonlinear optimization method combined with a Gauss-Newton iteration strategy to minimize the pose difference between adjacent nodes. Specifically, the change in the rotation angle of adjacent nodes should be less than 5°, and the change in the translation vector should be less than 0.1 meters to ensure that the local deformation of the spliced model is controlled within an acceptable range.

[0072] Further, the sequential road marker constraints utilize semantic recognition technology to sequentially match the hydraulic supports as natural road markers. Since the hydraulic supports have a fixed spacing (usually 1.5 to 2 meters) in design, the system identifies the three-dimensional positions of each support through image semantic segmentation algorithms (such as Mask R-CNN or PointPillars) and arranges them in sequence according to the design spacing. In the optimization process, the system introduces a distance constraint term to ensure that the actual distance between adjacent identified supports does not deviate from the design spacing by more than ±0.05 meters. At the same time, based on the known total number of supports, the system also imposes a sequence integrity constraint to prevent sequence breaks or misplacements caused by occlusion or misidentification.

[0073] This step is applicable to both distributed sensor synchronous acquisition and mobile scanning data acquisition modes in practical applications. In the static deployment mode, structural priors help eliminate splicing deviations caused by installation errors or local noise; in the mobile scanning mode, combined with closed-loop constraints, model drift caused by motion cumulative error can be significantly suppressed. By introducing structural priors, the system achieves higher geometric accuracy and semantic consistency in global modeling, providing a reliable foundation for subsequent spatial analysis and dimension solving.

[0074] S32, a closed-loop constraint is constructed, when a mobile continuous scanning mode is adopted, a constraint is established between data nodes collected at the same physical location but at different times through feature matching to eliminate the accumulated drift of a single-pass scan.

[0075] In the mobile continuous scanning mode, the core of constructing the closed-loop constraint is to identify data nodes of the same physical location scanned at different time points by the device through feature matching, and to establish a pose constraint relationship between these nodes. In a specific implementation, the system first extracts feature descriptors of point cloud data during mobile scanning, for example, using a three-dimensional feature description algorithm such as FPFH (Fast Point Feature Histograms) or SHOT (Signature of Histograms of Orientations) to perform feature extraction and matching on the point cloud in the key frame. When the device completes a reciprocating motion (such as moving from one end of the working surface to the other end and returning), the system identifies corresponding point pairs in the point cloud data of the start and end positions through feature matching to establish a closed-loop constraint. The constraint associates the end pose of the current scanning path with the start pose, forming a closed pose graph node for subsequent global optimization.

[0076] The construction of the closed-loop constraint depends on the accuracy and robustness of feature matching. In the present application, the threshold for feature matching is set to a similarity matching score of 0.75 to ensure the reliability of the matching point pairs. At the same time, to improve matching efficiency, the system uses an octree-based spatial indexing structure to divide the point cloud data into multiple local regions, and the dimension of the feature descriptor of each region is 32 or 64 dimensions to balance computational efficiency and feature expression ability. The frequency of closed-loop identification can be dynamically adjusted according to the device moving speed, and is usually set to perform closed-loop detection once every 100 meters or every 50 seconds. In addition, the weight of the closed-loop constraint is set to be higher than that of the odometry constraint in pose graph optimization to preferentially eliminate the accumulated deviation caused by IMU error or motion solution drift in a single-pass scan.

[0077] This closed-loop constraint mechanism is mainly applied to the scenario of continuous scanning of a handheld or mobile inspection device in a coal mining working surface. For example, during the advancing of the working surface, the inspection robot moves along the coal wall to collect point cloud and image data. Due to the long-term integral error of the IMU, a single-pass scan is prone to cause model deformation. Through the closed-loop constraint, the system can automatically identify and correct the accumulated error when the robot returns to the starting point, thereby ensuring the geometric consistency and spatial accuracy of the global three-dimensional model. This mechanism is particularly suitable for dynamic environment modeling in long distances and multiple cycles, such as periodic advancing of the coal mining working surface and device state monitoring.

[0078] The introduction of closed-loop constraints significantly improves the accuracy and stability of three-dimensional modeling in mobile scanning mode. By eliminating the cumulative drift in single-pass scanning, the system can generate a globally consistent model without internal deformation, providing a reliable foundation for subsequent spatial dimension measurement and equipment state recognition. This step plays a key role in error suppression and model correction in the overall technical solution, and is an important guarantee for achieving high-precision digital twin modeling and remote control of the coal mining face.

[0079] S4, identify and track the equipment instances in the digital twin model, based on the URDF kinematic model and principal component analysis algorithm, solve the motion parameters of key components, and real-time label the spatial dimensions and connection relationships between equipment to support remote control decisions.

[0080] Specifically, the system is based on the equipment instances in the digital twin model, combined with the URDF (Unified Robot Description Format) kinematic model and the principal component analysis (PCA) algorithm, to solve the motion parameters of key components, and real-time label the spatial dimensions and connection relationships between equipment to support remote control decisions. This step is the core link to realize the equipment state perception and spatial relationship modeling of the coal mining face.

[0081] In terms of technical implementation, the system first loads the corresponding URDF kinematic model from the pre-set digital twin model library through the equipment instance ID obtained by identification and tracking in step S104. The URDF model defines the geometric structure, joint type (such as rotary joint, translational joint) and its motion constraint relationship of the equipment. Then, the system performs PCA processing on the segmented key component point cloud subsets (such as the support plate, push rod), extracts their principal direction vectors, centroid coordinates and feature planes, etc. For translational joints, the system calculates the centroid vector between the connected components (such as the push rod and the support body), and projects it onto the motion axis defined in URDF, and determines the real-time extension of the joint through the vector length; for rotary joints, the real-time rotation angle is calculated through the angle between the feature plane normal vector or the principal direction vector. This process combines the geometric accuracy of point cloud and the structural semantics of URDF, achieving high-precision calculation of the equipment motion state.

[0082] In terms of parameter indicators, the PCA algorithm usually uses the singular value decomposition (SVD) method, retaining the first three principal components to cover more than 95% of the geometric variation information; the joint motion axis of the URDF model needs to be consistent with the mechanical structure of the actual equipment, with an error control within ±2°; the calculation of spatial dimensions is based on the global coordinate system after point cloud registration, with an accuracy of ±5mm, meeting the real-time and accuracy requirements of underground equipment control.

[0083] In the application scenario, this step can feedback the relative position between the coal mining machine and the hydraulic support, the expansion state of the support plate, the extension length of the push rod and other key parameters in real time, provide intuitive and accurate equipment state and spatial relationship information for the remote control personnel, and thus improve the response efficiency and operation safety of the unmanned mining system.

[0084] The technical effect of this step is that by fusing the point cloud geometric features and the URDF structure semantics, the dynamic parameter calculation and spatial relationship annotation of the key components of the coal mining equipment are realized, reliable data support is provided for subsequent remote control and intelligent decision-making, and the perception accuracy and automation level of the coal mining face are significantly improved.

[0085] The real-time perception and visualization method of the spatial size and equipment connection relationship of the coal mining face of the embodiment of the application realizes real-time and accurate perception and three-dimensional visualization display of the operation space size, equipment connection relationship and motion state of the coal mining face, and improves the remote control ability and operation safety of intelligent mining.

[0086] Also includes:

[0087] S5, GPU accelerated rendering buffer processing of the two-dimensional pixel coordinates input by the user on the visualization interface, instantaneous analysis of the unique instance identifier of the corresponding key component in the three-dimensional scene, and query of the spatial size and motion parameters calculated in real time according to the identifier, and the query result is rendered to the visualization interface in the form of text and graphics.

[0088] Specifically, in some implementations, this step involves GPU accelerated rendering buffer processing of the two-dimensional pixel coordinates input by the user on the visualization interface, and instantaneous analysis of the unique instance identifier (ID) of the corresponding key component in the three-dimensional scene, so as to realize the query and visualization superposition of the spatial size and motion parameters calculated in real time. This process is based on the graphics rendering pipeline and spatial mapping mechanism, combined with deep learning segmentation results and digital twin models, to realize high-precision, low-delay interactive visualization.

[0089] In terms of technical implementation, the system first constructs a digital twin model synchronized with the real coal mining face in a three-dimensional visualization engine, which is composed of multiple device instances, each having a unique ID identifier. When the user inputs two-dimensional pixel coordinates on the high-resolution display through mouse or touch operation, the system projects the coordinates back to the current frame's depth buffer (Depth Buffer) and instance ID buffer (Instance ID Buffer) through GPU-accelerated rendering buffer mechanism. The depth buffer records the Z value of each pixel in three-dimensional space, while the instance ID buffer maps each pixel to its three-dimensional object instance ID through frame buffer object (FBO). Through GPU parallel computing, the system can complete the mapping of pixel coordinates to three-dimensional instance ID within milliseconds.

[0090] In terms of parameter indicators, the accuracy of rendering buffer processing depends on the bit number of depth buffer (such as 24-bit or 32-bit depth buffer) and the encoding method of instance ID buffer (such as using 8-bit or 16-bit integer identifier). In this system, 32-bit floating-point depth buffer and 16-bit instance ID buffer are used to support high-precision spatial mapping and unique identification of large-scale device instances. In addition, the GPU rendering delay is controlled within 5ms, meeting the real-time interaction requirements.

[0091] In application scenarios, this step is widely used in remote monitoring and control systems of coal mining faces. The operator can click on a specific component (such as a support plate, a push rod, etc.) in the three-dimensional model, and the system will immediately return the real-time motion parameters (such as the extension amount, the rotation angle) of the component or the spatial distance between the component and other devices, and render the text box and auxiliary graphics (such as arrows, scales) to the visualization interface, achieving intuitive and accurate spatial perception.

[0092] The technical effect of this step is to achieve instantaneous response from user interaction to three-dimensional spatial information through GPU-accelerated rendering buffer mechanism, effectively improving the interaction efficiency and spatial perception accuracy of the system, and providing key visualization and data feedback support for intelligent unmanned mining of coal mining faces.

[0093] The real-time perception and visualization method of the spatial size and device connection relationship of the coal mining face according to the embodiment of the application maps the two-dimensional pixel coordinates input by the user to the unique instance identifier in the three-dimensional scene through GPU-accelerated rendering buffer processing, and real-time superimposes the spatial size and motion parameter information queried, significantly improving the interaction response speed and data correlation accuracy of the coal mining face visualization system, further enhancing the intuitiveness of remote control and the real-time nature of operation decision-making.

[0094] Embodiment 2

[0095] AsFigure 2 As shown, the coal mining face mining auxiliary method based on visual and point cloud data fusion modeling provided by the present application can be applied to a hardware device, which mainly comprises a visual-point cloud fusion sensor, a data storage module, a data analysis modeling module, a data display module, and a network base station.

[0096] As shown, the visual-point cloud fusion sensor mainly comprises a high-definition industrial camera, a point cloud sensor (taking a laser radar as an example), and an inertial sensor (IMU), wherein the high-definition industrial camera performs visual data acquisition, the laser radar performs point cloud data acquisition, and the IMU is responsible for monitoring the motion state. In order to ensure the physical consistency of the data source, the sensors are integrally packaged through a rigid support inside the sensor, and the industrial camera and the laser radar are strictly center-aligned and calibrated in the coordinate system, so as to ensure that they have a fixed relative pose and a relatively large overlapping field of view. In terms of time synchronization, the data output interfaces of the two are uniformly connected to an embedded micro data processor, which is responsible for applying accurate and uniform time stamps to the collected video frames and point cloud data, thereby realizing time synchronization. Figure 3

[0097] Exemplarily, the data storage module is mainly used for storing the data sensed by the visual-point cloud fusion sensor. The module can adopt a high-performance solid state disk array (SSD RAID) to meet the real-time and stable writing requirements of high-code rate video streams and massive point cloud data. The stored data is indexed according to the time sequence and the sensor ID, facilitating subsequent retrieval and calling.

[0098] Exemplarily, the data analysis modeling module is mainly used for modeling and splicing the stored data sensed by the visual-point cloud fusion sensor, so as to realize spatial size measurement and display of the coal mining face operation space, the device connection relationship, and the device motion position.

[0099] Exemplarily, the data display module is mainly used for displaying the coal mining face operation space, the device connection relationship, and the device motion position spatial model constructed by the data analysis modeling module. The module is usually one or more industrial-grade high-resolution displays running a specially developed visual interactive software. Users can rotate, scale, and translate the three-dimensional model on the screen through mouse or touch operations, and can observe the details of the working face from any angle. At the same time, the interface will superimpose the device recognition box, the motion trajectory, and the key size annotation information processed by the data analysis modeling module.

[0100] ​Specifically, the video image data model captured by a high-definition industrial camera can be displayed, the three-dimensional space point cloud data model constructed by a point cloud sensor can be displayed, and the video image data and three-dimensional space point cloud data fusion model can be displayed. In addition, the space size, space coordinates of the mining face operation space, equipment connection relationship and equipment movement position can be displayed, thereby providing intuitive visual perception and actual three-dimensional space size information for the mining face operator.

[0101] Embodiment 3

[0102] The hardware device provided by the application performs the method of sensing the mining face operation space, the equipment connection relationship and the equipment movement position, and specifically includes the following steps:

[0103] Among them, the visual and point cloud data fusion modeling mining face intelligent mining auxiliary device arranged on the mining face can be arranged in two different ways. One is to arrange one visual and point cloud data fusion modeling mining face intelligent mining auxiliary device on each interval of several hydraulic supports, and ensure that the sensing ranges of adjacent visual and point cloud data fusion modeling mining face intelligent mining auxiliary devices have a certain overlapping area, which is convenient for subsequent splicing. The other way is to manually hold (or install on a mobile inspection device) one visual and point cloud data fusion modeling mining face intelligent mining auxiliary device to scan the mining face operation space, the equipment connection relationship and the equipment movement position.

[0104] Among them, the visual and point cloud data fusion modeling: this step aims to upgrade single sensor data to multi-attribute data containing geometry, color and semantics.

[0105] First, multi-sensor joint calibration is performed. The camera internal parameter calibration of the visual sensor is performed to obtain its internal imaging parameters and correct the lens distortion. Then, the external parameter calibration of the visual sensor and the laser radar sensor is performed to accurately calculate the rotation matrix and translation vector between the coordinate systems of the two, and establish the rigid spatial transformation relationship. In addition to the external parameter calibration of the camera and the radar, joint calibration of the IMU and the laser radar / camera is also included.

[0106] Then, cross-modal data mapping is performed. Based on the internal and external parameter matrices obtained by joint calibration, the three-dimensional point cloud data in the laser radar coordinate system is projected to the two-dimensional image plane through coordinate system transformation and camera imaging model. The validity of the projected coordinates is judged, and the mapping relationship within the effective resolution range of the image is retained, thereby establishing a one-way sparse mapping from three-dimensional points to pixels.

[0107] Finally, multi-modal data fusion is performed to assign information in the two-dimensional image data to the corresponding three-dimensional point cloud using the mapping relationship. The fusion includes: assigning the RGB color information of the pixels to the three-dimensional points to generate a color point cloud; and / or assigning the pixel-level semantic labels obtained through the image semantic segmentation algorithm to the three-dimensional points to generate an attribute point cloud with high-level semantic information.

[0108] In the step of multi-sensor data stitching and global modeling, the plurality of locally three-dimensional data subsets that are discrete in space or time are accurately registered and fused into a globally consistent three-dimensional model, and the framework can uniformly process two different data acquisition modes of distributed synchronous acquisition and mobile continuous scanning.

[0109] First, local relative pose estimation is performed to calculate the relative spatial pose between any two local three-dimensional data subsets that have spatial overlap. In the distributed synchronous acquisition mode (e.g., sensors are installed on each hydraulic support), this step aims to calculate the accurate transformation relationship between two adjacent static acquisition units in space. Coarse registration is performed using the physical installation position, and then precise registration is performed by executing the iterative closest point (ICP) algorithm on the point cloud in the overlapping region of the field of view. In the mobile continuous scanning mode (e.g., handheld or airborne inspection), high-frequency motion data provided by the IMU are used for motion solving to provide an accurate initial pose guess value for point cloud registration. A key frame-based strategy is adopted to calculate the relative pose transformation between the current key frame and the previous key frame or local sub-map.

[0110] Then, global pose graph construction and optimization are performed. To eliminate the cumulative error at a long distance, this step constructs a global pose graph and performs joint optimization.

[0111] The pose of each independent local three-dimensional data subset (whether from a static unit or a key frame of mobile scanning) is abstracted as a node of a graph. Multi-element constraints are added between the nodes, including: 1) odometry constraints: relative pose constraints calculated in the above step, which ensure the local consistency of the model; 2) structure prior constraints, including local smoothness / rigidity constraints that require the surface constructed by locally adjacent key nodes to remain smooth or rigid to filter noise and suppress unreasonable local bending, and sequential road marker point constraints that use ordered hydraulic supports identified through semantic recognition as natural road markers, combined with distance constraints based on the known design distance of the supports and total number and sequence integrity constraints based on the known total number of supports; 3) loop closure constraints, which are specifically established between data nodes collected at the same physical location but at different times through feature matching in the mobile scanning mode, to eliminate the cumulative drift of a single trip;

[0112] A nonlinear optimization algorithm is used to solve the poses of all nodes in the graph in batch, and an optimal solution that satisfies all constraints is found.

[0113] Finally, according to the accurate global pose obtained after optimization, all fused data frames are transformed into a unified global coordinate system to generate a high-precision, internal deformation-free, and true global three-dimensional model reflecting the geometric shape of the working face.

[0114] Among them, image recognition, tracking and size marking include the following steps:

[0115] (1) Three-dimensional scene instance segmentation: Obtain the time-synchronized two-dimensional image and three-dimensional point cloud data of the coal mining face collected by the multi-sensor stitching system in real time. Using the pre-calibrated camera internal and external parameters, project each frame of three-dimensional point cloud to the corresponding two-dimensional image plane to establish the geometric mapping relationship between three-dimensional space points and two-dimensional pixels. Input the three-dimensional point cloud into the three-dimensional sparse convolution network to extract the local and global geometric structure features of the point cloud. Input the two-dimensional color image into the two-dimensional convolutional neural network to extract the semantic features containing color, texture, details, etc. Design a cross-modal fusion module. For each three-dimensional point, find its corresponding position on the two-dimensional feature map through the geometric mapping relationship between three-dimensional space points and two-dimensional pixels. Use the attention mechanism method to weight and fuse the geometric features of the three-dimensional point and its corresponding two-dimensional semantic features to generate multi-modal fusion features containing accurate spatial information and rich visual texture. Input the multi-modal fusion features into the instance prediction head, which outputs the class label, confidence and three-dimensional instance mask of each physical entity in the coal mining face scene in parallel. Perform non-maximum suppression on the instance prediction to eliminate redundancy, and associate the same physical entity with a unique instance identifier (ID) based on the timing information, finally generating a frame of accurately segmented and semantically rich three-dimensional data of the coal mining face.

[0116] (2) Key component motion association: According to the generated unique instance identifier, search and load the corresponding model from the preset digital twin model library containing high-precision CAD models of working face three-machine equipment and URDF kinematics definition files. Use a two-stage registration algorithm, first use global feature descriptors for coarse registration, and then use point-to-plane ICP algorithm for fine registration to achieve accurate alignment of the instance-based point cloud data generated in step one, which combines two-dimensional image texture and three-dimensional geometric features, where two-dimensional texture information helps to solve the registration ambiguity of similar geometric components. Through nearest neighbor query, the pre-defined function component labels on the CAD model are accurately migrated to the corresponding subset of the registered point cloud data, and the URDF kinematics model is bound to the equipment instance, thereby generating structured three-dimensional scene data with key functional components identified and associated with their kinematic constraints.

[0117] (3) Dynamic target pose solving: For mobile devices that need to determine the global pose (such as a coal mining machine), the point cloud instance segmented out in the current frame is registered with the corresponding point cloud instance of the previous frame through frame-to-frame ICP registration to solve the six-degree-of-freedom pose in the global coordinate system.

[0118] For the point cloud subset of each identified key component (such as the guard plate and the push rod), the main direction vector, centroid, geometric center, or feature plane of the key geometric parameters is fitted through principal component analysis or RANSAC algorithm. Based on the joint type defined by the URDF kinematic model, the geometric parameters are solved:

[0119] a) For a translational joint, the centroid vectors of the two components connected to the joint (such as the push rod and the support body) are calculated, and the vector is projected onto the motion axis defined by URDF. The length of the projection vector is the real-time extension of the joint.

[0120] b) For a rotational joint: the feature plane normal vector or the main direction vector of the two components connected to the joint is calculated, and the angle between them is the real-time rotation angle of the joint.

[0121] All joint state parameters (such as extension and rotation angle) and the overall pose of the device are integrated to output a set of accurate state parameters that describe the complete geometric shape and spatial position of the device at the current time.

[0122] (4) Real-time dimension information labeling:

[0123] In the three-dimensional visualization engine, the global pose of each device instance and the state of each joint inside it in the digital twin model are updated in real time according to the accurate state parameters. This step converts the geometric and kinematic state solved in the background into a dynamic three-dimensional virtual model that is synchronized with the real scene and can be observed and interacted with by users.

[0124] Receive the pick-up instruction input by the user on the dynamic three-dimensional virtual model through the visualization interface in the form of two-dimensional pixel coordinates; use GPU-accelerated rendering buffer technology to instantly analyze the two-dimensional pixel coordinates into the unique instance identifier of the corresponding key component in the three-dimensional scene; according to the unique instance identifier, query the real-time data of the component from the accurate state parameters, such as the calculated extension or the spatial distance between other components calculated through the real-time pose. According to the unique instance identifier, query the real-time data of the component from the accurate state parameters, such as the calculated extension or the spatial distance calculated through the real-time pose; the key dimensions queried and calculated are rendered to the area associated with the user's pick-up position on the visualization interface in the form of text and auxiliary graphics, and change in real time according to the update of the state parameters.

[0125] In summary, the application deploys multiple vision-point cloud fusion sensors on the coal mining face, synchronously collects high-definition vision images and three-dimensional point cloud data of the working face. The microprocessor inside the sensor stamps accurate time stamps on each frame of image and each set of point cloud data, and transmits them to the data storage module in real time through the network base station. The data analysis and modeling module reads the latest data from the data storage module. The vision image is corrected for distortion, and then the point cloud data is projected onto the image plane using the pre-calibrated internal and external parameters, giving high-precision three-dimensional point cloud data high-definition RGB color information, generating a color three-dimensional point cloud model with real texture. For data from multiple sensors, image stitching algorithms and point cloud registration algorithms are called. The vision images of each sensor are spliced into a panoramic image of the working face, and the local color point cloud models of each sensor are registered and spliced into a global three-dimensional scene model in a unified coordinate system. On the spliced panoramic image and global model, run the deep learning target detection and tracking algorithm. Real-time recognition of hydraulic support, coal mining machine, scraper conveyor and other main equipment, and continuous tracking of key moving parts such as the support plate and the cutting drum, obtain their real-time position and attitude in two-dimensional image and three-dimensional space. According to the three-dimensional coordinates of the tracked equipment and parts, the spatial position relationship and key dimensions between them are calculated in real time. For example, calculate and judge the retraction state of the hydraulic support plate, the extension length of the push rod, and the position of the coal mining machine relative to the support and conveyor, and output these structured analysis data. Send the global three-dimensional model, equipment recognition results, motion trajectory and calculated key dimension data to the data display module. Perform three-dimensional visualization on the display screen. The system also listens to the user's interactive operation, and when the user clicks on a specific device or part in the model, the object is immediately highlighted and detailed size and state information is automatically popped up.

[0126] Embodiment 4

[0127] To achieve the above-mentioned embodiments, as Figure 4 shown, the coal mining face mining auxiliary device 10 for vision and point cloud data fusion modeling is also provided in the embodiment, which comprises:

[0128] The multi-source data synchronous acquisition module 100 is used to deploy a multi-sensor fusion device, acquire vision image data of the coal mining face through a high-definition industrial camera, acquire three-dimensional point cloud data through a point cloud sensor, and acquire motion state information through an inertial sensor, to realize the spatio-temporal synchronous acquisition of vision image, point cloud data and motion state;

[0129] The cross-modal data fusion module 200 is configured to map the point cloud data to an image plane based on the jointly calibrated internal and external parameters and the IMU data, establish a geometric mapping relationship between three-dimensional points and pixels, and fuse image semantic information and point cloud geometric information through a cross-modal fusion algorithm to generate an attribute point cloud model containing geometry, color, and semantics.

[0130] The global pose optimization module 300 is configured to construct a global pose graph and introduce structure prior constraints and loop constraints, and use a nonlinear optimization algorithm to globally optimize the local three-dimensional data subsets collected by each sensor, eliminate accumulated errors, and generate a high-precision global three-dimensional digital twin model.

[0131] The equipment motion parameter solving and labeling module 400 is configured to identify and track equipment instances in the digital twin model, solve the motion parameters of key components based on a URDF kinematics model and a principal component analysis algorithm, and label the spatial dimensions and connection relationships between equipment in real time to support remote control decisions.

[0132] Further, the multi-source data synchronous acquisition module is further configured to:

[0133] The high-definition industrial camera, the point cloud sensor, and the inertial sensor are integrally packaged through a rigid support, and the industrial camera and the point cloud sensor are centrally aligned and calibrated in the coordinate system to ensure that they have a fixed relative pose and overlapping fields of view.

[0134] A built-in micro data processor is used to apply accurate timestamps to each frame of image and each set of point cloud data, achieving microsecond-level time synchronization of the visual image and the point cloud data.

[0135] Further, the cross-modal data fusion module is further configured to:

[0136] The point cloud data is projected to a two-dimensional image plane through a camera internal parameter and an external parameter matrix, and the projected point cloud coordinates are judged for effectiveness, and the mapping relationship within the effective resolution range of the image is retained.

[0137] The semantic features of the two-dimensional image are weighted and fused with the geometric features of the three-dimensional point cloud through an attention mechanism to generate multi-modal fusion features, which are input to an instance prediction head to output equipment category labels, confidence, and three-dimensional instance masks.

[0138] Further, the global pose optimization module is further configured to:

[0139] The structure prior constraints are introduced, including local smoothness / rigidity constraints and serialized road marker point constraints, wherein the serialized road marker point constraints use the ordered hydraulic support identified through semantic recognition as a natural road marker, and combine the distance constraint of the known design distance of the support and the sequence integrity constraint based on the known total number of the support.

[0140] Construct a closed loop constraint, when using mobile continuous scanning mode, identify the data nodes collected at the same physical location but different times through feature matching to establish a constraint between them, to eliminate the cumulative drift of single-pass scanning.

[0141] Further comprising:

[0142] A visualization interaction module is configured to perform GPU-accelerated rendering buffer processing on the two-dimensional pixel coordinates input by the user on the visualization interface, instantaneously resolve the unique instance identifier of the corresponding key component in the three-dimensional scene, and query the spatial size and motion parameters calculated in real time according to the identifier, and superimpose and render the query results in the form of text and graphics to the visualization interface.

[0143] The visual and point cloud data fusion modeling mining auxiliary device for the coal mining face according to the embodiment of the present application realizes high-precision real-time perception and three-dimensional visualization of the working space size, equipment connection relationship and motion position of the coal mining face, and significantly improves the remote control precision and response efficiency of unmanned mining.

[0144] In the description of the present specification, the description referring to the terms "one embodiment", "some embodiments", "an example", "a specific example", or "some examples" and the like means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In the present specification, the illustrative description of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any appropriate manner in any one or more embodiments or examples. In addition, the person skilled in the art can combine and combine the different embodiments or examples described in the present specification and the features of the different embodiments or examples without contradiction.

[0145] In addition, the terms "first", "second" are only for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of indicated technical features. Therefore, the features defined with "first", "second" can explicitly or implicitly include at least one of the features. In the description of the present application, the meaning of "a plurality of" is at least two, for example, two, three, etc., unless otherwise explicitly specified.

Claims

1. A mining auxiliary method for a coal mining face based on visual and point cloud data fusion modeling, characterized in that, Comprise: S1, deploy multi-sensor fusion device, collect visual image data of coal mining face through high-definition industrial camera, collect three-dimensional point cloud data through point cloud sensor, and obtain motion state information through inertial sensor, realize space-time synchronous collection of visual image, point cloud data and motion state; S2, based on joint calibration of internal and external parameters and IMU data, map point cloud data to image plane, establish geometric mapping relationship between three-dimensional points and pixels, and fuse image semantic information and point cloud geometric information through cross-modal fusion algorithm, generate attribute point cloud model containing geometry, color and semantics; S3, construct global pose graph and introduce structure priori constraint and closed loop constraint, use nonlinear optimization algorithm to optimize local three-dimensional data subset collected by each sensor, eliminate cumulative error, and generate high-precision global three-dimensional digital twin model; S4, identify and track equipment instances in digital twin model, solve motion parameters of key components based on URDF kinematics model and principal component analysis algorithm, and label space size and connection relationship between equipment in real time to support remote control decision.

2. The method of claim 1, wherein, The S1 further comprises: S11, integrate high-definition industrial camera, point cloud sensor and inertial sensor through rigid support, and center align and calibrate coordinate system of industrial camera and point cloud sensor to ensure fixed relative pose and overlapping field of view of them; S12, apply accurate time stamp to each frame of image and each group of point cloud data through built-in micro data processor, realize microsecond level time synchronization of visual image and point cloud data.

3. The method of claim 1, wherein, The S2 further comprises: S21, project point cloud data to two-dimensional image plane through camera internal and external parameter matrix, and judge effectiveness of projected point cloud coordinates, and retain mapping relationship within effective resolution range of image; S22, weight fuse semantic features of two-dimensional image and geometric features of three-dimensional point cloud through attention mechanism, generate multi-modal fusion features, and input to instance prediction head to output equipment category label, confidence and three-dimensional instance mask.

4. The method of claim 1, wherein, The S3 further comprises: S31, introduce structure priori constraint, including local smoothness / rigidity constraint and serialized road point constraint, wherein serialized road point constraint uses ordered hydraulic support recognized through semantic identification as natural road sign, combines distance constraint of known design distance of support and sequence integrity constraint based on known total number of support; S32, construct closed loop constraint, when mobile continuous scanning mode is adopted, establish constraint between data nodes collected at same physical position but different time through feature matching to eliminate cumulative drift of one-way scanning.

5. The method of claim 1, wherein, Further comprise: S5, perform GPU accelerated rendering buffer processing on two-dimensional pixel coordinates input by user on visualization interface, instantly analyze into unique instance identifier of corresponding key component in three-dimensional scene, query spatial size and motion parameters solved in real time according to the identifier, and superimpose and render query result to visualization interface in form of text and graphics.

6. A mining auxiliary device for a coal mining face based on visual and point cloud data fusion modeling, characterized in that, Comprise: The multi-source data synchronous acquisition module is used for deploying a multi-sensor fusion device, acquiring visual image data of a coal mining face through a high-definition industrial camera, acquiring three-dimensional point cloud data through a point cloud sensor, and acquiring motion state information through an inertial sensor, so as to realize the spatio-temporal synchronous acquisition of visual images, point cloud data and motion states; The cross-modal data fusion module is used for mapping the point cloud data to an image plane based on the joint calibrated internal and external parameters and IMU data, establishing a geometric mapping relationship between three-dimensional points and pixels, and fusing image semantic information and point cloud geometric information through a cross-modal fusion algorithm to generate an attribute point cloud model containing geometry, color and semantics; The global pose optimization module is used for constructing a global pose graph and introducing structure prior constraints and loop constraints, using a nonlinear optimization algorithm to globally optimize local three-dimensional data subsets collected by each sensor, eliminating cumulative errors, and generating a high-precision global three-dimensional digital twin model; The device motion parameter solving and labeling module is used for identifying and tracking device instances in the digital twin model, solving the motion parameters of key components based on a URDF kinematic model and a principal component analysis algorithm, and labeling the spatial dimensions and connection relationships between devices in real time to support remote control decisions.

7. The apparatus of claim 6, wherein, The multi-source data synchronous acquisition module is also used for: integrating the high-definition industrial camera, the point cloud sensor and the inertial sensor through a rigid support for integrated packaging, and centering and calibrating the coordinate system of the industrial camera and the point cloud sensor to ensure that they have a fixed relative pose and overlapping fields of view; applying accurate timestamps to each frame of image and each set of point cloud data through a built-in micro data processor to realize microsecond-level time synchronization of visual images and point cloud data.

8. The apparatus of claim 6, wherein, The cross-modal data fusion module is also used for: projecting the point cloud data to a two-dimensional image plane through camera internal and external parameter matrices, and judging the effectiveness of the projected point cloud coordinates to retain the mapping relationship within the effective resolution range of the image; weighting and fusing the semantic features of the two-dimensional image and the geometric features of the three-dimensional point cloud through an attention mechanism to generate multi-modal fusion features, and inputting the multi-modal fusion features into an instance prediction head to output device category labels, confidence and three-dimensional instance masks.

9. The apparatus of claim 6, wherein, The global pose optimization module is also used for: introducing structure prior constraints, including local smoothness / rigidity constraints and serialized road point constraints, wherein the serialized road point constraints use the ordered hydraulic supports recognized through semantic recognition as natural road markers, combined with the distance constraint of the known design distance of the supports and the sequence integrity constraint based on the known total number of supports; constructing loop constraints, when a mobile continuous scanning mode is used, establishing constraints between data nodes collected at the same physical location but different times through feature matching to eliminate the cumulative drift of single-pass scanning.

10. The apparatus of claim 6, wherein, Further comprising: The visual interaction module is used for GPU accelerated rendering buffer processing of two-dimensional pixel coordinates input by the user on the visual interface, instantaneous resolution into unique instance identifiers of corresponding key components in a three-dimensional scene, and query of spatial dimensions and motion parameters calculated in real time according to the identifiers, and the query results are superimposed and rendered to the visual interface in the form of text and graphics.

Citation Information

Cited By

  • Tunnel steel arch intelligent installation precision regulation and control method and system based on machine vision and computer readable storage medium

    CN121593831A

  • High-precision steel structure geometric dimension measuring and positioning method

    CN121782998A

  • Power channel lightweight semantic point cloud reconstruction and visualization method and system oriented to tree barrier analysis

    CN122289982A