A fully mechanized coal mining face roof bending deformation multi-modal intelligent identification method and system

By combining cameras and lidar in a multimodal intelligent recognition method, real-time and accurate identification of roof bending deformation in coal mines has been achieved, solving the problem of early identification of roof deformation in existing technologies and improving the accuracy and coverage of early warning.

CN121147643BActive Publication Date: 2026-02-10BEIJING HONGBO YATAI ELECTRICAL EQUIP CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511685851.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-18
Publication Date
2026-02-10
Estimated Expiration
2045-11-18

AI Technical Summary

Technical Problem

In fully mechanized coal mining faces, early detection and accurate monitoring of roof bending deformation are difficult, especially in complex environments with darkness, high dust levels, and limited equipment installation. Existing technologies cannot effectively identify minute cracks and precursors to delamination, making it difficult to predict potential safety hazards.

Method used

A multimodal intelligent recognition method is adopted. The panoramic image of the roof is acquired by a camera to identify two-dimensional deformation features. Three-dimensional point cloud data is acquired by LiDAR. Stable points are selected and clustered using information entropy and Euclidean distance. Point cloud registration is performed by combining ICP algorithm. Finally, multimodal data fusion and hierarchical early warning are carried out.

Benefits of technology

It enables real-time and accurate identification of roof deformation in coal mines, improves the accuracy and coverage of early warning, solves the dilemma of "not being able to see everything, not being able to identify accurately, and not being able to judge quickly" in existing technologies, and improves the real-time performance and accuracy of safety monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121147643B_ABST
    Figure CN121147643B_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of coal mine underground safety production monitoring, and discloses a fully mechanized coal face roof bending deformation multi-modal intelligent identification method and system, which realizes fully mechanized coal face roof deformation monitoring and early warning by fusing two-dimensional images and three-dimensional point cloud data. First, adjacent camera images are spliced to form a roof panorama, and an improved YOLOv8 model is used to identify two-dimensional deformation features such as cracks; simultaneously, the real-time collected point cloud data is processed, stable key points are obtained through information entropy screening and clustering, and the ICP algorithm is used for registration with the reference point cloud to obtain three-dimensional deformation features of the roof; then, the two-dimensional and three-dimensional deformation features are aligned and bidirectionally verified to obtain a fusion result; finally, the result is compared with a preset threshold to trigger a graded early warning. The method and system realize multi-modal accurate perception and safety early warning of the roof state.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of coal mine underground safety production monitoring technology, and in particular relates to a multimodal intelligent identification method and system for bending deformation of the roof of a fully mechanized mining face. Background Technology

[0002] In fully mechanized coal mining faces, roof bending and deformation is a major hidden danger leading to serious accidents such as roof falls and spalling. Achieving early and accurate identification of roof deformation, especially micro-cracks and precursors to delamination, has long been a key technical challenge that has remained unresolved in the field of underground safety monitoring. This challenge stems from the extremely unique working environment underground: continuous darkness, high concentrations of dust, severely limited equipment installation space, and dynamically changing mining conditions. These factors collectively pose a severe test to existing monitoring technology systems.

[0003] Therefore, underground roof safety monitoring has long faced the dilemma of "not being able to see everything, not being able to identify accurately, and not being able to judge quickly," and there is an urgent need for a multimodal intelligent identification method and system for the bending deformation of the fully mechanized mining face roof to solve the above problems. Summary of the Invention

[0004] Therefore, it is necessary to provide a multimodal intelligent identification method and system for bending deformation of the roof of a fully mechanized mining face to address the above-mentioned technical problems.

[0005] Firstly, this application provides a multimodal intelligent identification method for the bending deformation of the roof of a fully mechanized mining face, including:

[0006] A first region image and a second region image of the roof of the fully mechanized mining face are obtained. The overlapping region images of the first region image and the second region image are stitched together to form a panoramic image of the roof of the fully mechanized mining face. The bending deformation feature of the panoramic image of the roof of the fully mechanized mining face is identified by a detection model to obtain the two-dimensional deformation feature of the roof of the fully mechanized mining face.

[0007] The real-time point cloud of the fully mechanized mining face roof is obtained. The information entropy is used to filter out structurally stable low-entropy points from the real-time point cloud. The Euclidean distance is used to perform spatial clustering of the low-entropy points to form key points. Then, the ICP algorithm is used to perform point cloud registration with the key points in the reference point cloud to obtain the three-dimensional deformation features of the fully mechanized mining face roof.

[0008] Align the two-dimensional deformation features and the three-dimensional deformation features of the roof of the fully mechanized mining face, and perform bidirectional feature mutual verification to obtain the multimodal data fusion result.

[0009] The results of multimodal data fusion are compared with preset deformation thresholds, and a graded early warning is triggered based on the comparison results.

[0010] Preferably, the steps of acquiring a first region image and a second region image of the fully mechanized mining face roof, stitching together the overlapping region images of the first and second region images to form a panoramic image of the fully mechanized mining face roof, and using a detection model to identify bending deformation features of the panoramic image of the fully mechanized mining face roof to obtain the two-dimensional deformation features of the fully mechanized mining face roof include:

[0011] The Retinex algorithm is used to adjust the brightness and uniformity of images acquired by adjacent cameras, and bilateral filtering is performed to obtain the processed images acquired by adjacent cameras.

[0012] Using a brute-force matching algorithm and a k-nearest neighbor matching algorithm, images acquired by adjacent cameras are sequentially sorted to obtain a sequentially sorted image sequence of all cameras acquired at the same time. The image sequence includes at least a first region image and a second region image.

[0013] The image sequence is subjected to grayscale enhancement and noise suppression processing, and then input into a regression network to obtain coarsely aligned images at the same time.

[0014] SIFT feature points are extracted from the coarsely aligned image, and outliers in the SIFT feature point extraction are removed by random sampling consensus algorithm. Then, projection matrix transformation is performed to obtain a seamlessly stitched panoramic image.

[0015] The seamlessly stitched panoramic image is corrected and rectangularized to obtain a panoramic image of the roof of the fully mechanized mining face.

[0016] Preferably, the step of correcting and rectangularizing the seamlessly stitched panoramic image to obtain a panoramic image of the fully mechanized mining face roof includes the following:

[0017] The target YOLOv8 detection model is obtained by introducing the CBAM attention mechanism into the backbone layer of the YOLOv8 detection model and using the CIoU loss function in the YOLOv8 detection model.

[0018] Using the target YOLOv8 detection model, the bending deformation features of the panoramic image of the fully mechanized mining face roof are identified to obtain the two-dimensional deformation features of the fully mechanized mining face roof.

[0019] Preferably, the steps of acquiring the real-time point cloud of the fully mechanized mining face roof, filtering out structurally stable low-entropy points from the real-time point cloud using information entropy, spatially clustering the low-entropy points using Euclidean distance to form key points, and then using the ICP algorithm to perform point cloud registration with key points in the reference point cloud to obtain the three-dimensional deformation features of the fully mechanized mining face roof include:

[0020] The real-time point cloud of the roof of the fully mechanized mining face is obtained, and the real-time point cloud is downsampled by a voxel grid to obtain the downsampled real-time point cloud.

[0021] The distribution of the neighborhood distance of the downsampled real-time point cloud is statistically analyzed to form neighborhood points. Abnormal points whose average distance from the neighborhood points is greater than a preset multiple are removed to obtain real-time point cloud data after removing interference points.

[0022] Feature extraction is performed on the real-time point cloud data after removing interference points to obtain point cloud feature data;

[0023] Information entropy is used to filter out structurally stable low-entropy points from the point cloud feature data, and Euclidean distance is used to perform spatial clustering of the low-entropy points to form key points. Then, the ICP algorithm is used to perform point cloud registration with key points in the reference point cloud to obtain the three-dimensional deformation features of the roof of the fully mechanized mining face.

[0024] Preferably, the step of aligning the two-dimensional deformation features and the three-dimensional deformation features of the fully mechanized mining face roof, and performing bidirectional feature cross-verification to obtain the multimodal data fusion result includes:

[0025] Based on the intrinsic parameter matrix of the camera and the extrinsic parameter matrix of the explosion-proof LiDAR, coordinate alignment is performed to map the two-dimensional deformation features of the roof of the fully mechanized mining face to the three-dimensional deformation features of the roof of the fully mechanized mining face, and establish the correspondence between pixels and point clouds.

[0026] When the visual detection confidence level formed based on the two-dimensional deformation features and the three-dimensional deformation features of the fully mechanized mining face roof is greater than or equal to a preset two-dimensional value, and the corresponding point cloud deformation is greater than a preset three-dimensional value, a multimodal data fusion result is obtained.

[0027] Preferably, the step of comparing the multimodal data fusion result with a preset deformation threshold and triggering a graded early warning based on the comparison result further includes:

[0028] The multimodal data fusion result is compared with a preset deformation threshold to obtain a comparison result;

[0029] When the comparison result meets the preset warning conditions, a tiered warning is triggered;

[0030] If the comparison result shows that the preset warning conditions are not met, the monitoring of the roof of the fully mechanized mining face will continue.

[0031] Secondly, this application provides a multimodal intelligent identification system for the bending deformation of the roof of a fully mechanized mining face, applied to the aforementioned multimodal intelligent identification method for the bending deformation of the roof of a fully mechanized mining face, the system comprising:

[0032] The visual inspection unit is used to acquire a first region image and a second region image of the roof of the fully mechanized mining face, stitch the overlapping region image of the first region image and the second region image together to form a panoramic image of the roof of the fully mechanized mining face, and use the detection model to identify the bending deformation features of the panoramic image of the roof of the fully mechanized mining face to obtain the two-dimensional deformation features of the roof of the fully mechanized mining face.

[0033] The lidar detection unit is used to acquire the real-time point cloud of the fully mechanized mining face roof. It uses information entropy to filter out structurally stable low-entropy points from the real-time point cloud, and uses Euclidean distance to perform spatial clustering of the low-entropy points to form key points. Then, it uses the ICP algorithm to perform point cloud registration with key points in the reference point cloud to obtain the three-dimensional deformation features of the fully mechanized mining face roof.

[0034] The data fusion unit is used to align the two-dimensional deformation features and the three-dimensional deformation features of the roof of the fully mechanized mining face, and to perform bidirectional feature mutual verification to obtain multimodal data fusion results.

[0035] The early warning linkage unit is used to compare the multimodal data fusion results with the preset deformation threshold, and trigger graded early warnings based on the comparison results.

[0036] Preferably, the visual detection unit includes:

[0037] N explosion-proof mining cameras are deployed at one pre-set distance along the fully mechanized mining face, and the overlap rate of images acquired by adjacent cameras is greater than or equal to a pre-set overlap value. The adjacent cameras are configured to acquire images of the first region and the second region of the roof of the fully mechanized mining face at the same time.

[0038] The image preprocessing module, connected to the N mining explosion-proof cameras, is configured to: use the Retinex algorithm to adjust the brightness and uniformity of the images acquired by adjacent cameras, and perform bilateral filtering to obtain the processed images acquired by adjacent cameras;

[0039] The image stitching module, connected to the image preprocessing module, is configured to: use a brute-force matching algorithm and a k-nearest neighbor matching algorithm to sort the images acquired by adjacent cameras sequentially, obtaining a sequence of sequentially sorted images acquired by all cameras at the same time, wherein the image sequence includes at least a first region image and a second region image; perform grayscale enhancement and noise suppression processing on the image sequence, and input it into a regression network to obtain a coarsely aligned image at the same time; perform SIFT feature point extraction on the coarsely aligned image, and remove outliers in the SIFT feature point extraction using a random sampling consensus algorithm, and then perform projection matrix transformation to obtain a seamlessly stitched panoramic image; perform correction and rectangularization processing on the seamlessly stitched panoramic image to obtain a panoramic image of the fully mechanized mining face roof;

[0040] The YOLOv8 detection module, connected to the image stitching module, is configured to: introduce the CBAM attention mechanism into the Backbone layer of the YOLOv8 detection model, and use the CIoU loss function in the YOLOv8 detection model to obtain a target YOLOv8 detection model; and use the target YOLOv8 detection model to perform bending deformation feature recognition on the panoramic image of the fully mechanized mining face roof to obtain the two-dimensional deformation features of the fully mechanized mining face roof.

[0041] Preferably, the lidar detection unit includes:

[0042] M mining explosion-proof LiDARs, spaced at a second preset distance, are all installed on the top beam of the fully mechanized mining face support and are configured to scan the top plate of the fully mechanized mining face at a preset frequency to obtain real-time point clouds.

[0043] The point cloud preprocessing module, connected to the M mining explosion-proof LiDAR units, is configured to: receive the real-time point cloud and downsample the real-time point cloud using a voxel grid to form downsampled real-time point cloud data; statistically analyze the point cloud neighborhood distance distribution of the downsampled real-time point cloud data to form neighborhood points; remove outlier points whose average distance from the neighborhood points is greater than a preset multiple to obtain real-time point cloud data after removing interference points; and extract features from the real-time point cloud data after removing interference points to obtain point cloud feature data.

[0044] The point cloud registration module, connected to the point cloud preprocessing module, is configured to: use information entropy to filter out structurally stable low-entropy points from the point cloud feature data, and use Euclidean distance to perform spatial clustering of the low-entropy points to form key points; then use the ICP algorithm to perform point cloud registration with the key points in the reference point cloud to obtain the three-dimensional deformation features of the fully mechanized mining face roof.

[0045] Preferably, the data fusion unit is configured to: perform coordinate alignment based on the intrinsic parameter matrix of the camera and the extrinsic parameter matrix of the explosion-proof LiDAR, map the two-dimensional deformation features of the fully mechanized mining face roof to the three-dimensional deformation features of the fully mechanized mining face roof, and establish a correspondence between pixels and point clouds; when the visual detection confidence level formed based on the two-dimensional deformation features and the three-dimensional deformation features of the fully mechanized mining face roof is greater than or equal to a preset two-dimensional value, and the corresponding point cloud deformation is greater than a preset three-dimensional value, a multimodal data fusion result is obtained;

[0046] The early warning linkage unit includes an edge computing terminal and an audible and visual early warning device;

[0047] The edge computing terminal is configured to: compare the multimodal data fusion result with a preset deformation threshold to obtain a comparison result; obtain the multimodal data fusion result according to the triggering of hierarchical early warning; trigger hierarchical early warning when the comparison result meets the preset early warning conditions; and continue monitoring the roof of the fully mechanized mining face when the comparison result does not meet the preset early warning conditions.

[0048] Compared with the prior art, this application has the following beneficial effects:

[0049] The multimodal intelligent recognition method for bending deformation of the fully mechanized mining face roof described in this application includes: acquiring a first region image and a second region image of the fully mechanized mining face roof; stitching together the overlapping region image of the first region image and the second region image to form a panoramic image of the fully mechanized mining face roof; using a detection model to identify bending deformation features of the panoramic image of the fully mechanized mining face roof to obtain the two-dimensional deformation features of the fully mechanized mining face roof; acquiring a real-time point cloud of the fully mechanized mining face roof; using information entropy to filter out structurally stable low-entropy points from the real-time point cloud; using Euclidean distance to spatially cluster the low-entropy points to form key points; then using the ICP algorithm to perform point cloud registration with key points in the reference point cloud to obtain the three-dimensional deformation features of the fully mechanized mining face roof; aligning the two-dimensional deformation features and the three-dimensional deformation features of the fully mechanized mining face roof; performing bidirectional feature mutual verification to obtain a multimodal data fusion result; comparing the multimodal data fusion result with a preset deformation threshold; and triggering a graded early warning based on the comparison result. The above methods have enabled a combined technical approach of "visual panoramic detection, lidar three-dimensional monitoring and multimodal fusion", achieving a comprehensive upgrade in the real-time performance, accuracy and coverage of roof deformation identification in complex underground coal mine environments, and effectively improving the accuracy of early warning. Attached Figure Description

[0050] To more clearly illustrate the technical solutions in the embodiments of this application or the conventional technology, the drawings used in the description of the embodiments or the conventional technology will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0051] Figure 1 This is a flowchart of the multimodal intelligent identification method for bending deformation of the roof of the fully mechanized mining face in the embodiments of this application.

[0052] Figure 2 This is a schematic diagram of the multimodal intelligent recognition system for bending deformation of the roof of the fully mechanized mining face in an embodiment of this application. Detailed Implementation

[0053] To facilitate understanding of this application, a more complete description will be provided below with reference to the accompanying drawings, which illustrate embodiments of the present application. However, the present application can be implemented in many different forms and is not limited to the embodiments described herein. Rather, these embodiments are provided so that the disclosure of this application will be more thorough and complete.

[0054] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. The term "and / or" as used herein includes any and all couplings of one or more of the associated listed items.

[0055] It is understood that the terms “first,” “second,” etc., used in this application may be used herein to describe various elements, but these elements are not limited by these terms. These terms are only used to distinguish one element from another.

[0056] The following explanations of some terms used in this application are provided to aid in understanding the application:

[0057] The Retinex algorithm decomposes an image into illumination components and object reflection components. By estimating and removing the influence of the illumination components, it restores the inherent reflective properties of the object, thereby achieving image enhancement.

[0058] The brute-force matching algorithm is a method that uses exhaustive search to match feature points between two images. It uses lightweight feature blocks to perform fast coarse matching of adjacent images to determine their spatial order.

[0059] The k-nearest neighbor matching algorithm is a general algorithm that filters matching quality by comparing the similarity between the best and second-best matches, ensuring the reliability of the adjacency relationship of image sequences.

[0060] Regression networks are deep learning models used to learn complex mapping relationships between input data and continuous output values. They are pre-trained networks that can quickly transform a sorted sequence of images into a coarsely aligned panoramic image.

[0061] YOLOv8 is an advanced real-time detection algorithm that can complete target localization and classification in a single forward propagation.

[0062] CBAM (Convolutional Block Attention Module) is a lightweight, general-purpose attention mechanism that can be seamlessly integrated into CNNs. It allows the model to dynamically focus on more important features by inferring attention weights in both channel and spatial dimensions.

[0063] CIoU (Complete Intersection over Union) is a metric used to evaluate the similarity between predicted bounding boxes and ground truth bounding boxes. It can be directly used as a loss function for optimization. It considers both the center point distance and aspect ratio consistency on top of IoU.

[0064] Information entropy is a mathematical tool for measuring the "uncertainty" or "disorder" of information. Originating from information theory, it describes the amount of information contained in a system. The higher the information entropy, the greater the uncertainty and the richer the amount of information contained.

[0065] Euclidean distance is the straight-line distance between two points in space; Euclidean distance clustering is an unsupervised machine learning algorithm that automatically groups (clusters) points.

[0066] The ICP algorithm, Iterative Closest Point Algorithm, is a core algorithm used to align (register) two 3D point clouds with high precision.

[0067] KD-tree is a data structure used to efficiently organize points in space, with its core purpose being fast searching.

[0068] The loss function is a mathematical function that measures the difference between the predictions of a machine learning model and the actual situation.

[0069] The RANSAC algorithm robustly estimates optimal mathematical model parameters from a dataset containing a large amount of outlier data (noise, mismatches) through random sampling and iteration.

[0070] The homography matrix is ​​a 3x3 mathematical transformation matrix used to describe the perspective mapping relationship between two planes.

[0071] Fusion algorithms are a class of computational methods that integrate data or information from different sources to generate more consistent, accurate, and useful information than any single data source.

[0072] Zhang Zhengyou's calibration method is a method that uses multiple planar chessboard images from different orientations to solve for camera intrinsic parameters (focal length, principal point coordinates) and distortion coefficients.

[0073] Hand-eye calibration involves solving a fixed homogeneous transformation matrix (containing rotation and translation parameters). This matrix enables the system to map the two-dimensional deformation features detected by the vision unit to the three-dimensional point cloud coordinate system provided by the lidar unit, thereby achieving spatial alignment between "two-dimensional pixels" and "three-dimensional points" and providing a geometric basis for subsequent bidirectional feature verification.

[0074] LiDAR, short for Laser Detection and Ranging System, is an active remote sensing technology that calculates the distance to a target by emitting a laser beam and measuring its return time.

[0075] like Figure 1 As shown, this application provides a multimodal intelligent identification method for the bending deformation of the roof of a fully mechanized mining face, the method comprising:

[0076] S100: Acquire the first region image and the second region image of the fully mechanized mining face roof. Use the overlapping region image of the first region image and the second region image to stitch together to form a panoramic image of the fully mechanized mining face roof. Use the detection model to identify the bending deformation features of the panoramic image of the fully mechanized mining face roof to obtain the two-dimensional deformation features of the fully mechanized mining face roof.

[0077] The S100 step aims to collect images of a portion of the roof (such as images of the first and second areas) using multiple explosion-proof mining cameras (e.g., one camera every 3-5 meters), ensuring that the overlap rate of adjacent images is ≥60%. The images are then stitched together using image stitching technology to create a panoramic image covering the entire longwall face. Finally, a deep learning detection model (e.g., YOLOv8) is used to identify roof deformation features (such as cracks and depressions).

[0078] Specifically, the steps for obtaining the two-dimensional deformation characteristics of the roof of the fully mechanized mining face include:

[0079] S101, the Retinex algorithm is used to adjust the brightness and uniformity of the images acquired by adjacent cameras, and bilateral filtering is performed to obtain the processed images acquired by adjacent cameras.

[0080] After acquiring the first and second region images, the images are preprocessed. First, the Retinex algorithm can be used to enhance the uniformity of illumination, and bilateral filtering is used to remove dust noise, thereby improving the image clarity by 40%.

[0081] S102, using the brute-force matching algorithm and the k-nearest neighbor matching algorithm, the images acquired by adjacent cameras are sorted sequentially to obtain the sequentially sorted image sequence acquired by all cameras at the same time.

[0082] The image sequence includes at least a first region image and a second region image.

[0083] Specifically, the first region image and the second region image are for illustrative purposes only; in reality, multiple images may be used.

[0084] Coarse matching using a brute-force matching algorithm involves dividing the image into blocks, averaging the results, and then matching pixel blocks. For example, for a single image, only a few or dozens of pixel blocks might be selected. These blocks can be formed by downsampling to create simpler, lightweight features. Next, k-nearest neighbor matching is used to match pixel blocks from multiple images. This approach, by applying the "brute-force matching" idea (exhaustive search, traversing all image pairs to find associations) and the "k-nearest neighbor" idea (filtering matching quality), quickly calculates which images are adjacent, thus constructing an image sequence. In other words, the core goal of this step is "ranking," not "precise localization," solving the global topological problem of image relationships with relatively low computational cost.

[0085] S103, perform grayscale enhancement and noise suppression processing on the image sequence, and input it into the regression network to obtain coarsely aligned images at the same time.

[0086] It should be noted that the regression network (with ResNet-50 as its backbone, a 50-layer convolutional neural network architecture) in this step is able to map the ordered image sequence to a complex relationship in a unified panoramic coordinate system, thereby quickly and automatically generating an initial panoramic image that is structurally correct but has coarse details. This functionality requires that the regression network be a trained network that has undergone extensive learning of the transformation parameters of neighboring images, enabling it to quickly and automatically generate coarse panoramic images, with a coarse alignment error ≤ 2 pixels.

[0087] Specifically, using the sequence provided by S102, a roughly aligned panoramic image is generated through a regression network that learns global mapping relationships.

[0088] For example, a sorted sequence of images is input into the network, which learns how to quickly transform a set of overlapping images with a known order into a unified coordinate system. The output is a coarsely aligned image with correct structure, but the boundaries may have excessive or insufficient overlap. This action provides a good initial estimate for subsequent fine stitching, greatly reducing the difficulty of subsequent operations.

[0089] S104, SIFT feature point extraction is performed on the coarsely aligned image, and outliers in the SIFT feature point extraction are removed by random sampling consensus algorithm. Then, projection matrix transformation is performed to obtain a seamlessly stitched panoramic image.

[0090] Specifically, high-precision registration is performed on the roughly aligned images output by S103 to achieve seamless stitching.

[0091] For example, in the aforementioned steps, coarse alignment has already been performed. At this point, extracting SIFT features—which are computationally intensive but highly accurate—from the overlapping areas will result in a very high success rate and accuracy in matching. The search space for matching is reduced, leading to a seamlessly stitched panoramic image.

[0092] SIFT feature points are extracted (≥2000 per image). After outlier removal using brute-force matching and the RANSAC algorithm (retention rate ≥80%), a homography matrix is ​​calculated to achieve seamless stitching. Specifically, the homography matrix is ​​calculated using the matched SIFT feature point pairs. This homography matrix defines the projection transformation relationship between images, mapping the pixel coordinates of one image to the coordinate system of another. After applying this transformation, the images are aligned, and the overlapping areas are processed by a fusion algorithm to finally generate a seamless panoramic image. The fusion algorithm can calculate a weight for each pixel in the overlapping area. The weight is determined by the distance of the pixel to the image boundary; the closer the distance, the greater the weight, achieving a smooth transition.

[0093] S105, the seamlessly stitched panoramic image is corrected and rectangularized to obtain a panoramic image of the roof of the fully mechanized mining face.

[0094] Specifically, it eliminates geometric distortions generated during the preceding stitching process and converts the irregularly shaped preliminary panoramic image into a regular rectangular image without black borders, facilitating subsequent quantitative analysis and display.

[0095] It should be noted that step S105 is to correct the image distortion caused by the characteristics of the camera lens and the non-orthogonal shooting angle, so as to restore the true geometric proportions of the top plate shape in the image.

[0096] For example, the camera calibration process pre-acquires the camera's intrinsic parameter matrix (including parameters such as focal length and principal point coordinates) and lens distortion coefficients. Alternatively, the camera's intrinsic parameters can also be obtained using the Zhang Zhengyou calibration method.

[0097] Based on intrinsic parameters and distortion coefficients, a mapping relationship is calculated from the coordinates of the distorted image to the coordinates of the distortion-free ideal image. This mapping relationship can be understood as a detailed "correction lookup table". Alternatively, after obtaining the camera intrinsic parameters through Zhang Zhengyou's calibration method, perspective transformation is performed on the stitched image to eliminate edge distortion.

[0098] During the rectangularization process, each pixel in the seamlessly stitched panoramic image is repositioned to a new, correct location according to the calculated mapping relationship, thereby generating an image that eliminates barrel and pincushion lens distortions as well as perspective distortion. Next, on the perspective-corrected seamlessly stitched panoramic image, the largest inscribed rectangle representing the effective fully mechanized mining face roof area is found, and invalid black areas or irregular edges generated by image stitching and transformation are cropped.

[0099] For example, image processing techniques (such as thresholding and edge detection) are used to determine the precise boundary contour of the effective pixel area (i.e., the actual roof image) in the seamlessly stitched panoramic image after correction. The algorithm automatically finds the largest inscribed rectangle within this boundary contour. This rectangle represents the largest regular area that can be obtained from the current seamlessly stitched panoramic image. Finally, based on the coordinates of the calculated largest inscribed rectangle, the seamlessly stitched panoramic image is cropped, and a rectangular panoramic image with neat edges is finally output, which is the panoramic image of the fully mechanized mining face roof.

[0100] Using the above method, the image was processed step by step from coarse to fine, and finally a panoramic image of the roof of the fully mechanized mining face was obtained.

[0101] Following S105, the following steps are also included:

[0102] S106. The CBAM attention mechanism is introduced into the Backbone layer of the YOLOv8 detection model, and the CIoU loss function is used in the YOLOv8 detection model to obtain the target YOLOv8 detection model.

[0103] Specifically, the standard YOLOv8 detection model was enhanced to make it more suitable for detecting minute deformation features of the downhole roof.

[0104] For example, the CBAM attention mechanism is introduced by embedding CBAM (Convolutional Block Attention Module) into a specific layer (such as after the C3 module) of the YOLOv8 detection model's backbone. In this way, CBAM sequentially calculates channel attention and spatial attention on the feature map. Further, channel attention refers to the YOLOv8 detection model automatically learning the weights of each feature channel. For roof detection, the YOLOv8 detection model will therefore pay more attention to channel features sensitive to "cracks" and "depressions," suppressing interference from irrelevant channels (such as background rock textures). Spatial attention refers to the YOLOv8 detection model automatically learning the importance of each spatial location on the feature map. This allows the YOLOv8 detection model to focus more on local areas in the image that may contain anomalies (such as crack edges or depression centers), rather than processing the entire image uniformly. Essentially, channel attention tells the YOLOv8 detection model which features of the fully mechanized mining face roof to pay attention to, while spatial attention tells the YOLOv8 detection model which locations on the fully mechanized mining face roof to look at.

[0105] By introducing CBAM, the YOLOv8 detection model has significantly enhanced its ability to perceive small and fuzzy deformation features, especially improving its ability to detect small targets (such as cracks with a width ≥ 0.2 mm) in complex backgrounds.

[0106] Next, the CIoU loss function is used, considering the overlap area between the predicted bounding box and the ground truth bounding box, as well as the consistency of their center point distance and aspect ratio. In other words, during the optimization process, CIoU guides the YOLOv8 detection model to generate predicted boxes that not only overlap more with the ground truth in location, but also more accurately encompass deformed features in shape and center point. This is particularly suitable for scenarios requiring precise quantification of the size and location of deformed regions, effectively reducing localization errors.

[0107] Furthermore, the improved YOLOv8 detection model was trained with abundant real-world downhole data to achieve strong generalization capabilities and ensure that the bounding box positions, categories, and confidence scores of its outputs are accurate and reliable.

[0108] For example, the training dataset contains tens of thousands to hundreds of thousands of finely annotated downhole roof images. Annotation categories should include "fracture," "depression," and "precursor to delamination (crustation area)," etc. To improve the model's robustness, the training images can be augmented online to simulate complex downhole conditions; the main augmentation methods include:

[0109] Geometric transformations: random horizontal flipping and random small-angle rotation to simulate changes in viewpoint;

[0110] Pixel transformation: Adjust brightness, contrast, and saturation to simulate uneven lighting and differences in lighting fixtures;

[0111] Environmental simulation: Add simulated dust and fog effects and random occlusion to improve the model's recognition stability in high dust environments.

[0112] As another example, the improved YOLOv8 model training steps may include:

[0113] In the tens of thousands to hundreds of thousands of finely annotated downhole roof images, the collected roof images are divided into training set, validation set and test set in an 8:1:1 ratio, and data augmentation is performed by random flipping, brightness adjustment and dust simulation.

[0114] Model configuration: Load the pre-trained weights from yolov8n.pt, and set imgsz=640, batch=16, epochs=100, and initial learning rate=0.01.

[0115] Training and optimization: The SGD optimizer is used. Training is terminated when the validation set mAP@0.5 shows no improvement for 10 consecutive rounds, and the optimal weight file is saved.

[0116] Model deployment: Convert the trained model to ONNX format and deploy it to the downhole edge computing terminal with an inference speed of ≥20fps.

[0117] Finally, the improved YOLOv8 detection model is trained using the training dataset described above to obtain the target YOLOv8 detection model.

[0118] It should be noted that during the training of the YOLOv8 detection model, the loss function (usually the composite loss function built into YOLOv8) simultaneously optimizes bounding box regression, target classification, and confidence prediction. This means that the improved YOLOv8 detection model, by continuously reducing the difference between predictions and true values, not only learns how to locate and classify targets but also learns how to assess the confidence level of its own predictions. Thus, the trained YOLOv8 target detection model can not only identify anomalies in the panoramic image of the fully mechanized mining face roof but also output confidence scores.

[0119] S107. Using the target YOLOv8 detection model, anomaly identification is performed on the panoramic image of the fully mechanized mining face roof to obtain the two-dimensional deformation features of the fully mechanized mining face roof.

[0120] Specifically, after obtaining the target YOLOv8 detection model in the aforementioned steps, anomaly identification is performed on the panoramic image of the fully mechanized mining face roof. The YOLOv8 detection model performs forward propagation inference on the image and outputs the bounding box information of all identified abnormal targets.

[0121] Furthermore, the information within the bounding box is parsed, and the parsing results include:

[0122] Category: Indicate whether it is a crack, depression, or area of ​​chipping or spalling;

[0123] Bounding box coordinates: Used to define the location and extent of anomaly areas in a panoramic view;

[0124] Confidence level: The degree of certainty the YOLOv8 detection model has about the identification result. A threshold (e.g., ≥0.85) is usually set to filter out false positives with a low probability. The final result is a detection result containing all high-confidence deformation features, i.e., the two-dimensional deformation features of the fully mechanized mining face roof, which provides accurate two-dimensional information for subsequent fusion with three-dimensional point cloud data.

[0125] It should be noted that, in order to adapt to the detection of small targets on the top surface, the YOLOv8 detection model underwent triple optimization:

[0126] Feature enhancement: A CBAM attention mechanism was added after the C3 module (a regular module in YOLOv8), which improved the response value of crack edge features by 35%;

[0127] Loss optimization: By adopting the CIoU loss function, the bounding box regression error is reduced to 0.8 pixels;

[0128] Lightweight deployment: Redundant channels are removed through model pruning, increasing inference speed to 25fps to meet real-time requirements.

[0129] S200: Obtain the real-time point cloud of the fully mechanized mining face roof, use information entropy to filter out structurally stable low-entropy points from the real-time point cloud, then use Euclidean distance to perform spatial clustering on the low-entropy points to form key points, and perform point cloud registration with the key points in the reference point cloud to obtain the three-dimensional deformation features of the fully mechanized mining face roof.

[0130] The purpose of step S200 is to acquire three-dimensional point cloud data of the roof using lidar and to calculate the deformation of the roof by comparing it with the reference point cloud.

[0131] Specifically, obtaining the three-dimensional deformation characteristics of the roof of the fully mechanized mining face may include the following steps:

[0132] S201, acquire the real-time point cloud of the roof of the fully mechanized mining face, and downsample the real-time point cloud through a voxel grid to obtain downsampled real-time point cloud data.

[0133] Specifically, the purpose of step S201 is to significantly reduce the amount of point cloud data and improve the speed of subsequent processing while preserving the main geometric structure of the roof of the fully mechanized mining face.

[0134] For example, the three-dimensional point cloud space is divided into a series of uniform and tiny cubic grids, i.e. voxels. Then, all points falling within the same voxel are approximated by a representative point (such as the centroid of these points). By setting appropriate voxel sizes (such as 2cm x 2cm x 2cm), the amount of massive original point cloud data can be reduced by more than 70%, while effectively maintaining the macroscopic shape of the fully mechanized mining face roof, laying the foundation for subsequent calculations.

[0135] S202, statistically analyze the neighborhood distance distribution of the downsampled real-time point cloud data to form neighborhood points, and remove abnormal points whose average distance from the neighborhood points is greater than a preset multiple to obtain the real-time point cloud after removing interference points.

[0136] Specifically, the purpose of step S202 is to remove discrete noise points (such as abnormal points caused by dust or sensor errors) from the point cloud and retain valid top surface points.

[0137] For example, to statistically analyze the point cloud neighborhood distance distribution of the downsampled real-time point cloud data, the local distance is first calculated: for each point in the downsampled point cloud, the average distance to its K nearest neighbors is calculated. This distance reflects the "density" of the point in its local environment. Next, a global statistical analysis is performed: the mean of the dataset consisting of the "average distance" values ​​of all points is calculated. and standard deviation According to the normal distribution assumption, the average distance of the vast majority of normal points should be distributed within the range of \left [{\mu -n\sigma ,\mu +n\sigma} \right ] Within the interval (e.g., n=3); next, outlier removal is performed: set a threshold (e.g., ... If the average distance of a point exceeds this threshold, the point is considered an anomaly (such as a dust noise point) separated from the main structure and is removed. This method of anomaly removal effectively filters out most interfering points, thus reducing noise.

[0138] S203: Extract features from the real-time point cloud after removing interference points to obtain point cloud feature data.

[0139] Specifically, the purpose of feature extraction is to compute a mathematical vector for each point that describes the local three-dimensional shape features around it.

[0140] For example, for the denoised point cloud, the local surface properties formed by each point and its neighbors are calculated. For instance, the normal vector describes the orientation of the tangent plane at that point; the curvature describes the degree of surface curvature at that point; and the eigenvalues, calculated using principal component analysis (PCA), describe the distribution of the local point set along the three principal directions. After extracting the feature points, one or more eigenvalues ​​are assigned to each point, forming point cloud feature data. For example, the eigenvalues ​​of points in regions with obvious structural features (such as edges and planes) will differ significantly from those of points in regions with flat structures (such as noise).

[0141] S204. Using information entropy, structurally stable low-entropy points are selected from the point cloud feature data, and Euclidean distance is used to perform spatial clustering of the low-entropy points to form key points. Then, the ICP algorithm is used to perform point cloud registration with the key points in the reference point cloud to obtain the three-dimensional deformation features of the fully mechanized mining face roof.

[0142] Specifically, the purpose of step S204 is to select stable and unique "key points" from the point cloud and accurately calculate the three-dimensional deformation by comparing them with the benchmark point cloud.

[0143] For example, information entropy is used to filter low-entropy points. Here, entropy measures the structural uncertainty or disorder of a local area around a point. A point's feature vector has a low entropy value if it changes gently and predictably in its neighborhood (e.g., it lies on a flat surface), and a high entropy value if it changes drastically and is disordered (e.g., a noisy point).

[0144] Furthermore, the local feature information entropy H(i) of point i is calculated, which can be expressed by the formula:

[0145] ;

[0146] in, This represents the probability that a certain feature (such as curvature) of a point belongs to the k-th interval in its neighborhood. Points with low entropy values ​​(e.g., <0.3) may indicate that their local structure is stable and significant, making them suitable as key points representing the overall structure. In other words, the information entropy of all points is calculated, and points with entropy values ​​below a preset threshold are selected as low-entropy points or structurally stable key points.

[0147] Next, Euclidean distance is used to spatially cluster low-entropy points, thus avoiding overly dense keypoints and selecting representative points with even distribution. For the selected low-entropy points, spatial clustering is performed based on their Euclidean distances to form keypoints. Points that are too close are grouped into one class, and a center point is selected from each class as the final keypoint. This ensures that the keypoints are evenly distributed on the top surface and their number is controllable (e.g., 1000-2000).

[0148] Point cloud registration and deformation calculation employ the ICP algorithm. This involves first using the aforementioned key points to calculate an initial coarse registration transformation matrix by finding nearest neighbor pairs (KD-trees can be used to accelerate the search). Then, an iterative nearest-neighbor algorithm is used for registration. Further, the ICP process iteratively finds the correspondence between key points in the real-time point cloud and key points in the reference point cloud, continuously optimizing the rotation and translation matrices by minimizing the distance between corresponding points (e.g., the distance from a point to a surface) until the error is less than a threshold. This process is the registration. After registration, the coordinate difference between the real-time point cloud and the reference point cloud represents the three-dimensional deformation of the roof. By calculating the displacement of each point or region, the three-dimensional deformation characteristics of the fully mechanized mining face roof are obtained, such as subsidence and horizontal displacement.

[0149] S300, the two-dimensional deformation features and the three-dimensional deformation features of the roof of the fully mechanized mining face are aligned, and bidirectional feature mutual verification is performed to obtain the multimodal data fusion result.

[0150] Specifically, obtaining multimodal data fusion results may include the following steps:

[0151] S301, coordinate alignment is performed based on the intrinsic parameter matrix of the camera and the extrinsic parameter matrix of the explosion-proof LiDAR, and the two-dimensional deformation features of the roof of the fully mechanized mining face are mapped to the three-dimensional deformation features of the roof of the fully mechanized mining face, establishing the correspondence between pixels and point clouds.

[0152] Specifically, the purpose of step S301 is to establish a unified coordinate system so that two-dimensional anomalies found in the image can be accurately correlated with specific locations in the three-dimensional point cloud.

[0153] For example, calibration parameter preparation relies on sensor parameters that have been obtained in advance through hand-eye calibration.

[0154] The camera intrinsic parameter matrix describes how the camera projects three-dimensional spatial points onto a two-dimensional image plane, including focal length. , and principal point coordinates , The parameters are equal, and their matrix form is:

[0155] ;

[0156] extrinsic parameter matrix It describes the relative position and attitude relationship between the lidar coordinate system and the camera coordinate system, including a rotation matrix (R) and a translation vector (t).

[0157] After obtaining the intrinsic parameter matrix of the camera and the extrinsic parameter matrix of the explosion-proof LiDAR, coordinate mapping is performed. Any two-dimensional deformable region detected by the YOLOv8 detection model (represented by its bounding box) can be projected back into three-dimensional space using the aforementioned parameters and perspective geometry principles. Specifically, by combining the intrinsic parameter matrix K and the extrinsic parameter matrix... A ray can be calculated that passes through an image pixel from the camera's optical center. The area where this ray intersects with the three-dimensional top surface formed by the lidar point cloud is the corresponding position of the two-dimensional feature in three-dimensional space. This establishes the correspondence between each pixel in the image and a series of three-dimensional points in the point cloud, achieving spatial calibration between the two-dimensional detection results and the three-dimensional point cloud data.

[0158] S302, when the visual detection confidence level formed based on the two-dimensional deformation characteristics and the three-dimensional deformation characteristics of the fully mechanized mining face roof is greater than or equal to a preset two-dimensional value, and the corresponding point cloud deformation is greater than a preset three-dimensional value, a multimodal data fusion result is obtained.

[0159] Specifically, the purpose of step S302 is to formulate fusion decision rules to jointly judge the detection results of vision and lidar in order to generate the final multimodal data fusion result.

[0160] For example, the decision rule is as follows: when an identified target simultaneously meets the following two conditions, the system determines it to be a valid deformation event jointly verified by multimodal data:

[0161] To determine if a transformation is valid, both of the following conditions must be met simultaneously (using AND logic):

[0162] Condition 1 (Visual Verification): The YOLOv8 detection model's confidence level in detecting this deformable feature must be greater than or equal to a preset high threshold (preset two-dimensional value, e.g., ≥0.9). This indicates that the visual model is highly confident that the deformable feature exists in the image.

[0163] Condition 2 (3D Verification): Within the corresponding 3D point cloud region established in step S301, the calculated 3D deformation (displacement relative to the reference point cloud) must be greater than or equal to a preset minimum physical deformation threshold (preset 3D value, e.g., >5mm). This indicates that the deformation does indeed cause a measurable change in geometry.

[0164] Thus, the multimodal data fusion result refers to all valid deformation events that have passed the above two-way verification, which will be summarized and structurally integrated into a multimodal data fusion result. This multimodal data fusion result is the direct basis for the system to carry out hierarchical early warning and linkage control.

[0165] Finally, a list of multimodal data fusion results that have been verified by multimodal data fusion is generated. Each result in this list can include the type of deformation, two-dimensional location, three-dimensional spatial coordinates, deformation magnitude, and fusion confidence level, i.e., the multimodal data fusion result.

[0166] S400 compares the multimodal data fusion results with a preset deformation threshold and triggers a graded warning based on the comparison results.

[0167] The S400 step is the final decision-making stage of the system, which aims to automatically make early warning judgments based on valid deformation data that has been verified through multimodal analysis.

[0168] Specifically, triggering a tiered warning may include the following steps:

[0169] S401, compare the multimodal data fusion result with a preset deformation threshold to obtain a comparison result.

[0170] Specifically, a deformation threshold can be preset, and the results of multimodal data fusion can be compared using the deformation threshold.

[0171] For example, the key parameters of each "valid deformation event" in the multimodal data fusion result are compared one by one with the preset deformation threshold in the system. The key comparison parameters include the area of ​​the deformation region, the maximum depth of the deformation, or the amount of displacement. The comparison can be a comparison of the magnitude of the two data.

[0172] S402, when the comparison result meets the preset warning conditions, a graded warning is triggered.

[0173] Specifically, based on the comparison results, the corresponding early warning level and linkage control will be automatically activated.

[0174] Exemplarily, the warning condition: When the parameters in the multi-modal data fusion result meet or exceed (≥) the threshold conditions of any preset warning level, the corresponding level of warning is triggered. A multi-level threshold trigger mechanism is adopted, and warnings of different response levels are triggered according to the threshold conditions of different levels. For example:

[0175] Level 1 warning (general warning): Triggered when the deformed area of the effective deformation < A1 and the maximum deformation amount ≤ D1 (where A1 and D1 are the first-level thresholds, for example, A1 = 5㎡, D1 = 10mm).

[0176] Response measures: It may only be highlighted and logged on the monitoring system interface, and the local sound and light alarm is triggered to remind the on-site personnel to pay attention.

[0177] Level 2 warning (severe warning): Triggered when the deformed area of the effective deformation ≥ A2 or the maximum deformation amount > D2 (where A2 and D2 are the second-level thresholds, usually A2 = A1, D2 = D1 or higher).

[0178] Response measures: Immediately trigger the highest-level sound and light alarm, and automatically send a shutdown instruction to associated devices (such as coal mining machines and hydraulic supports) through the industrial Ethernet to force the operation to stop and ensure the safety of personnel.

[0179] In addition, to adapt to different mining conditions (such as different advancing speeds), the preset deformation thresholds can be dynamically adjusted according to preset rules. For example, when the mining progress accelerates (> 3m / h), the system can automatically reduce the thresholds (D1, D2) by 20% to achieve a more sensitive warning.

[0180] S403, when the comparison result does not meet the preset warning conditions, continue to monitor the roof of the fully-mechanized mining face.

[0181] Specifically, for minor deformations or suspected noises that do not reach the warning standard, the system maintains the monitoring state to avoid over-response.

[0182] If all parameters in the fusion result are lower (<) than the lowest first-level warning threshold, it is determined that the current roof state is within the safe range. The system will not trigger any warning, but continue to execute its normal monitoring cycle (that is, return to steps S100 and S200), collect new image and point cloud data, and perform a new round of processing, fusion and judgment to achieve 7x24-hour uninterrupted automated safety monitoring.

[0183] As Figure 2 shown, the present application also provides a multi-modal intelligent recognition system for roof bending deformation of a fully-mechanized mining face, which is applied to the multi-modal intelligent recognition method for roof bending deformation of the fully-mechanized mining face described above. The system includes:

[0184] The visual inspection unit is used to acquire a first region image and a second region image of the roof of the fully mechanized mining face, and to stitch together the overlapping region image of the first region image and the second region image to form a panoramic image of the roof of the fully mechanized mining face. The detection model is used to identify the bending deformation features of the panoramic image of the roof of the fully mechanized mining face to obtain the two-dimensional deformation features of the roof of the fully mechanized mining face.

[0185] The lidar detection unit is used to acquire the real-time point cloud of the fully mechanized mining face roof. It uses information entropy to filter out structurally stable low-entropy points from the real-time point cloud, and uses Euclidean distance to perform spatial clustering of the low-entropy points to form key points. Then, it uses the ICP algorithm to perform point cloud registration with key points in the reference point cloud to obtain the three-dimensional deformation features of the fully mechanized mining face roof.

[0186] The data fusion unit is used to align the two-dimensional deformation features and the three-dimensional deformation features of the roof of the fully mechanized mining face, and to perform bidirectional feature mutual verification to obtain multimodal data fusion results.

[0187] The early warning linkage unit is used to compare the multimodal data fusion results with the preset deformation threshold, and trigger graded early warnings based on the comparison results.

[0188] Preferably, the visual detection unit includes:

[0189] N explosion-proof mining cameras are deployed at one pre-set distance along the fully mechanized mining face, and the overlap rate of images acquired by adjacent cameras is greater than or equal to a pre-set overlap value. The adjacent cameras are configured to acquire images of the first region and the second region of the roof of the fully mechanized mining face at the same time.

[0190] The image preprocessing module, connected to the N mining explosion-proof cameras, is configured to: use the Retinex algorithm to adjust the brightness and uniformity of the images acquired by adjacent cameras, and perform bilateral filtering to obtain the processed images acquired by adjacent cameras;

[0191] The image stitching module, connected to the image preprocessing module, is configured to: use a brute-force matching algorithm and a k-nearest neighbor matching algorithm to sort the images acquired by adjacent cameras sequentially, obtaining a sequence of sequentially sorted images acquired by all cameras at the same time, wherein the image sequence includes at least a first region image and a second region image; perform grayscale enhancement and noise suppression processing on the image sequence, and input it into a regression network to obtain a coarsely aligned image at the same time; perform SIFT feature point extraction on the coarsely aligned image, and remove outliers in the SIFT feature point extraction using a random sampling consensus algorithm, and then perform projection matrix transformation to obtain a seamlessly stitched panoramic image; perform correction and rectangularization processing on the seamlessly stitched panoramic image to obtain a panoramic image of the fully mechanized mining face roof;

[0192] The YOLOv8 detection module, connected to the image stitching module, is configured to: introduce the CBAM attention mechanism into the Backbone layer of the YOLOv8 detection model, and use the CIoU loss function in the YOLOv8 detection model to obtain a target YOLOv8 detection model; and use the target YOLOv8 detection model to perform bending deformation feature recognition on the panoramic image of the fully mechanized mining face roof to obtain the two-dimensional deformation features of the fully mechanized mining face roof.

[0193] For example, to address downhole imaging defects, a two-stage stitching strategy of "unsupervised coarse stitching + key point fine stitching" is adopted:

[0194] Image preprocessing: The Retinex algorithm is used to enhance illumination uniformity, and bilateral filtering is used to remove dust noise, improving image clarity by 40%. This eliminates light spots and shadows caused by uneven illumination in the roof image of the fully mechanized mining face, enhances the visibility of surface details such as cracks, and provides high-quality input images for the subsequent YOLOv8 detection model.

[0195] Unsupervised coarse alignment: The input regression network (with ResNet-50 as the backbone) learns the transformation parameters of neighboring images, and the coarse alignment error is ≤2 pixels.

[0196] Key point fine stitching: Extract SIFT feature points (≥2000 per image), remove outliers by brute-force matching and RANSAC algorithm (retention rate ≥80%), and calculate homography matrix to achieve seamless stitching.

[0197] Perspective correction: The camera intrinsic parameters are obtained using the Zhang Zhengyou calibration method, and perspective transformation is performed on the stitched image to eliminate edge distortion.

[0198] To adapt to small target detection on the roof, the YOLOv8 detection model training dataset contains 120,000 downhole images, labeled with three types of deformable targets: cracks (0.2~50mm), depressions (1~30mm), and spalling areas (area ≥10cm²). Validated on the test set, the mAP@0.5 reached 97.3% (the model's average accuracy is 97.3% with an IoU (Intersection over Union) threshold of 0.5), and the small target detection rate was ≥92%.

[0199] In one embodiment, the lidar detection unit includes:

[0200] M mining explosion-proof LiDARs, spaced at a second preset distance, are all installed on the top beam of the fully mechanized mining face support and are configured to scan the top plate of the fully mechanized mining face at a preset frequency to obtain real-time point clouds.

[0201] The point cloud preprocessing module, connected to the M mining explosion-proof LiDAR units, is configured to: receive the real-time point cloud and downsample the real-time point cloud using a voxel grid to form downsampled real-time point cloud data; statistically analyze the point cloud neighborhood distance distribution of the downsampled real-time point cloud data to form neighborhood points; remove outlier points whose average distance from the neighborhood points is greater than a preset multiple to obtain real-time point cloud data after removing interference points; and extract features from the real-time point cloud data after removing interference points to obtain point cloud feature data.

[0202] The point cloud registration module, connected to the point cloud preprocessing module, is configured to: use information entropy to filter out structurally stable low-entropy points from the point cloud feature data, and use Euclidean distance to perform spatial clustering of the low-entropy points to form key points; then use the ICP algorithm to perform point cloud registration with the key points in the reference point cloud to obtain the three-dimensional deformation features of the fully mechanized mining face roof.

[0203] For example, the ICP registration algorithm based on information entropy clustering solves the noise interference problem:

[0204] Point cloud preprocessing: Voxel downsampling (2cm resolution) reduces data volume by 70%, and statistical filtering removes more than 95% of dust interference points;

[0205] Key point extraction: Calculate the information entropy of the feature value of each point, select stable points with entropy value < 0.25 as cluster centers, and extract 1500 core key points;

[0206] Fast registration: Point-to-point search is accelerated by bidirectional KD-tree (speed improvement of 40%), and the point-to-surface ICP algorithm is iteratively optimized, with registration RMSE ≤ 0.8mm and time < 0.1s / frame.

[0207] In one embodiment, the data fusion unit is configured to: perform coordinate alignment based on the intrinsic parameter matrix of the camera and the extrinsic parameter matrix of the explosion-proof LiDAR, map the two-dimensional deformation features of the fully mechanized mining face roof to the three-dimensional deformation features of the fully mechanized mining face roof, and establish a correspondence between pixels and point clouds; when the visual detection confidence level formed based on the two-dimensional deformation features and the three-dimensional deformation features of the fully mechanized mining face roof is greater than or equal to a preset two-dimensional value, and the corresponding point cloud deformation amount is greater than a preset three-dimensional value, a multimodal data fusion result is obtained;

[0208] The early warning linkage unit includes an edge computing terminal and an audible and visual early warning device;

[0209] The edge computing terminal is configured to: compare the multimodal data fusion result with a preset deformation threshold to obtain a comparison result; obtain the multimodal data fusion result according to the triggering of hierarchical early warning; trigger hierarchical early warning when the comparison result meets the preset early warning conditions; and continue monitoring the roof of the fully mechanized mining face when the comparison result does not meet the preset early warning conditions.

[0210] For example, the multimodal fusion unit can establish a two-dimensional-three-dimensional bidirectional calibration mechanism:

[0211] Spatial mapping: Based on hand-eye calibration, the camera and LiDAR extrinsic parameters (rotation matrix R, translation vector T) are obtained, and the two-dimensional bounding box (x1, y1, x2, y2) detected by YOLOv8 is mapped to the three-dimensional point cloud space to generate the region of interest (ROI).

[0212] Feature verification: When the point cloud deformation within the ROI is greater than 5mm and the visual detection confidence is greater than or equal to 0.9, it is considered a valid deformation; if deformation is detected only in a single modality, it is marked as a suspicious area and a secondary detection is triggered.

[0213] Dynamic threshold: The early warning threshold is dynamically adjusted according to the mining progress (the threshold is reduced by 20% when the advance speed is >3m / h) to adapt to different deformation rate scenarios.

[0214] Through the above-mentioned intelligent multimodal identification method and system for roof bending deformation in fully mechanized mining faces, a combined technical approach of "visual panoramic detection, lidar three-dimensional monitoring and multimodal fusion" has been realized, achieving a comprehensive upgrade in the real-time performance, accuracy and coverage of roof deformation identification in complex underground coal mine environments, and effectively improving the accuracy of early warning.

[0215] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0216] The various embodiments in this disclosure are described in a progressive manner. The same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on describing the differences from other embodiments.

[0217] The scope of protection of this disclosure is not limited to the embodiments described above. Obviously, those skilled in the art can make various modifications and variations to this disclosure without departing from its scope and spirit. If such modifications and variations fall within the scope of the claims of this disclosure and their equivalents, then the intent of this disclosure also includes such modifications and variations.

Claims

1. A multimodal intelligent identification method for bending deformation of the roof of a fully mechanized mining face, characterized in that, include: A first region image and a second region image of the roof of the fully mechanized mining face are obtained. The overlapping region images of the first region image and the second region image are stitched together to form a panoramic image of the roof of the fully mechanized mining face. The bending deformation feature of the panoramic image of the roof of the fully mechanized mining face is identified by a detection model to obtain the two-dimensional deformation feature of the roof of the fully mechanized mining face. The real-time point cloud of the fully mechanized mining face roof is obtained. The information entropy is used to filter out structurally stable low-entropy points from the real-time point cloud. The Euclidean distance is used to perform spatial clustering of the low-entropy points to form key points. Then, the ICP algorithm is used to perform point cloud registration with the key points in the reference point cloud to obtain the three-dimensional deformation features of the fully mechanized mining face roof. Align the two-dimensional deformation features and the three-dimensional deformation features of the roof of the fully mechanized mining face, and perform bidirectional feature mutual verification to obtain the multimodal data fusion result. The results of multimodal data fusion are compared with preset deformation thresholds, and graded early warnings are triggered based on the comparison results. The steps of acquiring a first region image and a second region image of the fully mechanized mining face roof, stitching together the overlapping region images of the first and second region images to form a panoramic image of the fully mechanized mining face roof, and using a detection model to identify bending deformation features of the panoramic image of the fully mechanized mining face roof to obtain the two-dimensional deformation features of the fully mechanized mining face roof include: The Retinex algorithm is used to adjust the brightness and uniformity of images acquired by adjacent cameras, and bilateral filtering is performed to obtain the processed images acquired by adjacent cameras. Using a brute-force matching algorithm and a k-nearest neighbor matching algorithm, images acquired by adjacent cameras are sequentially sorted to obtain a sequentially sorted image sequence of all cameras acquired at the same time. The image sequence includes at least a first region image and a second region image. The image sequence is subjected to grayscale enhancement and noise suppression processing, and then input into a regression network to obtain coarsely aligned images at the same time. SIFT feature points are extracted from the coarsely aligned image, and outliers in the SIFT feature point extraction are removed by random sampling consensus algorithm. Then, projection matrix transformation is performed to obtain a seamlessly stitched panoramic image. The seamlessly stitched panoramic image is corrected and rectangularized to obtain a panoramic image of the roof of the fully mechanized mining face. The target YOLOv8 detection model is obtained by introducing the CBAM attention mechanism into the backbone layer of the YOLOv8 detection model and using the CIoU loss function in the YOLOv8 detection model. Using the target YOLOv8 detection model, the bending deformation features of the panoramic image of the fully mechanized mining face roof are identified to obtain the two-dimensional deformation features of the fully mechanized mining face roof.

2. The intelligent multimodal identification method for bending deformation of the roof of a fully mechanized mining face according to claim 1, characterized in that, The steps of obtaining the real-time point cloud of the fully mechanized mining face roof, filtering out structurally stable low-entropy points from the real-time point cloud using information entropy, spatially clustering the low-entropy points using Euclidean distance to form key points, and then using the ICP algorithm to perform point cloud registration with key points in the reference point cloud to obtain the three-dimensional deformation features of the fully mechanized mining face roof include: The real-time point cloud of the roof of the fully mechanized mining face is obtained, and the real-time point cloud is downsampled by a voxel grid to obtain the downsampled real-time point cloud. The distribution of the neighborhood distance of the downsampled real-time point cloud is statistically analyzed to form neighborhood points. Abnormal points whose average distance from the neighborhood points is greater than a preset multiple are removed to obtain real-time point cloud data after removing interference points. Feature extraction is performed on the real-time point cloud data after removing interference points to obtain point cloud feature data; Information entropy is used to filter out structurally stable low-entropy points from the point cloud feature data, and Euclidean distance is used to perform spatial clustering of the low-entropy points to form key points. Then, the ICP algorithm is used to perform point cloud registration with key points in the reference point cloud to obtain the three-dimensional deformation features of the roof of the fully mechanized mining face.

3. The intelligent multimodal identification method for bending deformation of the roof of a fully mechanized mining face according to claim 1, characterized in that, The step of aligning the two-dimensional deformation features and the three-dimensional deformation features of the fully mechanized mining face roof, and performing bidirectional feature cross-verification to obtain the multimodal data fusion result includes: Based on the intrinsic parameter matrix of the camera and the extrinsic parameter matrix of the explosion-proof LiDAR, coordinate alignment is performed to map the two-dimensional deformation features of the roof of the fully mechanized mining face to the three-dimensional deformation features of the roof of the fully mechanized mining face, and establish the correspondence between pixels and point clouds. When the visual detection confidence level formed based on the two-dimensional deformation features and the three-dimensional deformation features of the fully mechanized mining face roof is greater than or equal to a preset two-dimensional value, and the corresponding point cloud deformation is greater than a preset three-dimensional value, a multimodal data fusion result is obtained.

4. The intelligent multimodal identification method for bending deformation of the roof of a fully mechanized mining face according to claim 3, characterized in that, The step of comparing the multimodal data fusion result with a preset deformation threshold and triggering a graded early warning based on the comparison result further includes: The multimodal data fusion result is compared with a preset deformation threshold to obtain a comparison result; When the comparison result meets the preset warning conditions, a tiered warning is triggered; If the comparison result shows that the preset warning conditions are not met, the monitoring of the roof of the fully mechanized mining face will continue.

5. A multimodal intelligent recognition system for bending deformation of the roof of a fully mechanized mining face, characterized in that, The intelligent multimodal identification method for bending deformation of the roof of a fully mechanized mining face, applicable to any one of claims 1-4, comprises: The visual inspection unit is used to acquire a first region image and a second region image of the roof of the fully mechanized mining face, stitch the overlapping region image of the first region image and the second region image together to form a panoramic image of the roof of the fully mechanized mining face, and use the detection model to identify the bending deformation features of the panoramic image of the roof of the fully mechanized mining face to obtain the two-dimensional deformation features of the roof of the fully mechanized mining face. The lidar detection unit is used to acquire the real-time point cloud of the fully mechanized mining face roof. It uses information entropy to filter out structurally stable low-entropy points from the real-time point cloud, and uses Euclidean distance to perform spatial clustering of the low-entropy points to form key points. Then, it uses the ICP algorithm to perform point cloud registration with key points in the reference point cloud to obtain the three-dimensional deformation features of the fully mechanized mining face roof. The data fusion unit is used to align the two-dimensional deformation features and the three-dimensional deformation features of the roof of the fully mechanized mining face, and to perform bidirectional feature mutual verification to obtain multimodal data fusion results. The early warning linkage unit is used to compare the multimodal data fusion results with the preset deformation threshold, and trigger graded early warnings based on the comparison results.

6. The multimodal intelligent recognition system for bending deformation of the roof of a fully mechanized mining face according to claim 5, characterized in that, The visual detection unit includes: N explosion-proof mining cameras are deployed at one pre-set distance along the fully mechanized mining face, and the overlap rate of images acquired by adjacent cameras is greater than or equal to a pre-set overlap value. The adjacent cameras are configured to acquire images of the first region and the second region of the roof of the fully mechanized mining face at the same time. The image preprocessing module, connected to the N mining explosion-proof cameras, is configured to: use the Retinex algorithm to adjust the brightness and uniformity of the images acquired by adjacent cameras, and perform bilateral filtering to obtain the processed images acquired by adjacent cameras; The image stitching module, connected to the image preprocessing module, is configured to: use a brute-force matching algorithm and a k-nearest neighbor matching algorithm to sort the images acquired by adjacent cameras sequentially, obtaining a sequence of sequentially sorted images acquired by all cameras at the same time, wherein the image sequence includes at least a first region image and a second region image; perform grayscale enhancement and noise suppression processing on the image sequence, and input it into a regression network to obtain a coarsely aligned image at the same time; perform SIFT feature point extraction on the coarsely aligned image, and remove outliers in the SIFT feature point extraction using a random sampling consensus algorithm, and then perform projection matrix transformation to obtain a seamlessly stitched panoramic image; perform correction and rectangularization processing on the seamlessly stitched panoramic image to obtain a panoramic image of the fully mechanized mining face roof; The YOLOv8 detection module, connected to the image stitching module, is configured to: introduce the CBAM attention mechanism into the Backbone layer of the YOLOv8 detection model, and use the CIoU loss function in the YOLOv8 detection model to obtain a target YOLOv8 detection model; and use the target YOLOv8 detection model to perform bending deformation feature recognition on the panoramic image of the fully mechanized mining face roof to obtain the two-dimensional deformation features of the fully mechanized mining face roof.

7. The intelligent multi-modal recognition system for bending deformation of the roof of a fully mechanized mining face according to claim 5, characterized in that, The lidar detection unit includes: M mining explosion-proof LiDARs, spaced at a second preset distance, are all installed on the top beam of the fully mechanized mining face support and are configured to scan the top plate of the fully mechanized mining face at a preset frequency to obtain real-time point clouds. The point cloud preprocessing module, connected to the M mining explosion-proof LiDAR units, is configured to: receive the real-time point cloud and downsample the real-time point cloud using a voxel grid to form downsampled real-time point cloud data; statistically analyze the point cloud neighborhood distance distribution of the downsampled real-time point cloud data to form neighborhood points; remove outlier points whose average distance from the neighborhood points is greater than a preset multiple to obtain real-time point cloud data after removing interference points; and extract features from the real-time point cloud data after removing interference points to obtain point cloud feature data. The point cloud registration module, connected to the point cloud preprocessing module, is configured to: use information entropy to filter out structurally stable low-entropy points from the point cloud feature data, and use Euclidean distance to perform spatial clustering of the low-entropy points to form key points; then use the ICP algorithm to perform point cloud registration with the key points in the reference point cloud to obtain the three-dimensional deformation features of the fully mechanized mining face roof.

8. The intelligent recognition system for multimodal bending deformation of the roof of a fully mechanized mining face according to claim 7, characterized in that: The data fusion unit is configured to: perform coordinate alignment based on the intrinsic parameter matrix of the camera and the extrinsic parameter matrix of the explosion-proof LiDAR, map the two-dimensional deformation features of the fully mechanized mining face roof to the three-dimensional deformation features of the fully mechanized mining face roof, and establish the correspondence between pixels and point clouds; when the visual detection confidence level formed based on the two-dimensional deformation features and the three-dimensional deformation features of the fully mechanized mining face roof is greater than or equal to a preset two-dimensional value, and the corresponding point cloud deformation is greater than a preset three-dimensional value, a multimodal data fusion result is obtained; The early warning linkage unit includes an edge computing terminal and an audible and visual early warning device; The edge computing terminal is configured to: compare the multimodal data fusion result with a preset deformation threshold to obtain a comparison result; obtain the multimodal data fusion result according to the triggering of hierarchical early warning; and trigger a hierarchical early warning when the comparison result meets the preset early warning conditions. If the comparison result shows that the preset warning conditions are not met, the monitoring of the roof of the fully mechanized mining face will continue.

Citation Information

Patent Citations

  • Traffic rail deformation detection method and device based on multi-modal three-dimensional point cloud fusion

    CN118485898A

  • Multi-view panoramic point cloud splicing method

    CN120219158A