Underground coal mine personnel behavior recognition system and method based on deep learning

By combining deep learning with 3D point cloud and 2D image data acquired by TOF and industrial cameras, and using deep convolutional neural networks for processing and feature fusion, the problem of low accuracy and poor reliability in the identification of personnel behavior in coal mines has been solved, achieving high-precision detection in complex environments.

CN120853255APending Publication Date: 2025-10-28TIANJIN UNIV OF SCI & TECH
View PDF 12 Cites 0 Cited by

Patent Information

Application Number
CN202510915437.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-03
Publication Date
2025-10-28

AI Technical Summary

Technical Problem

In the current technology for recognizing personnel behavior in underground coal mines, the accuracy and reliability of two-dimensional image information recognition are insufficient, the three-dimensional point cloud data suffers from severe noise interference, and the establishment of training datasets for diverse behavior categories consumes a lot of manpower and time, making it difficult to effectively combine point cloud and image information to improve detection accuracy.

Method used

Using deep learning, 3D point cloud data is acquired through a TOF camera and 2D image data is acquired through an industrial camera. A deep convolutional neural network is combined to design filtering layers, structural layers, and feature layers to process the point cloud data. A mapping map between the point cloud and the image is established using feature indexing, and feature fusion is performed to achieve behavior recognition.

Benefits of technology

It improves the accuracy and reliability of behavior detection in complex downhole environments, reduces data redundancy, maintains the stability and robustness of detection, and enhances the system's adaptability and generalization ability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120853255A_ABST
    Figure CN120853255A_ABST
Patent Text Reader

Abstract

The invention discloses an underground coal mine personnel behavior recognition system and method based on deep learning. The recognition system comprises a first data acquisition module, a second data acquisition module, a data feature extraction network, a data feature fusion module and a behavior recognition module. The feature data extraction network adopts a deep learning convolutional network; the deep learning convolutional network comprises a three-dimensional point cloud data processing unit, a two-dimensional image data processing unit and a feature index unit; the three-dimensional point cloud data processing unit is composed of a filtering layer, a structural layer and a feature layer; according to the method, point cloud neighborhood information is counted for three-dimensional point cloud data, a point cloud mass center is calculated, and neighborhood information of point cloud horizontal and vertical coordinates is divided for optimized three-dimensional point cloud data, so that a topological relation of a point cloud space is obtained; according to the invention, in a mine environment with complex environment and poor light, the measured three-dimensional point cloud information and the strong feature learning ability of the deep neural network are utilized to improve the accuracy and reliability of behavior detection and recognition of mine personnel.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention pertains to underground personnel detection technology in coal mines, and particularly relates to a deep learning-based system and method for recognizing the behavior of underground personnel in coal mines. Background Art

[0002] With the rapid development of deep learning technology, its application in fields such as face recognition and object detection has become increasingly widespread. In the field of underground pedestrian detection in coal mines, traditional recognition methods mainly rely on two-dimensional image information. However, in the complex underground environment, the light is dim and the dust interference is severe, resulting in insufficient recognition accuracy and reliability [1][2]. Although the technology for acquiring and processing three-dimensional point cloud data has made progress in recent years, in the underground environment, due to factors such as light changes and noise interference, the acquired point cloud data often has large noise and defects, which affects the accuracy of the recognition results [2][3].

[0003] To overcome these problems, methods for fusing three-dimensional point cloud data with two-dimensional image information were explored in order to improve the accuracy of underground personnel behavior detection through deep learning technology [4]. However, how to effectively combine point cloud and image information and design a deep convolutional neural network that adapts to the complex underground environment is still a technical problem that needs to be solved [5]. In addition, in order to ensure the generalization ability and adaptability of the detection algorithm, it is necessary to establish a training dataset containing diverse behavior categories, which not only requires a large amount of sample data, but also requires a lot of manpower and time for annotation work [6][7].

[0004] Therefore, developing a method for recognizing personnel behavior in coal mines that can make full use of three-dimensional point cloud data and two-dimensional image information and has good generalization ability is of great significance for improving the level of coal mine safety production management [2]. This requires in-depth research on data acquisition, feature extraction, and model training, and the development of deep convolutional neural networks that can adapt to complex underground environments in order to achieve more accurate and reliable personnel behavior detection and recognition [1][5].

[0005] The existing technology has the following drawbacks:

[0006] 1. Traditional methods for identifying the behavior of personnel in underground coal mines mainly rely on two-dimensional image information. In complex underground environments, the light is dim and dust interference is severe, resulting in insufficient accuracy and reliability of identification[8][9].

[0007] 2. Existing three-dimensional point cloud data processing methods often have large noise and defects in the point cloud data obtained in the underground environment due to factors such as light changes and noise interference, which affects the accuracy of the recognition results [7]

[10] .

[0008] 3. In the existing technology, it is difficult to effectively combine point cloud with image information, and there is a lack of a deep convolutional neural network that can adapt to complex downhole environments. It is impossible to make full use of three-dimensional point cloud data and two-dimensional image information to improve the accuracy of downhole personnel behavior detection [3]

[11] .

[0009] 4. In order to ensure the generalization ability and adaptability of the detection algorithm, it is necessary to establish a training dataset containing diverse behavioral categories. This requires not only a large amount of sample data, but also a lot of manpower and time for annotation work [9]

[12] .

[0010] Existing technologies suffer from low accuracy and poor reliability in recognizing the behavior of personnel underground in coal mines. Therefore, to address these issues, this invention provides a method for recognizing the behavior of personnel underground in coal mines based on deep learning-based 3D point cloud processing. Summary of the Invention

[0011] To address the technical problems existing in the prior art, this invention provides a deep learning-based system and method for recognizing the behavior of personnel in coal mines. This invention improves the accuracy and reliability of detecting and recognizing the behavior of personnel in mines in complex environments with poor lighting by utilizing measured three-dimensional point cloud information and the powerful feature learning capabilities of deep neural networks.

[0012] To address the technical problems existing in the prior art, the present invention adopts the following technical solution:

[0013] A deep learning-based underground coal mine personnel behavior recognition system, comprising a first data acquisition module, a second data acquisition module, a data feature extraction network, a data feature fusion module, and a behavior recognition module; the feature data extraction network employs a deep learning convolutional network; the deep learning convolutional network includes a 3D point cloud data processing unit, a 2D image data processing unit, and a feature indexing unit; the 3D point cloud data processing unit consists of a filtering layer, a structure layer, and a feature layer; wherein:

[0014] The first data acquisition module uses a 3D camera to detect the light reflected back from the target by detecting light pulses, calculates the distances of various targets inside the mine, and obtains three-dimensional point cloud data.

[0015] The second data acquisition module acquires target image data inside the mine using an industrial camera;

[0016] The three-dimensional point cloud data processing unit processes the three-dimensional point cloud data to obtain multi-scale point cloud feature data;

[0017] The two-dimensional image data processing unit extracts the features of the target image data in the mine to obtain two-dimensional image features;

[0018] The feature indexing unit indexes the topological relationships of the point cloud space to the pixel positions of the corresponding two-dimensional image features to construct a mapping map between the three-dimensional point cloud and the two-dimensional image.

[0019] The feature fusion module additively fuses multi-scale point cloud feature data with the mapping map to obtain the target feature fusion map.

[0020] The behavior recognition module categorizes personnel behavior scores in the target feature fusion map based on 3D target bounding boxes, and regresses the personnel coordinate positions in the target feature fusion map; finally, it outputs visualized recognition results.

[0021] Furthermore, the 3D point cloud data processing unit processes the 3D point cloud data to obtain multi-scale point cloud feature data; including:

[0022] By using a filtering layer to statistically analyze the neighborhood information of the 3D point cloud data and calculate the centroid of the point cloud, outliers are filtered out to obtain optimized 3D point cloud data.

[0023] The topological relationship of the point cloud space is obtained by dividing the neighborhood information of the horizontal and vertical coordinates of the optimized 3D point cloud data through the structural layer.

[0024] Multi-scale point cloud feature data is obtained by extracting the corresponding features of the mine detection target in the topological relationship of the point cloud space through the first feature layer.

[0025] Furthermore, the first data acquisition module uses a 3D camera to detect the light reflected back from the target by light pulse detection and calculates the distance of each target inside the mine to obtain three-dimensional point cloud data; that is, it uses a TOF camera to acquire three-dimensional point cloud data and infrared images inside the mine. The second data acquisition module uses an industrial camera to acquire target image data inside the mine, that is, it uses an industrial camera to acquire two-dimensional color images and grayscale image information inside the mine.

[0026] The present invention can also adopt the following technical solutions:

[0027] S1. Calculate the distances of various targets inside the mine by detecting the light reflected back from the target using a 3D camera light pulse detection system to obtain three-dimensional point cloud data;

[0028] S2. Acquire target image data inside the mine using an industrial camera;

[0029] S3. Process the 3D point cloud data to obtain multi-scale point cloud feature data; where:

[0030] 301. Calculate the centroid of the point cloud by statistically analyzing the neighborhood information of the 3D point cloud data, and filter out outliers to obtain optimized 3D point cloud data;

[0031] 302. Obtain the topological relationship of the point cloud space by dividing the neighborhood information of the horizontal and vertical coordinates of the optimized 3D point cloud data.

[0032] 303. Extract the corresponding features of the mine detection target from the topological relationship of the point cloud space to obtain multi-scale point cloud feature data;

[0033] S4. Extract the features of target image data in the mine to obtain two-dimensional image features;

[0034] S5. Index the topological relationships of the point cloud space to the pixel positions of the corresponding two-dimensional image features to construct a mapping map between the three-dimensional point cloud and the two-dimensional image.

[0035] S6. Perform additive fusion of multi-scale point cloud feature data and mapping map to obtain target feature fusion map;

[0036] S7. Classify the personnel behavior scores of the target feature fusion map based on the 3D target bounding box, and regress the personnel coordinate positions of the target feature fusion map; finally, output the visualized recognition results.

[0037] Beneficial effects

[0038] 1. This invention establishes a two-dimensional image and three-dimensional point cloud data acquisition system for mines, comprehensively utilizing the advantages of TOF cameras and industrial cameras. It can acquire clear three-dimensional point cloud data and two-dimensional image information in complex mine environments, overcoming the problem of low accuracy of traditional image detection in complex environments and improving the accuracy of mine personnel behavior detection. Through this collaborative acquisition of multi-source data, the details and changes in the mine environment can be better captured, ensuring the comprehensiveness and reliability of the detection process.

[0039] 2. The deep convolutional neural network for processing 3D point clouds in mines constructed in this invention, through filtering network modules and point cloud structured partitioning networks, can effectively process noise information in 3D point clouds, improve the quality of point cloud data, and lay a solid foundation for subsequent feature extraction and processing. This network structure design can not only reduce redundant information in the data, but also maintain the stability and consistency of the data under different environmental conditions, thereby providing more reliable data support for subsequent analysis.

[0040] 3. This invention employs a design that combines a feature extraction network and an attention mechanism layer, which can effectively extract features at different scales, fully explore deep clues in the data, and improve the comprehensiveness and accuracy of feature extraction. Through this design, key features can be identified in diverse data, and the ability to sensitively identify details in complex scenes can be maintained, ensuring the accuracy and reliability of detection.

[0041] 4. This invention effectively integrates three-dimensional point cloud data with two-dimensional images and establishes a mapping relationship between image features and point clouds using feature indexing, thereby achieving effective integration of multi-source data and improving the accuracy and reliability of mine personnel behavior detection. This fusion method not only improves the overall utilization rate of data, but also enables the system to perform consistent feature recognition under different perspectives and environmental conditions, enhancing the robustness and stability of detection.

[0042] 5. This invention improves the system's adaptability and generalization ability by labeling and training the collected data, combined with continuously updated datasets and model parameters, enabling it to maintain high detection accuracy in complex and ever-changing mining environments. Through self-learning and adaptive mechanisms, the system can quickly adjust and optimize detection strategies in constantly changing environments, ensuring accurate detection results under various complex conditions. Attached Figure Description

[0043] Figure 1 This is a schematic diagram of a mine personnel behavior recognition system based on deep learning in this invention;

[0044] Figure 2 This is a schematic diagram of the structure of the three-dimensional point cloud data processing unit for mines in this invention;

[0045] Figure 3 The feature fusion module structure for fusing three-dimensional point clouds and images in mines in this invention is practical;

[0046] Figure 4 This is the annotation of the three-dimensional point cloud dataset for the behavior detection and recognition of mine personnel by the behavior recognition module in this invention. DETAILED DESCRIPTION

[0047] The following is in conjunction with the appendix Figures 1-4 The present invention will be described in detail below:

[0048] This invention provides a deep learning-based system for recognizing the behavior of personnel in mines. The system includes a first data acquisition module, a second data acquisition module, a data feature extraction network, a data feature fusion module, and a behavior recognition module. The feature data extraction network employs a deep learning convolutional network. This deep learning convolutional network includes a 3D point cloud data processing unit, a 2D image data processing unit, and a feature indexing unit. Figure 1 As shown; the first data acquisition module uses a 3D camera to detect light reflected from the target using light pulses and calculates the distances to various targets inside the mine to obtain three-dimensional point cloud data; that is, it acquires three-dimensional point cloud data and infrared images inside the mine using a TOF camera. The second data acquisition module acquires target image data inside the mine using an industrial camera, that is, it acquires two-dimensional color images and grayscale image information inside the mine using an industrial camera. The feature data extraction network is constructed through software programs. Figure 2 The deep convolutional neural network shown is used for processing 3D point clouds in the mine. Figure 3 The shown example uses a deep convolutional neural network to fuse the 3D point cloud of the mine with an image, i.e., a data feature fusion module; the behavior recognition module performs the following processing on the collected 3D point cloud data of the mine interior: Figure 4 The diagram shows a labeled 3D point cloud dataset for detecting and recognizing personnel behavior in mines. Using this labeled dataset, deep convolutional neural networks (DCNNs) for processing 3D point clouds and for fusing 3D point clouds with images are trained to obtain system parameters. The dataset and system parameters are updated based on the test results. The trained system is then installed and deployed on a computer at the mine site. The collected 3D point cloud and 2D image data are analyzed using these DCNNs to detect and recognize personnel behavior within the mine. The optimal detection result is obtained based on the category score, and the result is displayed in 3D using different colored bounding boxes, thus completing the detection and recognition of personnel behavior in the mine. Specifically:

[0049] The 3D point cloud data processing unit processes the 3D point cloud data to obtain multi-scale point cloud feature data, including:

[0050] By using a filtering layer to statistically analyze the neighborhood information of the 3D point cloud data and calculate the centroid of the point cloud, outliers are filtered out to obtain optimized 3D point cloud data.

[0051] The topological relationship of the point cloud space is obtained by dividing the neighborhood information of the horizontal and vertical coordinates of the optimized 3D point cloud data through the structural layer.

[0052] Multi-scale point cloud feature data is obtained by extracting the corresponding features of the mine detection target in the topological relationship of the point cloud space through the first feature layer.

[0053] The two-dimensional image data processing unit extracts the features of the target image data in the mine to obtain two-dimensional image features;

[0054] The feature indexing unit indexes the topological relationships of the point cloud space to the pixel positions of the corresponding two-dimensional image features to construct a mapping map between the three-dimensional point cloud and the two-dimensional image.

[0055] The feature fusion module additively fuses multi-scale point cloud feature data with the mapping map to obtain the target feature fusion map.

[0056] The behavior recognition module identifies target features based on a 3D bounding box, fuses them into a graph, calculates personnel behavior category scores and location regression, and outputs visualized recognition results. Details are as follows:

[0057] 1. Composition of the mine interior 2D image and 3D point cloud data acquisition system

[0058] To achieve underground worker behavior recognition in coal mines based on deep learning-based 3D point cloud processing, it is first necessary to acquire 3D point cloud data of the mine interior using a 3D camera. For example... Figure 1 The system depicts a 2D image and 3D point cloud data acquisition system for the interior of a coal mine. An industrial camera acquires RGB color and grayscale images of the mine interior, while a Time-of-Flight (TOF) camera emits light pulses to detect light reflected back from targets within the mine, records the flight time, calculates the depth information of each target, and acquires 3D point cloud data. The TOF 3D camera uses an infrared laser as its light source, acquiring not only 3D data but also infrared grayscale images, providing supplementary infrared information for detecting personnel behavior within the mine. TOF emits infrared lasers, allowing for long-distance measurements (up to 100 meters), and is unaffected by surface grayscale and features, making it more suitable for the dim environment inside coal mines and improving the accuracy of personnel behavior recognition. The system uses a computer to acquire the collected images and point cloud data, processes and displays the measurement data using system processing software, and then uses a designed deep learning-based network for mine personnel behavior detection and recognition to process the point cloud data and identify personnel behavior within the mine.

[0059] 2. Mine Personnel Behavior Detection and Recognition Based on Deep Learning 3D Point Cloud Processing

[0060] 2.1 Construction of a Deep Convolutional Neural Network for 3D Point Cloud Processing in Mines

[0061] To detect and identify miners from 3D point clouds in a mine using deep learning methods, a deep convolutional neural network structure diagram is shown below. Figure 2As shown, 3D point cloud data of mines acquired using TOF cameras typically contains noisy information, requiring filtering. Therefore, the designed deep convolutional neural network structure for processing 3D point clouds of mines first includes a filtering network layer. This layer calculates the centroid of the point cloud by statistically analyzing its neighborhood information and filters out noise by removing outliers. Furthermore, 3D point cloud data contains depth information compared to 2D images, but it is usually unordered and cannot be directly processed by convolution. Therefore, a point cloud structuring network is used after the filtering network module. This structuring network mainly constructs the topological relationships of the point cloud space to determine the neighborhood information of the horizontal and vertical coordinates, resulting in a structured 3D point cloud of the mine. A feature extraction network is then used to extract features from the structured 3D point cloud. This feature extraction network mainly consists of convolutional layers and pooling layers. Since the size of personnel in the mine is relatively small compared to the mine itself, an attention mechanism layer is added to the feature extraction network to enhance the extraction of personnel information features and avoid information loss. Simultaneously, the feature extraction network extracts multi-scale features to distinguish information about personnel from other equipment in the mine. Finally, the extracted features are used to predict the location information and behavior of miners using 3D bounding boxes. The prediction of the 3D bounding boxes includes a score for the personnel behavior category within the bounding box, and regression adjustments are made to the bounding box position to obtain the regression coordinates (Reg) of the 3D bounding box position of the defective target. The predicted personnel behavior category and location coordinates are then displayed in a 3D point cloud using the color and position of the bounding box, thus realizing the detection and recognition of miner behavior.

[0062] 2.2 Construction of a Deep Convolutional Neural Network for Fusion of 3D Point Clouds and Images in a Mine

[0063] Three-dimensional point cloud data from mines contains depth information. When two-dimensional image information is insufficient or of poor quality, utilizing depth information can improve the accuracy of mine personnel behavior detection and recognition. However, two-dimensional image information possesses certain color and grayscale information that point clouds lack. Therefore, fusing two-dimensional image information with three-dimensional point cloud data can further improve the accuracy of mine personnel behavior detection and recognition. Especially inside mines, where the environment is complex, lighting is dim, monitoring video quality is poor, or interference from dust and spray is present, a three-dimensional defect detection network that fuses point clouds and images can still ensure the accuracy of mine personnel behavior detection and recognition. The structure diagram of the deep convolutional neural network for fusing three-dimensional point clouds and images in this solution is shown below. Figure 3As shown, the pseudo-image generated by the point cloud through the point cloud feature extraction network is subjected to multi-scale feature extraction through a convolutional network, and the feature map of the 2D image is extracted through a convolutional network. The key to image-point cloud fusion lies in establishing the mapping relationship between the image feature map and the features of the point cloud. In this invention, when the point cloud is structurally divided, the point cloud is indexed to the corresponding pixel position of the 2D image, thus facilitating the acquisition of the mapping relationship between the point cloud and the image. By fusing the image and the corresponding point cloud data through feature indexing, and using the fused features for 3D target bounding box prediction, the detection and recognition of mine personnel behavior can be guaranteed in complex mining environments.

[0064] 2.3 Creation and Annotation of 3D Point Cloud Dataset for Mine Personnel Behavior Detection and Recognition

[0065] This invention presents a method for detecting and recognizing mine personnel behavior based on deep learning-based 3D point cloud processing. It utilizes a deep convolutional neural network (DNN) to learn 3D point cloud features and fuses 3D and 2D image features to improve the accuracy of mine personnel behavior detection and recognition. The training effect of the DNN is crucial for ensuring the detection results. The DNN requires training with a large amount of mine point cloud and image data; therefore, it is necessary to create and annotate a 3D point cloud dataset for mine personnel behavior detection and recognition. The 3D point cloud dataset for mine personnel behavior detection and recognition is acquired through long-term shooting at the mine site using TOF and industrial cameras. The acquired 3D point cloud and image data are manually trained. The annotation of the 3D point cloud involves manually selecting 3D target boxes using 3D point cloud display and interactive software, and different colors are used to distinguish different behavior categories (e.g., smoking, wearing a safety helmet, falling, etc.). The types of annotations can be added or removed according to the actual needs within the mine, giving the mine personnel behavior detection good generalization ability and adaptability. The creation and annotation of the 3D point cloud dataset for mine personnel behavior detection and recognition is as follows: Figure 4 As shown.

[0066] Example 1:

[0067] This invention provides a method for detecting and recognizing the behavior of personnel in mines based on deep learning-based 3D point cloud processing. The specific implementation steps are as follows:

[0068] Step 1: Build a system for acquiring 2D images and 3D point cloud data inside the mine.

[0069] Step 101: Configure the TOF camera and industrial camera. The TOF camera model is Velodyne Puck LITE, and the industrial camera model is Hikvision DS-U3V09E6-I. Step 102: Set the camera acquisition parameters: TOF camera resolution is 1024×1024, frame rate is 20Hz, and exposure time is 100ns; industrial camera resolution is 2048×1536, frame rate is 25Hz, and exposure time is 1 / 1000s. Step 103: Establish a data transmission channel to transmit the acquired data to the central processing unit.

[0070] Step 2: Obtain 3D point cloud data and 2D image information of the mine interior.

[0071] Step 201: Emits an infrared laser with a wavelength of 905nm using a TOF camera, receives the reflected light, and calculates the target distance to obtain 3D point cloud data of the mine interior. Step 202: Acquires RGB color and grayscale images of the mine interior using an industrial camera. Step 203: Records the acquisition timestamp for subsequent data correlation.

[0072] Step 3: Preprocess the acquired 3D point cloud data.

[0073] Step 301: Use a statistical filtering algorithm to remove outliers from the point cloud. Step 302: Construct the topological relationships of the point cloud space and determine the neighborhood information of the horizontal and vertical coordinates of the point cloud. Step 303: Use a convolutional neural network to extract multi-scale features of the point cloud, including microscopic features, local features, and global features.

[0074] Step 4: Construct a deep convolutional neural network for processing 3D point clouds in the mine.

[0075] Step 401: Design a point cloud filtering network module, consisting of one 2D convolutional layer and one max pooling layer, with a 3×3 kernel and a 2×2 pooling window. Step 402: Design a point cloud structuring network, consisting of one 2D convolutional layer and one fully connected layer, with a 5×5 kernel. Step 403: Design a feature extraction network, consisting of two convolutional layers and two pooling layers, each with a 3×3 kernel. Step 404: Design an attention mechanism layer, employing a self-attention network structure to assign feature weights to different regions of the point cloud.

[0076] Step 5: Construct a deep convolutional neural network for fusing the 3D point cloud and image of the mine.

[0077] Step 501: Extract multi-scale features from point cloud and image using point cloud feature extraction networks and image feature extraction networks respectively. Step 502: Establish a feature index and implement the mapping relationship between point cloud features and image features using a convolutional neural network. Step 503: Design a feature fusion network, employing an additive fusion strategy to directly add point cloud and image features. Step 504: Design a 3D target bounding box prediction network, composed of multiple convolutional and regression layers, to predict the location and behavior category of mine personnel.

[0078] Step 6: Create a 3D point cloud dataset for detecting and recognizing the behavior of mine personnel.

[0079] Step 601: Acquire historical data, including 5 million 3D point cloud data points and 2 million image data points. Step 602: Preprocess the acquired data, including data normalization and data augmentation. Step 603: Manually annotate the 3D point cloud data using a 3D bounding box tool. Annotate the personnel behavior categories, including "normal walking," "stopped," "low posture," "high posture," "carrying items," "not wearing a helmet," and "wearing a helmet," totaling 14 categories. Step 604: Establish a dataset index to support random sampling and iterative training, with a training batch size of 64.

[0080] Step 7: Train the model and test the detection results.

[0081] Step 701: Initialize model parameters using the Xavier initialization method. Step 702: Train the point cloud processing network and fusion network using the labeled dataset, employing the Adam optimization algorithm with a learning rate of 0.001. Step 703: Use an iterative training strategy, training for 1000 epochs, with an early stopping strategy where training stops after 10 iterations if there is no improvement. Step 704: Evaluate the detection performance. If the average accuracy is below 95%, return to step 702 to continue training; otherwise, proceed to step 8.

[0082] Step 8: Deploy the system to detect and identify the behavior of mine personnel.

[0083] Step 801: Deploy the trained system on the mine's on-site computer, configuring it with an Intel i7-9700K processor and 16GB of memory. Step 802: Acquire 3D point cloud and 2D image data in real time at a frame rate of 10Hz. Step 803: Utilize the deployed network to detect and recognize the behavior of mine personnel, with a detection interval of 30 seconds. Step 804: Display the detection results using 3D bounding boxes, employing different colored 3D cubes to represent different behavior categories, thus achieving intelligent inspection.

[0084] Example 2:

[0085] This invention provides a method for detecting and recognizing the behavior of personnel in mines based on deep learning-based 3D point cloud processing. The specific implementation steps are as follows:

[0086] Step 1: Build a system for acquiring 2D images and 3D point cloud data inside the mine.

[0087] Step 101: Configure the TOF camera and industrial camera. The TOF camera model is Velodyne HDL-32E, and the industrial camera model is Hikvision DS-U3V09E6-I. Step 102: Set the camera acquisition parameters: TOF camera resolution is 1920×1080, frame rate is 30Hz, and exposure time is 200ns; industrial camera resolution is 2048×1536, frame rate is 30Hz, and exposure time is 1 / 1000s. Step 103: Establish a data transmission channel to transmit the acquired data to the central processing unit.

[0088] Step 2: Obtain 3D point cloud data and 2D image information of the mine interior.

[0089] Step 201: Emits an infrared laser with a wavelength of 905nm using a TOF camera, receives the reflected light, and calculates the target distance to obtain 3D point cloud data of the mine interior. Step 202: Acquires RGB color and grayscale images of the mine interior using an industrial camera. Step 203: Records the acquisition timestamp for subsequent data correlation.

[0090] Step 3: Preprocess the acquired 3D point cloud data.

[0091] Step 301: Use a statistical filtering algorithm to remove outliers from the point cloud. Step 302: Construct the topological relationships of the point cloud space and determine the neighborhood information of the horizontal and vertical coordinates of the point cloud. Step 303: Use a convolutional neural network to extract multi-scale features of the point cloud, including microscopic features, local features, and global features.

[0092] Step 4: Construct a deep convolutional neural network for processing 3D point clouds in the mine.

[0093] Step 401: Design a point cloud filtering network module, consisting of one 3D convolutional layer and one max pooling layer. The convolutional kernel size is 3×3×3, and the pooling window size is 2×2×2. Step 402: Design a point cloud structuring network, consisting of one 3D convolutional layer and one fully connected layer. The convolutional kernel size is 5×5×5. Step 403: Design a feature extraction network, consisting of three convolutional layers and two pooling layers. The convolutional kernel size is 3×3×3 for each layer. Step 404: Design an attention mechanism layer, employing a self-attention network structure to assign feature weights to different regions of the point cloud.

[0094] Step 5: Construct a deep convolutional neural network for fusing the 3D point cloud and image of the mine.

[0095] Step 501: Extract multi-scale features from point cloud and image using point cloud feature extraction networks and image feature extraction networks respectively. Step 502: Establish a feature index and implement the mapping relationship between point cloud features and image features using a convolutional neural network. Step 503: Design a feature fusion network, employing a multiplicative fusion strategy to multiply point cloud and image features. Step 504: Design a 3D target bounding box prediction network, composed of multiple convolutional and regression layers, to predict the location and behavior category of miners.

[0096] Step 6: Create a 3D point cloud dataset for detecting and recognizing the behavior of mine personnel.

[0097] Step 601: Acquire historical data, including 3 million 3D point cloud data and 1 million image data. Step 602: Preprocess the acquired data, including data normalization and data augmentation. Step 603: Manually annotate the 3D point cloud data using a 3D bounding box tool. Annotate the personnel behavior categories, including "normal walking," "stopped," "low posture," "high posture," "carrying items," "not wearing a helmet," and "wearing a helmet," totaling 14 categories. Step 604: Establish a dataset index, supporting random sampling and iterative training, with a training batch size of 32.

[0098] Step 7: Train the model and test the detection results.

[0099] Step 701: Initialize model parameters using the He initialization method. Step 702: Train the point cloud processing network and fusion network using the labeled dataset, employing the SGD optimization algorithm with a learning rate of 0.0001. Step 703: Use an iterative training strategy, training for 500 epochs, with an early stopping strategy: stopping training if there is no improvement after 5 epochs. Step 704: Evaluate the detection performance. If the average accuracy is below 90%, return to step 702 to continue training; otherwise, proceed to step 8.

[0100] Step 8: Deploy the model to detect and identify the behavior of mine personnel.

[0101] Step 801: Deploy the trained model on the mine's on-site computer, configuring it with an Intel Xeon E-2276M processor and 32GB of memory. Step 802: Real-time acquisition of 3D point cloud and 2D image data at a frame rate of 15Hz. Step 803: Utilize the deployed network to detect and recognize the behavior of mine personnel, with a detection interval of 60 seconds. Step 804: Display the detection results using 3D bounding boxes, employing different colored 3D cubes to represent different behavior categories, achieving intelligent inspection.

[0102] References

[0103] [1] CN113963420A, A 3D face recognition method based on deep learning, Xinxun Digital Technology (Hangzhou) Co., Ltd.

[0104] [2] CN115311241A, A method for detecting pedestrians in coal mines based on image fusion and feature enhancement, Tiandi (Changzhou) Automation Co., Ltd., +1

[0105] [3] CN111563923A, Method and related apparatus for obtaining dense depth maps, Zhejiang Dahua Technology Co., Ltd.

[0106] [4] CN113052066A, A multimodal fusion method based on multiple views and image segmentation in 3D target detection, University of Science and Technology of China

[0107] [5] CN112258631A, A method and system for three-dimensional target detection based on deep neural networks, Hohai University, Changzhou Campus

[0108] [6] CN112767391A, A method for locating defects in power grid line components by integrating 3D point cloud and 2D image, State Grid Fujian Electric Power Co., Ltd., +1

[0109] [7] CN112102472A, A method for densifying sparse 3D point clouds, Beijing University of Aeronautics and Astronautics

[0110] [8] CN110862033A, An intelligent early warning detection method applied to coal mine inclined shaft winches, CITIC Heavy Industries Kaicheng Intelligent Equipment Co., Ltd.

[0111] [9] CN112528979A, Obstacle Detection Method and System for Substation Inspection Robot, Chengdu University of Information Technology

[0112]

[10] CN117392192A, Image depth prediction method, apparatus, readable storage medium and electronic device, BYD Co., Ltd.

[0113]

[11] CN116597264A, A method for detecting 3D point cloud targets by integrating 2D image semantics, Nanjing University of Science and Technology

[0114]

[12] CN113052835A, A method and system for detecting medicine boxes based on the fusion of three-dimensional point cloud and image data, Jiangsu Xunjie Equipment Technology Co., Ltd.

[0115] Although the present invention has been described above, the present invention is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many modifications under the guidance of the present invention without departing from the spirit of the present invention, and these modifications are all within the protection scope of the present invention.

Claims

1. A deep learning-based system for recognizing the behavior of personnel in underground coal mines, characterized in that: The recognition system includes a first data acquisition module, a second data acquisition module, a data feature extraction network, a data feature fusion module, and a behavior recognition module; the feature data extraction network adopts a deep learning convolutional network; the deep learning convolutional network includes a 3D point cloud data processing unit, a 2D image data processing unit, and a feature indexing unit; the 3D point cloud data processing unit consists of a filtering layer, a structure layer, and a feature layer; wherein: The first data acquisition module uses a 3D camera to detect the light reflected back from the target by detecting light pulses, calculates the distances of various targets inside the mine, and obtains three-dimensional point cloud data. The second data acquisition module acquires target image data inside the mine using an industrial camera; The three-dimensional point cloud data processing unit processes the three-dimensional point cloud data to obtain multi-scale point cloud feature data; The two-dimensional image data processing unit extracts the features of the target image data in the mine to obtain two-dimensional image features; The feature indexing unit indexes the topological relationships of the point cloud space to the pixel positions of the corresponding two-dimensional image features to construct a mapping map between the three-dimensional point cloud and the two-dimensional image. The feature fusion module additively fuses multi-scale point cloud feature data with the mapping map to obtain the target feature fusion map. The behavior recognition module categorizes personnel behavior scores in the target feature fusion map based on 3D target bounding boxes, and regresses the personnel coordinate positions in the target feature fusion map; finally, it outputs visualized recognition results.

2. The deep learning-based underground personnel behavior recognition system for coal mines according to claim 1, characterized in that: The 3D point cloud data processing unit processes the 3D point cloud data to obtain multi-scale point cloud feature data, including: By using a filtering layer to statistically analyze the neighborhood information of the 3D point cloud data and calculate the centroid of the point cloud, outliers are filtered out to obtain optimized 3D point cloud data. The topological relationship of the point cloud space is obtained by dividing the neighborhood information of the horizontal and vertical coordinates of the optimized 3D point cloud data through the structural layer. Multi-scale point cloud feature data is obtained by extracting the corresponding features of the mine detection target in the topological relationship of the point cloud space through the first feature layer.

3. The deep learning-based underground personnel behavior recognition system for coal mines according to claim 1, characterized in that: The first data acquisition module uses a 3D camera to detect the light reflected back from the target by light pulse detection and calculates the distance of each target inside the mine to obtain three-dimensional point cloud data; that is, it uses a TOF camera to acquire three-dimensional point cloud data and infrared images inside the mine. The second data acquisition module uses an industrial camera to acquire target image data inside the mine, that is, it uses an industrial camera to acquire two-dimensional color images and grayscale image information inside the mine.

4. A deep learning-based method for recognizing the behavior of underground personnel in coal mines, characterized in that: The method, based on the system of any one of claims 1-3, includes the following steps: S1. Calculate the distances of various targets inside the mine by detecting the light reflected back from the target using a 3D camera light pulse detection system to obtain three-dimensional point cloud data; S2. Acquire target image data inside the mine using an industrial camera; S3. Process the 3D point cloud data to obtain multi-scale point cloud feature data; where: 301 pairs of 3D point cloud data were statistically analyzed to obtain point cloud centroids by calculating the neighborhood information of point clouds and filtering out outliers to obtain optimized 3D point cloud data. 302 pairs of optimized 3D point cloud data were used to divide the neighborhood information of the horizontal and vertical coordinates of the point cloud to obtain the topological relationship of the point cloud space; 303 Extract the corresponding features of the mine detection target from the topological relationship of the point cloud space to obtain multi-scale point cloud feature data; S4. Extract the features of target image data in the mine to obtain two-dimensional image features; S5. Index the topological relationships of the point cloud space to the pixel positions of the corresponding two-dimensional image features to construct a mapping map between the three-dimensional point cloud and the two-dimensional image. S6. Perform additive fusion of multi-scale point cloud feature data and mapping map to obtain target feature fusion map; S7. Classify the personnel behavior scores of the target feature fusion map based on the 3D target bounding box, and regress the personnel coordinate positions of the target feature fusion map; finally, output the visualized recognition results.

Citation Information

Patent Citations

  • Intelligent early warning detection method applied to coal mine inclined shaft winch

    CN110862033A

  • Method for obtaining dense depth map and related device

    CN111563923A

  • Sparse three-dimensional point cloud densification method

    CN112102472A

  • Three-dimensional target detection method and system based on deep neural network

    CN112258631A

  • Substation inspection robot obstacle discrimination method and system

    CN112528979A