Coal bunker material state monitoring method based on multi-modal data fusion
Through multimodal data fusion technology, combined with lidar and multi-spectral camera data, real-time and accurate monitoring of coal bin material status is achieved, solving the problems of insufficient real-time data processing and low recognition accuracy in the existing technology, and improving the reliability and stability of the system.
Patent Information
- Application Number
- CN202510299323.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-06-20
AI Technical Summary
The existing coal silo material detection methods have problems such as insufficient real-time data processing and low accuracy in identification of a single data source, making it difficult to achieve real-time and accurate coal level monitoring of coal silo.
The coal silo material status monitoring method based on multimodal data fusion is adopted, combined with lidar scanning and multispectral camera technology, and output is generated through data preprocessing, data downsampling and feature layering, global fusion, local fusion, location information encoding, and feature aggregation and residual connection output to generate 3D fusion images for detection.
It improves the real-time and accuracy of coal bin material status monitoring, enhances the reliability and stability of the system, and can more truly reflect the actual situation of the materials inside the coal bin.
Smart Images

Figure CN120182767A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for detecting materials inside a coal bunker, specifically a method for monitoring the state of coal bunker materials based on multi-modal data fusion. Background Art
[0002] China is a large energy country with rich coal resources, and coal mines are spread all over the country. The height of the coal level in the coal bunker directly affects the normal operation of coal mine safety production. The coal bunker needs to maintain a certain coal storage capacity. When the coal level is lower than the minimum limit value, the coal bunker becomes empty and the air flow is short-circuited; when the coal level exceeds the maximum limit value, situations such as coal piling in the coal bunker and the coal conveyor being stuck may occur, resulting in safety accidents. With the continuous improvement of the country's requirements for the intelligent level of coal mines, it has become an urgent requirement to monitor the actual state of the coal level in the coal bunker in real time and accurately. Preventing the occurrence of two phenomena, namely coal bunker overflow (choking) and emptying (empty bunker), is the key link, which is of great significance for the safe and efficient coal mining of coal mine enterprises and the steady development of China's economy.
[0003] Currently, the detection methods for materials inside the coal bunker are mainly divided into two categories: one is the contact measurement method, and the other is the non-contact measurement method. Contact measurements mainly include the heavy hammer type, electrode type, capacitance type, etc. However, these methods all have insurmountable defects in actual use. ① The heavy hammer type level gauge measures the coal level height by using a heavy hammer to track the rise and fall of the material. However, the particle size of raw coal varies greatly, and it is easy to bury the heavy hammer, resulting in the level gauge being unable to work. ② The electrode type level gauge measures the coal level at corresponding heights by installing electrodes at different heights in the raw coal bunker. However, the difference in the moisture content of raw coal leads to a very large change in the conductivity of raw coal, and false signals are extremely easy to generate. ③ The capacitance type level gauge forms a capacitance by contacting the sensor probe with the object to be measured and acts on the output relay, but continuous measurement cannot be achieved. Non-contact measurements mainly include ultrasonic type, laser type, nuclear type, etc. ① The reflection type and ultrasonic type are greatly affected by coal dust and humidity. When the ratio of the detection depth to the detection radius exceeds its range, the detection accuracy drops significantly, and even detection cannot be carried out. ② The laser type level gauge emits a laser beam from the sensor to irradiate the surface of the material to be measured and receives the emitted light. By measuring the time from emission to reception, it is converted into the depth of the material level. However, during the detection of different coal levels, it is easy to cause irregular scattering of the reflected wave, affecting the detection accuracy. ③ The nuclear type level gauge uses nuclear radiation to measure the material level. However, nuclear radiation instruments are dangerous, and the rays will be scattered and absorbed by the atoms of the material, and the measurement accuracy decays exponentially with the thickness of the material layer.
[0004] When processing the lidar point cloud data of the coal bunker above the well, the size of the data volume is the core indicator of the algorithm's time complexity. For the efficiency of point cloud data processing, it is required to reduce the number of point clouds while retaining the basic features of the point clouds. Liu et al. applied the improved voxel downsampling method to point cloud registration, and the performance of point cloud registration was improved. Zheng Jianhua et al. used the random downsampling algorithm in the improved random forest algorithm, and the obtained algorithm has better classification performance. In a three-dimensional scene, the directions of targets vary widely. To improve the robustness of the model to target transformations (such as rotation and symmetry), traditional detectors usually adopt methods such as data augmentation or test-time augmentation. However, the effects of these methods are limited and will increase a large amount of time cost. To solve this problem, Wu et al. introduced multi-channel transformed voxel features based on the Voxel R-CNN model, and this method successfully improved the prediction accuracy of the model for target directions.
[0005] Zhang et al. improved on the VoteNet model and proposed the IA-SSD model. This model designed two learnable sampling strategies, namely class-aware sampling and center-aware sampling, to better preserve foreground points. The IA-SSD model generates higher-quality seed point features by improving the voting mechanism and uses them for target center point prediction, bounding box prediction, and target classification. Chen et al. proposed the Semantic-Augmented Set Abstraction (SASA) module, which improves the traditional farthest point sampling method based on the semantic information and coordinate information of points. The SASA module attaches semantic weights to the distance metric, enabling the sampled points to be distributed over a larger range while retaining more foreground points. This improvement solves the problem of retaining the foreground points of simple targets while ignoring the foreground points of difficult-to-detect targets. Considering that the accuracy of point cloud reconstruction will decrease when the feature point matching is insufficient or incorrect, the MVSNet (Multi View Stereo Network) network model is applied to three-dimensional reconstruction. The multi-scale convolutional network is used to fully learn the high-level features of the depth image, reducing the dependence of three-dimensional reconstruction on feature points in adverse environments such as sparse or repetitive textures. Summary of the Invention
[0006] The present invention aims to simultaneously solve two major technical pain points in the field of coal bunker monitoring: improving the real-time performance of data processing while maintaining the integrity of the basic features of the lidar point cloud data of the coal bunker coal level; and breaking through the bottleneck of the recognition accuracy of a single data source in the monitoring of the coal bunker material state. Therefore, a method for monitoring the coal bunker material state based on multi-modal data fusion is provided.
[0007] The present invention is implemented by the following technical solutions: A method for monitoring the material state of a coal bunker based on multi-modal data fusion, including data preprocessing, data downsampling and feature stratification, global fusion, local fusion, position information encoding, and feature aggregation and residual connection output; in data preprocessing, the three-dimensional point cloud data is first voxelized, and then in data downsampling and feature stratification, multi-level voxel features are extracted from the 3D voxel data through a 3D Backbone, and at the same time, multi-level image features are extracted from the 2D image data through a 2D Backbone; then, region proposal networks are used for the voxel features to generate 3D bounding box candidates; in global fusion, the centroid of the voxel is calculated through a centroid dynamic fusion network module, the centroid of the voxel is projected onto the image feature plane, and the image features are sampled to achieve global fusion and generate global fusion features; then, in local fusion, each bounding box candidate is divided into uniform voxel grids, the grid center position is encoded through a position information encoding module in position information encoding to generate position information encoding features, and local fusion is achieved through a grid point dynamic fusion module to generate local fusion features; finally, the global fusion features, local fusion features, and position information encoding features are concatenated and aggregated to generate comprehensive features, and a 3D fusion image is finally generated through a self-attention mechanism and a residual connection block, and the real-time state in the coal bunker is detected by the 3D fusion image.
[0008] The above method for monitoring the material state of a coal bunker based on multi-modal data fusion uses lidar scanning to generate three-dimensional point cloud data of the coal bunker and uses a multi-spectral camera to generate 2D image data of the coal bunker.
[0009] In the above method for monitoring the material state of a coal bunker based on multi-modal data fusion, the 3D Backbone network uses a 3D sparse convolutional neural network to extract voxel features, and the 2D Backbone extracts image features through a 2D convolutional neural network.
[0010] In the above method for monitoring the material state of a coal bunker based on multi-modal data fusion, the specific process of the position information encoding module generating position information encoding features is as follows: The position information encoding module PIE divides each bounding box candidate into uniform grids, and each grid has a center point z j , encodes the center point position of each grid to generate position information encoding features
[0011] In the above method for monitoring the material state of a coal bunker based on multi-modal data fusion, the specific process of generating local fusion features in local fusion is as follows: First, project each grid center point z j onto the image feature plane to obtain the corresponding image coordinates o j , and on the image feature plane, according to the projected image coordinates o j, sample the image features around, and generate local sampled image features Then, through the grid point dynamic fusion module GDF, fuse the local sampled image features with the position information encoding features through the cross-attention module to generate local fused features
[0012] For the above coal bunker material state monitoring method based on multi-modal data fusion, the specific process of generating global fused features in global fusion is as follows: for each non-empty voxel in the voxel features, calculate its voxel centroid where, c i is the voxel centroid of the i-th voxel, V i is all the voxel points within the i-th voxel, |V i | is the number of voxel points; after calculating the voxel centroid, project the voxel centroid c i onto the image feature plane to obtain the corresponding image coordinates p i = M·c i , where M is the camera projection matrix; on the image feature plane, according to the projected voxel center position p i , sample the surrounding image features to obtain global sampled image features where, F I is the image feature, △p mik is the sampling offset; after obtaining the global sampled image features , fuse it with the voxel features through the centroid dynamic fusion network module CDF to generate global fused features
[0013] For the above coal bunker material state monitoring method based on multi-modal data fusion, accept the global fused features from the global fusion module the local fused features from the local fusion module and the position information encoding features from the position information encoding module Aggregate these three feature blocks to generate the comprehensive feature F S : Meanwhile, introduce a residual connection block. By adding the input features and the features processed by the self-attention mechanism, the final feature F final = RCB(Self-Attention(F S ))
[0014] The present invention combines lidar scanning technology with multispectral camera technology, making full use of the advantages of both technologies. Lidar scanning can generate accurate three-dimensional point cloud data, providing the three-dimensional spatial information of the materials inside the coal bunker; while the multispectral camera can collect two-dimensional image data containing rich color and texture information. By fusing these two types of data, not only the deficiencies of single-modal data are made up for, but also the complementarity and enhancement of information are achieved. The fused data is more comprehensive and accurate, and can more truly reflect the actual situation of the materials inside the coal bunker. This fusion technology not only improves the accuracy of detection, but also greatly enhances the reliability and stability of the system. Description of the Drawings
[0015] Figure 1 It is a flow chart of the present invention.
[0016] Figure 2 It is a structure diagram of the local fusion network.
[0017] Figure 3 It is a structure diagram of the global fusion network.
[0018] Figure 4 It is a structure diagram of the residual connection module. Detailed Implementation Manner
[0019] The method for monitoring the state of materials in a coal bunker based on multi-modal data fusion of the present invention mainly includes data preprocessing, data downsampling and feature stratification, global fusion, local fusion, position information encoding, feature aggregation and residual connection output. The present invention first voxelizes the three-dimensional point cloud data, and then extracts multi-level voxel features from the 3D voxel data through a 3D Backbone, and at the same time extracts multi-level image features from the 2D image data through a 2D Backbone. Then, a region proposal network (RPN) is used to generate 3D bounding box candidates, and the centroid of the voxel is calculated through a centroid dynamic fusion (CDF) module, and the centroid of the voxel is projected onto the image feature plane to sample the image features to achieve global fusion. Next, each bounding box candidate is divided into a uniform voxel grid, and the grid center position is encoded through a PIE module to generate position information encoding features, and local fusion is achieved through a grid point dynamic fusion (GDF) module. Finally, the global fusion features, local fusion features and position information encoding features are concatenated and aggregated to generate comprehensive features, and through a self-attention mechanism and a residual connection block (RCB), a 3D fusion image is finally generated.
[0020] Data Preprocessing
[0021] The method of the present invention inputs 3D point cloud data and 2D image data. First, the 3D point cloud data is voxelized, and the point cloud space is divided into voxel grids of a fixed size. Each voxel contains a certain number of points, and the voxelized point cloud data is used for subsequent feature extraction. Then, voxel feature extraction is performed through a 3D Backbone. Feature extraction is performed on the points within each non-empty voxel to generate voxel features. Specifically, information such as the mean of the point coordinates, the median of the reflectivity, and the point cloud density within each voxel is calculated, and this information is used as the feature representation of the voxel. These features will be used for subsequent 3D object detection. At the same time, the voxelization process of the present invention can effectively reduce the sparsity of the point cloud data and improve the efficiency of feature extraction. For the 2D image data, image features are extracted through a 2D Backbone network. These image features will be used for fusion with the voxel features.
[0022] Data Downsampling and Feature Hierarchy
[0023] The 3D Backbone network usually uses a 3D sparse convolutional neural network to extract voxel features, and these features can capture local and global information in the point cloud. The sparse convolutional layer only performs convolution operations on non-empty voxels, thus greatly reducing the computational amount and memory occupancy. The sparse convolutional layer converts the voxel features into a sparse tensor based on information such as the spatial shape, and only performs convolution operations on the positions where the features are non-zero. At the same time, the 3D Backbone network will also generate feature maps of different levels from the 3D voxel data, and these feature maps will be used for subsequent global fusion and local fusion modules. Feature hierarchy can capture feature information at different scales, which helps to improve the accuracy of object detection. The 2D image data extracts multi-level image features and generates feature maps of different levels through a 2D Backbone network, and these feature maps will be used for subsequent global fusion and local fusion modules. Feature hierarchy can capture feature information at different scales, which helps to improve the accuracy of object detection. These features contain semantic information and spatial information in the image, providing a basis for subsequent fusion. The 2D Backbone mainly extracts image features through a 2D convolutional neural network. The 2D Backbone receives the preprocessed 2D image data and extracts image features through multiple 2D convolutional layers. Each 2D convolutional layer contains convolutional kernels, and these convolutional kernels slide in the two-dimensional space to perform convolution operations on the image data and extract local features.
[0024] The features extracted by the 3D Backbone are input into the RPN to generate 3D bounding box candidates. The RPN generates candidate regions on the feature map by means of a sliding window, and these regions may contain target objects. The RPN module obtains region candidate boxes by performing regression and classification on the preset anchor boxes.
[0025] Position Information Encoding
[0026] The algorithm of the present invention encodes the position information of the point cloud data grid features in each boundary candidate box through a Position Information Encoder (PIE) module to generate position information encoded features. Specifically, the PIE module divides each boundary candidate box into uniform grids, and each grid has a center point z j . And encodes the position of the center point of each grid. To simplify the operation, even dimensions are taken to generate position information encoded features
[0027]
[0028] where d model is the dimension of the model, and i, j represent the encoded feature coordinates. The feature contains the position information of each grid, providing a basis for subsequent local fusion.
[0029] Local Fusion
[0030] The local fusion module is one of the key modules of the present invention, aiming to generate a richer local feature representation by performing fine-grained feature extraction and fusion on each 3D boundary candidate box. The present invention first projects each grid center point z j onto the image plane to obtain the corresponding image coordinates o j . On the image plane, according to the projected image coordinates o j , sample the surrounding image features to generate local sampled image features The present invention then passes the local sampled image features and the position information encoded features through a cross-attention module for fusion to generate local fusion features Such as Figure 2 .
[0031]
[0032] The GDF module can also reduce redundancy and enhance the complementarity between modalities by capturing the dynamics of context information in different semantic spaces. This dynamic interaction mechanism makes the fused features richer and more expressive. Through the fusion of the GDF module, the generated local fusion features can better capture the spatial information within the local area, thereby improving the accuracy of 3D object detection.
[0033] Global Fusion
[0034] The CDF module generates a richer multi-modal feature representation by performing global-level fusion on 3D voxel features and 2D image features, such as Figure 3。The present invention extracts the voxel centroid of the voxel feature. For each non-empty voxel, the voxel centroid c is calculated i :
[0035]
[0036] where c i is the voxel centroid of the i-th voxel, V i is all the voxel points within the i-th voxel, and |V i | is the number of voxel points.
[0037] After calculating the voxel centroid, the present invention projects the voxel center c i onto the image plane to obtain the corresponding image coordinates p i :
[0038] p i = M·c i
[0039] where M is the camera projection matrix.
[0040] On the image plane, the present invention samples the surrounding image features according to the projected voxel center position p i to obtain the global sampled image features
[0041]
[0042] where F I is the image feature, and △p mik is the sampling offset.
[0043] After obtaining the global sampled image features the present invention fuses them with the voxel features to generate the global fused features
[0044]
[0045] The global fusion module generates a richer multi-modal feature representation through the fusion of 3D point cloud features and 2D image features. Specifically, by calculating the voxel centroid, projecting onto the image plane, sampling the image features, and fusing the features, the global fusion module enhances the feature expression ability and captures the context information within the global range.
[0046] Feature aggregation
[0047] By concatenating and aggregating the global fused features, local fused features, and position information encoded features, a richer multi-modal feature representation is generated, receiving the global fused features from the global fusion module the local fused features of the local fusion module and the location information encoding features of the location information encoding module Aggregate these three feature blocks to generate the comprehensive feature F S :
[0048]
[0049] The feature aggregation module generates a richer multi-modal feature representation by splicing and aggregating the global fusion feature, local fusion feature, and location information encoding features.
[0050] Output the fusion result
[0051] The self-attention mechanism calculates the similarity between features to generate a weighted representation for each feature, thereby capturing global information. This process helps to enhance the feature representation and improve the model's ability to recognize targets.
[0052] The present invention also introduces a residual connection block, such as Figure 4 , by adding the input feature and the feature processed by the self-attention mechanism, the information of the original feature is retained, and at the same time, the expression ability of the feature is enhanced. The final fusion image F is generated through the residual connection block final .
[0053] F final = RCB(Self-Attention(F S ))
[0054] The aggregated features pass through the self-attention mechanism and the residual connection block to obtain the final fusion result. The fused data optimizes the feature representation and can improve the performance of object detection.
[0055] Experimental tests
[0056] 1. Dataset preparation
[0057] The dataset is obtained by deploying lidar and multi-spectral cameras inside the coal bunker for all-round and high-precision acquisition. This dataset is large in scale, including a total of 1500 pictures for training, 600 pictures for testing, and 300 pictures for validation. To ensure a comprehensive evaluation of the model performance, the present invention selects three key indicators: precision, robustness, and real-time performance as the evaluation criteria.
[0058] During the training process, these 1,500 training set images will be fully utilized to continuously iterate and optimize the model parameters in order to achieve the best recognition effect. At the same time, 600 test set images will be used to objectively and fairly test the model to verify its performance in actual applications. And 300 validation set images will mainly be used for model verification during the training process to detect and correct possible overfitting or underfitting problems.
[0059] 2. Parameter Settings
[0060] The voxel size is carefully set to (0.1m, 0.1m, 0.2m) to maximize the capture of spatial details while maintaining computational efficiency. To deeply explore the complex correlations between point clouds and images, 4 attention heads (M = 4) are configured, and each head can independently focus on different feature dimensions. In addition, the farthest point sampling strategy is adopted to carefully select 4 key sampling points (K = 4) to ensure the effective extraction of information. The global fusion module uses the last two voxel regions to fuse point cloud features and image features. To ensure the accuracy and effectiveness of the fusion process, the grid sizes of the global fusion module and the local fusion module are set to 6, and the number of training epochs is 150.
[0061] To improve the generalization ability and feature expression ability of the model, a data augmentation strategy of image rotation and scaling is implemented, which helps the model show stronger robustness when dealing with various perspective and scale changes. Finally, sufficient training is carried out for 150 training epochs to ensure that the model can fully learn the deep features of the data.
[0062] 3. Experimental Environment
[0063] The experimental environment is built based on the NVIDIA RTX 4090 GPU and AMD Ryzen Threadripper 7970X CPU, aiming to provide powerful computing capabilities to support the training of complex deep learning models. The operating system is selected as Ubuntu 20.04, with sufficient video memory and memory resources, as well as PyTorch 1.8.1 and CUDA 11.3 versions, laying a solid foundation for the smooth progress of the experiment.
[0064] 4. Experimental Results and Analysis
[0065] When evaluating the model performance, the precision will reflect the accuracy of the model to the actual situation inside the coal bunker; the robustness will test the stability and reliability of the model in the states of coal falling and not falling; and the real-time performance will measure the response speed and processing efficiency of the model when dealing with the dynamic changes inside the coal bunker. Through the comprehensive evaluation of these three indicators, it will be possible to comprehensively and objectively understand the performance of the model in the task of object recognition inside the coal bunker.
[0066] Table 1 Performance Comparison
[0067]
[0068]
[0069] Through the data comparison in Table 1, it can be clearly seen that compared with BEVFusion_TTA, the method of the present invention has achieved significant improvements in precision and robustness, increasing by 5% and 7.1% respectively, while the inference time has been reduced by 12 ms. Compared with DeepFusion_Ens, the method of the present invention has reduced the inference time by 35.3 ms, and also increased the precision and robustness by 2.1% and 4.1% respectively. In addition, in the comparison with AFDetV2_Ens, although the inference time of the method of the present invention has only increased by 1.1 ms, the precision and robustness have increased by 0.5% and 2.3% respectively.
Claims
1. A coal bunker material status monitoring method based on multimodal data fusion, characterized in that: The method includes data preprocessing, data downsampling and feature stratification, global fusion, local fusion, position information encoding, feature aggregation and residual connection output. In data preprocessing, the three-dimensional point cloud data is first voxelized, and then multi-level voxel features are extracted from the 3D voxel data through 3D Backbone in data downsampling and feature stratification, and multi-level image features are extracted from the 2D image data through 2D Backbone. Then, the voxel features are used to generate 3D boundary candidate boxes using the regional candidate network. In global fusion, the centroid is calculated through the centroid dynamic fusion network module, the centroid is projected onto the image feature plane, image features are sampled, global fusion is achieved, and global fusion features are generated. Then, in local fusion, each boundary candidate box is divided into a uniform voxel grid, and in position information encoding, the center position of the grid is encoded through the position information encoding module to generate position information encoding features, and local fusion is achieved through the grid point dynamic fusion module to generate local fusion features. Finally, the global fusion features, local fusion features and position information encoding features are aggregated to generate comprehensive features, and a 3D fused image is finally generated through the self-attention mechanism and the residual connection block.
2. The coal bunker material status monitoring method based on multimodal data fusion according to claim 1 is characterized in that: LiDAR scanning is used to generate the three-dimensional point cloud data of the coal bunker, and a multispectral camera is used to generate the 2D image data of the coal bunker.
3. The coal bunker material status monitoring method based on multimodal data fusion according to claim 1 or 2, characterized in that: The 3D Backbone network uses a 3D sparse convolutional neural network to extract voxel features, and the 2D Backbone uses a 2D convolutional neural network to extract image features.
4. The coal bunker material status monitoring method based on multimodal data fusion according to claim 3 is characterized in that: The specific process of the position information encoding module generating the position information encoding feature is as follows: the position information encoding module PIE divides each boundary candidate box into uniform grids, each grid has a center point z j , encode the center point position of each grid and generate the position information encoding feature 5. The coal bunker material status monitoring method based on multimodal data fusion according to claim 4 is characterized in that: The specific process of generating local fusion features in local fusion is as follows: First, each grid center point z j Project it onto the image feature plane and get the corresponding image coordinates o j , on the image feature plane, according to the projected image coordinates o j , sample the surrounding image features and generate local sampled image features Then the local sampled image features are merged into the grid point dynamic fusion module GDF. Encoding features with position information Fusion is performed through the cross-attention module to generate local fusion features 6. The coal bunker material status monitoring method based on multimodal data fusion according to claim 5 is characterized in that: The specific process of generating global fusion features in global fusion is: for each non-empty voxel in the voxel feature, calculate its body mass center Among them, c i is the centroid of the ith voxel, V i are all voxel points within the ith voxel, |V i | is the number of voxel points; after calculating the body mass centroid, the body mass centroid c i Project it onto the image feature plane and get the corresponding image coordinates p i =M·c i , where M is the camera projection matrix; on the image feature plane, according to the projected voxel center position p i , sample the surrounding image features to get the global sampled image features Among them, F I is the image feature, △p mik is the sampling offset; get the global sampling image features Then it is combined with the voxel feature Fusion is performed through the centroid dynamic fusion network module CDF to generate global fusion features 7. The coal bunker material status monitoring method based on multimodal data fusion according to claim 6 is characterized in that: Accept global fusion features from the global fusion module / Local fusion features of local fusion module and the position information encoding features of the position information encoding module Aggregate these three feature blocks to generate comprehensive features F S : The residual connection block is also introduced. By adding the input features and the features processed by the self-attention mechanism, the residual connection block generates the final feature F final =RCB(Self-Attention(F S )).