Thermal vibration heterogeneous data fusion conveyor belt edge tearing detection method
By employing a thermal-vibration heterogeneous data fusion method, utilizing line laser heating and improved neural network technology, efficient detection of edge tears in underground coal mine conveyor belts has been achieved. This solves the problem of insufficient edge region detection in existing technologies and improves detection accuracy and reliability.
Patent Information
- Application Number
- CN202511525115.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-24
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2045-10-24
AI Technical Summary
In existing technologies for detecting edge tears in underground conveyor belts in coal mines, the detection area focuses on the main body of the conveyor belt, ignoring the edge area. A single sensor cannot capture the transient mechanical impact and thermodynamic anomalies of edge tears, and the multi-source data fusion effect is poor, resulting in a high rate of missed detection.
A thermal-vibration heterogeneous data fusion method is adopted. The edge area of the conveyor belt is heated by a line laser, and infrared thermal images and vibration signals are collected simultaneously. Features are extracted using an improved ResNet-50 network and a multi-scale temporal convolutional network. Combined with dynamic modal balance coefficients and heterogeneous complementary attention mechanisms, efficient fusion of infrared images and vibration signals and hierarchical early warning are achieved.
It significantly improves the sensitivity and accuracy of edge tear detection, reduces false alarm and false negative rates, and is suitable for complex environments such as underground mines. It can accurately capture the edge tear contour and provide timely warnings.
Smart Images

Figure CN120986946A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of intelligent monitoring of industrial equipment, and particularly relates to a thermal vibration heterogeneous data fusion conveyor belt edge tear detection method. BACKGROUND
[0002] The underground environment of a coal mine has the characteristics of high humidity, low temperature, high concentration of coal dust, no light, high load and the like. As an indispensable transportation equipment in large-scale coal mining, the belt conveyor runs for a long time under complex working conditions. The edge area is frequently in contact and friction with the carrier roller and the baffle, and often bears the impact of materials, becoming a high-risk area of tearing. Conveyor belt edge tearing is often caused by factors such as scraping of hard materials such as anchor rods in coal blocks and blockage at the transfer point. The initial tearing range is small but expands rapidly. If not detected in time, it may lead to the scrap of the entire conveyor belt, causing production downtime, equipment damage and even safety risks.
[0003] The contact detection device disclosed in Chinese patent document CN219770980U detects tearing by touching the abnormality, but its detection range is concentrated in the middle of the conveyor belt, the response to the edge micro-tear is slow, and it is easily disturbed in the high-concentration coal dust environment, with a high false alarm rate. The visual detection method proposed in Chinese patent document CN119490042A relies on a single camera to collect images, which is seriously affected by the lack of light and dust in the mine. Moreover, the algorithm is mainly designed for the tearing features of the main body area of the conveyor belt, and the recognition accuracy of the tearing area with low contrast between the edge and the background is insufficient.
[0004] In summary, the existing technology has the following problems: the detection area focuses on the main body of the conveyor belt, ignoring the tearing state of the risk area of the edge; a single sensor cannot simultaneously capture the transient mechanical impact and thermal abnormality of the edge tear; and the multi-source data fusion only performs simple comprehensive judgment without optimizing the feature correlation of the edge area, resulting in a high undetected rate of edge tearing.
[0005] Therefore, it is necessary to invent a thermal vibration heterogeneous data fusion conveyor belt edge tear detection method to solve the above problems. SUMMARY
[0006] The present application provides a thermal vibration heterogeneous data fusion conveyor belt edge tear detection method to solve the problems of low sensitivity to the edge area, poor multi-source data fusion effect, and high undetected rate caused by dust and light interference in the complex environment of a coal mine of the existing conveyor belt tear detection method.
[0007] The present application is implemented by using the following technical solutions:
[0008] A thermal vibration heterogeneous data fusion conveyor belt edge tear detection method, comprising the following steps:
[0009] S10: During the operation of the belt conveyor, the surface of the conveyor belt is heated by a line laser, focusing on the area on both sides of the conveyor belt surface where the distance from the edge is less than L, where L is one-quarter of the belt width. Infrared thermal images and vibration signals of the conveyor belt in this area are collected simultaneously and timestamps are added. The vibration and impact signals are used as a spatiotemporal index. When the vibration sensor detects an abnormal impact signal, it triggers the infrared thermal imager to enter a high-speed sampling mode, increasing the frame rate, and locking the conveyor belt edge area corresponding to the impact position through a spatiotemporal mapping algorithm.
[0010] S20: Preprocess the acquired infrared thermal image and vibration signal. The preprocessing includes aligning the two in the time dimension, performing noise reduction, contrast enhancement, size adjustment and standardization on the acquired infrared thermal image in sequence to obtain the processed thermal image data, and performing noise reduction and normalization on the vibration signal to obtain the processed vibration data.
[0011] S30: Feature extraction is performed on the preprocessed thermal image data and vibration data. Feature extraction includes extracting infrared image feature vectors from the thermal image data using an improved ResNet-50 network. Vibration feature vectors are extracted from vibration data using a multi-scale temporal convolutional network. ;
[0012] S40: Fusion of infrared image feature vectors Vibration data feature vector Based on the fusion results, a tear probability is generated, and a graded early warning is executed according to the probability threshold.
[0013] Further, step S20 includes:
[0014] S21: The acquisition time of each frame of infrared thermal image is strictly aligned with the start time of the vibration signal segment through hardware pulse signals;
[0015] S22: Gaussian filtering is applied to the infrared thermal image for noise reduction, histogram equalization for enhancement, and size standardization. The pixel values are then converted to a distribution range with a mean of 0 and a standard deviation of 1 using the zero-mean normalization method to obtain the processed thermal image data.
[0016] S23: Perform wavelet threshold denoising on the vibration signal and normalize it to the [0,1] interval to obtain the processed vibration data.
[0017] Furthermore, in the improved ResNet-50 network, the initial layer is replaced with three 3×3 convolutional layers instead of the original 7×7 convolutional layers, the residual blocks are embedded with SE channel attention modules, and a spatial attention module is added at the end.
[0018] The calculation formula for the first 3×3 convolutional layer among the three 3×3 convolutional layers is:
[0019] ;
[0020] in: These are the parameters of the first 3×3 convolution kernel; For bias terms; This is the preprocessed infrared thermal image; The output feature map has a size of 224×224×64; ReLU is a linear rectified activation function used to introduce non-linear features; x and y represent the horizontal and vertical coordinates of the pixels in the output feature map; p and q correspond to the row and column indices of the convolution kernel, respectively.
[0021] The subsequent two 3×3 convolutional layers have the same structure, with the output feature map size remaining at 224×224 and the number of channels being 64 and 128 respectively.
[0022] Furthermore, the multi-scale temporal convolutional network includes three multi-scale convolutional branches:
[0023] First branch: 1D convolution kernel size is 3, dilation is set to 1, receptive field is 3, captures local high-frequency impact, and outputs feature dimension 1024×64;
[0024] Second branch: 1D convolution kernel size is 5, dilation is set to 2, receptive field is 9, captures mid-term vibration trend, and outputs feature dimension 1024×64;
[0025] The third branch: 1D convolution kernel size is 7, dilation is set to 4, receptive field is 25, long-term low-frequency components are captured, and the output feature dimension is 1024×64.
[0026] The three multi-scale convolutional branches are concatenated into a 192-channel feature map through a fusion layer, and then compressed to 128 channels through a 1×1 convolution, as shown in the formula:
[0027] ;
[0028] Where: n is the kernel size, and is 3, 5, or 7 respectively. For the i-th parameter of the 1D convolution kernel, The signal is the standardized vibration signal, and C represents the output characteristic. For bias terms, Indicates the time step index of the vibration signal. This represents the kernel index, and ReLU is the linear rectified activation function.
[0029] Further, step S40 includes:
[0030] S41: Real-time acquisition of environmental parameters, and dynamic calculation of modal equilibrium coefficients based on environmental parameters. Based on the modal balance coefficient, the feature vector of the infrared image is... Vibration data feature vector The weighted adjustment is performed using the following formula:
[0031] ;
[0032] in: For weighted , For weighted ;
[0033] right The channels are compressed to 512 dimensions using 1×1 convolution. The dimension is expanded to 512 by two fully connected layers, and the features are concatenated to form a 1×1152 joint feature vector. ;
[0034] S42: Combine the feature vectors Weighted fusion features are generated through a dual-channel attention mechanism. ;
[0035] S43: Weighted fusion features The input fully connected network outputs a tearing probability P∈[0,1]. The fully connected network includes a first fully connected layer and a second fully connected layer. The first fully connected layer outputs a 64-dimensional feature vector, and the second fully connected layer outputs a 1-dimensional probability value. The formula is as follows:
[0036] ;
[0037] ;
[0038] in: This is the output feature vector of the first fully connected layer. , Here are the weight matrices for the first and second fully connected layers. , For the bias terms of the first and second fully connected layers, P is the tearing probability, ReLU is the linear rectified activation function, and Sigmoid is the sigmoid activation function, which maps the output to the interval [0,1].
[0039] S44: Execute graded early warnings based on probability thresholds, adopting a three-level response mode.
[0040] Furthermore, the specific rules of the three-level response mode are as follows:
[0041] When 0 ≤ P < 0.2, it is considered a normal state and no intervention is required;
[0042] When 0.2≤P<0.6, it is judged as a Level 1 warning, that is, suspected tearing, triggering an audible and visual alarm, and inspection and investigation.
[0043] When 0.6≤P≤1, it is judged as a level 2 warning, that is, a severe tear, and the machine is shut down immediately.
[0044] Furthermore, the SE channel attention module performs:
[0045] Squeezing operation: Global average pooling is performed on the residual block output feature map to obtain channel-level statistics, as shown in the following formula:
[0046] ;
[0047] Where: F is the residual block output feature map, and the dimension of this output feature map is... H represents height, W represents width, C represents the number of channels, i and j are the row and column indices of the feature map pixels, and c represents the channel index. This represents the global feature of the c-th channel;
[0048] Activation operation: Channel weights are generated through two fully connected layers and a sigmoid activation function to perform weighted enhancement on the original feature map. The formula is as follows:
[0049] ;
[0050] ;
[0051] in: For the weights of the first fully connected layer, For the weights of the second fully connected layer, This is the channel statistics vector obtained from the compression operation. The channel weight vector is the output of the sigmoid function. For the sigmoid function, Let be the weight vector of the c-th channel. For weighted enhancement , Let F be the c-th channel, and ReLU be the linear rectified activation function.
[0052] Furthermore, the spatial attention module performs:
[0053] Perform max pooling and average pooling on the input features along the channel dimension;
[0054] The pooling results are concatenated and then convolved to generate a spatial weight matrix, as shown in the formula:
[0055] ;
[0056] in: The feature map is input to the spatial attention. This is the spatial weight matrix. For sigmoid function, MaxPool performs channel-based max pooling on the input feature map, AvgPool performs channel-based average pooling on the input feature map, and Conv performs convolution operation.
[0057] The formula is as follows: Weighted element-wise from the weight matrix.
[0058] ;
[0059] in: This is the feature map after being weighted by the spatial attention module. This indicates that the Hadamard product is an element-wise multiplication.
[0060] Furthermore, the dual-channel attention mechanism is executed as follows:
[0061] For joint eigenvectors Global average pooling is performed to generate channel-level statistical features. The mutual information (MI) between infrared and vibration sub-features is calculated, and associated channels with MI > 0.6 are selected. Basic channel weights are generated through a fully connected layer and a sigmoid function. The formula is:
[0062] ;
[0063] in: For the weights of the first fully connected layer, For the weights of the second fully connected layer, This is the channel statistics vector obtained from the compression operation. For the sigmoid function, , For the bias terms of the first and second fully connected layers;
[0064] A joint weight matrix is generated by fusing the correlation features between spatial coordinates and time points through a 3×3 convolutional layer. The formula is:
[0065] ;
[0066] in: This is the infrared image feature vector weighted by modal balance coefficients. This is the vibration data feature vector weighted by modal equilibrium coefficients. For the joint weight matrix, For sigmoid function, Conv is the convolution operation, MaxPool is the channel-dimensional max pooling operation on the input feature map, and PeakDetect is the peak detection operation on the vibration feature vector;
[0067] Basic channel weights With joint weight matrix Element-wise multiplication yields the dual-channel attention weights;
[0068] Joint feature vectors The weighted fusion feature is obtained by multiplying the dual-channel attention weights element-wise. The formula is:
[0069] .
[0070] A thermal shock heterogeneous data fusion conveyor belt edge tear detection system is provided. This system is used to execute the thermal shock heterogeneous data fusion conveyor belt edge tear detection method as described in this invention. The conveyor belt edge tear detection system includes:
[0071] Information acquisition module: includes a line laser, an infrared thermal imager, and a vibration sensor; the line laser emits a laser line that is perpendicularly irradiated onto the edge of the conveyor belt of the belt conveyor to actively heat the conveyor belt and form a temperature field; the infrared thermal imager is used to acquire infrared thermal images of the contour of the conveyor belt surface edge; and the vibration sensor is used to acquire vibration signals during the operation of the conveyor belt.
[0072] Preprocessing module: used to preprocess the acquired infrared thermal images and vibration signals;
[0073] Feature extraction module: used to extract features from preprocessed thermal image data and vibration data to obtain infrared image feature vectors and vibration data feature vectors;
[0074] The tear detection module includes a dynamic fusion unit, a dual-channel attention unit, and a hierarchical early warning unit. The dynamic fusion unit calculates the modal balance coefficient based on environmental parameters and uses it to weight and adjust the feature vectors of infrared images and vibration data and then concatenate them into a joint feature vector. The dual-channel attention unit is used to generate a weighted fusion feature from the joint feature vector. The hierarchical early warning unit is used to input the weighted fusion feature into a fully connected network to generate a tear probability and execute hierarchical early warning based on the probability threshold.
[0075] This invention provides a method for detecting edge tearing of conveyor belts using thermal shock heterogeneous data fusion, which has the following advantages compared with existing technologies:
[0076] 1. Innovatively, vibration and impact signals are introduced as a spatiotemporal index to dynamically drive the infrared module to focus on the edge area of the conveyor belt, realizing a closed-loop detection mechanism of "vibration triggering - infrared focusing - dual-mode fusion". This solves the problem of inaccurate feature capture in the edge area and greatly improves the sensitivity of edge tear detection.
[0077] 2. By fusing dual-modal data of infrared thermal imaging and vibration signals, and cross-validating the temperature field characteristics of infrared images with the dynamic characteristics of vibration signals, we innovatively introduce dynamic modal balance coefficients and heterogeneous complementary attention mechanisms to achieve adaptive balance and synergistic enhancement of the two types of data, thereby improving detection accuracy and reducing false alarm and false negative rates.
[0078] 3. This invention uses a line laser as an active excitation heat source and achieves uniform heating through the movement of a conveyor belt, upgrading infrared thermal imaging from "passive sensing" to "active controllability". It reduces the interference of ambient temperature on infrared data and can more clearly capture edge tear contours. Through the improved ResNet-50 network, it accurately captures temperature gradient differences in infrared images, making it particularly suitable for light-free scenarios such as underground mines and nighttime operations.
[0079] 4. The improved ResNet-50 network replaces the 7×7 layers with three consecutive 3×3 convolutional layers, preserving subtle temperature features of infrared images; it also embeds SE channel attention and spatial attention modules to enhance the response to high temperature anomalies in the torn region.
[0080] 5. An improved temporal convolutional network is adopted, which uses a three-branch parallel structure. Based on multi-scale convolution, it can adaptively capture tear-related temporal and frequency features through learnable convolutional kernels. At the same time, the parallel architecture can process features of different scales simultaneously, reducing detection latency. Attached Figure Description
[0081] Figure 1 This is the overall flowchart of the present invention.
[0082] Figure 2 A schematic diagram of the belt conveyor in this invention.
[0083] Figure 3 A schematic diagram of the conveyor belt edge position in this invention.
[0084] Figure 4 This is an image of the conveyor belt edge without tearing, acquired by an infrared thermal imager in this invention.
[0085] Figure 5 This is an image of the conveyor belt edge tear preprocessed by an infrared thermal imager acquired in this invention.
[0086] In the image: 1. Belt conveyor; 2. Infrared thermal imager; 3. Line laser; 4. Inverted U-shaped support; 5. Vibration sensor; 6. Left edge area of the conveyor belt; 7. Right edge area of the conveyor belt. Detailed Implementation
[0087] The relevant technical solutions will be clearly and completely described below. Obviously, the described embodiments are only some embodiments, not all embodiments. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this invention.
[0088] The present invention will be further explained and described below with reference to the accompanying drawings, embodiments, and comparative examples: Example 1
[0089] Implementation scenario: Underground transport roadways in coal mines, with environmental parameters of temperature 5-7℃, relative humidity 85-92%, and coal dust concentration of 30mg / L. In a dark environment, the conveyor belt of belt conveyor 1 runs at a speed of 2.5 m / s and a width of 1.2 m. The following data samples were collected for testing.
[0090] Test dataset: The dataset contains 3 types of samples, namely normal, mild tear and severe tear, with a total of 3000 sets of synchronized data. Each set contains one frame of infrared thermal image and one vibration signal.
[0091] Normal samples: 1500 groups, conveyor belt without tear; slightly torn samples: 750 groups, artificially pre-made tear 5cm-10cm long; severely torn samples: 750 groups, artificially pre-made tear longer than 10cm long.
[0092] A thermal shock heterogeneous data fusion conveyor belt tear edge detection system is deployed on a belt conveyor 1. The system includes:
[0093] Information acquisition module: includes a line laser 3, an infrared thermal imager 2, and a vibration sensor 5; as attached. Figure 2 As shown, the information acquisition modules are all focused on the edge area of the conveyor belt; the belt conveyor 1 is equipped with an inverted U-shaped bracket 4, and the line laser 3 is installed on the horizontal bar of the inverted U-shaped bracket 4. The laser line emitted by the laser is perpendicularly irradiated on the edge of the conveyor belt to actively heat the conveyor belt and form a temperature field; the infrared thermal imager 2 is installed on the vertical bar of the inverted U-shaped bracket 4 located on both sides of the conveyor belt to collect infrared thermal images of the edge contour of the conveyor belt after being heated by the laser; the vibration sensor 5 is installed on the idler bracket of the conveyor belt to collect vibration signals during the operation of the conveyor belt. Due to the tearing of the edge of the conveyor belt, a dynamic collision occurs instantaneously, releasing stress waves with vibration frequency and amplitude that are much higher than those in the normal operating state. The vibration sensor 5 receives these signals to determine the tearing state.
[0094] The infrared thermal imager 2 is a HIKMICRO K20 model, with a resolution of 640×480, a temperature measurement range of -20℃ to 400℃, a frame rate of 50Hz, and a temperature measurement accuracy of ±2℃; the vibration sensor 5 is a PCB352C33 model, a piezoelectric accelerometer, with a measurement range of ±50g, a sensitivity of 10mv / g, and a sampling frequency of 1000Hz.
[0095] Preprocessing module: Used to preprocess the acquired infrared thermal images and vibration signals.
[0096] Feature extraction module: used to extract features from preprocessed thermal image data and vibration data to obtain infrared image feature vectors and vibration data feature vectors.
[0097] The tear detection module includes a dynamic fusion unit, a dual-channel attention unit, and a hierarchical early warning unit. The dynamic fusion unit calculates the modal balance coefficient based on environmental parameters and uses it to weight and adjust the feature vectors of infrared images and vibration data and then concatenate them into a joint feature vector. The dual-channel attention unit is used to generate a weighted fusion feature from the joint feature vector. The hierarchical early warning unit is used to input the weighted fusion feature into a fully connected network to generate a tear probability and execute hierarchical early warning based on the probability threshold.
[0098] The preprocessing module, feature extraction module, and tear detection module are all integrated into a high-performance industrial computer. The high-performance industrial computer is an Advantech UNO-2484G model, equipped with an Intel i7-13700TE processor, an NVIDIA RTX A5000 GPU, and 64GB of DDR5 memory.
[0099] In the preprocessing module, the acquisition of infrared thermal images and vibration signals is aligned in the time dimension by a synchronization controller. The synchronization controller is independently deployed using an FPGA chip and achieves strict alignment of the acquisition time of the infrared thermal imager 2 and the vibration sensor 5 through hardware pulse signals.
[0100] It should be noted that, in specific implementations of the present invention, the line laser 3, infrared thermal imager 2, vibration sensor 5, high-performance industrial computer, and synchronization controller are selected according to actual needs, including but not limited to the selection types in this embodiment.
[0101] A method for detecting edge tears in conveyor belts based on a thermal shock heterogeneous data fusion conveyor belt tear detection system is presented below. Figure 1 As shown, it includes the following steps:
[0102] S10: Data Acquisition: The conveyor belt surface is continuously irradiated on both sides using a line laser 3; simultaneously, the infrared thermal imager 2 is controlled to collect infrared thermal images of areas less than 30cm from the edge on both sides of the conveyor belt surface at the set sampling frequency. The conveyor belt width is 120cm. (See attached image.) Figure 3 As shown, the left conveyor belt edge area 6 and the right conveyor belt edge area 7 are shown. The vibration sensor 5 synchronously collects the vibration signal of the conveyor belt and transmits the collected infrared thermal image and vibration signal to the preprocessing module. Each infrared thermal image and vibration signal has a corresponding timestamp information. The vibration impact signal serves as a spatiotemporal index. When the vibration sensor 5 detects an impact signal with a frequency > 500Hz and an amplitude > 3g, it triggers the infrared thermal imager 2 to enter the high-speed sampling mode, increasing the frame rate to 50Hz. The spatiotemporal mapping algorithm locks the conveyor belt edge area corresponding to the impact position and focuses on this area for continuous high-definition shooting of 3 frames.
[0103] S20: The preprocessing module performs preprocessing on the acquired infrared thermal image and vibration signal. The preprocessing includes aligning the two in the time dimension, performing noise reduction, contrast enhancement, size adjustment and zero-mean standardization on the acquired infrared thermal image in sequence to obtain the processed thermal image data, and performing noise reduction and normalization on the vibration signal to obtain the processed vibration data.
[0104] Step S20 specifically includes:
[0105] S21: A synchronization method combining hardware triggering and sliding window is adopted. The pulse signal issued by the synchronization controller provides a unified time reference for the thermal imager and vibration sensor 5. Based on the spatiotemporal index of the vibration and impact signal, a precise mapping between the impact moment and the infrared image frame is established to ensure that the infrared image within 0.1s after the impact occurs is strictly aligned with the start time of the vibration and impact segment.
[0106] S22: Preprocessing of infrared thermal images: For vibration-indexed edge region images, adaptive Gaussian filtering is used, with the kernel size dynamically adjusted from 3×3 to 7×7 according to the region complexity, to perform noise reduction, thereby smoothing the image and suppressing noise. The principle is to perform convolution operation on the image using a two-dimensional Gaussian kernel, so that each pixel value is replaced by the weighted average of its surrounding pixels. The formula is:
[0107] ;
[0108] in: The kernel function is a 5×5 Gaussian kernel. These are the original image pixel values. is the filtered pixel value; m and n are the row and column indices of the 5×5 Gaussian kernel; x and y represent the horizontal and vertical coordinates of the pixel.
[0109] The grayscale value distribution of infrared thermal images may be relatively concentrated, resulting in low image contrast and difficulty in distinguishing temperature differences on the conveyor belt surface, especially the temperature difference between torn and normal areas. Histogram equalization is used to enhance temperature details in edge areas and adjust the grayscale distribution of the image to make the grayscale values evenly distributed across the entire range, thereby enhancing image contrast. The above images are then standardized by cropping the vibration index region image into 112×112 sub-images and uniformly scaling the image size to 224×224 to meet the input requirements of deep learning models. The formula is:
[0110] ;
[0111] in: The mean gray level of the image after histogram equalization (normal sample) =128, torn sample =145), The standard deviation of gray levels in the histogram-equalized image (normal sample) =32, torn sample =48), These are the pixel values after histogram equalization.
[0112] As attached Figure 4 , 5 The images shown are pre-processed infrared thermal images of the conveyor belt in normal and torn states, respectively.
[0113] S23: Preprocessing of the vibration signal: The signal is converted to a time-spectrum using a short-time Fourier transform algorithm. Wavelet transform is then used for denoising. The signal is decomposed into sub-bands of different frequencies, and a 3-level decomposition is performed using the db4 wavelet basis. The high-frequency sub-band containing the vibration and impact signal (100-1000Hz) is prioritized. A threshold is applied to the noisy high-frequency sub-bands, retaining valid signals above the threshold and suppressing noise below the threshold. Finally, wavelet reconstruction yields the denoised signal. The min-max normalization method is used to map the signal values to the [0,1] interval. The standardized vibration signal can be better integrated with image features, improving the model's detection accuracy. The formula is:
[0114] ;
[0115] in: The minimum value of the vibration signal after denoising (normal sample) =-3.2g, tear-like =-8.5g), The maximum value of the vibration signal after denoising (normal sample) =2.8g, torn sample =10.2g), The signal value at the k-th time step of the standardized vibration signal. Let be the signal value at the k-th time step of the denoised vibration signal, where k is the time step index of the vibration signal.
[0116] S30: Feature extraction is performed on the preprocessed thermal image data and vibration data. Feature extraction includes extracting infrared image feature vectors from the thermal image data using an improved ResNet-50 network. Vibration feature vectors are extracted from vibration data using a multi-scale temporal convolutional network. .
[0117] Step S30 specifically includes:
[0118] S31: Optimize the initial convolutional layers of the original ResNet-50 network: Replace the original 7×7 convolutional layers with three consecutive 3×3 convolutional layers, each with a stride of 1, and remove the pooling operation after the first layer. The original large 7×7 convolutional kernels easily lead to the loss of subtle temperature difference features in infrared thermal images, while stacking three 3×3 convolutional layers can increase the nonlinear expressive power of the network while maintaining the same receptive field, and more accurately capture temperature gradient changes. The calculation formula for the first 3×3 convolutional layer is:
[0119] ;
[0120] in: These are the parameters of the first 3×3 convolution kernel; For bias terms; This is the preprocessed infrared thermal image; The output feature map has a size of 224×224×64; ReLU is a linear rectified activation function used to introduce non-linear features; x and y represent the horizontal and vertical coordinates of the pixels in the output feature map; p and q correspond to the row and column indices of the convolution kernel, respectively; the subsequent two 3×3 convolutional layers have the same structure, and the output feature map size remains 224×224, with 64 and 128 channels respectively.
[0121] An attention mechanism is incorporated into the residual blocks. Specifically, an SE channel attention module is embedded in each residual block from conv2_x to conv5_x to enhance the feature response to temperature anomaly regions. The SE channel attention module operation consists of two steps: squeezing and excitation.
[0122] Squeezing operation: Global average pooling is performed on the residual block output feature map to obtain channel-level statistics.
[0123] ;
[0124] Where: F is the residual block output feature map, and the dimension of this output feature map is... H represents height, W represents width, C represents the number of channels, i and j are the row and column indices of the feature map pixels, and c represents the channel index. This represents the global feature of the c-th channel;
[0125] Activation operation: Channel weights are generated through two fully connected layers and a sigmoid activation function to perform weighted enhancement on the original feature map:
[0126] ;
[0127] ;
[0128] in: For the weights of the first fully connected layer, For the weights of the second fully connected layer, This is the channel statistics vector obtained from the compression operation. The channel weight vector is the output of the sigmoid function. For the sigmoid function, Let be the weight vector of the c-th channel. For weighted enhancement , Let F be the c-th channel, and ReLU be the linear rectified activation function.
[0129] The downsampling strategy was also adjusted. The improved ResNet-50 performs downsampling only in the first residual block of conv3_x, conv4_x, and conv5_x, and adopts a parallel structure of "convolution + pooling": the main branch performs downsampling through a 3×3 convolution with a stride of 2, and the side branches use a 1×1 convolution with a stride of 2 combined with 2×2 average pooling to align the size. The outputs of the two are added together and then enter the SE channel attention module. This structure gradually reduces the feature map size from the initial 224×224 resolution to the final 7×7 while preserving the edge features of the torn region to the maximum extent.
[0130] At the end of the network, before global average pooling, a 1×1 convolutional layer is added to reduce the number of channels from 2048 to 1024, and a spatial attention module is introduced. The spatial attention module performs max pooling and average pooling on the channel dimensions of the feature maps, resulting in two 1×H×W feature maps. After convolutional fusion, a spatial weight matrix is generated, which spatially weights the feature maps. The formula is as follows:
[0131] ;
[0132] ;
[0133] in: The feature map is input to the spatial attention. This is the spatial weight matrix. For sigmoid function, MaxPool performs channel-based max pooling on the input feature map, AvgPool performs channel-based average pooling on the input feature map, and Conv performs convolution operation. This is the feature map after being weighted by the spatial attention module. This indicates that the Hadamard product is an element-wise multiplication.
[0134] Finally, the infrared image feature vector is obtained through global average pooling. , dimension 1×1024.
[0135] S32: Vibration signals exhibit significant temporal characteristics. A multi-scale temporal convolutional network is employed for feature extraction. The multi-scale convolutional branches consist of three branches: The first branch uses a 1D convolutional kernel of size 3, dilation of 1, and a receptive field of 3 to capture local high-frequency impacts, outputting a feature dimension of 1024×64; the second branch uses a 1D convolutional kernel of size 5, dilation of 2, and a receptive field of 9 to capture mid-term vibration trends, outputting a feature dimension of 1024×64; the third branch uses a 1D convolutional kernel of size 7, dilation of 4, and a receptive field of 25 to capture long-term low-frequency components, outputting a feature dimension of 1024×64. The fusion layer concatenates the outputs of the three multi-scale convolutional branches into a 192-channel feature map, which is then compressed to 128 channels using a 1×1 convolution. The formula is as follows:
[0136] ;
[0137] Where: n is the kernel size, and is 3, 5, or 7 respectively. For the i-th parameter of the 1D convolution kernel, The signal is the standardized vibration signal, and C represents the output characteristic. For bias terms, Indicates the time step index of the vibration signal. This represents the kernel index, and ReLU is the linear rectified activation function.
[0138] Feature compression and output: The 1024×128 multi-scale convolutional features are directly compressed into a 1×128 vibration feature vector through global pooling. .
[0139] S40: Fusion of infrared image feature vectors Vibration data feature vector Based on the fusion results, a tear probability is generated, and a graded early warning is executed according to the probability threshold.
[0140] Step S40 specifically includes:
[0141] S41: Perform dynamic modal equilibrium and feature preprocessing: Calculate the modal equilibrium coefficient based on real-time data collected by the environmental sensing module, including coal dust concentration, ambient temperature, infrared signal-to-noise ratio, and vibration noise levels. , Value selection rule: When the coal dust concentration is >50mg / m³ or the infrared image signal-to-noise ratio is <15dB, =0.3, meaning the vibration characteristic weight is increased; secondly, when the ambient temperature is <8℃ or the vibration signal noise is >5g, =0.7, meaning the infrared feature weight is increased; finally, under normal conditions =0.5, according to Dynamically adjust infrared and vibration weights:
[0142] ;
[0143] In this embodiment, the infrared and vibration weights α are dynamically adjusted to 0.7 according to the experimental scenario.
[0144] ;
[0145] in: For weighted , For weighted ;
[0146] right The channels are compressed to 512 dimensions using 1×1 convolution. The dimension is expanded to 512 by two fully connected layers, and the features are concatenated to form a 1×1152 joint feature vector. .
[0147] S42: Perform heterogeneous complementary attention weighting: on the joint feature vector Global average pooling is performed to obtain channel-level statistical features Z. The mutual information MI between the infrared sub-feature and the vibration sub-feature is calculated, where both the infrared and vibration sub-features are 512-dimensional. Associated channels with a mutual information MI > 0.6 are selected. A two-layer fully connected network structure is used to first reduce the feature dimension from 1024 to 256, then restore it to 1024. The sigmoid function is then used to generate basic channel weights, with an additional 20%-30% boost to the weights of associated channels. The formula is as follows:
[0148] ;
[0149] in: For the weights of the first fully connected layer, For the weights of the second fully connected layer, This is the channel statistics vector obtained from the compression operation. For the sigmoid function, , For the bias terms of the first and second fully connected layers;
[0150] By fusing the correlation features between spatial coordinates and time points through a 3×3 convolutional layer, a joint weight matrix is generated to enhance the weighted regions matching "spatial location - temporal time". The formula is as follows:
[0151] ;
[0152] in: This is the infrared image feature vector weighted by modal balance coefficients. This is the vibration data feature vector weighted by modal equilibrium coefficients. For the joint weight matrix, For sigmoid function, Conv is the convolution operation, MaxPool is the channel-dimensional max pooling operation on the input feature map, and PeakDetect is the peak detection operation on the vibration feature vector;
[0153] Final fusion features It is obtained by element-wise multiplication of the spliced features and the dual-channel attention weights, and the formula is:
[0154] ;
[0155] S43: Through the decision network, fully connected layer 1 outputs a 64-dimensional feature vector and fully connected layer 2 outputs a 1-dimensional probability value, ranging from [0,1], representing the probability of the conveyor belt existing. The formula is as follows:
[0156] ;
[0157] ;
[0158] in, This is the output feature vector of the first fully connected layer. , Here are the weight matrices for the first and second fully connected layers. , P represents the bias term for the first and second fully connected layers, where P is the tearing probability [0,1], ReLU is the linear rectified activation function, and Sigmoid is the sigmoid activation function that maps the output to the interval [0,1].
[0159] S44: The hierarchical early warning mechanism adopts a three-level response mode based on the tear probability P output by the decision network, as shown in Table 1. The specific rules are as follows:
[0160] When 0≤P<0.2, it is judged as a normal state, the conveyor belt is running normally, and the system does not output any warning signal at this time, but only records the normal operation status in the background;
[0161] When 0.2≤P<0.6, it is judged as a level one warning, that is, suspected tear. The system controls the sound and light alarm to emit a low frequency prompt sound, and at the same time the yellow warning light flashes, prompting the on-site inspection personnel to conduct a key inspection of the conveyor belt.
[0162] When 0.6≤P≤1, it is judged as a level 2 warning, that is, a severe tear. The system immediately activates the high-frequency alarm, the red warning light stays on, an emergency shutdown is executed, and the emergency braking mechanism is activated at the same time.
[0163] Table 1. Tear Fault Alarm Levels, Triggering Conditions, and Corresponding Measures for Belt Conveyors
[0164]
[0165] The following results were obtained from testing 3000 samples in the test set:
[0166] Normal samples (1500 groups): 1473 groups were detected as normal (0≤P<0.2), with a false alarm rate of 1.8%; Mild tear samples (750 groups): 718 groups were detected as Level 1 warning (0.2≤P<0.6), with an accuracy of 95.7%; Severe tear samples (750 groups): 738 groups were detected as Level 2 warning (0.6≤P≤1), with an accuracy of 98.4%; The overall mAP@0.5 was 93.5% (mild tear mAP 90.2%, severe tear mAP 95.7%), with a detection delay of 170ms. Comparative Example 1
[0167] The difference between this comparative example and Example 1 is that the vibration sensor 5 is not involved in the entire process. The infrared thermal imager 2 is used alone to detect the tear of the conveyor belt of the belt conveyor 1. The infrared thermal image data of the conveyor belt is obtained by the infrared thermal imager 2. After preprocessing, the data is processed by an improved ResNet-50 network, and the tear probability P is finally output.
[0168] The overall mAP of the output results was 65.2%, which was 28.3% lower than that in Example 1. Among them, the mAP for mild tearing was only 52.3%, which was very poor. Therefore, the single infrared mode has a weak ability to capture mild tearing and is significantly affected by image blurring caused by coal dust. Comparative Example 2
[0169] The difference between this comparative example and Example 1 is that the feature extraction of the image data adopts the traditional ResNet-50 network, which is an unoptimized convolutional layer and attention mechanism. The vibration data is still extracted using the same feature extraction method as in Example 1. Then, the infrared feature vector and the vibration feature vector are directly concatenated and fused without dual-channel attention, and the tearing probability P is finally output.
[0170] The overall mAP of the output results was 82.6%, which is a significant improvement over Comparative Example 1, but still 10.9% lower than Comparative Example 1. The detection of mild tearing was also poor, with an mAP of only 75.4%, while the mAP of severe tearing was 88.3%. Therefore, the unoptimized network and fusion strategy cannot effectively capture the subtle infrared temperature features, and the accuracy of mild tearing detection is insufficient.
[0171] Table 2 Comparison of Tear Detection Performance
[0172]
[0173] As shown in Table 2, mAP refers to mAP@0.5, meaning that the IoU threshold is set to 0.5 when calculating AP. Comparative Example 1 has a low overall mAP, especially in the detection of minor tears, where the mAP is less than 70%, indicating that the single-modal method has limited ability to capture subtle tear features. Although Comparative Example 2 improved the overall mAP to 82.6% through fusion, the mAP for minor tears was only 75.4% because the network was not optimized for infrared thermal images, resulting in insufficient sensitivity for early tear detection. The overall mAP of this invention reaches 93.5%, and the mAP for all levels of tears is higher than other methods, especially the mAP for severe tears, which reaches 95.7%, significantly better than Comparative Example 1 and Comparative Example 2.
[0174] In summary, this invention employs a thermal-vibration heterogeneous data fusion method for conveyor belt edge tear detection. This method detects and provides early warnings of tears on the edges of conveyor belts operating underground in coal mines. Linear lasers serve as the active excitation heat source, initiating conveyor belt movement to achieve uniform heating. Subsequently, a thermal imager captures and records changes in the conveyor belt surface temperature field. The torn area exhibits a different temperature distribution on the thermal image due to differences in thermal conductivity compared to the surrounding area. The torn area is then located based on semantic segmentation. Vibration data originates from sensors installed on the idlers. Based on changes in vibration signals, when a conveyor belt tears, the torn area may generate a unique and drastically changing high-frequency shock wave with vibration characteristics significantly different from normal operation. The vibration shock signal is introduced as a spatiotemporal index. When the vibration sensor detects a high-frequency shock signal matching tear characteristics, an infrared thermal imager is triggered in real-time to focus on capturing and analyzing the conveyor belt edge area corresponding to the impact location. By fusing these two types of heterogeneous data, feature complementarity and decision collaboration are achieved to determine whether the conveyor belt is torn, improving the accuracy of tear detection, reducing false alarm rates, and enhancing robustness in complex environments.
[0175] In the description of this invention, it should be understood that the indicated orientation or positional relationship is based on the orientation or positional relationship shown in the accompanying drawings, and is only for the convenience of describing this invention and simplifying the description, and is not intended to indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of this invention.
[0176] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A thermal-vibrational heterogeneous data fusion conveyor belt edge tear detection method, characterized by: The method comprises the following steps: S10: during the operation of the belt conveyor, heating the surface of the conveyor belt by a line laser, focusing on the area on both sides of the surface of the conveyor belt with a distance less than L from the edge, L being one fourth of the width of the conveyor belt, synchronously collecting an infrared thermal image and a vibration signal in the area of the conveyor belt and adding a time stamp; S20: preprocessing the collected infrared thermal image and vibration signal, the preprocessing including aligning the two in the time dimension, sequentially performing denoising, contrast enhancement, size adjustment and standardization processing on the collected infrared thermal image to obtain processed thermal image data, and performing denoising and normalization processing on the vibration signal to obtain processed vibration data; S30: feature extraction is performed on the preprocessed thermal image data and vibration data, the feature extraction including extracting an infrared image feature vector of the thermal image data through an improved ResNet-50 network , and extracting a vibration data feature vector of the vibration data through a multi-scale time sequence convolution network ; S40: fuse the infrared image feature vector with the vibration data feature vector , generate a tearing probability based on the fusion result, and perform a hierarchical early warning according to a probability threshold.
2. The method of claim 1, wherein the method is a thermal-vibrational heterogeneous data fusion conveyor belt edge tear detection method. Step S20 comprises: S21: strictly aligning the collection time of each frame of infrared thermal image with the starting time of the vibration signal segment through a hardware pulse signal; S22: performing Gaussian filter denoising, histogram equalization enhancement, size standardization on the infrared thermal image, and converting the pixel value of the infrared thermal image to a distribution range with a mean value of 0 and a standard deviation of 1 by using zero-mean normalization method to obtain the processed thermal image data; S23: performing wavelet threshold denoising on the vibration signal and normalizing it to the [0, 1] interval to obtain the processed vibration data.
3. The method of claim 1, wherein the method is a thermal-vibrational heterogeneous data fusion conveyor belt edge tear detection method. The improved ResNet-50 network replaces the original 7x7 convolutional layer with three 3x3 convolutional layers in the initial layer, embeds an SE channel attention module in the residual block, and adds a spatial attention module at the end. The calculation formula of the first 3x3 convolutional layer in the three 3x3 convolutional layers is: ; wherein: is a first 3x3 convolution kernel parameter; is a bias term; is a pre-processed infrared thermal image; is an output feature map, and the size of the output feature map is 224x224x64; ReLU is a linear rectification activation function, used to introduce nonlinear features; x, y represent the horizontal and vertical coordinates of the pixel points of the output feature map; p, q correspond to the row index and column index of the convolution kernel, respectively; The subsequent two 3x3 convolutional layers have the same structure, and the output feature map size is maintained at 224x224, and the channel number is 64, 128, respectively.
4. The method of claim 1, wherein the method is a thermal-vibrational heterogeneous data fusion conveyor belt edge tear detection method. The multi-scale time sequence convolutional network comprises three multi-scale convolutional branches: The first branch: the 1D convolution kernel size is 3, the dilation is set to 1, the receptive field is 3, the local high-frequency impact is captured, and the output feature dimension is 1024x64; The second branch: the 1D convolution kernel size is 5, the dilation is set to 2, the receptive field is 9, the medium-term vibration trend is captured, and the output feature dimension is 1024x64; The third branch: the 1D convolution kernel size is 7, the dilation is set to 4, the receptive field is 25, the long-term low-frequency component is captured, and the output feature dimension is 1024x64; The three multi-scale convolutional branches are spliced into a 192-channel feature map through a fusion layer, and compressed to 128 channels through a 1x1 convolution, and the formula is: ; where n is the size of the convolution kernel, and are 3, 5, 7, respectively, is the i-th parameter of the 1D convolution kernel, is the normalized vibration signal, C is the output feature, is the bias term, denotes the time step index of the vibration signal, denotes the convolution kernel index, and ReLU is the linear rectifier activation function.
5. The method of claim 3, wherein the method further comprises: Step S40 comprises: S41: Real-time acquisition of environmental parameters, dynamic calculation of modal balance coefficient according to environmental parameters , the infrared image feature vector and the vibration data feature vector are weighted and adjusted based on the modal balance coefficient, and the formula is: ; wherein: is the weighted , is the weighted ; To Channel compression to 512 dimensions by 1x1 convolution, to Expansion to 512 dimensions by two fully connected layers, feature concatenation, forming a joint feature vector of 1x1152 ; S42: generate a joint feature vector generate weighted fusion features through a dual-channel attention mechanism ; S43: outputting the weighted fusion feature inputting a tearing probability P e [0, 1] outputted by a full connection network, the full connection network comprising a first full connection layer and a second full connection layer, the first full connection layer outputting a 64-dimensional feature vector, and the second full connection layer outputting a 1-dimensional probability value, and the formula being: ; ; wherein: is an output feature vector of the first fully connected layer, , is a weight matrix of the first and second fully connected layers, , is a bias term of the first and second fully connected layers, P is a tearing probability, ReLU is a linear rectifier activation function, and Sigmoid is a sigmoid activation function that maps the output to the interval [0, 1]. S44: performing hierarchical early warning according to the probability threshold, and adopting a three-level response mode.
6. The method of claim 5, wherein the method further comprises: The specific rules of the three-level response mode are as follows: When 0≤P<0.2, it is determined to be a normal state, and no intervention is required; When 0.2≤P<0.6, it is determined to be a first-level warning, i.e., suspected tearing, triggering an audible and light alarm, and a patrol inspection is conducted; When 0.6≤P≤1, it is determined to be a second-level warning, i.e., serious tearing, and the machine is immediately stopped.
7. The method of claim 5, wherein the method further comprises: The SE channel attention module performs: Squeeze operation: performing global average pooling on the output feature map of the residual block to obtain channel-level statistical information, and the formula is as follows: ; wherein: F is a residual block output feature map, and the output feature map has a dimension of , H represents height, W represents width, C represents the number of channels, i, j are the row index and column index of the feature map pixel points, c represents the channel number index, is the global feature of the cth channel; The excitation operation: generate channel weight through two full connection layers and sigmoid activation function, and weight and enhance the original feature map, the formula is: ; ; wherein: is the first fully connected layer weight, is the second fully connected layer weight, is the channel statistic vector obtained by the squeezing operation, is the channel weight vector output by the sigmoid function, is the sigmoid function, is the weight vector of the cth channel, is the weighted enhanced , is the cth channel of F, and ReLU is a linear rectification activation function.
8. The method of claim 3, wherein the method is a thermal-vibrational heterogeneous data fusion conveyor belt edge tear detection method. The spatial attention module performs: Maximum and average pooling of the input features in the channel dimension; Concatenate the pooling results and generate a spatial weight matrix through convolution, the formula is: ; wherein: is a feature map input to spatial attention, is a spatial weight matrix, is a sigmoid function, MaxPool is a maximum pooling operation in the channel dimension of the input feature map, AvgPool is an average pooling operation in the channel dimension of the input feature map, and Conv is a convolution operation. Element-wise weighting according to the weight matrix, the formula is: ; wherein: is the feature map weighted by the spatial attention module, denotes the element-wise multiplication of the Hadamard product.
9. The method of claim 7, wherein the method further comprises: The dual-channel attention mechanism performs: The joint feature vector Global average pooling is performed to generate channel-level statistical features, mutual information MI of the infrared sub-feature and the vibration sub-feature is calculated, and the associated channels with MI>0.6 are screened, and the basic channel weight is generated through a fully connected layer and a Sigmoid function , and the formula is: ; wherein: is the first fully connected layer weight, is the second fully connected layer weight, is the squeezed channel statistics vector, is the sigmoid function, , are the bias terms for the first and second fully connected layers, respectively. The joint weight matrix is generated by fusing the spatial coordinates and the correlation features of the time points through a 3x3 convolution layer The formula is: ; wherein: is the infrared image feature vector weighted by the modal balance coefficient, is the vibration data feature vector weighted by the modal balance coefficient, is the joint weight matrix, is the sigmoid function, Conv is the convolution operation, MaxPool is the maximum pooling operation in the channel dimension of the input feature map, and PeakDetect is the peak detection operation on the vibration feature vector. The base channel weights are multiplied by the joint weight matrix with the joint weight matrix element-wise to obtain the dual-channel attention weights; The joint feature vector is combined with the attention weight vector to obtain a weighted fusion feature The weighted fusion feature is obtained by element-wise multiplication of the dual-channel attention weight The formula is: 。 10. A thermal-vibrational heterogeneous data fusion conveyor belt edge tear detection system characterized by: The system is used for performing a hot and heterogeneous data fusion conveyor belt edge tear detection method as claimed in any one of claims 1-9, and the conveyor belt edge tear detection system comprises: An information acquisition module: including a line laser, an infrared thermal imager and a vibration sensor; the laser line emitted by the line laser is vertically irradiated on the surface edge of the conveyor belt of the belt conveyor, for actively heating the conveyor belt to form a temperature field, the infrared thermal imager is used to acquire an infrared thermal image of the profile of the surface edge of the conveyor belt, and the vibration sensor is used to acquire a vibration signal in the running of the conveyor belt; A preprocessing module: used for preprocessing the acquired infrared thermal image and vibration signal; A feature extraction module: used for feature extraction on the preprocessed thermal image data and vibration data, to obtain an infrared image feature vector and a vibration data feature vector; A tear detection module: including a dynamic fusion unit, a dual-channel attention unit and a hierarchical early warning unit: the dynamic fusion unit is used for calculating a modal balance coefficient based on environmental parameters, for weighting and adjusting the infrared image feature vector and the vibration data feature vector, and then splicing them into a joint feature vector; the dual-channel attention unit is used for generating a weighted fusion feature from the joint feature vector; and the hierarchical early warning unit is used for inputting the weighted fusion feature into a full connection network to generate a tear probability, and performing hierarchical early warning according to a probability threshold.
Citation Information
Patent Citations
Conveying belt tearing detection device and method
CN119490042A
Contact type belt conveyor conveying belt tearing detection device
CN219770980U
Intelligent online monitoring system of belt conveyor
CN114275483A
Online detection method and system for edge tearing of conveying belt
CN115171051A
Output belt damage detection algorithm based on multiple modes
CN116543199A