Belt coal falling detection method for resource-constrained equipment
By combining lightweight convolutional neural network and event camera methods, the problem of large computing requirements and insufficient generalization capabilities of edge object detection on resource-constrained devices is solved, and efficient and accurate coal-fall detection is achieved.
Patent Information
- Application Number
- CN202510718015.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-05-30
AI Technical Summary
The existing edge target detection technology has problems such as large computing demand, insufficient generalization capability and poor real-time performance in resource-constrained equipment, making it difficult to effectively detect coal-fall failures in mine conveyor belts.
Combining the lightweight convolutional neural network structure and traditional object detection methods, an event camera is used to obtain asynchronous event data flow, and efficient coal drop detection is achieved through feature extraction and fusion.
It improves the accuracy and robustness of coal-fall detection, adapts to complex environments, reduces the computing resource requirements, and is suitable for edge devices with resource-constrained.
Smart Images

Figure CN120259636A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer vision target detection, and particularly relates to a method for detecting coal falling on a belt of a resource-constrained device. Background Art
[0002] Target detection is one of the core tasks of computer vision. In recent years, with the rapid development of deep learning and the improvement of the scale of data sets, target detection technology has made remarkable progress. The combination of artificial intelligence and the Internet of Things has promoted the development of target detection technology from the cloud to edge devices, making it widely used in fields such as intelligent cameras and mine robots. Traditional coal mine monitoring relies on manual inspections with poor real-time performance and high costs, as well as cloud computing. With the rapid development of intelligent mines, the coal industry is accelerating its transformation towards automation, unmanned operation, and intelligence. Edge target detection technology has played an important role in coal mine safety monitoring, equipment inspections, etc., not only reducing the cost of manual monitoring but also significantly improving production safety and efficiency. Among them, the mine conveyor belt, as a key device for coal transportation, operates for a long time in a harsh environment and is prone to faults such as tearing, deviation, and coal falling, which affect production efficiency and even cause safety accidents. The traditional manual inspection method has low efficiency and poor real-time performance, making it difficult to discover and handle problems in a timely manner. Therefore, edge target detection technology with high efficiency and real-time performance shows great research and application potential in the field of intelligent mine belt fault detection, providing strong support for the intelligent development of the coal industry.
[0003] The current edge target detection technologies are mainly divided into the following categories: edge target detection based on deep learning, edge target detection based on traditional computer vision, edge target detection based on sensor data, and target detection based on edge collaborative computing.
[0004] The main defects of current similar methods are as follows: 1. Edge target detection methods based on deep learning have high requirements for training data, limited generalization ability, and high energy consumption, making it difficult to perform target detection tasks on edge devices.
[0005] 2. Edge target detection methods based on traditional computer vision have poor adaptability, rely on manually set parameters, and have insufficient generalization ability, making it difficult to perform target detection tasks on edge devices.
[0006] 3. Edge target detection methods based on sensor data have poor adaptability to different environments, large sensor data noise, are easily affected by environmental interference, and have poor real-time performance in performing target detection tasks.
[0007] 4. Edge target detection methods based on edge system computing are limited by differences in bandwidth, power consumption, etc. of different devices, are difficult to schedule, and have great difficulty in performing target detection tasks. Summary of the Invention
[0008] In view of the problems existing in the above-mentioned prior art, the present invention provides a method for detecting coal falling on a belt of a resource-constrained device, which can combine a lightweight convolutional neural network structure with a traditional object detection method, reduce the demand for computing resources for the object detection task, improve the accuracy and robustness of coal falling detection, have strong adaptability to complex environments, have strong generalization ability, be applicable to the detection of coal falling on resource-constrained edge devices, and have high application value.
[0009] A method for detecting coal falling on a belt of a resource-constrained device according to the present invention includes steps S1, S2, S3, S4, and S5, wherein steps S2, S3, and S4 are performed synchronously, and the specific method is as follows: S1. Event data acquisition, and the specific method is as follows: S1.1. Configure the basic parameters of the event camera for focusing on the change area of the belt surface; S1.2. Use the event vision sensor to obtain an asynchronous event data stream; S1.3. Read the event stream in real time and aggregate it into event frames according to the set time window; S1.4. Cache the aggregated event frames into the processing queue for subsequent feature extraction modules to call; S2. Event feature preprocessing and construction, and the specific method is as follows: S2.1. Preprocess the event frames, including size normalization and channel standardization; S2.2. Use the lightweight RepNeXtBlock network structure to extract low-level spatial features and capture local structure information and edge contours; S2.3. Stack multiple layers of RepNeXtBlock deep modules to extract high-level semantic features, such as information on the position of the coal falling area, etc.; S2.4. Use FPN to fuse multi-scale features, enhance the detection effect of small target coal falling, and improve the semantic expression ability of features; S2.5. Input the extracted features into the detection head for classification and positioning of coal falling; S3. Compare the event frames before and after, and construct spatio-temporal feature fusion, and the specific method is as follows: S3.1. Perform spatial alignment and pixel normalization processing on two adjacent event maps; S3.2. Calculate the event change area between the front and back frames, and highlight the coal falling fault occurrence area through the frame difference method; S4. Perform fine recognition based on the event edge structure, and the specific method is as follows: S4.1. Grayscale the event frames and apply methods such as median filtering to remove high-frequency noise; S4.2. Calculate the gradient of the image using the Sobel operator; S4.3. Remove non-edge pixels using non-maximum suppression to obtain a refined edge; S4.4. Set double thresholds to achieve edge connection and output the result; S5. Perform spatio-temporal-edge feature fusion and output the coal dropping detection result.
[0010] Furthermore, in step S1.3, OpenCV is used to process the video stream, the camera data is read through cv2.VideoCapture, the interval is set to 1 second, and an image is saved every 1 second.
[0011] Furthermore, the specific method of step S2.1 is as follows: First, the collected image is converted to a standard size of 640×640 to avoid affecting the spatial region positioning result due to different sizes of the collected images; Then, color channel normalization is performed to normalize the collected RGB image to [-1, 1]; Finally, data augmentation techniques are adopted to perform rotation and color enhancement operations on some of the collected images to improve the generalization ability of the model.
[0012] Furthermore, the specific method of step S2.2 is as follows: 1) RepNeXtBlock adopts a classic depthwise separable convolution to reduce the computational complexity, making it suitable for resource-constrained edge devices. In a standard 3×3 convolution, each pixel calculation needs to consider all channels in the surrounding 3×3 positions, which will result in a large computational complexity. Its computational complexity is expressed as: ; Among them, , is used to represent the height and width of the feature map, represents the number of input channels, represents the number of output channels, represents the size of the convolution kernel. The computational complexity of the depthwise separable convolution is expressed as: ; Among them represents the number of channels obtained by reducing the dimension through a 1×1 convolution. It can be easily obtained that the depthwise separable convolution can effectively reduce the computational complexity and is suitable for edge devices.
[0013] 2) ReLU is used to introduce non-linear features, enabling the neural network to learn more complex patterns. Its mathematical definition is expressed as: ; This non-linear activation function enables the network to learn complex non-linear patterns such as image features, allowing the deep network to learn different feature representations. For example, it can learn edge features at the lower layers, complex shapes at the intermediate layers, and semantic features at the higher layers.
[0014] Furthermore, the specific method of step S2.3 is as follows: 1) RepNeXt introduces group convolution to reduce the number of parameters and enhance the feature representation ability. The computational operation of the common standard two-dimensional convolution is expressed as: ; where is the input feature map with a size of ; is the convolution kernel with a size of ; is the output feature map with a size of ; and are the number of input and output channels respectively. Such a traditional two-dimensional convolution operation has a high computational cost when the number of input and output channels is large, and is not suitable for edge devices with limited computing resources.
[0015] In group convolution, the input channels and the output channels are divided into G groups, and each group is calculated independently. Its computational operation is expressed as: ; where G represents the number of groups, and represent the input and convolution kernel of the th group. Such group convolution reduces the computational amount, is suitable for efficient deep network design, and is more suitable for edge devices with limited computing resources.
[0016] 2) The method of structural reparameterization is adopted to enhance the feature representation ability through multiple parallel paths during training, and simplify it to a single 3×3 convolution during inference to improve the computational efficiency. During the training phase, multiple paths are used to enhance the feature representation ability, and the computation at this stage is expressed as: ; where is a 3×3 group convolution used to enhance the local receptive field; is a 1×1 ordinary convolution used to enhance the feature representation ability; is the identity mapping used to improve the information flow. During the inference phase, the weights of these paths are merged to make the whole into an equivalent 3×3 convolution, reducing the computational amount and improving the inference speed. Its computation is expressed as: ; Among them means expanding the weights of the 1×1 convolution to a 3×3 size so that it can be added to additively; represents an identity matrix, which is added to the weights of the convolution kernel. In this way, the computational complexity of the inference operation is expressed as: ; Obviously, the computational complexity is reduced and the inference process is accelerated.
[0017] Furthermore, the specific method of the step S2.4 is as follows: Combine the high-level features and low-level features, and use FPN for feature fusion. According to steps 2.2 and 2.3, the low-level features from the shallow network and the high-level features from the deep network can already be obtained. The low-level features have high resolution, retain strong spatial details and weak semantic features, while the high-level features have strong semantic information but have low spatial resolution and insufficient spatial details during continuous downsampling operations. Therefore, FPN is used for feature fusion. FPN consists of three main parts, namely bottom-up, top-down, and lateral connections.
[0018] The bottom-up is the process of feature extraction by the Backbone. For a given input image , after passing through the backbone network RepNeXt, feature maps at multiple levels are extracted, and the process is expressed as: ; Among them represents the process of feature extraction by the Backbone; represents the shallowest feature, which has a small receptive field and mainly contains texture information; represents the deepest feature, which has a large receptive field and mainly contains semantic information.
[0019] The top-down uses transposed convolution for upsampling to restore high-resolution features. For the feature maps obtained in the bottom-up process , the top-down processing process is expressed as: ; Among them represents upsampling the previous layer of feature maps; represents using 1×1 convolution to reduce the dimension of the bottom-layer feature maps at the same level so that the number of channels is aligned with the upsampling result. The two are added to complete the lateral connection, enhancing the feature expression ability.
[0020] After the bottom-up and top-down processes, the output result can be obtained. These output features contain both the spatial information of the low-level features and the semantic information of the high-level features, and are suitable for the object detection scenario.
[0021] Further, in step S2.5, the extracted features are input into the detection head for the classification and positioning of coal dropping.
[0022] Further, the specific method of step S3.1 is as follows: 1) Align the front and rear frames. Since factors such as camera jitter affect the alignment between frames when collecting images according to the video stream, directly performing feature extraction in this case is likely to lead to comparison errors, thus affecting the detection results. Therefore, the front and rear frames are aligned first. During this process, global alignment based on optical flow and local alignment based on feature point matching need to be completed.
[0023] During global alignment, the global transformation matrix is calculated based on optical flow, and matrix transformation is performed on the entire input image. The optical flow method aligns the front and rear frames by estimating the pixel-level motion between adjacent frames, which is expressed as: ; where represents the optical flow field, which means the direction and speed of pixel moving in adjacent frames; are the optical flow components in the horizontal and vertical directions respectively.
[0024] The basic steps of local alignment include key point detection, feature point matching, and calculation of the transformation matrix. The present invention uses the KNN + cross-check method to complete feature point detection, that is, two nearest matching points are found for each key point, and matching tests are performed through their positioning. Setting the threshold to 0.5, this method is expressed as: ; where , represent the positions of two adjacent points. When their ratio is less than 0.5, it can be considered that this step of local alignment has been completed.
[0025] 2) The purpose of normalization is to eliminate the influence of different illuminations and contrasts on the feature extraction ability of adjacent frames and improve the generalization ability of the model. This step is very important in the field of coal industry target detection with complex and changeable environments. The present invention uses the Min-Max normalization method, which is expressed as: ; where represents the original pixel value, represents the maximum pixel value, represents the minimum pixel value, represents the pixel value after the normalization operation. This process compresses the pixel value to between [0, 1], reduces the differences in factors such as brightness, and improves the comparison accuracy.
[0026] Furthermore, the specific method of step S3.2 is as follows: The background modeling method is used to obtain the change region between the front and back frames, so as to obtain the region where the coal falling fault occurs. This process includes two parts: inter-frame difference and Gaussian mixture model. Inter-frame difference is to establish a stable background model using time series and detect foreground changes. This process is expressed as: ; where is the background model, is the current frame. The Gaussian mixture model not only considers the differences between adjacent frames, but also establishes a long-term stable background model through this model. The Gaussian mixture model is expressed as: ; where represents the weight of the th Gaussian distribution, represents the mean of the th Gaussian distribution, represents the variance of the th Gaussian distribution, represents the number of Gaussian mixture distributions.
[0027] Furthermore, the specific method of step S4.1 is as follows: 1) Grayscale conversion: Edge detection is based on grayscale images. The image is grayscale-converted, and the three-channel RGB color image is converted into a single-channel image. This operation is expressed as: ; where R, G, and B represent the pixel values of the red, green, and blue channels in the image respectively.
[0028] 2) Apply the median filtering method to remove high-frequency noise: The basic operation of median filtering is that for each pixel point in the image, the pixel values in its surrounding neighborhood are sorted, and the value of this pixel is replaced with the middle value after sorting. This operation can effectively remove salt-and-pepper noise and will not blur the edge information in the image. The steps of median filtering include selecting the filter window size, local sorting, calculating the median, and replacing the pixel value.
[0029] A sized filter window covers a pixel point in the image. The pixels in the neighborhood around this pixel point are represented as , and represents the number of pixels in the neighborhood. Then the median filtering operation is expressed as: ; where represents the pixel value after denoising, represents the neighborhood of the current pixel point represents selecting the middle value as the result after sorting the pixel values in the neighborhood
[0030] Through the above operations, an image processed by median filtering can be obtained, and the processing continues
[0031] Furthermore, the specific method of step S4.2 is as follows Calculate the gradient of the image using the Sobel operator. The process of calculating the gradient by the Sobel operator is expressed as ; ; ; where represents the horizontal direction gradient represents the vertical direction gradient
[0032] Furthermore, the specific method of step S4.3 is as follows The process of removing non-edge pixels can be divided into two steps, namely calculating the gradient direction and retaining the local maximum. The process of calculating the gradient direction is expressed as ; After obtaining the gradient direction compare the current pixel with adjacent pixels, retain the local maximum, and suppress non-maximum values to complete non-maximum suppression
[0033] Furthermore, the specific method of step S4.4 is as follows To reduce the influence of noise, set double thresholds, namely the high threshold and the low threshold , for pixel values greater than mark them as strong edges, and for pixels between and only retain them if they contain strong edges, otherwise suppress. Then continuously track the edge pixels. Retain them when they are connected to strong edges, otherwise delete them. The output result can be obtained through these steps
[0034] Furthermore, the specific method of step S5 is as follows Obtain the spatial feature through step S2.1 and step S2.5, obtain the temporal feature through step S3.1 and step S3.2, and obtain the edge feature through step S4.1 and step S4.4; use the weighted fusion method to combine the spatial feature , the temporal feature , Edge features Feature fusion is performed, and this process is expressed as: ; Among them represents the fused feature. Detection is performed on the trained model according to YOLOV12 with the backbone replaced by RepNeXtBlock, and this process is expressed as: ; Among them represents the result of coal falling detection.
[0035] The beneficial effects of the present invention are as follows: Compared with the prior art, the present invention introduces an event camera as the core of visual perception, and utilizes its characteristic of being highly sensitive to dynamic changes. Without the need for traditional inter-frame difference, high-precision spatio-temporal change information can be directly obtained, thereby significantly improving the detection effect of the coal falling fault area. The event camera has high temporal resolution and high dynamic range. In the face of complex working conditions such as strong light interference, low illuminance, and high-speed movement, it can maintain stable and clear edge contour information, effectively reducing the common motion blur and exposure distortion problems in traditional image processing methods, and further improving the robustness and detection accuracy of the system. In addition, the present invention combines the sparsity characteristic of event data, and uses a lightweight RepNeXt backbone network for feature extraction and target recognition, significantly reducing the consumption of computing resources and improving the real-time processing ability on edge devices. Through the event-driven data structure, the present invention naturally has the fusion characteristics of space, time, and edge information, avoiding the delay and redundancy caused by multi-module splicing, and ensuring that the detection system still has excellent response speed and stable performance under resource-constrained conditions. In summary, this method gives full play to the dynamic perception advantage of event vision, shows higher robustness, real-time performance, and engineering adaptability in the complex and changeable belt coal falling detection scenario, and has good application and promotion value. Description of the Drawings
[0036] Figure 1 is a flowchart of a method for detecting coal falling on a belt of a resource-constrained device according to the present invention. Detailed Embodiments
[0037] The following will elaborate on the technical content of its implementation plan in detail in combination with the specific drawings of the present invention. It should be noted that the embodiments listed herein are only partial examples of the present invention and do not exhaust all possibilities. Based on the technical solutions disclosed in the present invention, any other implementation solutions obtained by those skilled in the art without creative efforts shall fall within the protection scope of the present invention. Embodiment 1
[0038] As Figure 1As shown in the figure, a method for detecting coal falling on a belt of a resource-constrained device includes steps S1, S2, S3, S4, and S5, where steps S2, S3, and S4 are performed synchronously. The specific method is as follows: S1. Event data collection, and the specific method is as follows: 1.1 Configure the basic parameters of the event camera for focusing on the change of the belt surface area; 1.2 Use the event vision sensor to obtain an asynchronous event data stream; 1.3 Read the event stream in real time and aggregate it into event frames according to the set time window; 1.4 Cache the aggregated event frames into the processing queue for subsequent feature extraction modules to call.
[0039] S2. Use an improved lightweight model for feature extraction, and the specific method is as follows: 2.1 Preprocess the event frames, including size normalization and channel standardization. This step involves the following specific implementation details: First, convert the collected images to a standard size of 640×640 to avoid affecting the spatial region positioning results due to different sizes of the collected images; Then perform color channel normalization to normalize the collected RGB images to [-1, 1]; Finally, use data augmentation techniques to perform operations such as rotation and color enhancement on some of the collected images to improve the generalization ability of the model.
[0040] 2.2 Use the lightweight RepNeXtBlock network structure to extract low-level spatial features, capture local structure information and edge contours. A single RepNeXtBlock can be regarded as consisting of three parts, namely 1×1 convolution, 3×3 grouped convolution, and ReLU activation; RepNeXtBlock uses the classic depthwise separable convolution to reduce the computational amount, making it suitable for resource-constrained edge devices. In the standard 3×3 convolution, each pixel calculation needs to consider all channels in the surrounding 3×3 positions, resulting in a large computational amount. Its computational amount is expressed as: ; Among them, ,[[]]END]] is used to represent the height and width of the feature map, represents the number of input channels, represents the number of output channels, represents the size of the convolution kernel. The computational amount of the depthwise separable convolution is expressed as: ; Among them represents the number of channels obtained after dimensionality reduction by 1×1 convolution. As mentioned above, it is easy to see that depthwise separable convolution can effectively reduce the computational cost and is suitable for edge devices.
[0041] ReLU is used to introduce non-linear features, enabling the neural network to learn more complex patterns. Its mathematical definition is expressed as: ; Such non-linear activation functions enable the network to learn complex non-linear patterns such as image features, allowing the deep network to learn different feature representations, such as learning edge features at the low level, complex shapes at the middle level, and semantic features at the high level.
[0042] 2.3 Stack multiple layers of RepNeXtBlock deep modules to extract high-level semantic features, such as information about the position of the coal falling area; RepNeXt introduces group convolution to reduce the number of parameters and improve the feature representation ability. The common standard computational operation is expressed as: ; where is the input feature map with size ; is the convolution kernel with size ; is the output feature map with size ; and are the number of input and output channels respectively. Such traditional two-dimensional convolution operations have a high computational cost when the number of input and output channels is large and are not suitable for edge devices with limited computational resources.
[0043] In group convolution, the input channels and output channels are divided into G groups, and each group is calculated independently. Its computational operation is expressed as: ; where G represents the number of groups, and represent the input and convolution kernel of the th group. Such group convolution reduces the computational cost, is suitable for efficient deep network design, and is more suitable for edge devices with limited computational resources.
[0044] This embodiment also adopts the method of structural reparameterization to enhance the feature representation ability through multiple parallel paths during training and simplify it to a single 3×3 convolution during inference to improve the computational efficiency. During the training stage, multiple paths are used to enhance the feature representation ability. The computation in this stage is expressed as: ; Among them is a 3×3 group convolution used to enhance the local receptive field; is a 1×1 ordinary convolution used to enhance the feature expression ability; is an identity mapping used to improve the information flow. During the inference stage, the weights of these paths are combined so that the whole becomes an equivalent 3×3 convolution, reducing the computational amount and improving the inference speed. Its computational representation is: ; Among them means expanding the weights of the 1×1 convolution to a 3×3 size so that it can be added to ; represents an identity matrix, which is added to the weights of the convolution kernel. In this way, the computational amount of the inference operation is represented as: ; Obviously, the computational amount is reduced and the inference process is accelerated.
[0045] 2.4 Combine high-level features and low-level features and use FPN for feature fusion. After steps 2.2 and 2.3 of this embodiment, low-level features from the shallow network and high-level features from the deep network can already be obtained. Low-level features have high resolution, retain strong spatial details and weak semantic features, while high-level features have strong semantic information but have low spatial resolution and insufficient spatial details during continuous downsampling operations. Therefore, FPN is used for feature fusion. FPN consists of three main parts, namely bottom-up, top-down, and lateral connections.
[0046] The bottom-up is the process of feature extraction by the Backbone. For a given input image , the backbone network RepNeXt is used to complete the extraction of feature maps at multiple levels, and its process is represented as: ; Among them represents the process of feature extraction by the Backbone; represents the shallowest feature with a small receptive field, mainly containing texture information; represents the deepest feature with a large receptive field, mainly containing semantic information.
[0047] The top-down uses deconvolution for upsampling to restore high-resolution features. For the feature maps obtained during the bottom-up process , the top-down processing process is represented as: ; Among them represents upsampling the previous layer of feature maps; It means that a 1×1 convolution is used to reduce the dimension of the underlying feature map of the same level, so that the number of channels is aligned with the upsampling result. The two are added together to complete the lateral connection, enhancing the feature expression ability.
[0048] After the bottom-up and top-down processes, the output result can be obtained. These output features contain both the spatial information of low-level features and the semantic information of high-level features, and are suitable for the target detection scenario.
[0049] 2.5 Input the extracted features into the detection head for the classification and localization of coal falling.
[0050] The detection head of this embodiment adopts the Anchor-Free mechanism of the FCOS structure. The detection head consists of three parts, namely, a classification head for predicting the class probability of each pixel point, a regression head for directly regressing the offset of the target box, and a centerness branch for measuring whether a certain pixel point is the center of the target to improve the detection stability.
[0051] In addition, under the Anchor-Free mechanism of the FCOS structure, the loss function consists of three parts, namely, classification loss, regression loss, and centerness loss. The classification loss can solve the problem of class imbalance and reduce the interference of negative samples on training. The classification loss is expressed as: ; where is the probability of the predicted class, and are used to control the weights of easy and hard samples; the regression loss is used to optimize the target localization, and the regression loss is expressed as: ; where is the true bounding box, represents the predicted bounding box, represents the intersection over union between the two, which can measure the accuracy of the prediction. This process directly optimizes the IoU and improves the localization accuracy. The centerness loss is used to measure the reliability of the predicted box and prevent some boxes from being misjudged due to excessive deviation. The centerness loss is expressed as: ; where, represents the true centerness, and
[0052] S3. Combine the spatio-temporal feature map with the edge information to output the final result. The specific method is as follows 3.1 Align and normalize the front and rear frames. The specific method of this operation is as follows: Front and rear frame alignment. Since factors such as camera jitter during image acquisition according to the video stream affect the frame alignment, directly performing feature extraction in this case is likely to lead to comparison errors, thus affecting the detection results. Therefore, front and rear frame alignment is performed first. During this process, global alignment based on optical flow and local alignment based on feature point matching need to be completed.
[0053] During global alignment, a global transformation matrix is calculated based on optical flow, and matrix transformation is performed on the entire input image. The optical flow method aligns the front and rear frames by estimating the pixel-level motion between adjacent frames, which is expressed as: ; where represents the optical flow field, which means the direction and speed of pixel motion in adjacent frames; are the optical flow components in the horizontal and vertical directions respectively.
[0054] The basic steps of local alignment include key point detection, feature point matching, and calculation of the transformation matrix. The present invention uses the KNN + cross-check method to complete key point detection, that is, two nearest matching points are found for each key point, and matching tests are performed through their positioning. Setting the threshold to 0.5, this method is expressed as: ; where , represent the positions of two adjacent points. When their ratio is less than 0.5, it can be considered that this step of local alignment has been completed.
[0055] The purpose of normalization is to eliminate the ability of different illuminations and different contrasts to extract features from adjacent frames and improve the generalization ability of the model. This step is very important in the field of coal industry target detection where the environment is complex and changeable. The present invention uses the Min-Max normalization method, which is expressed as: ; where represents the original pixel value, represents the maximum pixel value, represents the minimum pixel value, represents the pixel value after the normalization operation. This process compresses the pixel value to between [0, 1], reduces the differences in factors such as brightness, and improves the comparison accuracy.
[0056] 3.2 Calculate the event change area between the front and rear frames, and highlight the area where the coal falling fault occurs through the frame difference method.
[0057] The present invention uses a background modeling method to obtain the change region between the front and rear frames, thereby obtaining the region where the coal dropping fault occurs. This process includes two parts: inter-frame difference and Gaussian mixture model. The inter-frame difference uses time series to establish a stable background model and detect foreground changes. This process is expressed as: ; where is the background model, is the current frame. The Gaussian mixture model not only considers the differences between adjacent frames, but also establishes a long-term stable background model through this model. The Gaussian mixture model is expressed as: ; where represents the weight of the th Gaussian distribution, represents the mean of the th Gaussian distribution, represents the variance of the th Gaussian distribution, represents the number of Gaussian mixture distributions.
[0058] S4. Perform fine recognition based on the event edge structure. The specific method is as follows: 4.1 Gray-scale the event frame and apply the median filtering method to remove high-frequency noise. The specific method is as follows: Gray-scale The edge detection in this embodiment is calculated based on the gray-scale image. The image is gray-scaled, and the color image with three RGB channels is converted into a single-channel image. This operation is expressed as: ; where R, G, and B respectively represent the pixel values of the red, green, and blue channels in the image.
[0059] Apply the median filtering method to remove high-frequency noise The basic operation of median filtering is that for each pixel point in the image, the pixel values in its surrounding neighborhood are sorted, and the value of this pixel is replaced with the median value after sorting. This operation can effectively remove salt-and-pepper noise and does not blur the edge information in the image. The steps of median filtering include selecting the filter window size, local sorting, calculating the median, and replacing the pixel value.
[0060] A sized filter window covers a pixel point in the image. The pixels in the surrounding neighborhood of this pixel point are represented as , and using to represent the number of pixels in the neighborhood, the median filtering operation is expressed as: ; where represents the pixel value after denoising, represents the neighborhood of the current pixel point, represents that the middle value is selected as the result after sorting the pixel values in the neighborhood.
[0061] Through the above operations, an image processed by median filtering can be obtained and further processed.
[0062] 4.2 Calculate the gradient of the image using the Sobel operator.
[0063] The process of calculating the gradient using the Sobel operator is expressed as: ; ; ; where represents the horizontal gradient, represents the vertical gradient.
[0064] 4.3 Use non-maximum suppression to remove non-edge pixels and obtain a refined edge.
[0065] The process of removing non-edge pixels can be divided into two steps, namely calculating the gradient direction and retaining the local maximum. The process of calculating the gradient direction is expressed as: ; After obtaining the gradient direction , the current pixel is compared with adjacent pixels, the local maximum is retained, and non-maximum values are suppressed to complete non-maximum suppression.
[0066] 4.4 Set double thresholds to achieve edge connection and output the result.
[0067] To reduce the influence of noise, the present invention sets double thresholds, namely a high threshold and a low threshold . For pixel values greater than , they are marked as strong edges. For pixels between and , they are retained only when they contain strong edges, otherwise they are suppressed. Then, if the edge pixels are continuously tracked, they are retained when connected to strong edges, otherwise they are deleted. Through these steps, the output result can be obtained.
[0068] S5. Spatial-temporal-edge feature fusion and output the coal falling detection result Spatial features can be obtained through S2 , temporal features can be obtained through S3 , and edge features can be obtained through S4 。This embodiment uses weighted fusion for feature fusion, and this process is expressed as: ; where represents the fused features. Detect the trained model according to YOLOV12 with the backbone replaced by RepNeXtBlock, and this process is expressed as: ; where represents the result of coal falling detection.
[0069] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.
Claims
1. A method for detecting coal falling on a belt of a resource-constrained device, characterized in that, It includes steps S1, S2, S3, S4, and S5. Among them, steps S2, S3, and S4 are carried out synchronously. The specific method is as follows: S1. Event data collection; S2. Event feature preprocessing and construction. The specific method is as follows: S2.
1. Preprocess the event frame, including size normalization and channel standardization; S2.
2. Use the lightweight RepNeXtBlock network structure to extract low-level spatial features and capture local structure information and edge contours; S2.
3. Stack multiple layers of RepNeXtBlock deep modules to extract high-level semantic features; S2.
4. Use FPN to fuse multi-scale features, enhance the detection effect of small target coal falling, and improve the semantic expression ability of features; S2.
5. Input the extracted features into the detection head for classification and localization of coal falling; S3. Compare the event frames before and after, and construct spatio-temporal feature fusion; S4. Perform fine recognition based on the event edge structure; S5. Spatial-temporal-edge feature fusion to output the coal falling detection result.
2. The belt coal falling detection method for a resource-constrained device according to claim 1, wherein The specific method of step S1 is as follows: S1.
1. Configure the basic parameters of the event camera for focusing on the change area of the belt surface; S1.
2. Use the event vision sensor to obtain the asynchronous event data stream; S1.
3. Read the event stream in real time, use OpenCV to process the video stream, and aggregate it into event frames according to the set time window; S1.
4. Cache the aggregated event frames into the processing queue for subsequent feature extraction modules to call; The specific method of step S3 is as follows: S3.
1. Perform spatial alignment and pixel normalization on two adjacent event images; S3.
2. Calculate the event change area between the front and back frames, and highlight the coal falling fault occurrence area through the frame difference method; The specific method of step S4 is as follows: S4.
1. Grayscale the event frame and apply the median filtering method to remove high-frequency noise; S4.
2. Use the Sobel operator to calculate the gradient of the image; S4.
3. Use non-maximum suppression to remove non-edge pixels and obtain a refined edge; S4.
4. Set double thresholds to achieve edge connection and output the result.
3. According to the belt coal falling detection method for a resource-constrained device described in claim 1, it is characterized in that The specific method of step S2.1 is as follows: First, convert the collected image to a standard size of 640×640 to avoid affecting the spatial region positioning result due to different sizes of the collected images; Then perform color channel normalization to normalize the collected RGB image to [-1, 1]; Finally, adopt data augmentation technology to perform rotation and color enhancement operations on some of the collected images to improve the generalization ability of the model.
4. A belt coal dropping detection method for a resource-constrained device according to claim 1, characterized in that, The specific method of step S2.2 is as follows: 1) RepNeXtBlock adopts the classic depthwise separable convolution to reduce the computational amount and make it applicable to resource-constrained edge devices; in the standard 3×3 convolution, the calculation for each pixel point needs to consider all channels in the surrounding 3×3 positions, and its computational amount is expressed as: ; Among them, , is used to represent the height and width of the feature map, represents the number of input channels, represents the number of output channels, represents the size of the convolutional kernel; and the computational cost of depthwise separable convolution is expressed as: ; Among them represents the number of channels obtained by dimensionality reduction through 1×1 convolution; 2) ReLU is used to introduce non-linear features to enable the neural network to learn more complex patterns, and its mathematical definition is expressed as: ; The specific method of step S2.3 is as follows: 1) RepNeXt introduces group convolution to reduce the number of parameters and improve the feature expression ability; the calculation operation of the common standard two-dimensional convolution is expressed as: ; Among them is the input feature map, with a size of ; is the convolutional kernel, with a size of ; is the output feature map, with a size of ; and are the number of input and output channels respectively; In group convolution, the input channels and output channels are divided into G groups, and each group is calculated independently. Its calculation operation is expressed as: ; where G represents the number of groups, and represent the input and convolutional kernel of the th group; 2) In the training stage, multiple paths are adopted to enhance the feature expression ability, and the calculation in this stage is expressed as: ; Among them is a 3×3 group convolution used to enhance the local receptive field; is a 1×1 ordinary convolution used to enhance the feature expression ability; is the identity mapping used to improve the information flow; in the inference stage, the weights of these paths are merged to make the whole an equivalent 3×3 convolution, reducing the computational amount and improving the inference speed, and its computational representation is: ; Among them means expanding the weights of the 1×1 convolution to a 3×3 size so that it can be added to addable; represents an identity matrix, which is added to the weights of the convolution kernel; in this way, the computational cost of the inference operation is expressed as: 。 5. The belt coal falling detection method for a resource-constrained device according to claim 1, characterized in that, The specific method of step S2.4 is as follows: Combine high-level features and low-level features, and use FPN for feature fusion; FPN includes three main parts, namely bottom-up, top-down, and lateral connection; Bottom-up is the process of feature extraction by the backbone. For a given input image , the backbone network RepNeXt completes the extraction of feature maps at multiple levels, and the process is expressed as: ; Among them represents the process of feature extraction by the Backbone; represents the shallowest features, mainly containing texture information; represents the deepest features, mainly containing semantic information; The top-down uses deconvolution for upsampling to restore high-resolution features; For the feature maps obtained in the bottom-up process , the top-down processing process is expressed as: ; Among them denotes upsampling the feature map of the previous layer; denotes using a 1×1 convolution to reduce the dimension of the underlying feature map at the same level, so that the number of channels is aligned with the upsampling result; The two are added to complete the lateral connection, enhancing the feature expression ability; Through the bottom-up and top-down processes, the output result can be obtained , and these output features contain both the spatial information of low-level features and the semantic information of high-level features, making them suitable for the target detection scenario; In step S2.5, the extracted features are input into the detection head for the classification and positioning of falling coal.
6. A belt coal dropping detection method for a resource-constrained device according to claim 2, characterized in that, The specific method of step S3.1 is as follows: 1) Align the front and rear frames; during the alignment of the front and rear frames, global alignment based on optical flow deflection and local alignment based on feature point matching need to be completed; During global alignment, the global transformation matrix is calculated based on optical flow, and matrix transformation is performed on the entire input image; the optical flow method aligns the front and rear frames by estimating the pixel-level motion between adjacent frames, and it is expressed as: ; Among them represents the optical flow field, which means the direction and speed of pixel motion in adjacent frames; are the optical flow components in the horizontal and vertical directions respectively; The basic steps of local alignment include key point detection, feature point matching, and calculation of the transformation matrix; the method of KNN + cross-check is used to complete feature point detection, that is, two nearest matching points are found for each key point, and matching tests are performed through their positioning, and the threshold is set to 0.
5. This method is expressed as: ; Among them and represent the positions of two adjacent points. When the ratio is less than 0.5, it is considered that the local alignment step has been completed; 2) The method of Min-Max normalization is adopted, which is expressed as: ; Among them represents the original pixel value, represents the maximum pixel value, represents the minimum pixel value, represents the pixel value after the normalization operation; this process compresses the pixel value between [0, 1], reduces the differences in factors such as brightness, and improves the contrast accuracy; The specific method of step S3.2 is as follows: The background modeling method is used to obtain the change region between the front and rear frames, so as to obtain the falling coal fault occurrence region. This process includes two parts: frame difference and Gaussian mixture model; frame difference is to use time series to establish a stable background model and detect foreground changes. This process is expressed as: ; Among them is the background model is the current frame; the Gaussian mixture model not only considers the differences between adjacent frames, but also establishes a long-term stable background model through this model. The Gaussian mixture model is expressed as: ; where represents the weight of the -th Gaussian distribution, represents the mean of the -th Gaussian distribution, represents the variance of the -th Gaussian distribution, represents the number of Gaussian mixture distributions.
7. A belt coal dropping detection method for a resource-constrained device according to claim 2, wherein The specific method of step S4.1 is as follows: 1) Grayscale conversion: Edge detection is calculated based on grayscale images. The image is grayscale converted, and the color image with three RGB channels is converted into a single-channel image. This operation is expressed as: ; Among them, R, G, and B respectively represent the pixel values of the red, green, and blue channels in the image; 2) Apply the median filtering method to remove high-frequency noise: The steps of median filtering include selecting the filter window size, local sorting, calculating the median, and replacing the pixel value; A -sized filtering window covers a pixel of the image, and the pixels in the neighborhood around this pixel are represented as . Using to represent the number of pixels in the neighborhood, the median filtering operation is expressed as: ; wherein represents the pixel value after denoising, represents the neighborhood of the current pixel point, represents selecting the middle value as the result after sorting the pixel values in the neighborhood; After the above operations, the image after median filtering processing can be obtained and continue to be processed; The specific method of step S4.2 is as follows: Use the Sobel operator to calculate the gradient of the image, and the process of calculating the gradient by the Sobel operator is expressed as: ; ; ; Among them represents the horizontal gradient represents the vertical gradient 8. A belt coal dropping detection method for a resource-constrained device according to claim 2, characterized in that The specific method of step S4.3 is as follows: The process of removing non-edge pixels is divided into two steps, namely calculating the gradient direction and retaining the local maximum; the process of calculating the gradient direction is expressed as: ; After obtaining the gradient direction compare the current pixel with adjacent pixels, retain the local maximum, suppress non-maxima, and complete non-maximum suppression.
9. A belt coal dropping detection method for a resource-constrained device according to claim 2, characterized in that, The specific method of step S4.4 is as follows: Set two thresholds, namely the high threshold and the low threshold . For pixel values greater than , mark them as strong edges. For pixels between and , only retain them if they contain strong edges, otherwise suppress them. Then continuously track the edge pixels. Retain them when they are connected to strong edges, otherwise delete them. After these steps, the output result can be obtained.
10. The belt coal falling detection method for a resource-constrained device according to claim 1, characterized in that, The specific method of step S5 is as follows: The spatial features are obtained through steps S2.1 and S2.5 The temporal features are obtained through steps S3.1 and S3.2 The edge features are obtained through steps S4.1 and S4.4 ; The spatial features , temporal features , and edge features are fused using weighted fusion, and this process is expressed as: ; Among them represents the fused features; the detection is performed on the trained model according to YOLOV12 with the backbone replaced by RepNeXtBlock, and this process is expressed as: ; Among them represents the result of coal falling detection.
Citation Information
Patent Citations
Coal blockage detection algorithm based on gray features and edge detection
CN115631191A
Multi-target tracking method and system for coal gangue and sundries
CN119850968A
Small target detection method based on multi-scale cavity fusion
CN119992390A