A Belt Coal-Falling Detection Method for Resource-Constrained Devices
By combining lightweight convolutional neural network and event camera, the problems of high computing resources, insufficient generalization capabilities and poor adaptability of edge object detection on resource-constrained devices are solved, and efficient and accurate detection of coal drop failures in mine conveyor belts is achieved.
Patent Information
- Application Number
- CN202510718015.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2045-05-30
AI Technical Summary
The existing edge target detection technology has problems such as high computing resource demand, insufficient generalization capability, poor adaptability and poor real-time performance in resource-constrained equipment, making it difficult to effectively detect coal-fall failures of mine conveyor belts.
Combining the lightweight convolutional neural network structure and traditional object detection methods, an event camera is used to collect asynchronous event data streams, and efficient detection of belt coal drops through feature extraction and fusion.
It improves the accuracy and robustness of coal-fall detection, adapts to complex environments, reduces the computing resource requirements, and is suitable for edge devices with resource-constrained.
Smart Images

Figure CN120259636B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer vision target detection, and particularly to a method for detecting coal falling on a belt of a resource-constrained device. Background Art
[0002] Target detection is one of the core tasks of computer vision. In recent years, with the rapid development of deep learning and the improvement of the scale of data sets, target detection technology has made remarkable progress. The combination of artificial intelligence and the Internet of Things has promoted the development of target detection technology from the cloud to edge devices, making it widely used in fields such as intelligent cameras and mine robots. Traditional coal mine monitoring relies on manual inspections with poor real-time performance and high costs, as well as cloud computing. With the rapid development of intelligent mines, the coal industry is accelerating its transformation towards automation, unmanned operation, and intelligence. Edge target detection technology has played an important role in coal mine safety monitoring, equipment inspections, etc., not only reducing the cost of manual monitoring but also significantly improving production safety and efficiency. Among them, the mine conveyor belt, as a key device for coal transportation, operates for a long time in a harsh environment and is prone to faults such as tearing, deviation, and coal falling, which affect production efficiency and even lead to safety accidents. The traditional manual inspection method has low efficiency and poor real-time performance, making it difficult to detect and handle problems in a timely manner. Therefore, edge target detection technology with high efficiency and real-time performance shows great research and application potential in the field of intelligent mine belt fault detection, providing strong support for the intelligent development of the coal industry.
[0003] Currently, the existing edge target detection technologies are mainly divided into the following categories: edge target detection based on deep learning, edge target detection based on traditional computer vision, edge target detection based on sensor data, and target detection based on edge collaborative computing.
[0004] The main defects of current similar methods are as follows:
[0005] 1. Edge target detection methods based on deep learning have high requirements for training data, limited generalization ability, and high energy consumption, making it difficult to perform target detection tasks on edge devices.
[0006] 2. Edge target detection methods based on traditional computer vision have poor adaptability, rely on manually set parameters, and have insufficient generalization ability, making it difficult to perform target detection tasks on edge devices.
[0007] 3. Edge target detection methods based on sensor data have poor adaptability to different environments, large sensor data noise, are easily affected by environmental interference, and have poor real-time performance in performing target detection tasks.
[0008] 4. Edge target detection methods based on edge system computing are limited by differences in bandwidth, power consumption, etc. of different devices, are difficult to schedule, and have great difficulty in performing target detection tasks. Summary of the Invention
[0009] Aiming at the problems existing in the above-mentioned prior art, the present invention provides a method for detecting coal falling on a belt of a resource-constrained device, which can combine a lightweight convolutional neural network structure with a traditional object detection method, reduce the demand for computing resources for the object detection task, improve the accuracy and robustness of coal falling detection, have strong adaptability to complex environments, and have strong generalization ability. It is applicable to the detection of coal falling on resource-constrained edge devices and has high application value.
[0010] A method for detecting coal falling on a belt of a resource-constrained device according to the present invention includes steps S1, S2, S3, S4, and S5, wherein steps S2, S3, and S4 are carried out synchronously, and the specific method is as follows:
[0011] S1. Event data collection, the specific method is as follows:
[0012] S1.1. Configure the basic parameters of the event camera for focusing on the change of the belt surface area;
[0013] S1.2. Use the event vision sensor to obtain an asynchronous event data stream;
[0014] S1.3. Read the event stream in real time and aggregate it into event frames according to the set time window;
[0015] S1.4. Cache the aggregated event frames into the processing queue for subsequent feature extraction modules to call;
[0016] S2. Event feature preprocessing and construction, the specific method is as follows:
[0017] S2.1. Preprocess the event frames, including size normalization and channel standardization;
[0018] S2.2. Use the lightweight RepNeXtBlock network structure to extract low-level spatial features and capture local structure information and edge contours;
[0019] S2.3. Stack multiple layers of RepNeXtBlock deep modules to extract high-level semantic features, such as information about the position of the coal falling area, etc.;
[0020] S2.4. Use FPN to fuse multi-scale features, enhance the detection effect of small target coal falling, and improve the semantic expression ability of features;
[0021] S2.5. Input the extracted features into the detection head for classification and positioning of coal falling;
[0022] S3. Compare the event frames before and after, and construct spatio-temporal feature fusion, the specific method is as follows:
[0023] S3.1. Align the adjacent two frames of event diagrams spatially and perform pixel normalization;
[0024] S3.2. Calculate the event change region between the front and back frames, and highlight the region where the coal falling fault occurs through the frame difference method;
[0025] S4. Based on the event edge structure, perform fine recognition. The specific method is as follows:
[0026] S4.1. Grayscale the event frame and apply methods such as median filtering to remove high-frequency noise;
[0027] S4.2. Use the Sobel operator to calculate the gradient of the image;
[0028] S4.3. Use non-maximum suppression to remove non-edge pixels and obtain a refined edge;
[0029] S4.4. Set double thresholds to achieve edge connection and output the result;
[0030] S5. Perform spatio-temporal-edge feature fusion and output the coal falling detection result.
[0031] Furthermore, in step S1.3, OpenCV is used to process the video stream, and the camera data is read through cv2.VideoCapture. The interval is set to 1 second, and an image is saved every 1 second.
[0032] Furthermore, the specific method of step S2.1 is as follows:
[0033] First, convert the collected image to a standard size of 640×640 to avoid affecting the spatial region positioning result due to different sizes of the collected images;
[0034] Then perform color channel normalization to normalize the collected RGB image to [-1, 1];
[0035] Finally, adopt data augmentation technology to perform operations of rotation and color enhancement on some of the collected images to improve the generalization ability of the model.
[0036] Furthermore, the specific method of step S2.2 is as follows:
[0037] 1) RepNeXtBlock adopts the classic depthwise separable convolution to reduce the computational complexity, making it suitable for resource-constrained edge devices. In the standard 3×3 convolution, each pixel calculation needs to consider all channels in the surrounding 3×3 positions, which will lead to a large computational complexity. Its computational complexity is expressed as:
[0038] ;
[0039] Among them, , For representing the height and width of the feature map, represents the number of input channels, represents the number of output channels, represents the size of the convolutional kernel. The computational cost of depthwise separable convolution is expressed as:
[0040] ;
[0041] where represents the number of channels obtained by dimension reduction through 1×1 convolution. It can be easily obtained from the above that depthwise separable convolution can effectively reduce the computational cost and is suitable for edge devices.
[0042] 2) ReLU is used to introduce non-linear features, enabling the neural network to learn more complex patterns. Its mathematical definition is expressed as:
[0043] ;
[0044] This non-linear activation function enables the network to learn complex non-linear patterns such as image features, allowing the deep network to learn different feature representations, such as learning edge features at the low layer, complex shapes at the middle layer, and semantic features at the high layer.
[0045] Furthermore, the specific method of the step S2.3 is as follows:
[0046] 1) RepNeXt introduces grouped convolution to reduce the number of parameters and enhance the feature representation ability. The computational operation of the common standard two-dimensional convolution is expressed as:
[0047] ;
[0048] where is the input feature map with a size of ; is the convolutional kernel with a size of ; is the output feature map with a size of ; and are the number of input and output channels respectively; such a traditional two-dimensional convolution operation has a high computational cost when the number of input and output channels is large and is not suitable for edge devices with limited computational resources.
[0049] In grouped convolution, the input channels and the output channels are divided into G groups and each group is calculated independently. Its computational operation is expressed as:
[0050] ;
[0051] Among them, G represents the number of groups, and represents the input and convolution kernel of the th group. Such group convolution reduces the computational load and is suitable for efficient deep network design and is more suitable for edge devices with limited computing resources.
[0052] 2) Adopt the method of structural reparameterization to enhance the feature expression ability of multiple parallel paths during training and simplify it to a single 3×3 convolution during inference to improve the computational efficiency; during the training stage, use multiple paths to enhance the feature expression ability, and the calculation in this stage is expressed as:
[0053] ;
[0054] Among them is a 3×3 group convolution used to enhance the local receptive field; is a 1×1 ordinary convolution used to enhance the feature expression ability; is an identity mapping used to improve the information flow. During the inference stage, the weights of these paths are merged so that the whole becomes an equivalent 3×3 convolution, reducing the computational load and improving the inference speed, and its calculation is expressed as:
[0055] ;
[0056] Among them represents expanding the weight of the 1×1 convolution to the 3×3 size so that it can be added to ; represents an identity matrix, which is added to the weight of the convolution kernel. In this way, the computational load of the inference operation is expressed as:
[0057] ;
[0058] Obviously, the computational load is reduced and the inference process is accelerated.
[0059] Furthermore, the specific method of the step S2.4 is as follows:
[0060] Combine the high-level features and low-level features and use FPN for feature fusion. According to steps 2.2 and 2.3, the low-level features from the shallow network and the high-level features from the deep network can already be obtained. The low-level features have high resolution, retain strong spatial details and weak semantic features, and the high-level features have strong semantic information, but during continuous downsampling operations, the spatial resolution is low and the spatial details are insufficient. Therefore, use FPN for feature fusion. FPN consists of three main parts, namely bottom-up, top-down, and lateral connections.
[0061] Bottom-up is the process of feature extraction by the Backbone. For a given input image , through the backbone network RepNeXt, feature maps of multiple levels are extracted, and the process is expressed as:
[0062] ;
[0063] Among them represents the process of feature extraction by the Backbone; represents the shallowest feature, with a small receptive field, mainly containing texture information; represents the deepest feature, with a large receptive field, mainly containing semantic information.
[0064] Top-down then uses transposed convolution for upsampling to restore high-resolution features. For the feature maps obtained in the bottom-up process , the top-down processing process is expressed as:
[0065] ;
[0066] Among them represents upsampling the previous layer of feature maps; represents using a 1×1 convolution to reduce the dimension of the bottom-layer feature map at the same level, so that the number of channels is aligned with the upsampling result. The two are added to complete the lateral connection, enhancing the feature expression ability.
[0067] After the bottom-up and top-down processes, the output result can be obtained. These output features contain both the spatial information of low-level features and the semantic information of high-level features, and are suitable for the target detection scenario.
[0068] Furthermore, in step S2.5, the extracted features are input into the detection head for the classification and localization of coal falling.
[0069] Furthermore, the specific method of step S3.1 is as follows:
[0070] 1) Align the front and rear frames. Since factors such as camera jitter affect the frame alignment when collecting images according to the video stream, directly performing feature extraction in this case is likely to cause comparison errors, thus affecting the detection result. Therefore, the front and rear frames are first aligned. During this process, global alignment based on optical flow and local alignment based on feature point matching need to be completed.
[0071] During global alignment, the global transformation matrix is calculated based on optical flow, and matrix transformation is performed on the entire input image. The optical flow method aligns the front and rear frames by estimating the pixel-level motion between adjacent frames, and it is expressed as:
[0072] ;
[0073] wherein represents the optical flow field, which means the direction and speed of pixel movement in adjacent frames; are the optical flow components in the horizontal and vertical directions respectively.
[0074] The basic steps of local alignment include key point detection, feature point matching, and calculation of the transformation matrix. The present invention uses the KNN + cross-check method to complete key point detection, that is, finding two nearest matching points for each key point, performing a matching test through their positioning, and setting the threshold to 0.5. Then this method is expressed as:
[0075] ;
[0076] wherein , represent the positions of two adjacent points. When their ratio is less than 0.5, it can be considered that this step of local alignment has been completed.
[0077] 2) The purpose of normalization is to eliminate the ability of different illuminations and different contrasts to extract features from adjacent frames and improve the generalization ability of the model. This step is very important in the field of coal industry target detection with complex and changeable environments. The present invention uses the Min-Max normalization method, which is expressed as:
[0078] ;
[0079] wherein represents the original pixel value, represents the maximum pixel value, represents the minimum pixel value, represents the pixel value after the normalization operation. This process compresses the pixel value to the range of [0, 1], reduces the differences in factors such as brightness, and improves the contrast accuracy.
[0080] Furthermore, the specific method of step S3.2 is as follows:
[0081] The background modeling method is used to obtain the change region between the front and rear frames, so as to obtain the region where the coal falling fault occurs. This process includes two parts: frame difference and Gaussian mixture model. Frame difference is to establish a stable background model using time series and detect foreground changes. This process is expressed as:
[0082] ;
[0083] wherein is the background model, is the current frame. The Gaussian mixture model not only considers the differences between adjacent frames, but also establishes a long-term stable background model through this model. The Gaussian mixture model is expressed as:
[0084] ;
[0085] where represents the weight of the -th Gaussian distribution, represents the mean of the -th Gaussian distribution, represents the variance of the -th Gaussian distribution, represents the number of Gaussian mixtures.
[0086] Furthermore, the specific method of step S4.1 is as follows:
[0087] 1) Grayscale conversion:
[0088] Edge detection is calculated based on grayscale images. The image is grayscale-converted, and the color image with RGB three channels is converted into a single-channel image. This operation is expressed as:
[0089] ;
[0090] where R, G, and B respectively represent the pixel values of the red, green, and blue channels in the image.
[0091] 2) Apply the median filtering method to remove high-frequency noise:
[0092] The basic operation of median filtering is that for each image pixel, the pixel values in its surrounding neighborhood are sorted, and the value of this pixel is replaced with the middle value after sorting. This operation can effectively remove salt-and-pepper noise and will not blur the edge information in the image. The steps of median filtering include selecting the filter window size, local sorting, calculating the median, and replacing the pixel value.
[0093] A -sized filter window covers a pixel point in the image. The pixels in the neighborhood of this pixel point are represented as , and using to represent the number of pixels in the neighborhood, then the median filtering operation is expressed as:
[0094] ;
[0095] where represents the pixel value after denoising, represents the neighborhood of the current pixel point, represents selecting the middle value as the result after sorting the pixel values in the neighborhood.
[0096] After the above operations, the image after median filtering processing can be obtained and continue to be processed.
[0097] Furthermore, the specific method of step S4.2 is as follows:
[0098] Calculate the gradient of the image using the Sobel operator. The process of calculating the gradient using the Sobel operator is expressed as:
[0099] ;
[0100] ;
[0101] ;
[0102] where represents the horizontal direction gradient, represents the vertical direction gradient.
[0103] Furthermore, the specific method of step S4.3 is as follows:
[0104] The process of removing non-edge pixels can be divided into two steps, namely calculating the gradient direction and retaining the local maximum. The process of calculating the gradient direction is expressed as:
[0105] ;
[0106] After obtaining the gradient direction , compare the current pixel with adjacent pixels, retain the local maximum, and suppress non-maximum values to complete non-maximum suppression.
[0107] Furthermore, the specific method of step S4.4 is as follows:
[0108] To reduce the influence of noise, set two thresholds, namely the high threshold and the low threshold . For pixel values greater than , mark them as strong edges. For pixels between and , only retain them if they contain strong edges, otherwise suppress them. Then continuously track the weak edge pixels. Retain them if they are connected to strong edges, otherwise delete them. After these steps, the output result can be obtained.
[0109] Furthermore, the specific method of step S5 is as follows:
[0110] The spatial feature is obtained through step S2.1 and step S2.5, the temporal feature is obtained through step S3.1 and step S3.2, and the edge feature is obtained through step S4.1 and step S4.4; use the weighted fusion method to fuse the spatial feature , the temporal feature , Edge features Feature fusion is performed, and this process is expressed as:
[0111] ;
[0112] where represents the fused feature. Detection is performed on the trained model according to YOLOV12 with the backbone replaced by RepNeXtBlock, and this process is expressed as:
[0113] ;
[0114] where represents the result of coal dropping detection.
[0115] The beneficial effects of the present invention are as follows: Compared with the prior art, the present invention introduces an event camera as the core of visual perception and utilizes its characteristic of being highly sensitive to dynamic changes. Without the need for traditional inter-frame difference, high-precision spatio-temporal change information can be directly obtained, thereby significantly improving the detection effect of the coal dropping fault area. The event camera has high temporal resolution and high dynamic range. In the face of complex working conditions such as strong light interference, low illuminance, and high-speed movement, it can maintain stable and clear edge contour information, effectively reducing the common motion blur and exposure distortion problems in traditional image processing methods, and further improving the robustness and detection accuracy of the system. In addition, the present invention combines the sparsity characteristic of event data and uses a lightweight RepNeXt backbone network for feature extraction and target recognition, significantly reducing the consumption of computing resources and improving the real-time processing ability on edge devices. Through the event-driven data structure, the present invention naturally has the fusion characteristics of space, time, and edge information, avoiding the delay and redundancy caused by multi-module splicing, and ensuring that the detection system still has excellent response speed and stable performance under resource-constrained conditions. In summary, this method gives full play to the dynamic perception advantage of event vision, shows higher robustness, real-time performance, and engineering adaptability in the complex and changeable belt coal dropping detection scenario, and has good application and promotion value. BRIEF DESCRIPTION OF THE DRAWINGS
[0116] Figure 1 is a flowchart of a method for detecting coal dropping on a belt of a resource-constrained device according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0117] The following will elaborate on the technical content of its implementation plan in detail in combination with the specific drawings of the present invention. It should be noted that the embodiments listed herein are only partial examples of the present invention and do not exhaust all possibilities. Based on the technical solutions disclosed in the present invention, any other implementation solutions obtained by any person skilled in the art without creative efforts shall fall within the protection scope of the present invention. Embodiment 1
[0118] As Figure 1 shown, a belt coal dropping detection method for resource-constrained devices includes steps S1, S2, S3, S4, and S5, where steps S2, S3, and S4 are performed synchronously. The specific method is as follows:
[0119] S1. Event data collection, and the specific method is as follows:
[0120] 1.1 Configure the basic parameters of the event camera for focusing on changes in the belt surface area;
[0121] 1.2 Use the event vision sensor to obtain an asynchronous event data stream;
[0122] 1.3 Read the event stream in real time and aggregate it into event frames according to the set time window;
[0123] 1.4 Cache the aggregated event frames into the processing queue for subsequent feature extraction modules to call.
[0124] S2. Use an improved lightweight model for feature extraction, and the specific method is as follows:
[0125] 2.1 Preprocess the event frames, including size normalization and channel standardization. This step involves the following specific implementation details:
[0126] First, convert the collected images to the standard size of 640×640 to avoid affecting the spatial region positioning results due to different sizes of the collected images;
[0127] Then perform color channel normalization to normalize the collected RGB images to [-1, 1];
[0128] Finally, use data augmentation techniques to perform operations such as rotation and color enhancement on some of the collected images to improve the generalization ability of the model.
[0129] 2.2 Use the lightweight RepNeXtBlock network structure to extract low-level spatial features, capture local structure information and edge contours. A single RepNeXtBlock can be regarded as consisting of three parts, namely 1×1 convolution, 3×3 grouped convolution, and ReLU activation;
[0130] RepNeXtBlock uses the classic depthwise separable convolution to reduce the computational complexity, making it suitable for resource-constrained edge devices. In the standard 3×3 convolution, each pixel calculation needs to consider all channels in the surrounding 3×3 positions, resulting in a large computational complexity. Its computational complexity is expressed as:
[0131] ;
[0132] Among them, , is used to represent the height and width of the feature map, represents the number of input channels, represents the number of output channels, represents the size of the convolutional kernel. The computational cost of depthwise separable convolution is expressed as:
[0133] ;
[0134] Among them represents the number of channels obtained after dimensionality reduction by 1×1 convolution. It can be easily obtained from the above that depthwise separable convolution can effectively reduce the computational cost and is suitable for edge devices.
[0135] ReLU is used to introduce non-linear features, enabling the neural network to learn more complex patterns. Its mathematical definition is expressed as:
[0136] ;
[0137] This non-linear activation function enables the network to learn complex non-linear patterns such as image features, allowing the deep network to learn different feature representations, such as learning edge features at the low layer, complex shapes at the middle layer, and semantic features at the high layer.
[0138] 2.3 Stack multiple layers of RepNeXtBlock deep modules to extract high-level semantic features, such as information about the position of the coal falling area, etc.;
[0139] RepNeXt introduces group convolution to reduce the number of parameters and improve the feature expression ability. The common standard computational operation is expressed as:
[0140] ;
[0141] Among them is the input feature map, with a size of ; is the convolutional kernel, with a size of ; is the output feature map, with a size of ; and are the number of input and output channels respectively. Such traditional two-dimensional convolution operations have a high computational cost when the number of input and output channels is large and are not suitable for edge devices with limited computational resources.
[0142] In group convolution, the input channels and the output channels are divided into G groups, and each group is calculated independently. Its computational operation is expressed as:
[0143] ;
[0144] Among them, G represents the number of groups, and represent the input and convolution kernel of the th group. Such group convolution reduces the computational amount and is suitable for efficient deep network design, and is more suitable for edge devices with limited computing resources.
[0145] This embodiment also adopts the method of structural reparameterization to enhance the feature expression ability of multiple parallel paths during training and simplify them into a single 3×3 convolution during inference, improving the computing efficiency. During the training phase, multiple paths are used to enhance the feature expression ability, and the calculation in this phase is expressed as:
[0146] ;
[0147] Among them is a 3×3 group convolution used to enhance the local receptive field; is a 1×1 ordinary convolution used to enhance the feature expression ability; is an identity mapping used to improve the information flow. During the inference phase, the weights of these paths are merged so that the whole becomes an equivalent 3×3 convolution, reducing the computational amount and improving the inference speed, and its calculation is expressed as:
[0148] ;
[0149] Among them represents expanding the weight of the 1×1 convolution to the 3×3 size so that it can be added to ; represents an identity matrix, which is added to the weight of the convolution kernel. In this way, the computational amount of the inference operation is expressed as:
[0150] ;
[0151] Obviously, the computational amount is reduced and the inference process is accelerated.
[0152] 2.4 Combine high-level features and low-level features and use FPN for feature fusion. After steps 2.2 and 2.3 of this embodiment, low-level features from the shallow network and high-level features from the deep network can already be obtained. The low-level features have high resolution, retain strong spatial details and weak semantic features, and the high-level features have strong semantic information, but during continuous downsampling operations, the spatial resolution is low and the spatial details are insufficient. Therefore, FPN is used for feature fusion. FPN consists of three main parts, namely bottom-up, top-down, and lateral connections.
[0153] The bottom-up is the process of Backbone extracting features. For a given input image After the backbone network RepNeXt completes the extraction of feature maps at multiple levels, the process is expressed as:
[0154] ;
[0155] Among them represents the process of the Backbone extracting features; represents the shallowest features, which have a smaller receptive field and mainly contain texture information; represents the deepest features, which have a larger receptive field and mainly contain semantic information.
[0156] In the top-down direction, deconvolution is used for upsampling to restore high-resolution features. For the feature maps obtained in the bottom-up process , the top-down processing process is expressed as:
[0157] ;
[0158] Among them represents upsampling the feature map of the previous layer; represents using a 1×1 convolution to reduce the dimension of the underlying feature map at the same level, so that the number of channels is aligned with the upsampling result. The addition of the two completes the lateral connection and enhances the feature expression ability.
[0159] After the bottom-up and top-down processes, the output result can be obtained. These output features contain both the spatial information of low-level features and the semantic information of high-level features, and are suitable for the target detection scenario.
[0160] 2.5 Input the extracted features into the detection head for the classification and localization of coal falling.
[0161] The detection head in this embodiment adopts the Anchor-Free mechanism of the FCOS structure. The detection head consists of three parts, namely a classification head for predicting the class probability of each pixel point, a regression head for directly regressing the offset of the target box, and a centerness branch for measuring whether a certain pixel point is the center of the target to improve the detection stability.
[0162] In addition, under the Anchor-Free mechanism of the FCOS structure, the loss function consists of three parts, namely classification loss, regression loss, and centerness loss. The classification loss can solve the problem of class imbalance and reduce the interference of negative samples to training. The classification loss is expressed as:
[0163] ;
[0164] Among them is the probability of the predicted class, and It is used to control the weights of easy and hard samples; the regression loss is used to optimize object localization, and the regression loss is expressed as:
[0165] ;
[0166] Where is the ground truth bounding box, represents the predicted bounding box, represents the intersection over union (IoU) between the two, which can measure the accuracy of the prediction. This process directly optimizes the IoU and improves the localization accuracy. The centerness loss is used to measure the reliability of the predicted bounding box and prevent some boxes from being misjudged due to excessive deviation. The centerness loss is expressed as:
[0167] ;
[0168] Wherein, represents the ground truth centerness, and
[0169] S3. The spatio-temporal feature map combines the edge information to output the final result. The specific method is as follows
[0170] 3.1 Align and normalize the front and rear frames. The specific method of this operation is as follows:
[0171] Align the front and rear frames. Since the frames are misaligned due to factors such as camera jitter when collecting images according to the video stream, directly performing feature extraction in this case is likely to lead to comparison errors, thus affecting the detection results. Therefore, the front and rear frames are aligned first. In this process, global alignment based on optical flow and local alignment based on feature point matching need to be completed.
[0172] For global alignment, the global transformation matrix is calculated based on optical flow, and the entire input image is subjected to matrix transformation. The optical flow method aligns the front and rear frames by estimating the pixel-level motion between adjacent frames, and it is expressed as:
[0173] ;
[0174] Where represents the optical flow field, which means the direction and speed of pixel motion in adjacent frames;
[0175] The basic steps of local alignment include key point detection, feature point matching, and calculation of the transformation matrix. The present invention uses the KNN + cross-check method to complete key point detection, that is, two nearest matching points are found for each key point, and a matching test is performed through its localization. The threshold is set to 0.5, then this method is expressed as:
[0176] ;
[0177] wherein and represent the positions of two adjacent points. When their ratio is less than 0.5, it can be considered that the local alignment step has been completed.
[0178] The purpose of normalization is to eliminate the ability of different illuminations and different contrasts to extract features from adjacent frames and improve the generalization ability of the model. This step is very important in the field of coal industry target detection where the environment is complex and changeable. The present invention adopts the Min-Max normalization method, which is expressed as:
[0179] ;
[0180] wherein represents the original pixel value, represents the maximum pixel value, represents the minimum pixel value, represents the pixel value after the normalization operation. This process compresses the pixel value between [0, 1], reduces the differences in factors such as brightness, and improves the contrast accuracy.
[0181] 3.2 Calculate the event change region between the front and rear frames, and highlight the region where the coal falling fault occurs through the frame difference method.
[0182] The present invention adopts the background modeling method to obtain the change region between the front and rear frames, thereby obtaining the region where the coal falling fault occurs. This process includes two parts: inter-frame difference and Gaussian mixture model. Inter-frame difference is to establish a stable background model using the time series and detect the foreground change. This process is expressed as:
[0183] ;
[0184] wherein is the background model, is the current frame. The Gaussian mixture model not only considers the differences between adjacent frames, but also establishes a long-term stable background model through this model. The Gaussian mixture model is expressed as:
[0185] ;
[0186] wherein represents the weight of the th Gaussian distribution, represents the mean of the th Gaussian distribution, represents the variance of the th Gaussian distribution, represents the number of Gaussian distributions in the mixture.
[0187] S4. Perform fine recognition based on the event edge structure. The specific method is as follows:
[0188] 4.1 Grayscale the event frames and apply the median filtering method to remove high-frequency noise. The specific method is as follows:
[0189] Grayscale conversion
[0190] The edge detection in this embodiment is based on the calculation of grayscale images. The image is grayscaled, and the color image with three RGB channels is converted into a single-channel image. This operation is expressed as:
[0191] ;
[0192] where R, G, and B represent the pixel values of the red, green, and blue channels in the image, respectively.
[0193] Apply the median filtering method to remove high-frequency noise
[0194] The basic operation of median filtering is that for each pixel point in the image, the pixel values in its surrounding neighborhood are sorted, and the value of this pixel is replaced with the median value after sorting. This operation can effectively remove salt-and-pepper noise and will not blur the edge information in the image. The steps of median filtering include selecting the filter window size, local sorting, calculating the median, and replacing the pixel value.
[0195] A sized filter window covers a pixel point in the image. The pixels in the neighborhood around this pixel point are represented as , using to represent the number of pixels in the neighborhood. Then the median filtering operation is expressed as:
[0196] ;
[0197] where represents the denoised pixel value, represents the neighborhood of the current pixel point, represents selecting the median value as the result after sorting the pixel values in the neighborhood.
[0198] After the above operations, an image processed by median filtering can be obtained and further processed.
[0199] 4.2 Use the Sobel operator to calculate the gradient of the image.
[0200] The process of calculating the gradient using the Sobel operator is expressed as:
[0201] ;
[0202] ;
[0203] ;
[0204] Among them represents the horizontal gradient, represents the vertical gradient.
[0205] 4.3 Use non-maximum suppression to remove non-edge pixels and obtain a refined edge.
[0206] The process of removing non-edge pixels can be divided into two steps, namely calculating the gradient direction and retaining the local maximum. The process of calculating the gradient direction is expressed as:
[0207] ;
[0208] After obtaining the gradient direction , compare the current pixel with adjacent pixels, retain the local maximum, suppress non-maxima, and complete non-maximum suppression.
[0209] 4.4 Set double thresholds to achieve edge connection and output the result.
[0210] To reduce the influence of noise, the present invention sets double thresholds, namely a high threshold and a low threshold . For pixel values greater than , mark them as strong edges. For pixels between and , only retain them if they contain strong edges, otherwise suppress them. Then continuously track the weak edge pixels. Retain them when they are connected to strong edges, otherwise delete them. Through these steps, the output result can be obtained.
[0211] S5, Spatial-Temporal-Edge Feature Fusion, Output the Coal-Falling Detection Result
[0212] Through S2, spatial features can be obtained, through S3, temporal features can be obtained, and through S4, edge features can be obtained. In this embodiment, weighted fusion is used for feature fusion, and this process is expressed as:
[0213] ;
[0214] Among them represents the fused features. Detect according to the YOLOV12 model trained with the backbone replaced by RepNeXtBlock, and this process is expressed as:
[0215] ;
[0216] Among them represents the result of coal-falling detection.
[0217] Obviously, those skilled in the art can make various modifications and variations to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention also intends to include these modifications and variations.
Claims
1. A method for detecting coal falling on a belt of a resource-constrained device, characterized in that, It includes steps S1, S2, S3, S4, and S5, where steps S2, S3, and S4 are carried out synchronously. The specific methods are as follows: S1. Event data acquisition. The specific methods are as follows: S1.
1. Configure the basic parameters of the event camera for focusing on the change area of the belt surface; S1.
2. Use the event vision sensor to obtain an asynchronous event data stream; S1.
3. Read the event stream in real time, process the video stream using OpenCV, and aggregate it into event frames according to the set time window; S1.
4. Cache the aggregated event frames into the processing queue for subsequent calls by the feature extraction module; S2. Event feature preprocessing and construction. The specific methods are as follows: S2.
1. Preprocess the event frames, including size normalization and channel standardization; S2.
2. Use the lightweight RepNeXtBlock network structure to extract low-level spatial features and capture local structure information and edge contours; S2.
3. Stack multiple layers of RepNeXtBlock deep modules to extract high-level semantic features; S2.
4. Use FPN to fuse multi-scale features, enhance the detection effect of small target coal falling, and improve the semantic expression ability of features; S2.
5. Input the extracted features into the detection head for classification and localization of coal falling; S3. Compare the event frames before and after, and construct spatio-temporal feature fusion. The specific methods are as follows: S3.
1. Perform spatial alignment and pixel normalization on two adjacent event images; S3.
2. Calculate the event change area between the front and back frames, and highlight the area where the coal falling fault occurs through the frame difference method; S4. Perform fine recognition based on the event edge structure. The specific methods are as follows: S4.
1. Grayscale the event frames and apply the median filtering method to remove high-frequency noise; S4.
2. Use the Sobel operator to calculate the gradient of the image; S4.
3. Use non-maximum suppression to remove non-edge pixels and obtain a refined edge; S4.
4. Set double thresholds to achieve edge connection and output the results; S5. Spatial-temporal-edge feature fusion to output the coal falling detection result. The specific methods are as follows: The spatial features are obtained through steps S2.1 and S2.5 The temporal features are obtained through steps S3.1 and S3.2 The edge features are obtained through steps S4.1 and S4.4 ; The spatial features , temporal features , and edge features are fused using weighted fusion, and this process is expressed as: ; Among them represents the fused features; the detection is performed on the trained model according to YOLOV12 with the backbone replaced by RepNeXtBlock, and this process is expressed as: ; Among them represents the result of coal falling detection.
2. The belt coal falling detection method for a resource-constrained device according to claim 1, characterized in that The specific method of step S2.1 is as follows: First, convert the collected image to a standard size of 640×640 to avoid affecting the spatial region positioning result due to different sizes of the collected images; Then, perform color channel normalization to normalize the collected RGB image to [-1,1]; Finally, adopt data augmentation techniques to perform rotation and color enhancement operations on some of the collected images to improve the generalization ability of the model.
3. The belt coal falling detection method for a resource-constrained device according to claim 1, characterized in that, The specific method of step S2.2 is as follows: RepNeXtBlock uses classical depthwise separable convolutions to reduce the computational amount, making it suitable for edge devices with limited resources; in a standard 3×3 convolution, the calculation for each pixel needs to consider all channels in the surrounding 3×3 positions, and its computational amount is expressed as: ; Among them, , is used to represent the height and width of the feature map, represents the number of input channels, represents the number of output channels, represents the size of the convolutional kernel; and the computational cost of depthwise separable convolution is expressed as: ; Among them represents the number of channels obtained by dimensionality reduction through 1×1 convolution; ReLU is used to introduce non-linear features, enabling the neural network to learn more complex patterns, and its mathematical definition is expressed as: ; The specific method of step S2.3 is as follows: 1) RepNeXt introduces group convolution to reduce the number of parameters and enhance the feature representation ability; the calculation operation of the common standard two-dimensional convolution is expressed as: ; Among them is the input feature map with a size of ; is the convolution kernel with a size of ; is the output feature map with a size of ; and are the number of input and output channels respectively; In group convolution, the input channels and output channels are divided into G groups, and each group is calculated independently. Its calculation operation is expressed as: ; where G represents the number of groups, and represent the input and convolution kernel of the th group; 2) In the training stage, multiple paths are adopted to enhance the feature representation ability, and the calculation in this stage is expressed as: ; Among them is a 3×3 group convolution used to enhance the local receptive field; is a 1×1 ordinary convolution used to enhance the feature expression ability; is an identity mapping used to improve the information flow; in the inference stage, the weights of these paths are merged to make the whole an equivalent 3×3 convolution, reducing the computational amount and improving the inference speed. Its computational representation is: ; Among them means expanding the weights of the 1×1 convolution to a 3×3 size so that it can be added to addable; represents an identity matrix, which is added to the weights of the convolution kernel; thus, the computational complexity of the inference operation is expressed as: 。 4. The belt coal falling detection method for a resource-constrained device according to claim 1, characterized in that, The specific method of the step S2.4 is as follows: Combine the high-level features and low-level features, and use FPN for feature fusion; FPN includes three main parts, namely bottom-up, top-down, and lateral connections; Bottom-up is the process of feature extraction by the backbone. For a given input image , the backbone network RepNeXt completes the extraction of feature maps at multiple levels, and its process is expressed as: ; Among them represents the process of feature extraction by the Backbone; represents the shallowest features, mainly containing texture information; represents the deepest features, mainly containing semantic information; The top-down uses deconvolution for upsampling to restore high-resolution features; For the feature map obtained in the bottom-up process , the top-down processing process is expressed as: ; Among them Indicates upsampling the upper-layer feature map; Indicates using 1×1 convolution to reduce the dimension of the underlying feature map at the same level, so that the number of channels is aligned with the upsampling result; The two are added to complete the lateral connection, enhancing the feature representation ability; Through the bottom-up and top-down processes, the output result can be obtained , and these output features contain both the spatial information of low-level features and the semantic information of high-level features, making them suitable for object detection scenarios; In the step S2.5, the extracted features are input into the detection head for the classification and positioning of coal dropping.
5. A method for detecting coal falling from a belt of a resource-constrained device according to claim 1, characterized in that, The specific method of the step S3.1 is as follows: Align the front and rear frames; during the alignment of the front and rear frames, global alignment based on optical flow and local alignment based on feature point matching need to be completed; During global alignment, calculate the global transformation matrix based on optical flow and perform matrix transformation on the entire input image; the optical flow method aligns the front and rear frames by estimating the pixel-level motion between adjacent frames, and it is expressed as: ; where represents the optical flow field, which means the direction and speed of pixel motion in adjacent frames; are the optical flow components in the horizontal and vertical directions respectively; The basic steps of local alignment include key point detection, feature point matching, and calculation of the transformation matrix; the method of KNN + cross-check is used to complete feature point detection, that is, find two nearest matching points at each key point, and perform a matching test through their positioning, and set the threshold to 0.
5. This method is expressed as: ; Among them and represent the positions of two adjacent points. When their ratio is less than 0.5, it is considered that the step of local alignment has been completed; Adopt the method of Min-Max normalization, which is expressed as: ; Among them represents the original pixel value, represents the maximum pixel value, represents the minimum pixel value, represents the pixel value after the normalization operation; this process compresses the pixel value to between [0, 1], reduces the difference in brightness factors, and improves the contrast accuracy; The specific method of the step S3.2 is as follows: Adopt the background modeling method to obtain the changed area between the front and rear frames, so as to obtain the area where the coal dropping fault occurs. This process includes two parts: frame difference and Gaussian mixture model; frame difference is to establish a stable background model using time series and detect foreground changes. This process is expressed as: ; wherein is the background model, is the current frame; the Gaussian mixture model not only considers the differences between adjacent frames, but also establishes a long-term stable background model through this model. The Gaussian mixture model is expressed as: ; wherein represents the weight of the th Gaussian distribution, represents the mean of the th Gaussian distribution, represents the variance of the th Gaussian distribution, represents the number of Gaussian mixture distributions.
6. A belt coal dropping detection method for a resource-constrained device according to claim 1, characterized in that The specific method of the step S4.1 is as follows: Grayscale conversion: Edge detection is calculated based on the grayscale image. The image is grayscale-converted, and the color image with three RGB channels is converted into a single-channel image. This operation is expressed as: ; Among them, R, G, and B respectively represent the pixel values of the red, green, and blue channels in the image; Apply the median filter method to remove high-frequency noise: The steps of median filtering include selecting the filter window size, local sorting, calculating the median, and replacing the pixel value; A -sized filtering window covers a pixel of the image, and the pixels in the neighborhood around this pixel are represented as . Using to represent the number of pixels in the neighborhood, the median filtering operation is expressed as: ; wherein represents the pixel value after denoising, represents the neighborhood of the current pixel point, represents selecting the middle value as the result after sorting the pixel values in the neighborhood; After the above operations, an image processed by median filtering can be obtained and continue to be processed; The specific method of the step S4.2 is as follows: Use the Sobel operator to calculate the gradient of the image. The process of calculating the gradient by the Sobel operator is expressed as: ; ; ; Among them represents the horizontal gradient, represents the vertical gradient.
7. A belt coal dropping detection method for a resource-constrained device according to claim 1, characterized in that, The specific method of the step S4.3 is as follows: The process of removing non-edge pixels is divided into two steps, namely calculating the gradient direction and retaining the local maximum; the process of calculating the gradient direction is expressed as: ; After obtaining the gradient direction then compare the current pixel with adjacent pixels, retain the local maximum, suppress non-maxima, and complete non-maximum suppression.
8. A belt coal dropping detection method for a resource-constrained device according to claim 1, characterized in that, The specific method of the step S4.4 is as follows: Set two thresholds, namely the high threshold and the low threshold . For pixel values greater than , mark them as strong edges. For pixels between and , only retain them if they contain strong edges, otherwise suppress them. Then continuously track the edge pixels. Retain them when they are connected to strong edges, otherwise delete them. After these steps, the output result is obtained.
Citation Information
Patent Citations
Coal blockage detection algorithm based on gray features and edge detection
CN115631191A
Multi-target tracking method and system for coal gangue and sundries
CN119850968A