Carrier roller rotation detection method based on video understanding
Through the roller rotation detection method based on video understanding, the directional enclosure box version of the yolov7 network model and the multi-channel fusion algorithm are used to solve the efficiency and accuracy of roller rotation state detection in the prior art, and efficient and safe roller rotation state monitoring is achieved.
Patent Information
- Application Number
- CN202510116978.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-27
AI Technical Summary
The existing roller rotation state detection technology has problems such as single equipment functions, high deployment cost, complex structure and high failure rate. The optical flow method has low detection efficiency, is susceptible to light interference, is low robustness, and has low personnel patrol efficiency and safety.
Using a roller rotation detection method based on video understanding, using machine vision and artificial intelligence technology, real-time detection of roller rotation state is achieved through the directional enclosure box version of the yolov7 network model and multi-channel fusion algorithm.
It improves the accuracy and efficiency of the roller rotation state detection, reduces hardware requirements, reduces the calculation amount and time, enhances the robustness and safety of the detection, and effectively prevents safety accidents caused by the roller not rotating.
Smart Images

Figure CN120047869A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of industrial conveyor equipment detection, and specifically, it is a method for detecting the rotation of idlers based on video understanding. Background Art
[0002] The belt conveyor is an important transportation equipment in modern industrial production and is widely used in scenarios such as mines, ports, and steel mills. In a belt conveyor, the idlers bear the function of supporting the conveyor belt and are the core components of the belt conveyor. Their normal rotation is of great significance for ensuring the smooth operation of the conveyor belt, extending the service life of the conveyor belt, and ensuring the transportation efficiency. However, due to the complexity of its working environment and the transported objects, the idlers may experience phenomena such as being locked or slipping, which will not only affect the normal operation of the conveyor belt but may even lead to equipment failures and production safety accidents in severe cases. Therefore, the real-time monitoring of the rotation state of the idlers is particularly important.
[0003] However, there are certain problems in the existing detection technologies for the rotation state of the idlers of belt conveyors. For example, in the fixed contact method, there are problems such as single equipment function, high deployment cost, complex structure, and high failure rate; the method of using optical flow to detect the rotation of idlers has problems such as low detection efficiency, susceptibility to light interference, and low robustness, resulting in low detection accuracy; if the method of manual inspection is adopted, there are problems of low efficiency and low safety. Summary of the Invention
[0004] The purpose of the present invention is to solve the problems existing in the prior art, and provide a method for detecting the rotation of idlers based on video understanding, which uses machine vision and artificial intelligence technologies to achieve real-time mobile detection of the rotation state of idlers, reduce the occurrence of safety accidents of belt conveyors caused by non-rotating idlers, and improve the safety and reliability of belt conveyors.
[0005] To achieve the above object, the present invention is realized through the following technical solutions:
[0006] A method for detecting the rotation of idlers based on video understanding includes the steps:
[0007] S1. Generate a target detection model dataset;
[0008] S2. Build and train a yolov7 network model with an oriented bounding box version;
[0009] S3. Perform forward calculation on the time-series images of a single idler group using the trained network model in step S2 to obtain the confidence of each frame of image and the coordinates of the idler object area. When the confidence of the idler object is greater than the confidence threshold, the current idler object is valid; conversely, when the confidence of the idler object is less than the confidence threshold, the current idler object is invalid.
[0010] S4. Use the network model in step S2 to identify the frame images under each group of idlers, and obtain the directional bounding box information of each frame image under each group of idlers;
[0011] S5. According to the obtained directional bounding box information, perform multi-channel fusion processing on the frame images under each group of idlers;
[0012] S6. After obtaining the multi-channel fusion detection data of each group of idlers, label the data according to the rotation state of the idlers. Each group of idlers has one and only one state label, and the state label is rotation or stop, and generate a video understanding dataset;
[0013] S7. Build and train a video understanding model;
[0014] S8. First, use the network model in step S2 to perform idler target detection on multiple frame images to obtain the position information of the directional bounding box of the idler object. Then, perform multi-channel fusion processing on the multiple frame images according to the position information, fuse the effective information of the multiple frame images together, perform forward calculation using the video understanding model to obtain the rotation state of the idler and the corresponding confidence level, and compare the obtained confidence level with the idler rotation threshold. When the confidence level is greater than the threshold, the current result is considered valid, otherwise the current result is invalid.
[0015] Preferably, in step S1, through the inspection system of the belt conveyor, collect the working video of the idlers, extract the video into frames as images, label the idler objects in the images, establish a target detection model dataset, and randomly divide the target detection model dataset into a training set and a test set.
[0016] Preferably, step S2 includes the steps:
[0017] S21. Build a directional bounding box yolov7 network model, which consists of a backbone feature extraction network, an enhanced feature extraction network, and a head. Among them, the picture is subjected to feature extraction in the backbone feature extraction network to obtain the effective feature layer of the picture; in the enhanced feature extraction network, the different scale information in the effective feature layer is fused to further extract features; the head judges the enhanced effective feature layer to determine whether there are detection targets and the types of detection targets;
[0018] S22. After the network is built, use the target detection model dataset to iteratively train the network model, optimize the parameters in the network model, obtain the target detection network model with the best detection effect, and save the weight file of the optimal model object for subsequent model deployment;
[0019] Preferably, in step S5, the multi-channel fusion processing steps include the steps:
[0020] S51. Screen valid frame data from each group of idler frame images according to the oriented bounding boxes and confidence information obtained after the forward calculation of the target detection model for each frame of image;
[0021] S52. According to the oriented bounding box information, filter out the invalid region information from the valid data in each group of idler data, reduce the invalid and even interfering information when judging the rotation state of the idler, and perform unified scale processing on the frame data information;
[0022] S53. For the screened data, select the first K frames for fusion operation processing to obtain the idler rotation judgment data of the current idler group.
[0023] Preferably, step S6 includes the steps:
[0024] S61. Build a video understanding model, adopt multiple convolutional and pooling layers, and add activation functions after each convolutional layer to extract features in the image; subsequently, classify the extracted features through a fully connected layer;
[0025] S62. Use the video understanding dataset to perform multiple iterative trainings on the video understanding network model to continuously optimize the parameters in the network model; through training, obtain the model with the best target detection effect; finally, save the weight file of the model for use in subsequent model deployment.
[0026] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0027] 1. The present invention builds an idler rotation detection method including the yolov7 network model, multi-channel fusion algorithm and video understanding classification model. The yolov7 network model is used to obtain the oriented bounding box information of the idler object in the idler group data, and the multi-channel fusion algorithm is used to perform fusion processing on the idler group data according to the oriented bounding box information. The multi-channel fusion data is input into the video understanding classification model. By extracting high-dimensional feature information including basic texture features and high-level semantic information of the time series, the prediction accuracy of the model can be improved. At the same time, it also overcomes the disadvantages of long time consumption, high hardware requirements and large amount of calculation when directly inferring the frame data after target detection using the video understanding model in the prior art, thereby reducing the application cost and improving the operation speed.
[0028] 2. The yolov7 target detection network with the oriented bounding box version adopted by the present invention can obtain the target box with rotation of the idler target. Compared with the axis-aligned bounding box obtained by the ordinary target detection model, it can obtain more accurate idler target contour information, thereby reducing the interference of the invalid region in the data and achieving better detection effects; the multi-channel fusion algorithm in the method can perform fusion processing on multiple frames of data, reduce the calculation amount of the video understanding model, and improve the operation speed of the algorithm.
[0029] 3. Compared with the existing methods, the method of the present invention reduces the requirements for hardware devices, which plays a positive role in the popularization and application of the method in reality; effectively improves the prediction accuracy of the rotating state of the idler, reduces the amount of calculation, and improves the calculation speed, so as to be able to detect the rotating state of the idler more efficiently and accurately. When using the device deployed with this method to detect the rotating state of the idler, it effectively improves the driving speed and detection accuracy, so as to be able to timely detect the abnormal rotating state of the idler and make targeted treatments in a timely manner to prevent the occurrence of safety accidents. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 is the flowchart of the present invention;
[0031] Figure 2 is the calculation flowchart of the multi-data fusion algorithm;
[0032] Figure 3 is one of the schematic structural diagrams of the oriented bounding box yolov7 target detection network model;
[0033] Figure 4 is the second schematic structural diagram of the oriented bounding box yolov7 target detection network model. DETAILED DESCRIPTION OF THE INVENTION
[0034] The present invention will be further described below in conjunction with specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. In addition, it should be understood that after reading the content taught by the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent forms also fall within the scope defined by this application.
[0035] Embodiment 1: The present invention relates to a method for detecting the rotation of idlers based on video understanding. First, continuous videos of each idler during the operation of a belt conveyor are collected, and the videos are frame-sampled to obtain sequential frame images. The yolov7 object detection model with an oriented bounding box version is used to perform forward calculation on the sequential frame images to obtain the confidence level and regional coordinates of the idler objects in each frame image. By comparing with a set target detection confidence threshold, it is confirmed whether the detected idler target is valid. When the confidence level is greater than the target detection confidence threshold, the idler target is valid; when the confidence level is less than the target detection confidence threshold, the idler target is invalid. Further, part of the data is selected from the valid data of the target detection for fusion processing, and the multi-frame image data is fused into a single data to reduce interference from invalid information and reduce the amount of information calculation. The fused information is used for forward calculation using a video understanding model to obtain the result and confidence level of the idler rotation state. By comparing with the rotation state confidence threshold, when the confidence level is greater than the rotation state confidence threshold, the idler state is valid; otherwise, it is invalid.
[0036] Specifically, it includes:
[0037] Step 1: Obtain the working video of the idlers of the belt conveyor, perform frame-sampling on the working video, and perform data annotation on the images with idler targets according to the idler position categories to obtain a single-frame training data set;
[0038] Step 2: Construct a yolov7 object detection network with an oriented bounding box version, and use the single-frame training data set to train this network to obtain a generalization model that can detect idler targets, which can perform static recognition on the idler targets in the belt conveyor video;
[0039] Step 3: Perform temporal sampling processing on the working video and divide it according to idler groups to obtain a training data set of idler groups;
[0040] Step 4: Use the object detection model in Step S2 to detect the data in the idler group data set to obtain the oriented bounding boxes of each idler target in the data;
[0041] Step 5: Perform multi-channel fusion processing on the first K-frame data according to the idler oriented bounding boxes of the first K-frame data in each group of idler data;
[0042] Step 6: According to the actual rotation state of the idler targets in the idler group data, perform annotation on whether the idler rotates or does not rotate on the fused data to obtain a video understanding data set;
[0043] Step 7: Construct a video understanding classification model based on multi-frame fusion data. The backbone network of this model performs continuous convolution and downsampling on the K-frame multi-channel fusion data to extract high-dimensional features after fusing the basic texture features of the data and the high-level semantic information in the time series. Finally, a classification head based on a fully connected layer is used to classify the multi-channel fusion information to determine whether the rotation state of the idler when the belt conveyor starts is rotation or slipping.
[0044] Step 8: Use the video understanding dataset in Step S6 to train the video understanding classification model to obtain a video understanding generalization model that can determine whether the idler rotates.
[0045] Step 9: Read the weight file of the trained video understanding generalization model and call the video understanding model inference program.
[0046] Step 10: Infer the frame data of the idler group using the generalization model of object detection to obtain the target-oriented bounding box data of the idler, generate multi-channel fusion data, and call the video understanding model for forward calculation to obtain the confidence of the idler state.
[0047] Step 11: Compare the confidence obtained in Step 10 with the confidence threshold to determine the true state of the idler.
[0048] Step 12: When performing dynamic inspection on the belt conveyor, the video data can be analyzed in real time according to the idler group. The detection model detects whether there is an idler object in each frame of data. If there is, the target-oriented bounding box of the idler is obtained at the same time; the data fusion algorithm screens the optimal data for multi-channel data fusion processing; the video understanding model analyzes the idler state from the fused data.
[0049] Example 2: As shown in Figure 1 This paper describes the method for detecting the rotation of an idler based on video understanding of the present invention. After obtaining the working video of the idler of the belt conveyor, the video is parsed into time-series frame images, and the forward calculation is performed on the time-series frame images using the yolov7 network model with an oriented bounding box version to obtain the position area and confidence of the idler object in the image. Then, it is judged whether there is an idler object in the image according to the confidence. The main part of this method lies in the generation of the object detection dataset and the construction and training of the yolov7 network model with an oriented bounding box version.
[0050] According to Figure 1 the dashed box part in, the process of dataset generation and construction and training of the network model is as follows:
[0051] S1: Through the inspection system of the belt conveyor, collect the working video of the idler, extract the video into images, label the idler objects in the images, establish an object detection dataset, and randomly divide the dataset into a training set and a test set.
[0052] Through the inspection system of the belt conveyor, the working video of the idler object can be conveniently recorded. In order to ensure that the sample data has sufficient representativeness, when sampling the idler object samples, various factors such as weather, light, time period, camera angle, lighting, idler frame, and idler need to be considered, so that the belt conveyor idler samples can cover different belt conveyor idler detection scenarios.
[0053] When annotating the target detection image, different from the general target detection network that annotates the axis-aligned bounding box, the oriented bounding box target detection network requires the data to be annotated according to the shape and direction of the target, reducing the interference of invalid information on the learning of the network model.
[0054] S2. Build an oriented bounding box yolov7 network model. The model consists of three parts: the backbone feature extraction network, the enhanced feature extraction network, and the head. The picture extracts features in the backbone network to obtain the effective feature layer of the picture; in the enhanced feature extraction network, the different scale information in the effective feature layer is fused to further extract features; the head part judges the enhanced effective feature layer to determine whether there are detection targets and the types of detection targets.
[0055] S3. After the network is built, according to factors such as the size of the dataset, hardware performance, and software framework, select appropriate training parameters such as the learning rate, number of iterations, and optimizer type, and use the target detection dataset to iteratively train the network model to optimize the parameters in the network model, obtain the target detection network model with the best detection effect, and save the weight file of the optimal model object for subsequent model deployment.
[0056] S4. After the target detection model is trained, the trained target detection model can be used for idler object detection. In actual applications, the trained target detection model is used for forward calculation of the sequential images of a single idler group to obtain the confidence level and the coordinate of the idler object area of each frame image.
[0057] To reduce misjudgment, a detection target confidence threshold needs to be set. Compare the confidence level of the idler object detected in the frame image with the confidence threshold. When the confidence level of the idler object is greater than the confidence threshold, the current idler object is valid; otherwise, when the confidence level of the idler object is less than the confidence threshold, the current idler object is invalid.
[0058] Combined with Figure 1 As shown, if the rotation state of the idler of the belt conveyor is judged, after the training of the target detection network is completed, data fusion processing of the data and the video understanding model are also required to judge the rotation state of the idler. The steps are as follows:
[0059] The working video of the belt conveyor idlers collected by the inspection system is processed by extracting frames in time sequence for each group of idlers to obtain the time-sequence frame images under each group of idlers. The object detection model is used to detect and identify the idlers in the frame images under each group of idlers, and the oriented bounding box information of each frame image under each group of idlers is obtained.
[0060] According to the obtained oriented bounding box information, multi-channel fusion processing is performed on the frame images under each group of idlers. The purpose of multi-channel fusion processing is to fuse multiple frame images under a single group of idlers into a single detection data while retaining the valid information in each frame image.
[0061] First, according to the oriented bounding box and confidence information obtained after the forward calculation of the object detection model for each frame image, the valid frame data is screened from the frame images of each group of idlers; second, according to the oriented bounding box information, for the valid data in each group of idler data, the invalid regional information can be filtered out, reducing the invalid and even interfering information when judging the rotation state of the idlers, and performing unified scale processing on the frame data information; finally, for the data screened above, the first K frames are selected for fusion operation processing to obtain the idler rotation judgment data of the current idler group. After multi-channel fusion processing, the idler rotation information in each group of idlers is retained, and at the same time, the interference of the predicted invalid data of the number of frames is reduced, which has great positive significance for reducing the detection time, improving the detection speed, reducing the hardware requirements, and improving the detection accuracy.
[0062] After obtaining the multi-channel fusion detection data of each group of idlers, according to the rotation state of the idlers, the data is labeled, and each group of idler fusion data has and only has one status label: rotating or stopped. A video understanding training dataset is generated.
[0063] A video understanding model is built. This network model uses multiple convolutional and pooling layers, and an activation function is added after each convolutional layer to extract the features in the image. Subsequently, the extracted features are classified through a fully connected layer. The entire network has a deeper and wider hierarchical structure, enabling it to learn more complex image features and patterns.
[0064] Before training the video understanding model, some factors need to be considered, such as the scale of the dataset, hardware performance, and software framework. According to these factors, appropriate training parameters such as the learning rate, number of iterations, and optimizer type are selected. Next, the video understanding dataset is used to perform multiple iterative trainings on the video understanding network model to continuously optimize the parameters in the network model. Through training, the model with the best object detection effect can be obtained. Finally, the weight file of the model is saved for use in subsequent model deployment.
[0065] After the video understanding model is trained, the trained model can be used for the detection of the rotating state of the idler. In practical applications, for the sequential frame images of the idler group, first use the object detection network to perform idler object detection on multiple frames of images to obtain the position information of the oriented bounding box of the idler object. According to the position information, perform multi-channel fusion processing on multiple frames of images, fuse the effective information of multiple frames of images together, and use the video understanding model for forward calculation to obtain the rotating state of the idler and the corresponding confidence. Similar to the object detection model, judging the true state of the idler also needs to be judged according to the confidence of the rotating state. Compare the obtained confidence with the idler rotation threshold. When the confidence is greater than the threshold, the current result is considered valid; otherwise, the current result is invalid.
Claims
1. A roller rotation detection method based on video understanding, characterized in that: Includes steps: S1, generate target detection model dataset; S2. Build and train the oriented bounding box version of the yolov7 network model; S3, using the network model trained in step S2 to perform forward calculation on the time series image of a single roller group, to obtain the confidence and roller object area coordinates of each frame image, when the roller object confidence is greater than the confidence threshold, the current roller object is valid, otherwise, when the roller object confidence is less than the confidence threshold, the current roller object is invalid; S4, using the network model in step S2 to identify the frame image under each group of rollers, and obtain the directional bounding box information of each frame image under each group of rollers; S5. Perform multi-channel fusion processing on the frame images under each group of rollers according to the obtained directional bounding box information; S6. After obtaining the multi-channel fusion detection data of each group of rollers, the data is labeled according to the rotation state of the rollers. The fusion data of each group of rollers has only one state label, which is rotation or stop, and a video understanding data set is generated; S7. Build and train a video understanding model; S8. First, use the network model in step S2 to perform roller target detection on multiple frames of images to obtain the directional bounding box position information of the roller object, and then perform multi-channel fusion processing on the multiple frames of images according to the position information, fuse the effective information of the multiple frames of images together, use the video understanding model to perform forward calculation, obtain the rotation state of the roller and the corresponding confidence, compare the obtained confidence with the roller rotation threshold, and when the confidence is greater than the threshold, the current result is considered valid, otherwise the current result is invalid.
2. The method for detecting roller rotation based on video understanding according to claim 1, characterized in that: In step S1, the inspection system of the belt conveyor is used to collect roller working videos, extract frames of the videos into images, annotate the roller objects in the images, establish a target detection model data set, and randomly divide the target detection model data set into a training set and a test set.
3. The method for detecting roller rotation based on video understanding according to claim 1, characterized in that: Step S2 comprises the steps of: S21. Build a directional bounding box yolov7 network model, which consists of a backbone feature extraction network, an enhanced feature extraction network and a head. The image is feature extracted in the backbone feature extraction network to obtain an effective feature layer of the image. In the enhanced feature extraction network, different scale information in the effective feature layer is fused to further extract features. The head determines whether there is a detection target and the type of detection target in the enhanced effective feature layer. S22. After the network is built, the target detection model data set is used to iteratively train the network model, optimize the parameters in the network model, obtain the target detection network model with the best detection effect, and save the weight file of the optimal model object for subsequent model deployment.
4. The method for detecting roller rotation based on video understanding according to claim 1, characterized in that: In step S5, the multi-channel fusion processing step includes the following steps: S51, filtering valid frame data from each group of roller frame images according to the directional bounding box and confidence information obtained after the forward calculation of the target detection model for each frame image; S52, according to the directional bounding box information, for the valid data in each group of roller data, filter out invalid area information, reduce invalid or even interference information when judging the rotation state of the roller, and perform unified scale processing on the frame data information; S53, for the filtered data, select the first K frames for fusion operation processing to obtain the roller rotation judgment data of the current roller group.
5. The method for detecting roller rotation based on video understanding according to claim 1, characterized in that: Step S6 comprises the steps of: S61. Build a video understanding model using multiple layers of convolution and pooling layers, and add activation functions after each convolution layer to extract features from the image. Subsequently, the extracted features are classified through a fully connected layer; S62. Use the video understanding dataset to perform multiple iterative training on the video understanding network model to continuously optimize the parameters in the network model; through training, obtain the model with the best target detection effect; finally, save the model weight file for use in subsequent model deployment.