Belt deviation detection method based on multi-modal data

The YOLO model is trained through the self-supervised learning of light guide mask autoencoder, and belt deviation detection is combined with multimodal data, which solves the detection accuracy and response delay problems of traditional detection technology under complex operating conditions, and realizes real-time and reliable detection of belt deviation.

CN120355698APending Publication Date: 2025-07-22滨沅国科(秦皇岛)智能科技股份有限公司
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510735390.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-04
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

Traditional belt deviation detection technology has problems such as attenuation of detection accuracy, high hardware cost, large computing resource consumption, high false alarm rate and response delay in complex working conditions, making it difficult to achieve real-time and reliable detection.

Method used

The YOLO model is trained using the IG-MAE light guide mask autoencoder based on self-supervised representation learning, combining multimodal data (visible light video and thermal imaging video) for belt deviation detection, and using the dual verification mechanism and dynamic threshold decisions, the millisecond-level real-time detection of belt and roller positions is achieved.

Benefits of technology

Maintain good robustness in low light or complex backgrounds, significantly improving the generalization ability and accuracy of detection, reducing false alarm rates, and realizing blind spot monitoring and major accident prevention capabilities throughout the life cycle of the belt conveyor.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120355698A_ABST
    Figure CN120355698A_ABST
Patent Text Reader

Abstract

The invention provides a belt deviation detection method based on multi-modal data, and relates to the technical field of fault detection, and the method comprises the steps: firstly constructing a no-label data set and a label data set based on a historical data set; then, based on the label-free data set and the label data set, training a target detection model by adopting an illumination-guided mask auto-encoder IG-MAE of self-supervised representation learning as an auxiliary agent task; and based on the trained target detection model, analyzing by combining the visible light video stream and the thermal imaging video stream when the to-be-detected belt runs, and obtaining a belt deviation detection result. According to the method, the illumination-guided mask auto-encoder IG-MAE based on self-supervised representation learning is used as an auxiliary agent task for training, and good robustness can be kept under low illumination or complex backgrounds.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of fault detection, and particularly to a belt deviation detection method based on multimodal data. Background Art

[0002] With the acceleration of the industrial automation process, the belt conveyor system, as the core equipment in scenarios such as bulk cargo ports and coal mines, its operating safety and efficiency directly affect production benefits. Traditional belt deviation detection technologies mainly rely on mechanical limit devices or single-sensor solutions, and have the following obvious defects under complex working conditions: 1) Traditional mechanical detection devices achieve limit measurement through physical contact, and the detection accuracy is likely to decay due to mechanical wear during long-term use; 2) Contact detection devices based on color detection methods are sensitive to dust, vibration, and temperature changes, and are significantly affected by light and the cleanliness of the conveyor belt; 3) Non-contact solutions based on vision or laser require complex algorithm processing and it is difficult to achieve millisecond-level response during dynamic operation; 4) Multi-sensor fusion systems require multi-module collaborative work, resulting in a significant increase in hardware costs and maintenance difficulties; 5) Traditional vision algorithms are prone to detection errors under low light or complex backgrounds, affecting the reliability of the results; 6) Multi-source information fusion methods such as D-S evidence theory consume a large amount of computing resources and it is difficult to achieve real-time inference on edge devices. These technical bottlenecks lead to frequent belt deviation accidents, seriously restricting the safety and efficiency of industrial production. Summary of the Invention

[0003] The purpose of the present invention is to provide a belt deviation detection method based on multimodal data, which is trained with an illumination-guided masked autoencoder IG-MAE based on self-supervised representation learning as an auxiliary proxy task, and can maintain good robustness under low light or complex backgrounds.

[0004] A belt deviation detection method based on multimodal data, the method comprising: S1, obtaining a historical data set, and selecting key frames in the historical data set as an unlabeled data set, and labeling the unlabeled data set to obtain a labeled data set; the historical data set includes a left visible light video and a right visible light video during the operation of the belt conveyor; S2, constructing an object detection model, and training the object detection model based on the historical data set to obtain the trained object detection model; The target detection model is the YOLO model. During the training process of the YOLO model, an illumination-guided masked autoencoder IG-MAE based on self-supervised representation learning is used as an auxiliary proxy task for training. The IG-MAE adopts an asymmetric encoding and decoding structure, including an encoder based on the Transformer architecture and a lightweight decoder; the unlabeled dataset is used as the input of the encoder, and the labeled dataset is used as the input of the YOLO model; the encoder adopts an illumination-guided random masking strategy; The illumination-guided random masking strategy is as follows: The input image is divided into N image patches of a set size, and each of the image patches is converted into an embedding vector through a linear embedding layer; the N image patches are converted from the RGB space to the YUV space, and the luminance channel Y is extracted to obtain N Y-channel data; the average value of each of the Y-channel data is calculated to obtain the initial luminance weight of each of the image patches, each of the initial luminance weights is normalized to obtain N normalized weights, and each of the normalized weights is subjected to non-linear enhancement processing to obtain N enhanced luminance weights; a noise sequence with a shape of (B, N) is randomly generated, where B is the batch size set during training; the noise sequence is weighted based on each of the enhanced luminance weights to obtain a weighted noise sequence; the weighted noise sequence is sorted in descending order to obtain a sorting index; based on the masking ratio ratio, the indexes of the first N×(1 - ratio) image patches are retained, and the remaining image patches will be masked to obtain a masking index; a masked image is obtained based on each of the embedding vectors and the masking index; The masked image is used as the input of the VIT structure of the encoder. Based on the MLP, n first features output by n intermediate layers of the VIT structure are aligned with n second features output by n C3k2 modules in the backbone network of the YOLO model, and the first features and the second features of the same level are feature fused based on the cross-attention mechanism to obtain n fused features; the n fused features are used as the input of the next layer of the n C3k2 modules; The reconstruction loss is obtained based on the output of the decoder, the detection loss is obtained based on the output of the YOLO model, n InfoNCE losses are obtained based on the n fused features, the reconstruction loss, the detection loss, and the n InfoNCE losses are weighted and summed to obtain the total training loss, and the weights of the YOLO model are updated based on the total training loss to obtain the trained target detection model; S3. Obtain the visible light video stream and the thermal imaging video stream when the belt to be detected is running. Input the visible light video stream into the trained target detection model for inference to obtain the belt detection frames and the idler detection frames in each frame of the visible light image. Based on each of the belt detection frames and each of the idler detection frames, obtain several overlapping areas between the idlers and the belt. According to each of the overlapping areas, combined with multi-frame motion analysis, obtain the first deviation detection result. Each of the belt detection frames and each of the idler detection frames are tracked based on the IoU Tracker in each frame of the visible light image. S4. Based on each of the belt detection frames and each of the idler detection frames, obtain several pre-labeled visible light ROIs. Map the coordinates of each of the pre-labeled visible light ROIs to the thermal imaging image of its corresponding frame to obtain several thermal imaging ROIs. S5. Perform temperature jump analysis on each of the thermal imaging ROIs to obtain the second deviation detection result. When the first deviation detection result is that no deviation occurs and the second deviation detection result is that no deviation occurs, determine that the belt has not deviated. When the first deviation detection result or the second deviation detection result is that the belt to be detected has deviated, send a warning signal. When both the first deviation detection result and the second deviation detection result are that the belt to be detected has deviated, determine that the belt has deviated, send an alarm signal, and trigger a chain shutdown.

[0005] Optionally, obtain the visible light video stream and the thermal imaging video stream when the belt to be detected is running based on a quadruped bionic robot. The quadruped bionic robot includes a robot body, an autonomous charging module, a battery power management module, a fuzzy logic decision-making module, a visible light video module, and a thermal imaging video module. The autonomous charging module, the battery power management module, and the fuzzy logic decision-making module are all arranged on the robot body. The visible light video module and the thermal imaging video module are arranged on the robot body based on an attitude adjuster. The autonomous charging module is used to support a closed-loop operation starting from the charging room and ending at the charging room for the inspection task. The fuzzy logic decision-making module is used to achieve the classification and avoidance of dynamic obstacles and static obstacles. The visible light video module is used to obtain the visible light video stream. The thermal imaging video module is used to obtain the thermal imaging video stream. The battery power management module is used to monitor the remaining power in real time and predict the remaining cruising range. When the remaining power < 10%, start the shortest path return. When the remaining cruising range is less than the mileage required for the current inspection task, perform an emergency brake and send a fault code and the last position information to the control center. The mechanical structure design of the robot body supports a maximum climbing angle of 35°, a maximum roll angle of 25°, and an obstacle crossing height of 15 cm. The quadruped bionic robot has an IP67 protection rating.

[0006] Optionally, n is 4, and the neck structure of the YOLO model adopts a bidirectional feature pyramid structure; The bidirectional feature pyramid structure includes an upsampling fusion path and a downsampling fusion path; The upsampling fusion path includes a first convolutional layer, a second convolutional layer, a third convolutional layer, a fourth convolutional layer, a first upsampling fusion module, and a second upsampling fusion module; the first upsampling fusion module includes a first upsampling layer, a first CA attention fusion layer, and a first C3k2 layer; the second upsampling fusion module includes a second upsampling layer, a second CA attention fusion layer, and a second C3k2 layer; The downsampling fusion path includes a first downsampling fusion module and a second downsampling fusion module; the first downsampling fusion module includes a fifth convolutional layer, a third CA attention fusion layer, and a third C3k2 layer; the second downsampling fusion module includes a sixth convolutional layer, a fourth CA attention fusion layer, and a fourth C3k2 layer; The third convolutional layer, the first upsampling layer, the first CA attention fusion layer, the first C3k2 layer, the fourth convolutional layer, the second upsampling layer, the second CA attention fusion layer, and the second C3k2 layer are connected in sequence; The second C3k2 layer is respectively connected to the fifth convolutional layer and the first detection head; the first convolutional layer is respectively connected to the first CA attention fusion layer and the third C3k2 module of the backbone network of the YOLO model; the second convolutional layer is respectively connected to the second CA attention fusion layer and the second C3k2 module of the backbone network of the YOLO model; Both the third convolutional layer and the fourth CA attention fusion layer are connected to the C2PSA module of the backbone network of the YOLO model; The fifth convolutional layer, the third CA attention fusion layer, the third C3k2 layer, the sixth convolutional layer, the fourth CA attention fusion layer, and the fourth C3k2 layer are connected in sequence; the third CA attention fusion layer is respectively connected to the first C3k2 layer and the third C3k2 module of the backbone network of the YOLO model; The third C3k2 layer is connected to the second detection head, and the fourth C3k2 layer is connected to the third detection head.

[0007] Optionally, the convolutional channels of the first convolutional layer, the second convolutional layer, the third convolutional layer, the fourth convolutional layer, the fifth convolutional layer, and the sixth convolutional layer are all 1×1.

[0008] Optionally, the method further includes: collecting historical detection data based on the port edge cloud platform, and fine-tuning and updating the trained target detection model according to the set period based on the historical detection data.

[0009] Optionally, S3 - S5 are implemented based on an edge computing unit, which is disposed on the robot body.

[0010] Optionally, S3 is specifically as follows: Assign a unique ID to each idler, and perform cross - frame matching on each belt detection frame and each idler detection frame based on the IoU Tracker to obtain the overlapping area between each idler and the belt in each visible - light image frame. Suppose the i - th idler first disappears in the (j + 1) - th visible - light image frame and first appears in the (j - n) - th visible - light image frame; obtain the overlapping area between the i - th idler and the belt from the (j - n) - th visible - light image frame to the j - th visible - light image frame. Calculate the change rate of the overlapping area between the i - th idler in the j - th visible - light image frame and the overlapping area between the i - th idler and the belt from the (j - n) - th visible - light image frame to the (j - 1) - th visible - light image frame respectively. If there is a continuous set number of overlapping area change rates greater than the first set area threshold or there is any overlapping area change rate greater than the second set area threshold and less than the third set area threshold, then the relative position of the i - th idler and the belt is abnormal; otherwise, the relative position of the i - th idler and the belt is normal; the third set area threshold is greater than the second set area threshold, and the second set area threshold is greater than the first set area threshold. If the relative positions of a continuous set number of idlers are all abnormal, then the first belt deviation detection result is that the belt to be detected is deviated; otherwise, the first belt deviation detection result is that the belt to be detected is not deviated.

[0011] Optionally, S4 is specifically as follows: Pre - mark the ROI of the contact area between the visible - light idler and the belt based on each belt detection frame and each idler detection frame; track the ROI based on the IoU Tracker, and map each pre - marked visible - light ROI to the thermal - imaging video stream based on the coordinate mapping relationship between the visible - light video stream and the thermal - imaging video stream to obtain several thermal - imaging ROIs, and obtain the temperature values of each thermal - imaging ROI; make a judgment based on the temperature values corresponding to the set temperature number of frames, calculate the temperature change values of the temperature values corresponding to the set temperature number of frames. If there is a continuous set number of temperature change values greater than the first temperature set value or there is any temperature change value greater than the second temperature set value, then the second belt deviation detection result is that the belt to be detected is deviated; otherwise, the second belt deviation detection result is that the belt to be detected is not deviated; the second temperature set value is greater than the first temperature set value.

[0012] Optionally, the method further includes: If the overlapping area of consecutive set frames is less than the set ratio of the area of the i-th idler detection frame or the change rate of the overlapping area is greater than the third set area threshold, a calibration signal is triggered; Taking the set detection frames as a period, obtaining the total number of idlers included in each detection period. If the total number of idlers in the current detection period is less than the set ratio of the total number of idlers in the previous detection period, the calibration signal is triggered; Obtaining the distance from the center of the belt detection frame in each frame of visible light image to the center of the visible light image, obtaining a number of distance values. If the distance values corresponding to consecutive set frames are all greater than the set ratio of the width of the visible light image, the calibration signal is triggered.

[0013] Optionally, when triggering the decoupling of the calibration signal, the following process is executed: Obtaining the point cloud data of the belt to be detected and the surrounding environment, and the attitude information of the attitude adjuster; Based on the point cloud data, the attitude information of the attitude adjuster, and combining with the six-degree-of-freedom kinematic model, obtaining the attitude adjustment parameters of the attitude adjuster; Adjusting the attitude of the attitude adjuster based on the attitude adjustment parameters.

[0014] The effects of the present invention are as follows: The belt deviation detection method based on multi-modal data of the present invention realizes the dual verification of belt deviation based on the multi-modal perception fusion strategy; at the same time, the illumination-guided masked auto-encoder IG-MAE based on self-supervised representation learning is used as an auxiliary proxy task for training, which can maintain good robustness under low light or complex backgrounds, significantly improving the detection generalization ability in complex scenarios; and using the dynamic threshold decision mechanism to effectively distinguish between working condition fluctuations and real faults, greatly reducing the false alarm rate; finally, realizing the full-life-cycle non-blind-spot monitoring of the belt conveyor, significantly improving the major accident pre-control ability.

[0015] The belt deviation detection method based on multi-modal data of the present invention uses an improved YOLO object detection model combined with a dynamic threshold discrimination and multi-frame motion analysis algorithm to realize the millisecond-level real-time detection of the positions of the belt and the idlers, breaking through the response delay bottleneck of traditional vision solutions.

[0016] The belt deviation detection method based on multi-modal data of the present invention significantly reduces the complexity and maintenance cost compared with the traditional multi-sensor fixed deployment scheme; the introduced thermal imaging temperature verification technology and self-supervised learning optimization strategy not only maintain the detection robustness under low light or complex backgrounds, but also continuously improve the algorithm reliability through the dynamic model update mechanism, solving the problems of large computational resource consumption and difficult edge device deployment in traditional fusion methods. Description of the Drawings

[0017] Figure 1 It is the flow chart of the belt deviation detection method based on multi-modal data of the present invention; Figure 2 It is the architecture diagram of the YOLO model and the IG-MAE network; Figure 3 It is the schematic diagram of dataset annotation; Figure 4 It is the schematic diagram of the training result of the YOLO model; Figure 5 It is the schematic diagram of the detection result of the trained YOLO model; Figure 6 It is the schematic diagram of thermal imaging ROI monitoring; Figure 7 It is the schematic diagram of the attitude adjustment process. Specific embodiments

[0018] Hereinafter, the embodiments of the present invention will be described with reference to the accompanying drawings.

[0019] Figure 1 It is the flow chart of the belt deviation detection method based on multi-modal data of the present invention. As Figure 1 shown, the present invention provides a belt deviation detection method based on multi-modal data, and the method includes: S1. Obtain a historical dataset, select key frames in the historical dataset as an unlabeled dataset, and annotate the unlabeled dataset to obtain a labeled dataset. The historical dataset includes the left visible light video and the right visible light video during the operation of the belt conveyor. The data annotation process is as Figure 3 shown.

[0020] S2. Construct a target detection model, train the target detection model based on the historical dataset, and obtain a trained target detection model.

[0021] Preferably, the network architecture diagram of the target detection model is as Figure 2 shown. The target detection model is the YOLO model, specifically the YOLOv11-nano model. During the training process of the YOLO model, the illumination-guided masked autoencoder IG-MAE based on self-supervised representation learning is used as an auxiliary proxy task for training. The IG-MAE adopts an asymmetric encoding and decoding structure, including an encoder based on the Transformer architecture and a lightweight decoder; the unlabeled dataset is used as the input of the encoder, and the labeled dataset is used as the input of the YOLO model; the encoder adopts an illumination-guided random masking strategy.

[0022] The light-guided random masking strategy is as follows: divide the input image into N image patches of a set size, and convert each image patch into an embedding vector through a linear embedding layer; convert the N image patches from the RGB space to the YUV space, and extract the luminance channel Y to obtain N Y-channel data; calculate the average value of each Y-channel data to obtain the initial luminance weight of each image patch, normalize each initial luminance weight to obtain N normalized weights, and perform non-linear enhancement processing on each normalized weight to obtain N enhanced luminance weights; randomly generate a noise sequence with a shape of (B, N), where B is the batch size set during training; weight the noise sequence based on each enhanced luminance weight to obtain a weighted noise sequence; perform descending sorting on the weighted noise sequence to obtain sorting indices; based on the masking ratio ratio, retain the indices of the first N×(1 - ratio) image patches, and the remaining image patches will be masked to obtain masking indices; obtain the masked image based on each embedding vector and the masking indices.

[0023] Preferably, N is 16 and ratio is 0.75. As Figure 2 shown, the input of the VIT structure of the encoder is the masked image, where the masked image patches are completely ignored during the encoding stage, and the unmasked image patches extract high-level semantic features through the VIT structure; the output of the VIT structure of the encoder is combined with the learnable vector at the masking position as the input of the decoder, and the original image is restored through the decoder, and only the reconstruction loss of the masked area is calculated.

[0024] Specifically, the VIT structure includes 12 ViT layers connected in sequence. n is 4.

[0025] The neck structure of the YOLO model adopts a bidirectional feature pyramid structure, and through bidirectional feature flow and coordinate attention weighted fusion, the detection ability of the model for relatively small idler rollers is improved. The backbone network of the YOLO model includes 5 convolutional modules, 4 C3k2 modules, 1 spatial pyramid pooling module, and 1 C2PSA module.

[0026] The 5 convolutional modules are respectively defined as the first convolutional module, the second convolutional module, the third convolutional module, the fourth convolutional module, and the fifth convolutional module; the 4 C3k2 modules are respectively defined as the first C3k2 module, the second C3k2 module, the third C3k2 module, and the fourth C3k2 module; The first convolutional module, the second convolutional module, the first C3k2 module, the third convolutional module, the second C3k2 module, the fourth convolutional module, the third C3k2 module, the fifth convolutional module, the fourth C3k2 module, the spatial pyramid pooling module, and the C2PSA module are connected in sequence.

[0027] The bidirectional feature pyramid structure includes an upsampling fusion path and a downsampling fusion path, which realizes cross-stage fusion of high-level and low-level features.

[0028] The upsampling fusion path includes a first convolutional layer, a second convolutional layer, a third convolutional layer, a fourth convolutional layer, a first upsampling fusion module, and a second upsampling fusion module; the first upsampling fusion module includes a first upsampling layer, a first CA attention fusion layer, and a first C3k2 layer; the second upsampling fusion module includes a second upsampling layer, a second CA attention fusion layer, and a second C3k2 layer.

[0029] The downsampling fusion path includes a first downsampling fusion module and a second downsampling fusion module; the first downsampling fusion module includes a fifth convolutional layer, a third CA attention fusion layer, and a third C3k2 layer; the second downsampling fusion module includes a sixth convolutional layer, a fourth CA attention fusion layer, and a fourth C3k2 layer.

[0030] The third convolutional layer, the first upsampling layer, the first CA attention fusion layer, the first C3k2 layer, the fourth convolutional layer, the second upsampling layer, the second CA attention fusion layer, and the second C3k2 layer are connected in sequence.

[0031] The second C3k2 layer is respectively connected to the fifth convolutional layer and the first detection head; the first convolutional layer is respectively connected to the first CA attention fusion layer and the third C3k2 module; the second convolutional layer is respectively connected to the second CA attention fusion layer and the second C3k2 module.

[0032] Both the third convolutional layer and the fourth CA attention fusion layer are connected to the C2PSA module.

[0033] The fifth convolutional layer, the third CA attention fusion layer, the third C3k2 layer, the sixth convolutional layer, the fourth CA attention fusion layer, and the fourth C3k2 layer are connected in sequence; the third CA attention fusion layer is respectively connected to the first C3k2 layer and the third C3k2 module of the backbone network of the YOLO model.

[0034] The third C3k2 layer is connected to the second detection head, and the fourth C3k2 layer is connected to the third detection head.

[0035] Preferably, the convolutional channels of the first convolutional layer, the second convolutional layer, the third convolutional layer, the fourth convolutional layer, the fifth convolutional layer, and the sixth convolutional layer are all 1×1.

[0036] Specifically, the CA attention is coordinate attention. Coordinate attention captures long-range dependencies by explicitly encoding the spatial coordinate information of the feature map. Coordinate attention not only focuses on the information interaction between channels but also retains the information in the spatial direction. For the horizontal belt and vertical idler in the task, coordinate attention is more suitable for processing directional targets; the C3k2 layer is a basic building block in the YOLO model, which is essentially a variant of the residual structure. By stacking multiple convolutional layers and residual connections, it enhances the feature expression ability, enabling the model to capture the position information of the belt and idler more accurately while maintaining light weight, thus improving the accuracy of deviation detection. As Figure 2 shown, the detection head is an important part of the YOLO model, responsible for the final processing of features to output the position and category information of the target. The detection head uses the multi-parameter distance intersection over union (MPDIoU) as the loss function to optimize the positioning accuracy of the detection box. MPDIoU not only considers the overlapping area between the detection box and the ground truth box but also introduces a geometric distance penalty, improving the positioning accuracy of the belt and idler.

[0037] Based on the MLP, align the n first features output by the n intermediate layers of the VIT structure with the second features output by the n C3k2 modules in the backbone network of the YOLO model, and based on the cross-attention mechanism, fuse the first features and the second features at the same level to obtain n fused features; use the n fused features as the input for the next layer of the n C3k2 modules.

[0038] The 4 intermediate layers selected from the VIT structure are the 3rd layer, the 6th layer, the 9th layer, and the 12th layer. The 3rd layer is defined as the first VIT layer, the 6th layer is defined as the second VIT layer, the 9th layer is defined as the third VIT layer, and the 12th layer is defined as the fourth VIT layer.

[0039] Align the first feature output by the first VIT layer with the second feature output by the first C3k2 module, and perform feature fusion based on the cross-attention mechanism to obtain the 1st fused feature. The 1st fused feature is used as the input for the third convolutional module; align the first feature output by the second VIT layer with the second feature output by the second C3k2 module, and perform feature fusion based on the cross-attention mechanism to obtain the 2nd fused feature. The 2nd fused feature is used as the input for the fourth convolutional module and the second convolutional layer; align the first feature output by the third VIT layer with the second feature output by the third C3k2 module, and perform feature fusion based on the cross-attention mechanism to obtain the 3rd fused feature. The 3rd fused feature is used as the input for the fifth convolutional module and the first convolutional layer; align the first feature output by the fourth VIT layer with the second feature output by the fourth C3k2 module, and perform feature fusion based on the cross-attention mechanism to obtain the 4th fused feature. The 4th fused feature is used as the input for the spatial pyramid pooling module.

[0040] Obtain the reconstruction loss based on the output of the decoder, obtain the detection loss based on the output of the YOLO model, obtain n InfoNCE losses based on n fusion features, perform weighted summation on the reconstruction loss, the detection loss and the n InfoNCE losses to obtain the total training loss, and update the weights of the YOLO model based on the total training loss to obtain a trained object detection model.

[0041] S3. Obtain the visible light video stream and the thermal imaging video stream when the belt to be detected is running, input the visible light video stream into the trained object detection model for inference to obtain the belt detection frames and the idler detection frames in each frame of visible light image, and obtain a number of overlapping areas between the idlers and the belt based on the belt detection frames and the idler detection frames; according to the overlapping areas, combined with multi-frame motion analysis, obtain the first deviation detection result. The belt detection frames and the idler detection frames are tracked based on the IoU Tracker in each frame of visible light image.

[0042] Preferably, obtain the visible light video stream and the thermal imaging video stream when the belt to be detected is running based on the quadruped bionic robot. The resolution of the visible light video stream is 1920×1080, the resolution of the thermal imaging video stream is 1280×720, the visible light video stream and the thermal imaging video stream are synchronously triggered at an interval of 10 ms, and the spatio-temporal alignment algorithm is used to ensure the consistency of multi-modal data.

[0043] The quadruped bionic robot includes a robot body, an autonomous charging module, a battery power management module, a fuzzy logic decision-making module, a visible light video module and a thermal imaging video module; the autonomous charging module, the battery power management module and the fuzzy logic decision-making module are all arranged on the robot body; the visible light video module and the thermal imaging video module are arranged on the robot body based on an attitude adjuster.

[0044] The autonomous charging module is used to support the closed-loop operation starting from the charging room and ending at the charging room; the fuzzy logic decision-making module is used to realize the classification and avoidance of dynamic obstacles and static obstacles; the visible light video module is used to obtain the visible light video stream; the thermal imaging video module is used to obtain the thermal imaging video stream; the battery power management module is used to monitor the remaining power in real time and predict the remaining cruising range, start the shortest path return when the remaining power < 10%, and perform emergency braking and send the fault code and the last position information to the control center when the remaining cruising range is less than the mileage required for the current inspection task.

[0045] The mechanical structure design of the robot body supports a maximum climbing angle of 35°, a maximum roll angle of 25° and an obstacle crossing height of 15 cm; the quadruped bionic robot has an IP67 protection level.

[0046] Specifically, a unique ID is assigned to each idler, and cross-frame matching is performed on each belt detection box and each idler detection box based on the IoU Tracker to obtain the overlapping area between each idler and the belt in each frame of visible light image.

[0047] Suppose the i-th idler first disappears in the (j + 1)-th frame of visible light image, and the i-th idler first appears in the (j - n)-th frame of visible light image; obtain the overlapping area between the i-th idler and the belt in the visible light images from the (j - n)-th frame to the j-th frame.

[0048] Calculate the change rate of the overlapping area between the i-th idler in the j-th frame of visible light image and the overlapping areas between the i-th idler and the belt in the visible light images from the (j - n)-th frame to the (j - 1)-th frame respectively.

[0049] If the change rate of the overlapping area for consecutive set frames is greater than the first set area threshold or there is any change rate of the overlapping area greater than the second set area threshold and less than the third set area threshold, then the relative position of the i-th idler and the belt is abnormal; otherwise, the relative position of the i-th idler and the belt is normal; the third set area threshold is greater than the second set area threshold, and the second set area threshold is greater than the first set area threshold.

[0050] If it is satisfied that the relative positions of the idlers in consecutive set frames are all abnormal, then the first belt deviation detection result is that the belt to be detected is deviated; otherwise, the first belt deviation detection result is that the belt to be detected is not deviated.

[0051] Preferably, the number of set frames is 5, the first set area threshold is 15%, the second set area threshold is 25%, and the third set area threshold is 70%. The exponential weighted moving average method is used to smooth the data, suppress high-frequency noise, and improve the discrimination accuracy.

[0052] S4. Obtain a number of pre-labeled visible light ROIs based on each belt detection box and each idler detection box; map the coordinates of each pre-labeled visible light ROI to the thermal imaging image of its corresponding frame to obtain a number of thermal imaging ROIs.

[0053] Furthermore, based on each belt detection frame and each roller detection frame, a visible light roller and belt contact area ROI is pre-marked; based on the IoU Tracker, the ROI is tracked, and based on the coordinate mapping relationship between the visible light video stream and the thermal imaging video stream, each pre-marked visible light ROI is mapped to the thermal imaging video stream to obtain a number of thermal imaging ROIs, and the temperature value of each thermal imaging ROI is obtained; based on the temperature value corresponding to the set temperature frame number, a judgment is made, and the temperature change value of each temperature value corresponding to the set temperature frame number is calculated. If there are a continuous set number of temperature change values greater than the first temperature setting value or there is any temperature change value greater than the second temperature setting value, then the second deviation detection result is that the belt to be detected has deviated, otherwise the second deviation detection result is that the belt to be detected has not deviated; the second temperature setting value is greater than the first temperature setting value.

[0054] Preferably, the setting quantity is 5, the first temperature setting value is 5° C., the second temperature setting value is 8° C., and Gaussian filtering is used to smooth the temperature data to improve the accuracy of judgment.

[0055] S5, perform temperature jump analysis on each thermal imaging ROI to obtain a second deviation detection result. When the first deviation detection result is that no deviation has occurred and the second deviation detection result is that no deviation has occurred, it is determined that the belt has not deviated; when the first deviation detection result or the second deviation detection result is that the belt to be detected has deviated, a warning signal is issued; when the first deviation detection result and the second deviation detection result are both that the belt to be detected has deviated, it is determined that the belt has deviated, an alarm signal is issued, and the machine is shut down.

[0056] Preferably, the method also includes: collecting historical detection data based on the port edge cloud platform, regularly collecting belt running images under different seasons and working conditions, and fine-tuning and updating the trained target detection model based on the historical detection data according to the set cycle. Specifically, after manually annotating the newly collected unlabeled data, it is merged with the original annotated data set to form a hybrid data set, the model loads the existing pre-trained weights, and the cosine annealing learning rate strategy is used to fine-tune the hybrid data set. During the training process, the backbone network parameters are kept frozen, and only the classification layer and the detection head are updated. The model retains the ability to recognize basic features and adapts to the detection needs under new working conditions.

[0057] S3-S5 are implemented based on an edge computing unit, which is set on the robot body. Preferably, the edge computing unit is based on the NVIDIA Jetson AGX Orin platform, integrating a 5G communication module and 1TB SSD storage.

[0058] Preferably, the method further comprises: If the overlapping area of consecutive z frames is less than the set ratio r of the area of the i-th idler detection frame or the change rate of the overlapping area is greater than the third set area threshold, a calibration signal is triggered.

[0059] Taking the set number of detection frames as a period, obtain the total number of idlers included in each detection period. If the total number of idlers in the current detection period is less than the set ratio of the total number of idlers in the previous detection period, a calibration signal is triggered.

[0060] Obtain the distance from the center of the belt detection frame in each frame of visible light image to the center of the visible light image, and obtain a number of distance values. If there are consecutive z frames corresponding to distance values that are all greater than the set ratio r of the width of the visible light image, a calibration signal is triggered.

[0061] Trigger the decoupling of the calibration signal and perform the following process: Obtain the point cloud data of the belt to be detected and the surrounding environment, and the attitude information of the attitude adjuster.

[0062] Based on the point cloud data, the attitude information of the attitude adjuster, and combined with the six-degree-of-freedom kinematic model, obtain the attitude adjustment parameters of the attitude adjuster.

[0063] Adjust the attitude of the attitude adjuster based on the attitude adjustment parameters.

[0064] Preferably, z is 3, r is 20%, and the third set area threshold is 70%; as Figure 7 shown, the six-degree-of-freedom kinematic model constructs a rotation matrix according to the point cloud data and the attitude information of the attitude adjuster , where is the roll angle, is the pitch angle, is the yaw angle; specifically, solve the gimbal target angle through the inverse kinematic equation , where is the horizontal rotation angle, is the pitch angle, is the roll angle, and the error calculation is .

[0065] Specifically, taking a bulk cargo transportation system in a certain port as an example, the deployment and operation process of the detection method of the present invention are as follows: The quadruped robot performs three full-line inspections along the belt line every day. Each single inspection covers about 1.6 kilometers of transportation, and the path takes 45 minutes. Build a three-dimensional map through laser SLAM, set transition points at the conveyor belt transfer tower, and set the gait to the climbing mode at the upper and lower stairs. When it is determined that the belt is running off track, the off-track frame is pushed to the central control room in real time through the 5G network, and an audible and visual alarm is triggered.

[0066] Select 5,000 images with different lighting scenarios and different dust concentration scenarios from the historical dataset to obtain an unlabeled dataset, and divide it into a training set, a validation set, and a test set according to the ratio of 8:1:1.

[0067] When training the object detection model, the initial learning rate is 0.001, the batch size is 4, and the training period is 200. The cosine annealing strategy is used to dynamically adjust the learning rate; after the model training is completed, the model is finally evaluated. As Figure 4 shown is the P-R curve obtained by training, and the performance of the model is evaluated according to indicators such as the accuracy, recall rate, and F1 value of the model. Figure 4 In it, belt represents the average precision (Average Precision, AP) of the belt category, roller represents the AP of the idler category, and mAP@0.5 represents the average precision when calculating IoU≥0.5.

[0068] The detection results of using the trained YOLO model for deviation detection are as Figure 5 shown, and the thermal imaging ROI monitoring is as Figure 6 shown. Based on the trained object detection model, the average detection error of the deviation amount of 0 - 300 mm is ≤5 mm, and the false alarm rate is <0.3%, which is significantly better than traditional mechanical detection equipment; the end-to-end delay from image acquisition to alarm trigger is ≤80 ms, meeting the industrial real-time requirements; the detection accuracy remains above 90% under complex working conditions.

[0069] The above embodiments only describe the preferred embodiments of the present invention, and do not limit the scope of the present invention. Without departing from the design spirit of the present invention, various deformations and improvements made by those of ordinary skill in the art to the technical solutions of the present invention shall fall within the protection scope determined by the claims of the present invention.

Claims

1. A belt deviation detection method based on multimodal data, characterized in that, The method includes: S1. Obtain a historical dataset, select key frames in the historical dataset as an unlabeled dataset, and label the unlabeled dataset to obtain a labeled dataset; the historical dataset includes a left visible light video and a right visible light video during the operation of the belt conveyor; S2. Construct a target detection model, and train the target detection model based on the historical dataset to obtain the trained target detection model; The target detection model is a YOLO model. During the training process of the YOLO model, an illumination-guided masked autoencoder IG-MAE based on self-supervised representation learning is used as an auxiliary proxy task for training. The IG-MAE adopts an asymmetric encoding and decoding structure, including an encoder based on the Transformer architecture and a lightweight decoder; the unlabeled dataset is used as the input of the encoder, and the labeled dataset is used as the input of the YOLO model; the encoder adopts an illumination-guided random masking strategy; The illumination-guided random masking strategy is as follows: Divide the input image into N image patches of a set size, and convert each of the image patches into an embedding vector through a linear embedding layer; convert the N image patches from the RGB space to the YUV space, and extract the luminance channel Y to obtain N Y-channel data; calculate the average value of each of the Y-channel data to obtain the initial luminance weight of each of the image patches, normalize each of the initial luminance weights to obtain N normalized weights, and perform non-linear enhancement processing on each of the normalized weights to obtain N enhanced luminance weights; randomly generate a noise sequence with a shape of (B, N), where B is the batch size set during training; weight the noise sequence based on each of the enhanced luminance weights to obtain a weighted noise sequence; sort the weighted noise sequence in descending order to obtain a sorting index; based on the masking ratio ratio, retain the indexes of the first N×(1 - ratio) image patches, and the remaining image patches will be masked to obtain a masking index; obtain a masked image based on each of the embedding vectors and the masking index; The masked image is used as the input of the VIT structure of the encoder. Based on the MLP, align n first features output by n intermediate layers of the VIT structure with second features output by n C3k2 modules in the backbone network of the YOLO model, and perform feature fusion on the first features and the second features of the same level based on the cross-attention mechanism to obtain n fused features; use the n fused features as the input of the next layer of the n C3k2 modules; Obtain a reconstruction loss based on the output of the decoder, obtain a detection loss based on the output of the YOLO model, obtain n InfoNCE losses based on the n fused features, perform weighted summation on the reconstruction loss, the detection loss, and the n InfoNCE losses to obtain a total training loss, and update the weights of the YOLO model based on the total training loss to obtain the trained target detection model; S3. Obtain the visible light video stream and the thermal imaging video stream when the belt to be detected is running. Input the visible light video stream into the trained target detection model for inference to obtain the belt detection frames and idler detection frames in each frame of visible light image. Based on each of the belt detection frames and each of the idler detection frames, obtain several overlapping areas between the idlers and the belt; according to each of the overlapping areas, combined with multi-frame motion analysis, obtain the first deviation detection result; each of the belt detection frames and each of the idler detection frames are tracked based on the IoU Tracker in each frame of visible light image; S4. Based on each of the belt detection frames and each of the idler detection frames, obtain several pre-labeled visible light ROIs; map the coordinates of each of the pre-labeled visible light ROIs to the thermal imaging image of its corresponding frame to obtain several thermal imaging ROIs; S5. Perform temperature jump analysis on each of the thermal imaging ROIs to obtain the second deviation detection result. When the first deviation detection result is that no deviation occurs and the second deviation detection result is that no deviation occurs, determine that the belt does not deviate; when the first deviation detection result or the second deviation detection result is that the belt to be detected deviates, send a warning signal; when both the first deviation detection result and the second deviation detection result are that the belt to be detected deviates, determine that the belt deviates, send an alarm signal, and interlock for shutdown.

2. The belt deviation detection method based on multimodal data according to claim 1, wherein Obtain the visible light video stream and the thermal imaging video stream when the belt to be detected is running based on the quadruped bionic robot; The quadruped bionic robot includes a robot body, an autonomous charging module, a battery power management module, a fuzzy logic decision-making module, a visible light video module, and a thermal imaging video module; the autonomous charging module, the battery power management module, and the fuzzy logic decision-making module are all arranged on the robot body; the visible light video module and the thermal imaging video module are arranged on the robot body based on an attitude adjuster; The autonomous charging module is used to support the closed-loop operation starting from the charging room and ending at the charging room for the inspection task; The fuzzy logic decision-making module is used to achieve the classification and avoidance of dynamic obstacles and static obstacles; the visible light video module is used to obtain the visible light video stream; the thermal imaging video module is used to obtain the thermal imaging video stream; the battery power management module is used to monitor the remaining power in real time and predict the remaining battery life. When the remaining power < 10%, start the shortest path return. When the remaining battery life is less than the mileage required for the current inspection task, perform an emergency brake and send the fault code and the last position information to the control center; The mechanical structure design of the robot body supports a maximum climbing angle of 35°, a side tilt of 25°, and an obstacle crossing height of 15 cm; the quadruped bionic robot has an IP67 protection level.

3. The belt deviation detection method based on multi-modal data according to claim 1, characterized in that, n is 4, and the neck structure of the YOLO model adopts a bidirectional feature pyramid structure; The bidirectional feature pyramid structure includes an upsampling fusion path and a downsampling fusion path; The upsampling fusion path includes a first convolutional layer, a second convolutional layer, a third convolutional layer, a fourth convolutional layer, a first upsampling fusion module, and a second upsampling fusion module; the first upsampling fusion module includes a first upsampling layer, a first CA attention fusion layer, and a first C3k2 layer; the second upsampling fusion module includes a second upsampling layer, a second CA attention fusion layer, and a second C3k2 layer; The downsampling fusion path includes a first downsampling fusion module and a second downsampling fusion module; the first downsampling fusion module includes a fifth convolutional layer, a third CA attention fusion layer, and a third C3k2 layer; The second downsampling fusion module includes a sixth convolutional layer, a fourth CA attention fusion layer, and a fourth C3k2 layer; The third convolutional layer, the first upsampling layer, the first CA attention fusion layer, the first C3k2 layer, the fourth convolutional layer, the second upsampling layer, the second CA attention fusion layer, and the second C3k2 layer are connected in sequence; The second C3k2 layer is respectively connected to the fifth convolutional layer and the first detection head; The first convolutional layer is respectively connected to the first CA attention fusion layer and the third C3k2 module of the backbone network of the YOLO model; the second convolutional layer is respectively connected to the second CA attention fusion layer and the second C3k2 module of the backbone network of the YOLO model; Both the third convolutional layer and the fourth CA attention fusion layer are connected to the C2PSA module of the backbone network of the YOLO model; The fifth convolutional layer, the third CA attention fusion layer, the third C3k2 layer, the sixth convolutional layer, the fourth CA attention fusion layer, and the fourth C3k2 layer are connected in sequence; the third CA attention fusion layer is respectively connected to the first C3k2 layer and the third C3k2 module of the backbone network of the YOLO model; The third C3k2 layer is connected to the second detection head, and the fourth C3k2 layer is connected to the third detection head.

4. The belt deviation detection method based on multimodal data according to claim 3, characterized in that The convolutional channels of the first convolutional layer, the second convolutional layer, the third convolutional layer, the fourth convolutional layer, the fifth convolutional layer, and the sixth convolutional layer are all 1×1.

5. The belt deviation detection method based on multimodal data according to claim 1, wherein The method further includes: collecting historical detection data based on the port edge cloud platform, and fine-tuning and updating the trained target detection model according to the set period based on the historical detection data.

6. The belt deviation detection method based on multimodal data according to claim 2, wherein, S3 - S5 are implemented based on the edge computing unit, and the edge computing unit is arranged on the robot body.

7. The belt deviation detection method based on multimodal data according to claim 2, wherein S3 is specifically: Assign a unique ID to each idler, and perform cross-frame matching on each belt detection box and each idler detection box based on the IoU Tracker to obtain the overlapping area between each idler and the belt in each frame of visible light image; Suppose the i-th idler first disappears in the (j + 1)-th frame of visible light image, and the i-th idler first appears in the (j - n)-th frame of visible light image; Obtain the overlapping area between the i-th idler and the belt from the (j - n)-th frame of visible light image to the j-th frame of visible light image; Calculate the change rate of the overlapping area between the $i$-th idler in the $j$-th visible light image and the $i$-th idler and the belt in the visible light images from the $(j - n)$-th frame to the $(j - 1)$-th frame respectively; If the change rate of the overlapping area for a continuous set number of frames is greater than the first set area threshold or if there is any change rate of the overlapping area that is greater than the second set area threshold and less than the third set area threshold, then the relative position of the $i$-th idler and the belt is abnormal; otherwise, the relative position of the $i$-th idler and the belt is normal; the third set area threshold is greater than the second set area threshold, and the second set area threshold is greater than the first set area threshold; If it is satisfied that the relative positions of a continuous set number of idlers are all abnormal, then the first belt deviation detection result is that the belt to be detected is deviated; otherwise, the first belt deviation detection result is that the belt to be detected is not deviated.

8. The belt deviation detection method based on multimodal data according to claim 1, wherein S4 is specifically as follows: Pre-mark the ROI of the contact area between the visible light idler and the belt based on each of the belt detection frames and each of the idler detection frames; track the ROI based on the IoU Tracker, and map each of the pre-marked visible light ROIs to the thermal imaging video stream based on the coordinate mapping relationship between the visible light video stream and the thermal imaging video stream to obtain a number of thermal imaging ROIs, and obtain the temperature values of each of the thermal imaging ROIs; Judge based on the temperature values corresponding to the set temperature frames, calculate the temperature change values of each of the temperature values corresponding to the set temperature frames. If there is a continuous set number of temperature change values greater than the first temperature set value or if there is any temperature change value greater than the second temperature set value, then the second belt deviation detection result is that the belt to be detected is deviated; otherwise, the second belt deviation detection result is that the belt to be detected is not deviated; the second temperature set value is greater than the first temperature set value.

9. The belt deviation detection method based on multimodal data according to claim 7, characterized in that The method further includes: If there is a continuous set number of frames with an overlapping area less than the set proportion of the area of the $i$-th idler detection frame or if the change rate of the overlapping area is greater than the third set area threshold, then trigger a calibration signal; Taking the set detection frames as a period, obtain the total number of idlers included in each detection period. If the total number of idlers in the current detection period is less than the set proportion of the total number of idlers in the previous detection period, then trigger the calibration signal; Obtain the distance from the center of the belt detection frame in each visible light image to the center of the visible light image, obtaining a number of distance values. If there are a continuous set number of frames corresponding to the distance values that are all greater than the set proportion of the width of the visible light image, then trigger the calibration signal.

10. The belt deviation detection method based on multi-modal data according to claim 9, characterized in that, Trigger the decoupling of the calibration signal and execute the following process: Obtain the point cloud data of the belt to be detected and the surrounding environment and the attitude information of the attitude adjuster; Based on the point cloud data and the attitude information of the attitude adjuster, and combining with the six-degree-of-freedom kinematic model, obtain the attitude adjustment parameters of the attitude adjuster; Adjust the attitude of the attitude adjuster based on the attitude adjustment parameters.

Citation Information

Cited By

  • Multi-mode monitoring and anomaly detection method, system and equipment of belt conveying system and medium

    CN120558607A