Conveyor belt longitudinal tear detection system based on binocular image fusion
Through binocular image fusion and improved YOLOv5 model, the problem of external environmental interference in longitudinal tear detection of conveyor belts is solved, and the stability and robustness of the detection results are achieved, ensuring the continuity of production and the stability of the equipment.
Patent Information
- Application Number
- CN202510495981.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-08-12
AI Technical Summary
The prior art is susceptible to external environmental factors in the longitudinal tear detection of conveyor belts, resulting in unstable detection results.
The longitudinal tear detection system of conveyor belt based on binocular image fusion is adopted, and the visible light and infrared images of the conveyor belt are collected using a binocular camera, and decomposed into high-frequency and low-frequency subband signals through non-downsampled contour wave transformation, combining the maximum entropy threshold segmentation and weighted average method to fuse the images, and tear detection is used using the improved YOLOv5 model to enhance the stability of the detection.
Effectively respond to the influence of factors such as light changes, noise interference and materials conveyed, improve the stability and robustness of the detection results, and ensure the continuity of production and the stable operation of the equipment.
Smart Images

Figure CN120471837A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of machine vision, and in particular relates to a conveyor belt longitudinal tear detection system based on binocular image fusion. Background Art
[0002] Conveyor belts are a widely used type of conveying equipment, primarily used to move materials from one location to another. They are highly adaptable and can be automated, significantly improving production efficiency.
[0003] During operation, conveyor belts may experience a number of problems, including belt breakage or wear, drive system failure, tensioning device failure, deviation, cleaning device malfunction, and electrical failure. Belt breakage is the most serious, halting the flow of materials during conveyance, leading to production halts and equipment damage. Belt breakage occurs when cracks develop and gradually expand on the belt surface. These cracks accelerate belt wear, causing faster aging and damage, increasing repair costs and impacting equipment operation. Cracks can also weaken the belt, impacting conveying capacity. Belt breakage during operation can cause material accumulation, blockages, and even safety incidents.
[0004] Existing technologies generally use traditional machine vision methods or deep learning models combined with sound signals for fault detection. During the actual detection process, external environmental factors such as lighting changes, noise interference, and conveyed materials will affect the detection effect, resulting in unstable detection results. Summary of the Invention
[0005] The purpose of the present invention is to provide a conveyor belt longitudinal tear detection system based on binocular image fusion, to solve the problem that the detection results of the existing technology are easily affected by the external environment during the detection process, to overcome the interference of the existing technology in factors such as lighting changes, noise interference, and conveyed materials, and to improve the stability of the detection results.
[0006] To achieve the above object, the technical solution adopted by the present invention is: A conveyor belt longitudinal tear detection system based on binocular image fusion, the system comprising: Original image acquisition module: uses a binocular camera to collect real-time images of the conveyor belt; Original image processing module: performing denoising, filtering and registration processing on the acquired image; Image decomposition module: uses non-subsampled contourlet transform to decompose the image into high-frequency and low-frequency sub-band signals; Image fusion module: Use the maximum entropy threshold segmentation method to fuse the low-frequency sub-band signals, and use the weighted average method to fuse the high-frequency sub-band signals, and finally generate the fused image through the inverse non-subsampled contourlet transform; Detection module: tearing detection based on improved YOLOv5.
[0007] The images collected by the binocular camera are used to obtain three-dimensional information of the conveyor belt.
[0008] The registration process in the original image processing module includes adjusting the sizes of the two images to be consistent and making the corresponding belt areas the same.
[0009] The improved YOLOv5 model includes: Backbone network: extracts image features through convolutional layers, depth-wise separable convolutions, and spatial attention mechanisms; Feature Enhancement Neck Network: Enhance the feature map through upsampling and multi-layer convolution; Detection head: responsible for outputting the detection results of longitudinal tearing of the conveyor belt.
[0010] The backbone network is mainly composed of Conv+BN+Hardswish, MoV3, and SPFF modules; Conv+BN+Hardswish: a convolutional layer combination used for the input part of the network to improve the network's nonlinear expression ability and training stability; MoV3: consists of three submodules: MoV31, MoV32, and MoV33. Each submodule extracts features through convolution and depthwise separable convolution blocks, and adds Batch Normalization (BN) layers and different activation functions to improve model performance. SPFF: An important module in the network that fuses multi-layer feature maps to improve the detection capability of the model.
[0011] The feature enhancement neck network adopts the dual-tower structure of "FPN+PAN" for feature extraction, integrating features at different levels to achieve the aggregation of feature parameters and the extraction of target position information at different scales; The feature enhancement neck network introduces the BiFPN framework to enrich the semantic information in the deep features; The feature enhancement neck network includes a C3 module, which performs feature processing through multiple Conv+BN+SiLu layers and fuses feature maps through Add and Concat operations to enhance the representation capability of feature maps and the network's perception capability of multi-scale and multi-level features.
[0012] The Conv+BN+Hardswish is a combination of convolutional layers, where Conv is mainly used to extract local features in an image or feature map. The convolution operation slides on the input feature map through a filter and performs a dot product operation to obtain an output feature map, which can effectively capture local spatial structures and information such as edges and textures in the image. Its function is to extract higher-level features by applying the convolution operation to the input feature map; BN standardizes the input of each layer to have zero mean and unit variance, thereby accelerating model convergence and reducing the gradient vanishing / exploding problem. Its main function is to standardize the input data during each training process to keep the activation value distribution of each layer stable and avoid excessive numerical fluctuations in the network. Hardswish is an efficient activation function that uses a simpler calculation process than the Swish activation function during inference, with less computational effort. It is particularly suitable for devices with limited computing resources, such as mobile terminals. Its formula is: .
[0013] The MoV31 module performs dimensionality increase through 1×1 convolution operations, increasing the number of feature map channels to help the network process more complex information. It applies a batch normalization layer for normalization, adds nonlinear factors through the ReLU activation function, and uses a 3×3 depthwise separable convolution for feature extraction. Finally, the MoV31 module inputs the extracted features into the SE module, which adaptively adjusts the weight of each channel. The MoV32 module increases the number of channels based on the MoV31 module to extract more layers of feature information and enhance the diversity of feature extraction. It uses 3×3 depthwise separable convolution to process images and applies the ReLU activation function to reduce accuracy loss. The MoV33 module uses 3×3 depthwise separable convolution to extract features, applies BN and HardSwish activation functions, and also passes the features to the SE module to further optimize the feature expression of each channel, ensuring that the network can more accurately capture the tear information on the conveyor belt.
[0014] The SPFF part includes: Conv+BN+SiLu: Convolutional layers extract features, and the BN layer normalizes the convolution output. SiLu is a new activation function similar to Swish that improves the nonlinear expression ability during training. MaxPool: Use the maximum pooling operation to reduce the size of the feature map while retaining the most important information, reducing the amount of computation and improving the translation invariance of the model; Concat: Concatenates the MaxPool feature map with other feature maps to fuse feature information of different scales, enhance the model's expressiveness, and enable the network to focus on multi-scale information simultaneously. Conv + BN+SiLu: The concatenated feature map is processed again through convolution, BN and SiLu activation functions to further fuse features and improve the performance of the model.
[0015] The improved YOLOv5 model performance verification part uses the accuracy P , recall rate R , average precision mean m AP and frame rate FPS As the evaluation index, m AP is the mean of the average precision of all categories m AP Calculated based on the precision-recall curve.
[0016] The accuracy used to evaluate the performance of the improved YOLOv5 model P , recall rate R And the mean average precision m AP The calculation formula is as follows:
[0017]
[0018]
[0019] in, T P The number of correct positive samples to be predicted; F P is the number of negative samples predicted incorrectly; F N is the number of positive samples predicted incorrectly; AP i is the average precision of category i; n is the total number of categories.
[0020] Using the above solution, the system employs a binocular camera and an improved YOLOv5 deep learning model for real-time conveyor belt monitoring and tear detection. The binocular camera captures visible and infrared images of the conveyor belt. Denoising, filtering, and registration processes produce clear images. The non-subsampled contourlet transform (NSCT) is used to decompose the images into high- and low-frequency subband signals. The image fusion module fuses the subband signals from different frequency bands using maximum entropy threshold segmentation and weighted averaging to generate a fused image. The improved YOLOv5 model is used in the detection module for accurate tear detection. By adding depthwise separable convolution (DWConv) and the SE attention mechanism, the model improves the accuracy of tear feature detection. The system effectively handles factors such as lighting variations, noise interference, and conveyed materials, improving the stability and robustness of detection results. This technology can be widely applied in mining, logistics, manufacturing, and other fields, effectively improving conveyor belt safety, reducing maintenance costs, and ensuring production continuity and stable equipment operation. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 This is the flow chart of the conveyor belt longitudinal tear detection system; Figure 2 This is a block diagram of the non-subsampled contourlet transform fusion process; Figure 3 This is the architecture diagram of the improved YOLOv5 conveyor belt longitudinal tear detection algorithm; Figure 4 This is the MoV3 module diagram; Figure 5 This is a partial structural diagram of SPFF; Figure 6 This is the structural diagram of the C3 part. DETAILED DESCRIPTION
[0022] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0023] See Figure 1 The present invention discloses a conveyor belt longitudinal tear detection system, the system comprising: Original image acquisition module: uses a binocular camera to collect real-time images of the conveyor belt; Original image processing module: performing denoising, filtering and registration processing on the acquired image; Image decomposition module: uses non-subsampled contourlet transform (NSCT) to decompose the image into high-frequency and low-frequency sub-band image signals; Image fusion module: Use the maximum entropy threshold segmentation method to fuse the low-frequency sub-band image signals, and use the weighted average method to fuse the high-frequency sub-band image signals, and finally generate the fused image through the inverse non-subsampled contourlet transform; Detection module: tearing detection based on improved YOLOv5.
[0024] Furthermore, the binocular camera integrates a visible light camera and an infrared camera into one device, and the collected images are used to obtain three-dimensional information of the conveyor belt, thereby enhancing detection accuracy and robustness.
[0025] See Figure 2 , the non-subsampled contourlet transform fusion process block diagram, including: Original image acquisition module: uses the binocular camera to collect real-time images of the conveyor belt, and transmits the collected real-time images to the original image processing module.
[0026] Original image processing module: performs pre-processing operations such as denoising, filtering and registration on the image acquired by the original image acquisition module to facilitate the subsequent non-subsampled contourlet transform fusion process; Image Decomposition Module: The Non-Subsampled Contourlet Transform (NSCT) accurately characterizes the edge signals and texture features of an image. First, the image pre-processed by the original image processing module undergoes an NSCT transform to obtain low-frequency and high-frequency sub-band image signals for the infrared and visible light signals, respectively.
[0027] Image Fusion Module: Low-frequency subband image signals largely contain useful information and display the image's texture features. The quality of the fused image is significantly affected by the low-frequency component. In infrared images, the torn areas are significantly different from the background, and the target information is more prominent. Therefore, the selection of low-frequency subband image signals should primarily come from the infrared image. The maximum entropy threshold segmentation method can relatively completely preserve the image's texture information. This scheme uses the maximum entropy threshold segmentation method to fuse the two low-frequency subband image signals, ultimately obtaining the low-frequency subband signal of the fused image. High-frequency subband signals primarily include image edges and detail information, while visible light images generally provide more scene information and detail information. Therefore, the selection of high-frequency subband signals is generally determined by the visible light signal. This scheme uses a weighted average method to fuse the two high-frequency subband signals.
[0028] Finally, the fused image is obtained by inverse NSCT transformation.
[0029] The present invention adopts a target detection model structure consisting of a backbone network, a neck structure and a detection head, wherein the backbone network is used to extract image features, the neck structure realizes multi-scale feature fusion through feature splicing (Concat) and upsampling (Upsampling), and finally outputs the detection results through multiple two-dimensional convolutional layers (Conv2d). Figure 3 This is the architecture diagram of the improved YOLOv5 conveyor belt longitudinal tear detection algorithm. It consists of three main components: the backbone network (Backbone), the neck structure (Neck), and the detection head (Head). The meanings of the modules in the diagram are as follows: Conv+BN+Activation Function: represents the combination of convolution layer (Convolution), batch normalization (Batch Normalization), and activation function; HardSwish / SiLU: HardSwish and SiLU activation functions, respectively; C3: Residual structure module; Concat: Feature concatenation operation; Upsampling: Upsampling operation; SPPF: Spatial Pyramid PoolingFast; MoV3: Improved MobileNetV3 module; Conv2d: Two-dimensional convolution layer, used to output detection results. The improved YOLOv5 conveyor belt longitudinal tear detection algorithm architecture diagram includes: Backbone network: extracts image features through convolutional layers, depthwise separable convolution, and spatial attention mechanisms. It mainly consists of modules such as Conv+BN+Hardswish, MoV3, and SPFF. Neck (Feature Enhancement Neck Network): Adopts a "dual-tower structure" feature extraction method, combined with FPN and PAN modules, to fuse features at different levels, achieve feature parameter integration and multi-scale target location extraction. The FPN module upsamples deep semantic features and fuses them with shallow features to generate feature maps containing rich semantic information. The PAN module transfers shallow location information to deep layers to improve the model's multi-scale positioning capabilities. By combining the FPN and PAN modules, the prediction layer can simultaneously utilize feature information from the top and bottom layers to complete feature aggregation and multi-scale positioning feature extraction, thereby improving detection accuracy. Head (detection head): responsible for outputting the detection results of longitudinal tearing of the conveyor belt.
[0030] See Figure 4 , MoV3 module diagram, the Backbone part includes: Conv+BN+Hardswish: a convolutional layer combination used for the input part of the network to improve the network's nonlinear expression ability and training stability; MoV3: consists of three submodules (MoV31, MoV32, and MoV33). Each submodule extracts features through convolution and depthwise separable convolution blocks, and adds Batch Normalization (BN) layers and different activation functions to improve model performance. SPFF: An important module in the network that fuses multi-layer feature maps to improve the detection capability of the model.
[0031] Furthermore, the Conv+BN+Hardswish is a common convolutional layer combination, in which Conv is mainly used to extract local features in images or feature maps. The convolution operation slides on the input feature map through a filter (i.e., convolution kernel) to perform dot product operations to obtain an output feature map, which can effectively capture local spatial structures and information such as edges and textures in the image. Its function is to extract higher-level features by applying convolution operations to the input feature map; BN standardizes the input of each layer to have zero mean and unit variance, thereby accelerating model convergence and reducing gradient vanishing / explosion problems. Its main function is to standardize the input data during each training process in order to keep the activation value distribution of each layer stable and avoid excessive numerical fluctuations in the network; Hardswish is an efficient activation function that uses a simpler calculation process than the Swish activation function during inference, with less computational effort, and is particularly suitable for devices with limited computing resources such as mobile terminals. Its formula is: .
[0032] Furthermore, the MoV3 module is the core of this network, performing feature extraction through multiple submodules (MoV31, MoV32, and MoV33). Each submodule uses a different combination of convolution operations and activation functions to improve the network's representation capabilities and computational efficiency. Figure 4 The article presents three MoV3 (Modified MobileNetV3) module structures, each used to construct different layers of the backbone network. Each module consists of the following components: Conv+BN+ReLU (or HardSwish): represents a combination of a convolutional layer, batch normalization, and an activation function (ReLU or HardSwish); DWConv+BN+Activation Function: represents depthwise separable convolution and its subsequent operations; SE Attention: represents the introduction of the Squeeze-and-Excitation Attention module; and HardSwish: is a lightweight activation function used in MobileNetV3. Different variants adapt the network's feature extraction requirements by adjusting the activation function, attention mechanism, or structural order. The MoV31 module performs dimensionality increase through 1×1 convolution operations, increasing the number of feature map channels to help the network process more complex information. The BN layer is applied for standardization to eliminate the distribution bias of different batches of data and improve the stability of training. After the ReLU activation function, nonlinear factors are added to enable the network to learn more complex features. 3×3 depth-wise separable convolution is used for feature extraction. Finally, the MoV31 module inputs the extracted features into the SE module, which adaptively adjusts the weight of each channel to increase the network's attention to important features and enhance the expressiveness of features. The MoV32 module increases the number of channels based on the MoV31 module to extract more layers of feature information and enhance the diversity of feature extraction. It uses 3×3 depthwise separable convolution to process images and applies the ReLU activation function to reduce accuracy loss. The MoV33 module uses 3×3 depthwise separable convolution to extract features, applies BN and HardSwish activation functions, and also passes the features to the SE module to further optimize the feature expression of each channel, ensuring that the network can more accurately capture the tear information on the conveyor belt.
[0033] See Figure 5 After the SPFF part passes through one convolution (Conv+BN+SiLU), it performs three maximum pooling (MaxPool) operations in sequence, and concatenates the pooling results of different scales. Finally, it uses a layer of convolution to fuse the features, thereby realizing the compression and expression of spatial multi-scale features and improving the model's perception of different target sizes.
[0034] Specifically, the SPFF section includes: Conv+BN+SiLu: Convolutional layers extract features, and the BN layer normalizes the convolution output. SiLu (Sigmoid Linear Unit) is a new activation function similar to Swish that improves nonlinear expression capabilities during training. MaxPool: Use the maximum pooling operation to reduce the size of the feature map while retaining the most important information, reducing the amount of computation and improving the translation invariance of the model; Concat: Concatenates the MaxPool feature map with other feature maps to fuse feature information of different scales, enhance the model's expressiveness, and enable the network to focus on multi-scale information simultaneously. Conv+BN+SiLu: The concatenated feature map is processed again through convolution, BN and SiLu activation functions to further fuse features and improve the performance of the model.
[0035] read Figure 6 The C3 module in the improved YOLOv5 model's Neck section performs feature processing through multiple Conv+BN+SiLu layers and fuses feature maps through Add and Concat operations, enhancing the representational power of the feature maps and the network's ability to perceive multi-scale and multi-layer features. Specifically, the C3 module consists of two paths: one that undergoes three convolutions (Conv+BN+SiLU) followed by feature addition (Add); the other that undergoes only a single convolution (Conv+BN+SiLU). The outputs of the two paths are then concatenated (Concat) and finally subjected to a single convolution (Conv+BN+SiLU), completing the C3 module's feature processing. Deep convolutions capture more abstract and semantically rich features in the image, while shallow convolutions preserve basic image details. This combination improves the extraction of longitudinal tear features in conveyor belt images.
[0036] The improved YOLOv5 model performance verification part uses the precision rate (Precision, P ), Recall, R ), mean Average Precision, m AP ) and frame rate (Frames Per Second, FPS ) as the evaluation index. Among them, m AP is the average precision of all categories ( AP ) AP Based on the calculation of the Precision-Recall curve, the accuracy and robustness of the model can be comprehensively evaluated.
[0037] Accuracy P , recall rate R And the mean average precision m AP The calculation formula is as follows:
[0038]
[0039]
[0040] in, T P The number of correct positive samples to be predicted; F P is the number of negative samples predicted incorrectly; F N is the number of positive samples predicted incorrectly; APi is the average precision of category i; n is the total number of categories.
[0041] The present invention's conveyor belt longitudinal tear detection system, based on binocular image fusion, has broad application prospects, particularly in conveyor belt monitoring within industrial automation. As crucial conveying equipment, conveyor belts are widely used in a variety of industries, including mining, logistics, and manufacturing. In these industries, belt damage, particularly longitudinal tears, can lead to production halts, equipment damage, and even accidents. Therefore, timely detection of conveyor belt cracks or tears is crucial for ensuring production continuity, extending equipment life, and ensuring safety.
[0042] Specific application scenarios include but are not limited to: Mining Industry: Conveyor belts are particularly widely used in mines to transport materials such as ore and coal. Tearing of conveyor belts can not only cause production stagnation but also impact safe production in mines. This invention can be used to detect cracks in conveyor belts in real time, issuing timely alarms and reducing accidents.
[0043] Logistics Industry: In logistics centers, conveyor belts are used to sort and transport materials. A tear in a conveyor belt can cause material blockages, delay sorting, and impact logistics efficiency. The detection system of this invention enables automatic monitoring of conveyor belts and early warning of failures.
[0044] Manufacturing: Conveyor belts are often used in automated production lines to transport materials. Cracks or tears in these belts can not only affect production efficiency but also disrupt subsequent processes. This system ensures reliable conveyor belt operation and prevents losses caused by unexpected failures.
[0045] Through the detection system of the present invention, problems can be discovered in time at the early stage of cracks or tears in the conveyor belt, and repairs and replacements can be carried out to avoid equipment shutdown, damage and safety hazards, thereby effectively improving production efficiency, reducing maintenance costs, and ensuring continuous and stable production operation.
[0046] The above description is merely an embodiment of the present invention and does not limit the technical scope of the present invention. Therefore, any minor modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of the present invention are still within the scope of the technical solution of the present invention.
Claims
1. A conveyor belt longitudinal tear detection system based on binocular image fusion, characterized in that: The system includes: Original image acquisition module: uses a binocular camera to collect real-time images of the conveyor belt; Original image processing module: performing denoising, filtering and registration processing on the acquired image; Image decomposition module: uses non-subsampled contourlet transform to decompose the image into high-frequency and low-frequency sub-band signals; Image fusion module: Use the maximum entropy threshold segmentation method to fuse the low-frequency sub-band signals, and use the weighted average method to fuse the high-frequency sub-band signals, and finally generate the fused image through the inverse non-subsampled contourlet transform; Detection module: tearing detection based on improved YOLOv5.
2. A conveyor belt longitudinal tear detection system based on binocular image fusion according to claim 1, characterized in that: The images collected by the binocular camera are used to obtain three-dimensional information of the conveyor belt.
3. The conveyor belt longitudinal tear detection system based on binocular image fusion according to claim 1 is characterized in that: The registration process in the original image processing module includes adjusting the sizes of the two images to be consistent and making the corresponding belt areas the same.
4. The conveyor belt longitudinal tear detection system based on binocular image fusion according to claim 1 is characterized in that: The improved YOLOv5 model includes: Backbone network: extracts image features through convolutional layers, depth-wise separable convolutions, and spatial attention mechanisms; Feature Enhancement Neck Network: Enhance the feature map through upsampling and multi-layer convolution; Detection head: responsible for outputting the detection results of longitudinal tearing of the conveyor belt.
5. A conveyor belt longitudinal tear detection system based on binocular image fusion according to claim 4, characterized in that: The backbone network is mainly composed of Conv+BN+Hardswish, MoV3, and SPFF modules; Conv+BN+Hardswish: a convolutional layer combination used for the input part of the network to improve the network's nonlinear expression ability and training stability; MoV3: consists of three submodules: MoV31, MoV32, and MoV33. Each submodule extracts features through convolution and depthwise separable convolution blocks, and adds Batch Normalization (BN) layers and different activation functions to improve model performance. SPFF: An important module in the network that fuses multi-layer feature maps to improve the detection capability of the model.
6. A conveyor belt longitudinal tear detection system based on binocular image fusion according to claim 4, characterized in that: The feature enhancement neck network adopts the "FPN+PAN" dual-tower structure for feature extraction, integrating features at different levels to achieve the aggregation of feature parameters and the extraction of target position information at different scales; The feature enhancement neck network introduces the BiFPN framework to enrich the semantic information in the deep features; The feature enhancement neck network includes a C3 module, which performs feature processing through multiple Conv+BN+SiLu layers and fuses feature maps through Add and Concat operations to enhance the representation capability of feature maps and the network's perception capability of multi-scale and multi-level features.
7. The conveyor belt longitudinal tear detection system based on binocular image fusion according to claim 5, characterized in that: The Conv+BN+Hardswish is a combination of convolutional layers, where Conv is mainly used to extract local features in an image or feature map. The convolution operation slides on the input feature map through a filter and performs a dot product operation to obtain an output feature map, which can effectively capture local spatial structures and information such as edges and textures in the image. Its function is to extract higher-level features by applying the convolution operation to the input feature map; BN standardizes the input of each layer to have zero mean and unit variance, thereby accelerating model convergence and reducing the gradient vanishing / exploding problem. Its main function is to standardize the input data during each training process to keep the activation value distribution of each layer stable and avoid excessive numerical fluctuations in the network. Hardswish is an efficient activation function that uses a simpler calculation process than the Swish activation function during inference, with less computational effort. It is particularly suitable for devices with limited computing resources, such as mobile terminals. Its formula is: 。 8. The conveyor belt longitudinal tear detection system based on binocular image fusion according to claim 5, characterized in that: The MoV31 module performs dimensionality increase through 1×1 convolution operations, increasing the number of feature map channels to help the network process more complex information. A batch normalization layer is applied for normalization. After the ReLU activation function, nonlinear factors are added and a 3×3 depthwise separable convolution is used for feature extraction. Finally, the MoV31 module inputs the extracted features into the SE module, which adaptively adjusts the weight of each channel. The MoV32 module increases the number of channels based on the MoV31 module to extract more layers of feature information and enhance the diversity of feature extraction. It uses 3×3 depthwise separable convolution to process images and applies the ReLU activation function to reduce accuracy loss. The MoV33 module uses 3×3 depthwise separable convolution to extract features, applies BN and HardSwish activation functions, and also passes the features to the SE module to further optimize the feature expression of each channel, ensuring that the network can more accurately capture the tear information on the conveyor belt.
9. The conveyor belt longitudinal tear detection system based on binocular image fusion according to claim 5, characterized in that: The SPFF part includes: Conv+BN+SiLu: Convolutional layers extract features, and the BN layer normalizes the convolution output. SiLu is a new activation function similar to Swish that improves the nonlinear expression ability during training. MaxPool: Use the maximum pooling operation to reduce the size of the feature map while retaining the most important information, reducing the amount of computation and improving the translation invariance of the model; Concat: Concatenates the MaxPool feature map with other feature maps to fuse feature information of different scales, enhance the model's expressiveness, and enable the network to focus on multi-scale information simultaneously. Conv + BN+SiLu: The concatenated feature map is processed again through convolution, BN and SiLu activation functions to further fuse features and improve the performance of the model.
10. The conveyor belt longitudinal tear detection system based on binocular image fusion according to claim 1, characterized in that: The improved YOLOv5 model performance verification part uses the accuracy P , recall rate R , average precision mean m AP and frame rate FPS As the evaluation index, m AP is the mean of the average precision of all categories m AP Calculated based on the precision-recall curve; The accuracy used to evaluate the performance of the improved YOLOv5 model P , recall rate R And the mean average precision m AP The calculation formula is as follows: in, T P The number of correct positive samples to be predicted; F P is the number of negative samples predicted incorrectly; F N is the number of positive samples predicted incorrectly; AP i is the average precision of category i; n is the total number of categories.