A Tearing Detection Method Based on Encoding-Decoding Neural Network

Image preprocessing and semantic segmentation are performed through linear laser emitters and industrial cameras combined with codec neural networks, which solves the problem of insufficient light underground in coal mines, and achieves high-accuracy tear detection, reducing costs and maintenance difficulties, and improving the adaptability of detection.

CN116485710BActive Publication Date: 2025-07-25CHINA COAL TECH & ENG GRP CHONGQING RES INST CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310158684.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-24
Publication Date
2025-07-25
Estimated Expiration
2043-02-24

AI Technical Summary

Technical Problem

Poor lighting conditions under coal mines lead to poor visible light image quality, affecting the accuracy and generalization ability of tear detection of belt conveyors. The existing technology requires additional fill light equipment to increase cost and maintenance complexity.

Method used

A linear laser emitter and industrial camera are used to combine a codec neural network to perform image preprocessing and semantic segmentation, detect tear external rectangles, and use a codec neural network to perform tear detection, avoid additional fill-up equipment and reduce cost and complexity.

Benefits of technology

It realizes high-accurate tear detection under low light conditions, reduces equipment costs and maintenance difficulties, and improves the generalization ability and adaptability of detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116485710B_ABST
    Figure CN116485710B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of belt conveyor fault detection and relates to a tearing detection method based on an encoder-decoder neural network. It mainly includes a line laser emitter, an industrial camera, and an embedded development board. The line laser emitter emits a line laser, which reaches the lower surface of the belt and is reflected, and then the industrial camera captures the image; the industrial camera transmits the image to the embedded development board, and the embedded development board uses semantic segmentation based on the encoder-decoder neural network to diagnose tearing faults. The main steps of the tearing detection method are divided into image preprocessing, encoder-decoder neural network, and image postprocessing. This solution combines the line structured light technology with the encoder-decoder neural network, solves the problem of poor lighting conditions in coal mines and poor quality of visible light images obtained, does not require setting up additional lighting equipment for supplementary lighting, has a simple structure, low cost, and is convenient for installation and maintenance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of belt conveyor transportation monitoring, and relates to a tearing detection method based on an encoding and decoding neural network. Background Technique

[0002] Belt conveyors are important mechanical equipment in the coal mine transportation process, with the advantages of low transportation energy consumption, long distance, and high efficiency. However, the tearing of the conveyor belt is one of the common faults of belt conveyors. The main reasons for the occurrence of conveyor belt tearing are as follows: 1) Hard foreign objects mixed in the transported coal press, smash, and scratch the conveyor belt; 2) The belt conveyor has design defects or incorrect installation, resulting in other sharp objects damaging the conveyor belt; 3) The conveyor belt runs off track, causing the idler or steel frame to scratch the conveyor belt. When tearing occurs, it will lead to the accumulation of floating coal. At this time, transportation must be stopped and processed in time, otherwise the entire conveyor belt will break, resulting in the spillage of coal materials, damage to transportation equipment, and threatening the lives of underground workers.

[0003] At present, the detection technologies for belt conveyor tearing faults can be divided into three categories: sensor detection, multispectral detection, and computer vision detection. With the rapid development of artificial intelligence and computer hardware, the longitudinal tearing detection method based on computer vision not only has higher detection speed and accuracy than other methods, but also has less installation and maintenance workload. Therefore, using computer vision technology to solve the tearing detection problem is one of the development directions of coal mine intelligentization. This type of method usually relies on industrial cameras to collect tearing image data and transmit it to the processing end for image recognition to determine whether there is tearing and complete the diagnosis of tearing faults.

[0004] However, the computer vision detection method has the following disadvantages: 1) The lighting conditions in the coal mine are poor, and the quality of visible light images obtained by the image acquisition system is poor, resulting in a decrease in the accuracy of subsequent image processing; 2) The image features of tearing are constructed by manual statistics. When the on-site conditions in the coal mine change, the previous tearing features are no longer applicable, and researchers need to re-extract the tearing features, with poor generalization ability. Summary of the Invention

[0005] In view of this, the purpose of the present invention is to provide a tearing detection method based on an encoding and decoding neural network, which solves the problem of poor lighting conditions in the coal mine and poor quality of visible light images obtained, does not require setting additional lighting equipment for supplementary lighting, has a simple structure, low cost, and is convenient for installation and maintenance.

[0006] To achieve the above purpose, the present invention provides the following technical solutions:

[0007] A tearing detection method based on an encoding and decoding neural network, based on an image acquisition system, the image acquisition system includes a line laser emitter, an industrial camera, and an embedded development board; the method includes the following steps:

[0008] S1. The line laser emitter emits a line laser, which irradiates the lower surface of the belt. The industrial camera captures an image I(y, x) and transmits the captured image I(y, x) to the embedded development board for processing;

[0009] S2. Use the image I(y, x) as the input for image preprocessing to obtain the output feature F(x);

[0010] The method of image preprocessing is as follows:

[0011] Traverse x and calculate the first k maximum gray values and the y index values of the maximum values along the y direction: [V, D] = Topk(I(y, x)), where the scales of V and D are both k×w, V represents the first k maximum gray values, D represents the y index values of the first k maximum values, and w is the image width;

[0012] Obtain the key stripe coordinates and stripe gray levels through the weighted average sum algorithm:

[0013] Calculate the size of F(x) through the preset resolution of the industrial camera;

[0014] S3. Use F(x) as the input of the encoding and decoding neural network for semantic segmentation to obtain the output data P;

[0015] S4. Perform post-image processing and calculate the circumscribed rectangle R of the tear in the image according to P obtained in S2;

[0016] The specific method is as follows:

[0017]

[0018] Among them, when P(1, x) ≥ P(0, x), it is determined that x belongs to the torn pixel, and the continuous torn pixels are found to form a torn interval [x s , x e .

[0019] Furthermore, in S1, k is 16, and the image bilinear interpolation algorithm is used to calculate w.

[0020] Furthermore, in S3, the size of P is 2×w.

[0021] Furthermore, in S3, the semantic segmentation includes the following steps:

[0022] Input feature map F 1×w , continuously perform two c×3 convolution operations to obtain the output feature map

[0023] Input feature map First, perform pooling with a stride of 2, and then continuously perform two convolutional operations of 2c×3 to obtain the output feature map

[0024] Input feature map First, perform pooling with a stride of 2, and then continuously perform two convolutional operations of 4c×3 to obtain the output feature map

[0025] Input feature map First, perform pooling with a stride of 2, then continuously perform two convolutional operations of 8c×3, and finally perform a transposed convolutional operation with a stride of 2 and 4c×3 to obtain the output feature map

[0026] Input feature map and Concatenate them along the channels to form a feature map

[0027] Input feature map First, continuously perform two convolutional operations of 4c×3, and then perform a transposed convolutional operation with a stride of 2 and 2c×3 to obtain the output feature map

[0028] Input feature map and Concatenate them along the channels to

[0029] Input feature map First, continuously perform two convolutional operations of 2c×3, and then perform a transposed convolutional operation with a stride of 2 and c×3 to obtain the output feature map

[0030] Input feature map and Concatenate them along the channels to

[0031] Input feature map First, continuously perform two convolutional operations of c×3, and then perform a convolutional operation of 2×3 to obtain the output feature map P 2×w ;

[0032] Among them, c represents the channel parameter

[0033] Furthermore, in the S3, w = 1024 and c = 64. Parameter description: w is the width of the image, and it should be divisible by 8 as much as possible. w should not be too large, as it will consume a large amount of hardware. It is recommended that w = 1024. The image bilinear interpolation algorithm can also be used to adjust w to the appropriate range; c represents the channel parameter, and its value range is between [32, 128]. If it is too large, it will consume a large amount of hardware. It is recommended that c = 64

[0034] The beneficial effects of the present invention are as follows:

[0035] In this solution, the line-structured light image is subjected to neural network operations to detect the circumscribed rectangle of the tear, and then determine whether a belt tear fault has occurred, so as to realize the timely detection and alarm of the tear fault. Compared with the prior art, this solution combines the line-structured light technology with the encoding and decoding neural network, solves the problem of poor lighting conditions in coal mines and poor quality of visible light images obtained, overcomes the technical prejudice of the prior art that additional lighting is required for the phase-structured light technology underground, ensures the accuracy of judgment, does not require setting up additional lighting equipment for supplementary lighting, has a simple structure, low cost, and is convenient for installation and maintenance.

[0036] At the same time, this solution adopts a semantic segmentation network and performs image post-processing, which can complete training with only a small number of samples, has low requirements for the quantity of the data sample library, strong generalization ability, and strong adaptability to changes in the on-site conditions in coal mines; and the model scale of this solution is small, the method calculation is simple, and the requirements for hardware devices are low, further reducing the cost.

[0037] Other advantages, objectives and features of the present invention will be described to some extent in the subsequent specification, and to some extent, will be obvious to those skilled in the art based on the study of the following text, or can be taught from the practice of the present invention. The objectives and other advantages of the present invention can be realized and obtained through the following specification. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be described in detail preferably with reference to the accompanying drawings, where:

[0039] Figure 1 is the system structure diagram of a tear detection method based on an encoding and decoding neural network of the present invention;

[0040] Figure 2 is the schematic diagram of the input image of a tear detection method based on an encoding and decoding neural network of the present invention;

[0041] Figure 3 is the encoding and decoding neural network structure of a tear detection method based on an encoding and decoding neural network of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0042] The following describes the implementation manners of the present invention through specific examples. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the drawings provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.

[0043] Among them, the drawings are only for illustrative purposes, showing only schematic diagrams, not physical diagrams, and should not be construed as a limitation to the present invention; in order to better illustrate the embodiments of the present invention, some components in the drawings will be omitted, enlarged or reduced, which does not represent the size of the actual product; for those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted.

[0044] In the drawings of the embodiments of the present invention, the same or similar reference numerals correspond to the same or similar components; in the description of the present invention, it should be understood that if there are terms such as "upper", "lower", "left", "right", "front", "rear", etc. indicating the orientation or positional relationship, they are based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, the terms describing the positional relationship in the drawings are only for illustrative purposes and should not be construed as a limitation to the present invention. For those of ordinary skill in the art, the specific meanings of the above terms can be understood according to specific circumstances.

[0045] Please refer to Figures 1 to 3 , which is a tearing detection method based on an encoding and decoding neural network. The image acquisition system is composed of a line structured light, mainly including a line laser emitter, an industrial camera, and an embedded development board.

[0046] S1. The line laser emitter emits a line laser to the lower surface of the belt. After reflection, the industrial camera captures an image; the industrial camera transmits the captured image to the embedded development board.

[0047] S2. Image preprocessing

[0048] Taking the image I(y, x) captured by the industrial camera as the output, perform image preprocessing to obtain the output feature F(x).

[0049] Specifically, it includes:

[0050] First, traverse x and calculate the top k maximum grayscale values and the y-index values of the maximum values along the y direction: [V, D] = Topk(I(y, x)), where the scales of V and D are both k×w, V represents the top k maximum grayscale values, D represents the y-index values of the top k maximum values, and w is the image width;

[0051] Then, use the weighted average sum algorithm to obtain the key stripe coordinates and stripe grayscales: x ∈ [0, w - 1];

[0052] Among them, the image width w is obtained from the preset resolution of the industrial camera and is brought into the above formula together with the k value to calculate the size of F(x). In this embodiment, k is preferably 16, and the image width is preferably w = 1024. Then the size of F(x) is w, which is expanded to 1×w.

[0053] S3. Encoder-Decoder Neural Network

[0054] Perform semantic segmentation according to the encoder-decoder neural network structure as Figure 3 shown. The input data for this step is F(x), which enters the encoder-decoder neural network for semantic segmentation, and finally the output data P is obtained. In this embodiment,

[0055] L1 module: Input feature map F 1×w , perform two consecutive c×3 convolution operations to obtain the output feature map

[0056] L2 module: Input feature map First perform pooling with a stride of 2, and then perform two consecutive 2c×3 convolution operations to obtain the output feature map

[0057] L3 module: Input feature map First perform pooling with a stride of 2, and then perform two consecutive 4c×3 convolution operations to obtain the output feature map

[0058] L4 module: Input feature map First perform pooling with a stride of 2, then perform two consecutive 8c×3 convolution operations, and finally perform a transposed convolution operation with a stride of 2 and 4c×3 to obtain the output feature map

[0059] Cat module: There are three modules, and their functions are respectively: (1) Input feature maps and are concatenated along the channels as (2) Input feature maps and are concatenated along the channels as (3) Input feature map and are concatenated by channel to form a feature map

[0060] R3 module: Input feature map First, perform two consecutive convolution operations of 4c×3, and then perform a transposed convolution operation with a stride of 2 and 2c×3 to obtain the output feature map

[0061] R2 module: Input feature map First, perform two consecutive convolution operations of 2c×3, and then perform a transposed convolution operation with a stride of 2 and c×3 to obtain the output feature map

[0062] R1 module: Input feature map First, perform two consecutive convolution operations of c×3, and then perform a convolution operation of 2×3 to obtain the output feature map P 2×w .

[0063] Parameter description: w is the width of the image. It is recommended to ensure that it can be divisible by 8 as much as possible. w should not be too large, as it will consume a large amount of hardware. It is recommended that w = 1024. You can also use the image bilinear interpolation algorithm to adjust w to the appropriate range; c represents the channel parameter, and its value range is between [32, 128]. Too large a value will consume a large amount of hardware. It is recommended that c = 64.

[0064] The key points are:

[0065] Quadrilateral: Represents the input and output data. F is the output of image preprocessing, with a size of 1×w, representing the number of channels 1 and the length w respectively; P is the output of the neural network, with a size of 2×w, representing the number of channels 2 and the length w respectively;

[0066] Rectangle: Represents the operation module. The first row is the input size of the feature map, and the mode is "number of channels × length";

[0067] Conv: Represents the convolution operation. The following number represents the operation size, and the mode is "number of channels × convolution kernel length"; The number of channels determines the number of channels of the output feature map of the convolution operation, and the length of the feature map remains unchanged; After convolution, a nonlinear activation operation Mish is performed by default;

[0068] Maxpool: Represents the max pooling operation. The following number is the operation stride, which shrinks the feature map times, and the number of channels remains unchanged;

[0069] Deconv: Represents the transposed convolution operation. The number following it is the operation size, and the pattern is "number of channels × convolution kernel length". The number of channels determines the number of channels of the output feature map, and the length of the feature map remains unchanged; stride represents the operation step size, and the number 2 following it means that the length of the feature map after the operation is doubled.

[0070] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the present technical solution, and they should all be covered within the scope of the claims of the present invention.

Claims

1. A tearing detection method based on an encoding and decoding neural network, based on an image acquisition system, the image acquisition system comprising a line laser emitter, an industrial camera, and an embedded development board; characterized in that: It includes the following steps: S1. The line laser emitter emits a line laser, which irradiates the lower surface of the belt. The industrial camera acquires the image I(y, x) and transmits the acquired image I(y, x) to the embedded development board for processing; S2. Use the image I(y, x) as the input for image preprocessing to obtain the output feature F(x); The method of image preprocessing is as follows: Traverse x, and calculate the first k maximum gray values and the y index values of the maximum values along the y direction: [V, D]=Topk(I(y, x)), where the scales of V and D are both k×w, V represents the first k maximum gray values, D represents the y index values of the first k maximum values, and w is the image width; Obtain the key fringe coordinates and fringe grayscale through the weighted average sum algorithm: x ∈ [0, w - 1]; Calculate the size of F(x) through the preset resolution of the industrial camera; S3. Use F(x) as the input of the encoding and decoding neural network for semantic segmentation to obtain the output data P; S4. Perform image postprocessing, and calculate the circumscribed rectangle R of the tear in the image according to P obtained in S3; The specific method is as follows: Among them, when P(1, x) ≥ P(0, x), it is determined that x belongs to the torn pixel, and continuous torn pixels are found to form a torn interval [x s , x e .

2. The tear detection method based on an encoding and decoding neural network according to claim 1, wherein: In S2, k is 16, and the image bilinear interpolation algorithm is used to calculate w.

3. The tear detection method based on an encoding and decoding neural network according to claim 1, characterized in that: In S3, the size of P is 2×w.

4. The tear detection method based on an encoding and decoding neural network according to claim 3, wherein: In S3, the semantic segmentation includes the following steps: Input feature map F 1×w , perform two consecutive c×3 convolution operations to obtain the output feature map Input feature map First, perform pooling with a stride of 2, and then continuously perform two convolutional operations of 2c×3 to obtain the output feature map Input feature map First, perform pooling with a stride of 2, and then continuously perform two convolutional operations of 4c×3 to obtain the output feature map Input feature map First, perform pooling with a stride of 2, then continuously perform two convolutional operations of 8c×3, and finally perform a transposed convolutional operation of 4c×3 with a stride of 2 to obtain the output feature map Input feature map and are concatenated into a feature map by channel Input feature map First, perform two consecutive 4c×3 convolution operations, and then perform a transposed convolution operation with a stride of 2 and 2c×3 to obtain the output feature map Input feature map and are concatenated by channel to form Input feature map First, perform two consecutive 2c×3 convolutional operations, and then perform a transposed convolutional operation with a stride of 2 and c×3 to obtain the output feature map Input feature map and are concatenated by channel to form Input feature map First, perform two consecutive c×3 convolution operations, and then perform a 2×3 convolution operation to obtain the output feature map P 2×w ; Where c represents the channel parameter.

5. A tearing detection method based on an encoding and decoding neural network according to claim 4, characterized in that: In S3, w = 1024 and c = 64.

Citation Information

Patent Citations

  • Device and method for detecting longitudinal tearing of large conveyor belt

    CN109335575A

  • Conveyor belt inspection system and method

    US20030168317A1