Belt Tear Detection Method and System Based on Dual Visual State Space Model
By adopting a dual visual state space model in belt tear detection, global and local features are extracted and multi-scale features are fusion, the problems of large information loss, background interference and computing resource occupation in the existing technology are solved, and efficient and accurate early small-target defect detection is achieved.
Patent Information
- Application Number
- CN202411201826.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-29
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2044-08-29
AI Technical Summary
In the detection of small target defects in the early stage of belt tear, the problems of information loss, complex background interference and large computing resource occupation caused by image downsampling, making it difficult to achieve efficient and accurate detection.
The belt tear detection method based on the dual visual state space model is adopted. Through the image serialization module, the dual visual state space belt tear image feature extraction network, the multi-scale feature fusion module and the dual-branch object detection output module, global and local features are extracted, and multi-scale feature fusion is performed to output the final object detection result.
This method avoids information loss caused by image downsampling, improves detection accuracy in complex backgrounds, and reduces the use of computing resources through the state space modeling method of linear complexity, and realizes efficient and accurate detection of small target defects in early belt tear.
Smart Images

Figure CN119006434B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of belt tear detection, and particularly to a belt tear detection method and system based on a dual-vision state space model. Background Art
[0002] In the scenario of industrial automation, the stable operation of the belt conveyor system is a key factor in ensuring the continuity of the production process. However, the belt may be torn due to factors such as material aging, overload, and foreign object jamming. Such tearing phenomena may not only interrupt the production process, but also cause irreversible damage to the equipment and even trigger safety accidents. These tears often become more serious over time. Therefore, timely and effective detection and maintenance of early and minor defects can effectively reduce the occurrence of the above-mentioned equipment damage and safety accidents. Therefore, developing an efficient and accurate belt tear small target defect detection technology to accurately identify and detect early defects on the belt surface is of crucial significance for ensuring production safety and the stable operation of equipment.
[0003] Traditional belt tear small target defect detection methods mainly rely on manual inspections and physical sensors. Although manual inspections are intuitive, for small target defects, due to limitations such as human concentration and eyesight, their efficiency and accuracy are limited, and there are also safety hazards. Although physical sensors can, in principle, monitor the belt in real time, due to strict requirements for installation positions and environments, they require high maintenance costs. Moreover, if the sensor signal processing is not accurate enough, environmental disturbances will seriously affect the detection of early small target defects.
[0004] In recent years, with the rapid development of deep learning technology, belt tear small target defect detection methods based on image processing technology have gradually been widely studied and applied. Deep learning models can automatically extract rich semantic features from images through training and use these features to achieve various downstream tasks, such as object detection and semantic segmentation. However, current mainstream deep learning algorithms still have limitations in the detection of early small target defects in high-resolution belt tears:
[0005] Firstly, in terms of small target feature extraction, existing methods often downsample the belt tear image, resulting in serious loss of image information, especially the loss of local small target information during the downsampling process;
[0006] Secondly, affected by the operating scenario of the belt, the collected belt tear images often have complex image backgrounds, which contain a large number of background objects similar to early small target defects, such as scratches, textures, water stains, and spots. These background interferences lead to a high missed detection rate in small target detection;
[0007] Finally, when applying the early small target detection technology for belt tearing to real-time online detection, the consumption of computing resources and the inference efficiency of the model are also important considerations. When the existing Transformer models are used for long sequence modeling, due to the limitations of the computational principle of the core self-attention mechanism, there are two major problems: large computational resource occupation and high computational complexity during model inference, making it difficult to be implemented. Summary of the Invention
[0008] The present invention provides a belt tearing detection method and system based on a dual visual state space model to solve the problems existing in the above-mentioned prior art. The technical solutions are as follows:
[0009] On the one hand, a belt tearing detection method based on a dual visual state space model is provided, including:
[0010] S1. Obtain the surface image of the belt to be detected;
[0011] S2. Input the surface image of the belt to be detected into the trained backbone network for detecting belt tearing image defects, and detect and output the target detection result. The backbone network for detecting belt tearing image defects is composed of an image serialization module, a dual visual state space belt tearing image feature extraction network, a multi-scale feature fusion module, and a dual-branch target detection output module;
[0012] Among them, the image serialization module serializes the surface image of the belt to be detected and outputs an image sequence and an image micro-sequence;
[0013] The dual visual state space belt tearing image feature extraction network uses a dual visual state space feature extraction module to extract the global semantic category information and local detail information of small target defects in belt tearing;
[0014] The multi-scale feature fusion module uses a feature pyramid structure based on the visual state space feature extraction module to perform multi-scale feature fusion and obtain image fusion features of three sizes;
[0015] The dual-branch target detection output module uses two branches, utilizes the image fusion features of the three sizes, and respectively performs target recognition and classification tasks, and outputs the final target detection result.
[0016] Optionally, the image serialization module first cuts the surface image of the belt to be detected into multiple image blocks;
[0017] Then, each image block is locally cut into multiple local image blocks;
[0018] Then, using the convolutional downsampling operation, the cut image patches and local image patches are respectively converted into two image sequence pairs, namely an image sequence and an image micro-sequence, both of which are discrete two-dimensional image sequences.
[0019] Optionally, the dual-visual-state-space belt tear image feature extraction network uses the dual-visual-state-space feature extraction module to extract the global semantic category information and local detail information of the small target defects of belt tears in four stages. In the four stages, N1, N2, N3, and N4 dual-visual-state-space feature extraction modules are cascaded. In different stages, convolution with a stride of 2 is used for downsampling to reduce the size of the image features, and image fusion features F1, F2, F3, and F4 of four sizes are obtained respectively. The sizes are in a relationship of two-fold downsampling in sequence. Among them, F1 contains more noise information and is not used to generate the subsequent target detection results; through continuous parameter iteration in the training stage, downsampling effectively filters out the invalid information in the feature map of the previous stage, reduces the computational amount, and does not lose the effective features of the image.
[0020] The local detail information extracted by the dual-visual-state-space feature extraction module does not directly participate in the output of the target detection result. Instead, in the dual-visual-state-space feature extraction module, through the method of feature fusion, the local detail information is integrated into the image fusion feature, indirectly playing a role in accurately detecting the output result.
[0021] Optionally, the first dual-visual-state-space feature extraction module in the dual-visual-state-space feature extraction module takes the image sequence and the image micro-sequence as inputs. First, it uses a local visual state space branch composed of visual state space feature extraction modules to extract the local image features inside the image micro-sequence, and performs feature fusion with the input image sequence by addition. Then, the fused sequence features are sent to another global visual state space branch composed of visual state space feature extraction modules to extract the long-distance global semantic category information between the image sequences, and at the same time model the relationship between the global features of the image sequence and the local features of the image micro-sequence to obtain the image fusion feature. The two parts, namely the local image features extracted by the local visual state space branch and the image fusion features extracted by the global visual state space branch, are sent to the cascaded second dual-visual-state-space feature extraction module to continue extracting deeper image features.
[0022] Optionally, the visual state space feature extraction module is responsible for modeling the discrete two-dimensional image sequence to obtain the relationship between sequences. When the discrete two-dimensional image sequence is used as the input, in addition to the front and back sequences, there are positional relationships in the four directions of up, down, left, and right among the image sequences. According to the cross-scanning methods in four different directions, all the image matrices composed of several image sequences are rearranged to form four long image sequences composed of image sequences. Then, each long image sequence is respectively sent into the visual state space feature extraction module to extract the corresponding semantic information in the sequence, obtaining four feature sequences. The four feature sequences are added and averaged to obtain a new image matrix, and then the new image matrix is flattened into a one-dimensional vector to obtain the serialized image vector;
[0023] Among them, the visual state space feature extraction module sequentially extracts image features from each long image sequence through layer normalization, linear layer, activation function, depthwise separable convolution, discrete state space model, and layer normalization operations, and then multiplies with the weight matrix obtained by passing through the linear layer and activation function in another branch of the long image sequence. Then, the result after multiplication passes through a linear layer to obtain the image feature sequence. A residual connection is introduced, and the image feature sequence is added to the long image sequence to obtain the feature sequence.
[0024] Optionally, the discrete state space model takes pixel intensity, edge, texture, color histogram, object position, shape, or motion as states, and these states are quantized into discrete values to form a state space. An intermediate implicit state is introduced to extract features in the image, and the complex dynamics and semantic relationships in the image sequence are accurately modeled through the discretized state transition and observation processes.
[0025] Optionally, the multi-scale feature fusion module layer-by-layer fuses the features of the current layer and the previous layer through two paths of top-down and bottom-up to generate a new feature map, and uses the feature pyramid structure based on the visual state space feature extraction module to fuse and extract the feature map to generate a new feature map;
[0026] Among them, top-down means that the image fusion feature F4 with the smallest size is upsampled to expand the resolution to be the same as F3 and then concatenated with F3. The concatenated feature is sent into the visual state space feature extraction module to fuse the features of two different scales. The fused feature is then upsampled and concatenated with the larger-sized F2, and the concatenated feature is sent into the visual state space feature extraction module to fuse the features of two different scales to obtain the multi-scale fusion feature P2;
[0027] For bottom-up fusion, downsample P2 to the same size as F3, then concatenate it with the corresponding features in the top-down fusion, and send them into the visual state space feature extraction module to fuse features of two different scales, obtaining the multi-scale fusion feature P3;
[0028] Downsample P3 to the same size as F4, then concatenate it with the corresponding features in the top-down fusion, and send them into the visual state space feature extraction module to fuse features of two different scales, obtaining the multi-scale fusion feature P4.
[0029] Optionally, the dual-branch object detection output module upsamples the multi-scale fusion features P2, P3, and P4 to the original image size, concatenates them, and then connects a 1×1 convolution to adjust the feature map dimension;
[0030] Then, the classification task branch consists of 2 3×3 convolutions and 1 1×1 convolution. Assuming the image size is H×W and C is the number of object detection categories, it outputs a classification result of H×W×C;
[0031] The object recognition branch consists of 2 3×3 convolutions and 2 1×1 convolutions. The 2 1×1 convolutions are respectively used to output the position detection result of H×W×4 and the intersection over union score of H×W×1;
[0032] And finally, it outputs the object detection result in the format of (c o ,x o ,y o ,h o ,w o ), where the subscript o represents the output, c represents the defect category number, (x,y) represents the center point coordinates of the object detection box, and (h,w) represents the height and width of the object detection box.
[0033] Optionally, the method splits the small target detection task of belt tearing into two tasks: a classification task and a rectangular box regression task. The loss function consists of a classification task loss function and a rectangular box regression task loss function. The classification task loss function is the binary cross-entropy loss function, and the rectangular box regression task loss function consists of a distribution focal loss function and a complete intersection over union loss function. These three loss functions are weighted proportionally to form the loss function, which is used to supervise the learning process of the model;
[0034] Among them, the binary cross-entropy loss function judges the quality of the classification model prediction result by calculating the difference between the class prediction probability and the class label. The specific formula is as follows:
[0035]
[0036] Among them, y icDenote the c-th label of the i-th sample output by the model, p ic Denote the probability that the output belongs to the label, and N denotes the number of groups of objects predicted by the model;
[0037] The said distribution focal loss function can be used to make the network quickly focus on the values near the label, make the probability density at the label as large as possible, and the calculation is based on KL divergence, which is used to compare the difference between the predicted distribution and the true distribution. For each predicted bounding box, the true class and the predicted class are respectively encoded as distribution vectors, and the KL divergence between the two distribution vectors is calculated to judge the quality of the model output. Suppose the distribution probabilities of the predicted bounding box positions are p0, p1, …, p n , then for the position y of each true bounding box, find the two predicted bounding box positions y i and y i+1 that are closest to y, and then calculate the losses corresponding to the probabilities p i and p i+1 at these two positions. By accumulating the losses at all these positions, the final DFL loss is obtained, and the specific formula is as follows:
[0038] DFL_Loss(p i , p i+1 ) = -((y i+1 - y) log(p i ) + (y - y i ) log(p i+1 ))
[0039] where p i and p i+1 are the distribution probabilities of the two predicted bounding box positions output by the model, y is the position of the true bounding box, and y i and y i+1 represent the predicted bounding box positions closest to y;
[0040] Intersection over Union (IoU) is a way to describe the overlap degree between the predicted bounding box and the true bounding box in object detection. The regression degree of the box is measured by the IoU between the predicted bounding box and the true bounding box. The complete IoU loss function is based on this principle and adds two measurement criteria: the consistency of the center point distance and the aspect ratio. Using the relative position information and shape difference between the predicted bounding box and the detected bounding box to judge the fitting situation of the model output rectangular box, the specific formula is as follows:
[0041]
[0042] where ρ 2 (b, b gt) represents the Euclidean distance between the center points of the predicted bounding box and the ground truth bounding box. d represents the length of the diagonal of the smallest closed region that can simultaneously contain the predicted bounding box and the ground truth bounding box. w and h respectively represent the width and height of the predicted bounding box, and w gt and h gt respectively represent the width and height of the ground truth bounding box;
[0043] Finally, the total loss function is the weighted sum of the above three loss functions, and the specific formula is as follows:
[0044] Loss=w bce BCE_Loss+w dfl DFL_Loss+w ciou CIoU_Loss
[0045] Among them, w bce 、w dfl 、w ciou respectively represent the corresponding weights.
[0046] On the other hand, a belt tear detection system based on a dual visual state space model is provided. The system includes:
[0047] An acquisition module for acquiring an image of the surface of the belt to be detected;
[0048] A detection module for inputting the image of the surface of the belt to be detected into a trained backbone network for detecting belt tear image defects, and detecting and outputting a target detection result. The backbone network for detecting belt tear image defects is composed of an image serialization module, a dual visual state space belt tear image feature extraction network, a multi-scale feature fusion module, and a dual-branch target detection output module;
[0049] Among them, the image serialization module serializes the image of the surface of the belt to be detected and outputs an image sequence and an image micro-sequence;
[0050] The dual visual state space belt tear image feature extraction network uses a dual visual state space feature extraction module to extract the global semantic category information and local detail information of small target defects of belt tears;
[0051] The multi-scale feature fusion module uses a feature pyramid structure based on a visual state space feature extraction module to perform multi-scale feature fusion and obtain image fusion features of three sizes;
[0052] The dual-branch target detection output module uses two branches, utilizes the image fusion features of the three sizes, respectively performs target recognition and classification tasks, and outputs the final target detection result.
[0053] On the other hand, an electronic device is provided. The electronic device includes a processor and a memory. At least one instruction is stored in the memory, and the at least one instruction is loaded and executed by the processor to implement the above-mentioned belt tear detection method based on the dual visual state space model.
[0054] On the other hand, a computer-readable storage medium is provided. At least one instruction is stored in the storage medium, and the at least one instruction is loaded and executed by a processor to implement the above-mentioned belt tear detection method based on the dual visual state space model.
[0055] The beneficial effects brought by the technical solution provided by the present invention at least include:
[0056] By improving the deep learning framework of the visual state space model, the present invention proposes a belt tear small target defect detection technology based on the dual visual state space model. This technology omits the image downsampling process, avoids information loss in the processing of high-resolution belt tear images, and improves the detection accuracy of the model for early small target defects in complex belt backgrounds through feature extraction means such as global and local feature fusion and multi-scale feature fusion. Moreover, the state space modeling method with linear complexity is used to ensure high inference efficiency. That is, the present invention comprehensively considers the deficiencies of the prior art in three aspects: small target feature extraction, anti-interference, and implementation and deployment, and provides a more accurate and efficient method and system for detecting early small target defects of belt tears, which is applicable to multiple fields such as industrial manufacturing, mine transportation, power production, and port logistics. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0058] Figure 1 is a flowchart of a belt tear detection method based on the dual visual state space model provided by an embodiment of the present invention;
[0059] Figure 2 is the overall flowchart of the belt tear detection method provided by an embodiment of the present invention;
[0060] Figure 3 is a flowchart of dataset production provided by an embodiment of the present invention;
[0061] Figure 4 is a structural diagram of an image serialization module and a dual visual state space belt tear image feature extraction network provided by an embodiment of the present invention;
[0062] Figure 5 It is the structural diagram of the dual visual state space feature extraction module provided by the embodiments of the present invention;
[0063] Figure 6 It is the schematic diagram of the visual state space branch scanning provided by the embodiments of the present invention, with local branches (upper - image micro - sequence / local feature sequence) and global branches (lower - image sequence / fusion feature sequence);
[0064] Figure 7 It is the structural diagram of the visual state space feature extraction module provided by the embodiments of the present invention;
[0065] Figure 8 It is the structural diagram of the multi - scale feature fusion of the feature pyramid structure based on the visual state space feature extraction module provided by the embodiments of the present invention;
[0066] Figure 9 It is the structural diagram of the dual - branch target detection output module provided by the embodiments of the present invention;
[0067] Figure 10 It is the block diagram of the belt tear detection system based on the dual visual state space model provided by the embodiments of the present invention;
[0068] Figure 11 It is the schematic diagram of the structure of an electronic device provided by the embodiments of the present invention. Detailed implementation manners
[0069] To make the technical problems, technical solutions and advantages to be solved by the present invention clearer, the following will be described in detail with reference to the accompanying drawings and specific embodiments.
[0070] The embodiments of the present invention provide a belt tear detection method based on a dual visual state space model. This method can be implemented by an electronic device, which can be a terminal or a server. As Figure 1 shown in the flowchart of a belt tear detection method based on a dual visual state space model, the processing flow of this method can include the following steps:
[0071] S1. Obtain the surface image of the belt to be detected;
[0072] S2. Input the surface image of the belt to be detected into the trained backbone network for belt tear image defect detection, and detect and output the target detection result. The backbone network for belt tear image defect detection is composed of an image serialization module, a dual visual state space belt tear image feature extraction network, a multi - scale feature fusion module, and a dual - branch target detection output module;
[0073] Among them, the image serialization module serializes the surface image of the belt to be detected and outputs an image sequence and an image micro - sequence;
[0074] The dual-vision state space belt tear image feature extraction network uses a dual-vision state space feature extraction module to extract the global semantic category information and local detail information of small target defects of belt tears.
[0075] The multi-scale feature fusion module uses a feature pyramid structure based on the vision state space feature extraction module to perform multi-scale feature fusion and obtain image fusion features of three sizes.
[0076] The dual-branch object detection output module uses two branches, utilizes the image fusion features of the three sizes, respectively performs object recognition and classification tasks, and outputs the final object detection result.
[0077] The embodiment of the present invention provides a belt tear detection method based on a dual-vision state space model. First, aiming at the problem of missing features of small target defects caused by image downsampling in the prior art, the technical solution of the embodiment of the present invention removes the downsampling operation to retain the original features of the image, and designs a local state space feature extraction branch, so that the model can more effectively extract local detail features crucial for small target defect detection; Second, aiming at the problem of background target interference in complex backgrounds, the technical solution of the embodiment of the present invention uses a global vision state space branch to model the global information between long image sequences, fuse the above local features, establish the relationship between global and local features, and combine the feature pyramid structure based on the vision state space feature extraction module to perform multi-scale feature fusion to enhance the model's image content understanding ability; Finally, aiming at the problem of difficult landing deployment due to computational complexity and computational efficiency in the prior art, the technical solution of the embodiment of the present invention utilizes the linear computational complexity and parallel computing ability of the state space model to make the model meet the requirements of actual applications in terms of real-time performance and resource occupancy, as Figure 2 shown, the following is the complete technical solution:
[0078] S1. Obtain an image of the surface of the belt to be detected;
[0079] In the belt working scenario of the embodiment of the present invention, a 2048-resolution linear array camera (the acquired image resolution is 2048×1) is used to capture images of the running belt, and 2000 consecutive frames are taken to form a high-resolution belt tear image to be detected with a size of 2048×2000.
[0080] S2. Input the image of the surface of the belt to be detected into the trained main network for detecting belt tear image defects, and detect and output the object detection result. The main network for detecting belt tear image defects is composed of an image serialization module, a dual-vision state space belt tear image feature extraction network, a multi-scale feature fusion module, and a dual-branch object detection output module;
[0081] Among them, the image serialization module serializes the surface image of the belt to be detected, and outputs an image sequence and an image micro-sequence;
[0082] The dual-visual-state-space belt tear image feature extraction network uses a dual-visual-state-space feature extraction module to extract the global semantic category information and local detail information of small target defects of belt tears;
[0083] The multi-scale feature fusion module uses a feature pyramid structure based on the visual state space feature extraction module to perform multi-scale feature fusion and obtain image fusion features of three sizes;
[0084] The dual-branch object detection output module uses two branches, utilizes the image fusion features of the three sizes, performs object recognition and classification tasks respectively, and outputs the final object detection result.
[0085] The training of the belt tear image defect detection backbone network described in the embodiments of the present invention is carried out using the image data in the training set, as Figure 3 shown, including:
[0086] 1) Belt tear image data collection.
[0087] In the belt working scenario, a 2048-resolution line array camera (the acquired image resolution is 2048×1) is used to capture the images of the running belt, and 2000 consecutive frames are taken to form a belt tear image with a high resolution of 2048×2000. Multiple 2048×2000 high-resolution belt tear images are formed in this way, and the defective and non-defective images among them are screened and a data set is constructed according to a certain ratio.
[0088] 2) Belt tear image data annotation.
[0089] The defects in the collected images are divided into C different defect categories according to different shape texture features and severity levels. The label img tool is used to annotate the belt images, and the annotation format uses the YOLO object detection annotation format (c t ,x t ,y t ,h t ,w t ), where the subscript t represents the label truth value, c represents the defect category number, the range is [0, C), (x, y) represents the center point coordinates of the object detection box, (h, w) represents the height and width of the object detection box, x±w / 2 should be within the range of 0 to 2048, and y±h / 2 should be within the range of 0 to 2000, and the output is stored in txt format.
[0090] 3) Belt tear detection data set division.
[0091] The image-label pairs with and without defects are respectively divided into a training set and a validation set according to a ratio of 4:1. The two parts with and without defects are combined together to form the training set and the validation set of the final belt tear detection data set. Among them, the training set is used for model training, and the validation set is used to evaluate the performance during the model training process. Finally, the optimal weights on the validation set are selected for model inference and application.
[0092] 4) Image preprocessing of the belt tear detection data set.
[0093] All image-label pairs are subjected to data augmentation by means of random horizontal mirror flipping, random scale scaling, random size cropping, contrast enhancement, etc. to expand the training data set, and the input images and labels are obtained. Moreover, additional data augmentation is performed for some defects with a small sample size to make the quantity balanced.
[0094] In order to extract global features and local features subsequently (the visual state space feature extraction module can only process serialized data), the embodiments of the present invention need to perform a slicing serialization operation on the images, and use a convolutional downsampling operation to convert the sliced image blocks and local image blocks into an image sequence and an image micro-sequence respectively, and send the generated pair of image sequences into the dual visual state space belt tear image feature extraction network.
[0095] Optionally, as Figure 4 shown, the image serialization module first slices the surface image of the belt to be detected to obtain a plurality of image blocks;
[0096] Then, each image block is locally sliced to obtain a plurality of local image blocks;
[0097] Then, using the convolutional downsampling operation, the sliced image blocks and local image blocks are respectively converted into two image sequence pairs, namely an image sequence and an image micro-sequence, and both the image sequence and the image micro-sequence are discrete two-dimensional image sequences.
[0098] Optionally, as Figure 4As shown, the dual-visual-state-space belt tear image feature extraction network uses the dual-visual-state-space feature extraction module to extract the global semantic category information and local detail information of small target defects in belt tears in four stages. In the four stages, N1, N2, N3, and N4 dual-visual-state-space feature extraction modules are cascaded. In different stages, convolution with a stride of 2 is used for downsampling to reduce the size of image features, and image fusion features F1, F2, F3, and F4 of four sizes (image local features L1, L2, L3, and L4 of four sizes will also be obtained) are obtained respectively. The sizes are in a relationship of being downsampled by a factor of two in sequence. Among them, F1 contains more noise information and is not used to generate the subsequent target detection results; through continuous parameter iteration in the training stage, downsampling effectively filters out the invalid information in the feature map of the previous stage, reduces the computational amount, and does not lose the effective features of the image.
[0099] The local detail information extracted by the dual-visual-state-space feature extraction module does not directly participate in the output of the target detection result. Instead, in the dual-visual-state-space feature extraction module, through the method of feature fusion, the local detail information is incorporated into the image fusion feature, indirectly playing the role of accurately detecting the output result.
[0100] The state space model is a mathematical model that describes the behavior of a dynamic system. It uses a set of first-order differential equations (continuous-time system) or difference equations (discrete-time system) to represent the evolution of the internal state of the system, and another set of equations to describe the relationship between the system state and the output. The state space model has been widely used in the visual field due to its high model efficiency, less graphics processor memory usage, and better long-range dependence modeling ability. In the field of image processing, an image sequence (feature sequence) is regarded as different states, and the model extracts image feature information by learning the transformation relationship between these states. However, directly applying the state space model to the detection of early small target defects in belt tears will lead to inaccurate positioning or even missed detection because it explores less local features (which are very important for the detection of early small target defects). In the embodiments of the present invention, in order to utilize the advantages of the state space model in the field of detecting small target defects in belt tears, a dual-visual-state-space feature extraction module is proposed to replace the visual state space module in the current visual state space model, enabling it to utilize the long-range modeling ability of the state space model to obtain the global features between image sequences while also obtaining the local detail features between pixels within the image sequence.
[0101] Optionally, the first dual-visual-state-space feature extraction module in the dual-visual-state-space feature extraction module, such as Figure 5As shown, taking the image sequence and the image micro-sequence as inputs, first, a local visual state space branch composed of a visual state space feature extraction module is used to extract the local image features within the image micro-sequence, and perform feature fusion with the input image sequence by addition. Then, the fused sequence features are fed into another global visual state space branch composed of a visual state space feature extraction module to extract the long-distance global semantic category information between image sequences, and simultaneously model the relationship between the global features of the image sequence and the local features of the image micro-sequence to obtain image fusion features. The two parts, namely the local image features extracted by the local visual state space branch and the image fusion features extracted by the global visual state space branch, are fed into the second cascaded dual visual state space feature extraction module to continue extracting deeper image features.
[0102] Optionally, the visual state space feature extraction module is responsible for modeling the discrete two-dimensional image sequence to obtain the relationship between sequences. When taking the discrete two-dimensional image sequence as input, in addition to the front and back sequences, there are positional relationships in the four directions of up, down, left, and right between image sequences, as Figure 6 shown. In the embodiment of the present invention, according to the cross-scanning method in four different directions (when state space modeling was initially used in natural language image processing, there was only temporal information between the front and back sequences, so it was only necessary to scan once from front to back to model the semantics), all image matrices composed of several image sequences are rearranged to form four long image sequences composed of image sequences. Then, each long image sequence is respectively fed into the visual state space feature extraction module to extract the corresponding semantic information in the sequence, obtaining four feature sequences. The four feature sequences are added and averaged to obtain a new image matrix, and then the new image matrix is flattened into a one-dimensional vector to obtain a serialized image vector;
[0103] By adopting this complementary traversal path, the model can integrate the information of each pixel in the image from all other pixels in different directions.
[0104] Among them, as Figure 7 shown, the visual state space feature extraction module extracts image features by passing each long image sequence through layer normalization, a linear layer, an activation function, depthwise separable convolution, a discrete state space model, and layer normalization operations in sequence, and then multiplies with the weight matrix obtained from another branch of the long image sequence through a linear layer and an activation function. Then, the multiplied result passes through a linear layer to obtain an image feature sequence. A residual connection is introduced (to increase the stability of model training), and the image feature sequence is added to the long image sequence to obtain the feature sequence.
[0105] The state space model is usually applied to continuous systems. In the small target detection task of belt tearing, the data obtained through the above steps are discrete serialized images or features. Therefore, the embodiments of the present invention use a discrete state space model to model the image sequence.
[0106] Optionally, in the discrete state space model, pixel intensity, edge, texture, color histogram, object position, shape, or motion is used as the state. These states are quantized into discrete values to form a state space, and intermediate implicit states are introduced to extract features from the image. Through the discretized state transition and observation processes, the complex dynamics and semantic relationships in the image sequence are accurately modeled.
[0107] Compared with the existing modeling method based on the self-attention mechanism, the discrete state space modeling also has the advantage of linear complexity. Combined with the parallel computing design, it can achieve high accuracy while having a fast inference speed.
[0108] Optionally, as Figure 8 shown, the multi-scale feature fusion module fuses the features of the current layer and the previous layer layer by layer through two paths: top-down and bottom-up to generate a new feature map, and uses the feature pyramid structure based on the visual state space feature extraction module (different from the design of using convolution in the traditional pyramid feature fusion module) to fuse and extract the feature map to generate a new feature map;
[0109] Among them, top-down means that the image fusion feature F4 with the smallest size is upsampled to expand the resolution, and after being the same as F3, it is concatenated with F3. The concatenated feature is sent to the visual state space feature extraction module to fuse two different scales of features. The fused feature is then upsampled and concatenated with the larger-sized F2, and the concatenated feature is sent to the visual state space feature extraction module to fuse two different scales of features to obtain the multi-scale fusion feature P2;
[0110] For bottom-up fusion, P2 is downsampled to the same size as F3, and then concatenated with the corresponding feature in the top-down fusion and sent to the visual state space feature extraction module to fuse two different scales of features to obtain the multi-scale fusion feature P3;
[0111] P3 is downsampled to the same size as F4, and then concatenated with the corresponding feature in the top-down fusion and sent to the visual state space feature extraction module to fuse two different scales of features to obtain the multi-scale fusion feature P4.
[0112] Optionally, as Figure 9 shown, the dual-branch object detection output module upsamples the multi-scale fusion features P2, P3, and P4 to the original image size, concatenates them, and then connects a 1×1 convolution to adjust the dimension of the feature map;
[0113] Then, the classification task branch consists of 2 3×3 convolutions and 1 1×1 convolution. Assuming the image size is H×W and C is the number of object detection categories, the classification result of H×W×C is output;
[0114] The object recognition branch consists of 2 3×3 convolutions and 2 1×1 convolutions. The 2 1×1 convolutions are used to output the position detection result of H×W×4 and the intersection over union score of H×W×1 respectively;
[0115] And finally, the object detection result in the format of (c o ,x o ,y o ,h o ,w o ) is output, where the subscript o represents the output, c represents the defect category number, (x,y) represents the center point coordinates of the object detection box, and (h,w) represents the height and width of the object detection box.
[0116] This design allows each part of the model to focus on its specific task, which can improve the accuracy of classification and the precision of detection
[0117] Optionally, the method splits the small object detection task of belt tearing into two tasks: a classification task and a rectangle box regression task. The loss function consists of a classification task loss function and a rectangle box regression task loss function. The classification task loss function is a binary cross-entropy loss function, and the rectangle box regression task loss function consists of a distribution focal loss function and a complete intersection over union loss function. These three loss functions are weighted proportionally to form the loss function, which is used to supervise the learning process of the model;
[0118] Among them, the binary cross-entropy loss function judges the quality of the classification model prediction result by calculating the difference between the class prediction probability and the class label (for the case where the label is 1, if the predicted value approaches 1, then the value of the loss function should approach 0; conversely, if the predicted value approaches 0 at this time, then the value of the loss function is very large). The specific formula is as follows:
[0119]
[0120] Among them, y ic represents the c-th label of the i-th sample output by the model, p ic represents the probability that the output belongs to the label, and N represents the number of groups of model prediction objects;
[0121] The distribution focal loss function can be used to make the network quickly focus on the values near the label, maximize the probability density at the label, and is calculated based on the KL divergence, which is used to compare the difference between the predicted distribution and the true distribution. For each predicted bounding box, the true class and the predicted class are respectively encoded as distribution vectors, and the KL divergence between the two distribution vectors is calculated to judge the quality of the module output. Assume that the distribution probabilities of the predicted bounding box positions are p0, p1, …, p n , then for the position y of each true bounding box, find the two predicted bounding box positions y i and y i+1 that are closest to y, and then calculate the losses of the probabilities p i and p i+1 corresponding to these two positions. By accumulating the losses of all these positions, the final DFL loss is obtained. The specific formula is as follows:
[0122] DFL_Loss(p i , p i+1 ) = -((y i+1 - y) log(p i ) + (y - y i ) log(p i+1 ))
[0123] where p i and p i+1 are the distribution probabilities of the two predicted bounding box positions output by the model, y is the position of the true bounding box, and y i and y i+1 represent the predicted bounding box positions closest to y;
[0124] The intersection over union is a way to describe the overlap degree between the predicted bounding box and the true bounding box in object detection. The regression degree of the box is measured by the intersection over union of the predicted bounding box and the true bounding box. The complete intersection over union loss function is based on this principle and adds two measurement criteria: the consistency of the center point distance and the aspect ratio. It uses the relative position information and shape difference between the predicted bounding box and the detected bounding box to judge the fitting situation of the model output rectangular box. The specific formula is as follows:
[0125]
[0126] where ρ 2 (b, b gt ) represents the Euclidean distance between the center points of the predicted bounding box and the true bounding box, d represents the length of the diagonal of the smallest closed region that can simultaneously contain the predicted bounding box and the true bounding box, w and h respectively represent the width and height of the predicted bounding box, and w gt and h gt respectively represent the width and height of the true bounding box;
[0127] Finally, the total loss function is the weighted sum of the above three loss functions, and the specific formula is as follows:
[0128] Loss=w bce BCE_Loss+w dfl DFL_Loss+w ciou CIoU_Loss
[0129] Among them, w bce 、w dfl 、w ciou respectively represent the corresponding weights (range 0-1).
[0130] As Figure 10 shown, the embodiment of the present invention also provides a belt tear detection system based on a dual visual state space model, and the system includes:
[0131] An acquisition module 1010, configured to acquire an image of the surface of the belt to be detected;
[0132] A detection module 1020, configured to input the image of the surface of the belt to be detected into a trained backbone network for detecting belt tear image defects, and detect and output a target detection result. The backbone network for detecting belt tear image defects is composed of an image serialization module 10201, a dual visual state space belt tear image feature extraction network 10202, a multi-scale feature fusion module 10203, and a dual-branch target detection output module 10204;
[0133] Among them, the image serialization module 10201 serializes the image of the surface of the belt to be detected and outputs an image sequence and an image micro-sequence;
[0134] The dual visual state space belt tear image feature extraction network 10202 uses a dual visual state space feature extraction module to extract the global semantic category information and local detail information of small target defects of belt tears;
[0135] The multi-scale feature fusion module 10203 uses a feature pyramid structure based on a visual state space feature extraction module to perform multi-scale feature fusion and obtain image fusion features of three sizes;
[0136] The dual-branch target detection output module 10204 uses two branches, utilizes the image fusion features of the three sizes, respectively performs target recognition and classification tasks, and outputs the final target detection result.
[0137] A belt tearing detection system based on a dual-vision state space model provided by an embodiment of the present invention has a functional structure corresponding to a belt tearing detection method based on a dual-vision state space model provided by an embodiment of the present invention, which will not be elaborated herein.
[0138] Figure 11 FIG. 4 is a schematic structural diagram of an electronic device 1100 provided by an embodiment of the present invention. The electronic device 1100 may vary greatly due to different configurations or performances, and may include one or more central processing units (CPUs) 1101 and one or more memories 1102. Among them, at least one instruction is stored in the memory 302, and the at least one instruction is loaded and executed by the processor 1101 to implement the steps of the above-mentioned belt tearing detection method based on a dual-vision state space model.
[0139] In an exemplary embodiment, a computer-readable storage medium is further provided, such as a memory including instructions. The above instructions can be executed by a processor in a terminal to complete the above-mentioned belt tearing detection method based on a dual-vision state space model. For example, the computer-readable storage medium may be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.
[0140] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above embodiments can be completed by hardware, or can be completed by a program instructing relevant hardware. The program can be stored in a computer-readable storage medium, and the above-mentioned storage medium may be a read-only memory, a magnetic disk or an optical disc, etc.
[0141] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A belt tear detection method based on a dual visual state space model, characterized in that: The method comprises: S1, obtaining the surface image of the belt to be detected; S2, inputting the belt surface image to be detected into the trained belt tear image defect detection backbone network, detecting and outputting the target detection result, wherein the belt tear image defect detection backbone network is composed of an image serialization module, a dual visual state space belt tear image feature extraction network, a multi-scale feature fusion module and a dual-branch target detection output module; The image serialization module serializes the belt surface image to be detected and outputs an image sequence and an image microsequence; The dual visual state space belt tear image feature extraction network uses a dual visual state space feature extraction module to extract global semantic category information and local detail information of the belt tear small target defect; The multi-scale feature fusion module uses a feature pyramid structure based on a visual state space feature extraction module to perform multi-scale feature fusion to obtain image fusion features of three sizes; The dual-branch target detection output module uses two branches and utilizes the image fusion features of the three sizes to perform target recognition and classification tasks respectively, and outputs the final target detection result; The image serialization module first slices the belt surface image to be detected into blocks to obtain a plurality of image blocks; Then, each image block is locally cut into blocks to obtain multiple local image blocks; Then, by using the convolution down sampling operation, the sliced image blocks and the local image blocks are converted into two image sequence pairs, namely, an image sequence and an image microsequence, wherein both the image sequence and the image microsequence are discrete two-dimensional image sequences; The first dual visual state space feature extraction module in the dual visual state space feature extraction module takes the image sequence and the image micro-sequence as input, first uses a local visual state space branch composed of visual state space feature extraction modules to extract the local image features inside the image micro-sequence, and performs feature fusion with the input image sequence by addition, and then sends the fused sequence features to another global visual state space branch composed of visual state space feature extraction modules to extract long-distance global semantic category information between image sequences, and at the same time models the relationship between the global features of the image sequence and the local features of the image micro-sequence to obtain image fusion features, and sends the image local features extracted by the local visual state space branch and the image fusion features extracted by the global visual state space branch to the second cascaded dual visual state space feature extraction module to continue to extract deeper image features.
2. The method according to claim 1, characterized in that: The dual vision state space belt tear image feature extraction network uses the dual vision state space feature extraction module to extract the global semantic category information and local detail information of the belt tear small target defect in four stages. The four stages use N1, N2, N3, and N4 dual vision state space feature extraction modules in cascade, and down-sample using convolution with a step size of 2 in different stages to reduce the image feature size, and obtain image fusion features F1, F2, F3, and F4 of four sizes, respectively, and the sizes are two times the down-sampling relationship, among which F1 is not used to generate target detection results later because it contains more noise information; down-sampling is through continuous parameter iteration in the training stage, effectively filtering out invalid information in the feature map of the previous stage, reducing the amount of calculation, and will not lose effective image features; The local detail information extracted by the dual vision state space feature extraction module does not directly participate in the output of the target detection result. Instead, the local detail information is integrated into the image fusion feature through feature fusion in the dual vision state space feature extraction module, thereby indirectly playing a role in accurately detecting the output result.
3. The method according to claim 1, characterized in that The visual state space feature extraction module is responsible for modeling discrete two-dimensional image sequences and obtaining the relationship between sequences. When the discrete two-dimensional image sequence is used as input, in addition to the front and back sequences, there are positional relationships in four directions, namely, up, down, left, and right, between the image sequences. According to the cross-scanning method in four different directions, all image matrices composed of several image sequences are rearranged to form four long image sequences composed of image sequences. Each long image sequence is then sent to the visual state space feature extraction module to extract the corresponding semantic information in the sequence to obtain four feature sequences. The four feature sequences are added and averaged to obtain a new image matrix, and then the new image matrix is flattened into a one-dimensional vector to obtain a serialized image vector. Among them, the visual state space feature extraction module extracts image features from each long image sequence through layer normalization, linear layer, activation function, depthwise separable convolution, discrete state space model, and layer normalization operations in sequence, and then multiplies the image features with the weight matrix obtained by the other branch of the long image sequence through the linear layer and the activation function, and then passes the multiplication result through a linear layer to obtain an image feature sequence, introduces a residual connection, and adds the image feature sequence to the long image sequence to obtain the feature sequence.
4. The method according to claim 3, characterized in that The discrete state space model takes pixel intensity, edge, texture, color histogram, object position, shape or motion as states. These states are quantified into discrete values to form a state space. Intermediate implicit states are introduced to extract features in the image. Through discretized state transfer and observation processes, the complex dynamic and semantic relationships in the image sequence are accurately modeled.
5. The method according to claim 2, characterized in that: The multi-scale feature fusion module fuses the features of the current layer and the previous layer layer by layer through two paths, top-down and bottom-up, to generate a new feature map, and uses the feature pyramid structure based on the visual state space feature extraction module to fuse the extracted feature map to generate a new feature map; Among them, top-down means that the image fusion feature F4 with the smallest size is upsampled to expand the resolution, and then spliced with F3 like F3, and the spliced features are sent to the visual state space feature extraction module to fuse the features of two different scales, and the fused features are upsampled and spliced with F2 of a larger size, and the spliced features are sent to the visual state space feature extraction module to fuse the features of two different scales, and obtain the multi-scale fusion feature P2; For bottom-up fusion, P2 is downsampled to the same size as F3, and then concatenated with the corresponding features in top-down fusion, and sent to the visual state space feature extraction module to fuse the features of two different scales to obtain the multi-scale fusion feature P3; P3 is downsampled to the same size as F4, and then concatenated with the corresponding features in the top-down fusion, and sent to the visual state space feature extraction module to fuse the features of two different scales to obtain the multi-scale fusion feature P4.
6. The method according to claim 1, characterized in that The dual-branch target detection output module adjusts the feature map dimension by upsampling the multi-scale fusion features P2, P3, and P4 to the original image size and concatenating them followed by a 1×1 convolution; Then, the classification task branch consists of 2 3×3 convolutions and 1 1×1 convolution. Assuming the image size is H×W, C is the number of target detection categories, and outputs the classification result of H×W×C; The target recognition branch consists of two 3×3 convolutions and two 1×1 convolutions. The two 1×1 convolutions are used to output H×W×4 position detection results and H×W×1 intersection-over-union scores respectively. And the final output format is (c o ,x o ,y o ,h o ,w o ), where subscript o represents output, c represents defect category number, (x, y) represents the center point coordinates of the target detection box, and (h, w) represents the height and width of the target detection box.
7. The method according to claim 1, characterized in that The method splits the belt tearing small target detection task into two tasks: a classification task and a rectangular box regression task. The loss function consists of a classification task loss function and a rectangular box regression task loss function. The classification task loss function is a binary cross entropy loss function. The rectangular box regression task loss function consists of a distribution focus loss function and a complete intersection-over-union loss function. These three loss functions are weighted proportionally to form the loss function, which is used to supervise the learning process of the model. The binary cross entropy loss function is used to judge the quality of the classification model prediction results by calculating the difference between the category prediction probability and the category label. The specific formula is as follows: Among them, y ic represents the cth label of the i-th sample output by the model, p ic Represents the probability that the output belongs to the label, and N represents the number of groups of objects predicted by the model; The distribution focus loss function can be used to allow the network to quickly focus on the value near the label, making the probability density at the label as large as possible. The calculation is based on the KL divergence, which is used to compare the difference between the predicted distribution and the true distribution. For each predicted bounding box, the true category and the predicted category are encoded as distribution vectors respectively. The KL divergence between the two distribution vectors is calculated to judge the quality of the model output. Assume that the distribution probability of the predicted bounding box position is p0, p1, ..., p n , then for each true bounding box position y, find the two predicted bounding box positions y closest to y i and i+1 , and then calculate the probability p corresponding to these two positions i and p i+1 The loss of all these positions is accumulated to get the final DFL loss. The specific formula is as follows: DFL_Loss(p i ,p i+1 )=-((y i+1 -y)log(p i )+(y-y i )log(p i+1 )) Among them, p i and p i+1 is the distribution probability of the two predicted bounding box positions output by the model, y is the position of the true bounding box, and y i and i+1 represents the predicted bounding box position closest to y; The intersection-over-union (IoU) ratio is a way to describe the overlap between the predicted bounding box and the true bounding box of target detection. The IoU ratio between the predicted bounding box and the true bounding box is used to measure the degree of regression of the box. The complete IoU loss function is based on this principle and adds two metrics: the consistency of the center point distance and the aspect ratio. The relative position information and shape difference between the predicted bounding box and the detected bounding box are used to judge the fitting of the model output rectangular box. The specific formula is as follows: Among them, ρ 2 (b,b gt ) represents the Euclidean distance between the center points of the predicted bounding box and the true bounding box, d represents the diagonal length of the minimum closed area that can contain both the predicted bounding box and the true bounding box, w and h represent the width and height of the predicted bounding box, respectively. gt and h gt Represent the width and height of the real bounding box respectively; Finally, the total loss function is the weighted sum of the above three loss functions. The specific formula is as follows: Loss=w bce BCE_Loss+w dfl DFL_Loss+w ciou CIoU_Loss Among them, w bce 、w dfl 、w ciou Represent the corresponding weights respectively.
8. A belt tear detection system based on a dual visual state space model, characterized in that: The system comprises: An acquisition module, used for acquiring the surface image of the belt to be detected; A detection module is used to input the belt surface image to be detected into a trained belt tear image defect detection backbone network, and detect and output a target detection result. The belt tear image defect detection backbone network consists of an image serialization module, a dual visual state space belt tear image feature extraction network, a multi-scale feature fusion module, and a dual-branch target detection output module; The image serialization module serializes the belt surface image to be detected and outputs an image sequence and an image microsequence; The dual visual state space belt tear image feature extraction network uses a dual visual state space feature extraction module to extract global semantic category information and local detail information of the belt tear small target defect; The multi-scale feature fusion module uses a feature pyramid structure based on a visual state space feature extraction module to perform multi-scale feature fusion to obtain image fusion features of three sizes; The dual-branch target detection output module uses two branches and utilizes the image fusion features of the three sizes to perform target recognition and classification tasks respectively, and outputs the final target detection result; The image serialization module first slices the belt surface image to be detected into blocks to obtain a plurality of image blocks; Then, each image block is locally cut into blocks to obtain multiple local image blocks; Then, by using the convolution down sampling operation, the sliced image blocks and the local image blocks are converted into two image sequence pairs, namely, an image sequence and an image microsequence, wherein both the image sequence and the image microsequence are discrete two-dimensional image sequences; The first dual visual state space feature extraction module in the dual visual state space feature extraction module takes the image sequence and the image micro-sequence as input, first uses a local visual state space branch composed of visual state space feature extraction modules to extract the local image features inside the image micro-sequence, and performs feature fusion with the input image sequence by addition, and then sends the fused sequence features to another global visual state space branch composed of visual state space feature extraction modules to extract long-distance global semantic category information between image sequences, and at the same time models the relationship between the global features of the image sequence and the local features of the image micro-sequence to obtain image fusion features, and sends the image local features extracted by the local visual state space branch and the image fusion features extracted by the global visual state space branch to the second cascaded dual visual state space feature extraction module to continue to extract deeper image features.
Citation Information
Patent Citations
Infrared small target detection method based on YOLOv4 multi-scale feature fusion
CN115546502A
Belt tearing detection method and system based on discrete state selectable space model
CN118333979A