A belt tearing detection method and system based on key feature fusion
By adopting a dual-stream fusion network and a dual-branch perceived attention mechanism in belt tear detection, the problems of false detection and missed detection in the prior art are solved, achieving higher detection accuracy and system practicality.
Patent Information
- Application Number
- CN202410945217.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-15
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2044-07-15
AI Technical Summary
Existing belt tear detection technology is difficult to effectively distinguish between tear and distractor, resulting in missed detection and missed detection problems, especially in small target tear categories with small data samples and small morphology.
A dual-stream fusion network based on key feature fusion is adopted, and a two-way feature pyramid fusion is performed through a shared feature extraction layer, combining the dual-branch perceived attention mechanism and false detection feature enhancement module to enhance the model's perception of small targets and small sample tearing categories, and reduce false detection.
It improves the accuracy of belt tear detection, reduces false detection and missed detection rates, and ensures the practicality and accuracy of the detection system.
Smart Images

Figure CN118887183B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of belt tear detection, and particularly to a belt tear detection method and system based on key feature fusion. Background Art
[0002] Belt tear detection technology aims to analyze feature information such as texture and shape in the imaging of the belt surface, and accurately locate and classify the defects existing on the belt. Common tear types include edge breakage, belt stacking, patch scratching, large-area breakage, and small target tears, etc. In actual industrial production, it usually involves large equipment, high-speed operating mechanical systems, and harsh operating conditions. Various environmental factors such as abnormal materials, equipment wear, and overloading greatly increase the risk of belt tear. Belt tear detection technology can quickly obtain the tear type and the location of the tear on the belt surface by collecting high-resolution belt surface image data. It can not only stop the machine in time when a serious tear breakage is detected on the belt surface, but also quickly identify the risk when a slight tear occurs, preventing production interruption after a major tear. However, the surface breakage caused by belt tear is easily confused with interference objects such as water stains, textures, and dust particles existing on the belt during the recognition process of the detection model, resulting in false detection. At the same time, tear categories with a small number of samples and small target tears with small shapes in the data are extremely easy to be missed, which poses a great challenge to meeting the high-precision requirements of detection. Summary of the Invention
[0003] The present invention provides a belt tear detection method and system based on key feature fusion to solve the problems existing in the above-mentioned prior art. The technical solutions are as follows:
[0004] On the one hand, a belt tear detection method based on key feature fusion is provided, including:
[0005] S1. Collect and preprocess multiple belt surface images, respectively perform labeling for the target detection task and the segmentation task on the multiple belt surface images, and divide the training set and the validation set;
[0006] S2. Input the images in the training set into the shared feature layer of the dual-stream fusion network for training, extract general feature information, and perform feature fusion based on the bidirectional feature pyramid to output a fused feature image for sharing by subsequent dual-branch tasks;
[0007] S3, input the fused feature image and train the lightweight target detection branch of the dual-stream fusion network for missed detection problems, and closely combine the global features with the key features of the missed detection category through the dual-branch perception attention mechanism, so as to enhance the model's sensitivity to the feature information of small targets and small sample tearing categories, output the positioning detection anchor frame of the target detection and the category prediction result of the target, and output the confidence C corresponding to the predicted category. d ;
[0008] S4, input the fused feature image and train the lightweight segmentation decision branch of the dual-stream fusion network for the false detection problem, and improve the model's recognition ability of the texture features of the interference object through the false detection feature enhancement module, accurately distinguish the interference object in the image background from the real torn area, and output the prediction result of whether the target is really torn and the corresponding confidence C s ;
[0009] S5. C d and C s Input and train the confidence fuser of the two-stream fusion network, output the final classification of the detected target, and determine the detection anchor frame position where the target is finally marked according to the predicted frame position of the detected target of the lightweight target detection branch;
[0010] S6. Use the trained two-stream fusion network to perform belt tear detection on the belt surface image to be detected.
[0011] Optionally, after the shared feature layer performs data enhancement for the belt surface image that is common to the dual-branch task, part of the image is subjected to feature extraction through 9 convolution layers with a convolution kernel size of 5×5, and each convolution operation is followed by a batch normalization process to normalize each feature channel to a zero-mean distribution with unit variance, and then a nonlinear transformation is introduced through a ReLU activation function, and finally the resolution of each layer is reduced by half to obtain common feature information of different scales; another part of the image is subjected to 3 maximum pooling layers that will reduce the image resolution, thereby reducing the spatial dimension and the amount of model calculation, while improving the generalization ability of the model and capturing higher-level and more abstract global feature information;
[0012] After merging the common feature information of different scales with the pooled global feature information, efficient bidirectional cross-scale feature fusion is performed through the bidirectional feature pyramid structure in the shared feature layer. Not only are high-level abstract features upsampled layer by layer along a top-down path to match the scale of bottom-level features and fuse with each other, but a bottom-up feedback mechanism is also introduced to gradually downsample the bottom-level features containing rich details and interact with high-level features. After that, a fused feature image with the same dimension as the input feature is output through a 1x1 convolutional layer.
[0013] Optionally, for the lightweight object detection branch, the fused feature image is passed through an initial convolutional layer of 3×3 to extract preliminary feature information, and then fed into a batch normalization layer for normalization processing. Subsequently, it undergoes non-linear transformation through the h-swish activation function to obtain the basic features provided to the lightweight detection module.
[0014] For the lightweight detection module, the basic features are input into a depthwise separable convolutional layer of 3×3 with a stride of 2 to extract depth convolutional features. Then, pointwise convolution of 1×1 is used to integrate the depth convolutional features and reduce the number of feature channels. Subsequently, the feature map output by the pointwise convolution is again subjected to non-linear transformation through the h-swish activation function to enhance the expression ability of the model. The above steps are repeated 11 times to obtain 11 feature maps F i , where i = 1, 2, …, 11. A top-down method in a feature pyramid is adopted to gradually merge the feature maps of adjacent levels until a fused feature map F_fusion is formed.
[0015] The fused feature map F_fusion is input into the dual-branch perceptual attention mechanism module. The global feature perceptual attention sub-module of the dual-branch perceptual attention mechanism module outputs the global feature weighted vector X global , and the missed detection class feature perceptual attention sub-module of the dual-branch perceptual attention mechanism module outputs the missed detection class feature weighted vector X local . X global and X local are concatenated, and after the channel size is converted through a fully connected layer, the dual-branch perceptual attention output X output is obtained.
[0016] Based on X output , the localization loss function, confidence loss function, and classification loss function are calculated for the object detection anchor boxes and object category predictions. At the same time, a class balance loss function is added to specifically address the problem of unbalanced data class ratios. The calculation formula of the class balance loss function is as follows:
[0017] F(p t ) = -(1 - p t ) γ log(p t )
[0018] where p tis the predicted probability of the model for a sample to be a positive class. γ is used as a modulation parameter to control the relative loss contribution of easy-to-classify samples and difficult samples. When γ > 0, difficult samples will obtain a larger loss weight, enabling the model to pay more attention to the sample categories that are easily misclassified. After this calculation process, the model outputs the localization detection anchor boxes for object detection and the class prediction results of the objects, and simultaneously outputs the confidence C corresponding to the predicted class d 。
[0019] Optionally, the dual-branch perception attention mechanism module multiplies the input fused feature map F_fusion by different weight parameters to obtain the inputs of the Q, K, and V branches, and then respectively passes them through three 1×1 convolutional layers Query_Conv, Key_Conv, and Value_Conv. At the same time, the view function in the Pytorch environment is used to perform dimensional transformation on the Query and Key feature vectors, and the permute function is used to invert the Query vector to obtain the output of three different-dimensional feature vectors Q, K, and V
[0020] The global feature perception attention sub-module of the dual-branch perception attention mechanism module passes through a fully connected layer to extract global information to obtain Q_global, K_global, and V_global, and then downsamples K_global and V_global through global pooling. Multiply Q_global by the downsampled K_global matrix and perform Softmax normalization processing, and then multiply the obtained result by the downsampled V_global matrix to finally extract the low-frequency global information and obtain the global feature weighted vector Xglobal. Its formula is expressed as follows
[0021] X global =Softmax(Q_global·Pool(K_global))·Pool(V_globa)
[0022] The missed detection class feature perception attention sub-module of the dual-branch perception attention mechanism module specifically captures the high-frequency feature information corresponding to the missed detection class for perception enhancement. For the three different-dimensional feature vectors Q, K, and V, three depthwise separable convolutions with a stride of 1×1 are respectively used to extract local information to obtain Q_local, K_local, and V_local. At the same time, the weights of the depthwise separable convolutional layer are globally shared. Then, calculate the Hadamard product of Q_local and K_local to merge the two to generate a context-aware feature matrix M context , and then through a fully connected layer and the Swish activation function, and then through another fully connected layer and the Tanh activation function, perform feature integration and non-linear transformation on this result to obtain a context-aware feature W between -1 and 1 contex, the relevant formula is expressed as follows:
[0023] M context = Q_local ⊙ k_local
[0024]
[0025] where n represents the number of feature channels, the symbol ⊙ represents the Hadamard product, and FC represents the fully connected layer;
[0026] Then use W contex and the V_local after integrating the 1×1 depthwise separable convolution feature information to perform the Hadamard product again to obtain the feature weighted vector X after enhancing the perception of the undetected category features local , X local = W context ⊙ V_local;
[0027] Concatenate X global and X local The two are concatenated, and after passing through the fully connected layer to convert the channel size, the dual-branch perception attention output X output , X output = FC(Contact(X global , X local )) is obtained, enabling the model to effectively perceive high-frequency and low-frequency information simultaneously.
[0028] Optionally, for the lightweight segmentation decision branch, first pass the fused feature image through the lightweight segmentation module, and use the depthwise separable convolution network for lightweight improvement, specifically including:
[0029] Use a depthwise separable convolution with a size of 15×15 and a stride of 1 to extract depth convolution features, and then use a 1×1 pointwise convolution to integrate the depth convolution features and reduce the number of feature channels. First, use the non-linear activation function ReLU to increase the information volume of the feature map, and then pass through a 1×1 convolution layer for reducing the number of output channels, and perform batch normalization and use the non-linear activation function ReLU again to obtain a single-channel feature map;
[0030] While obtaining the single-channel feature map, input the feature map output after the first use of the non-linear activation function ReLU into the misdetection region feature enhancement module. The axial global attention sub-module of the misdetection region feature enhancement module outputs enhanced global feature information, and the misdetection region detail feature enhancement sub-module of the misdetection region feature enhancement module outputs misdetection detail enhancement features. Multiply and fuse the enhanced global feature information and the misdetection detail enhancement features to obtain a feature map with higher semantic information, and splice it with the single-channel feature map to obtain the feature map output of the lightweight segmentation module as the input of the decision module;
[0031] The decision-making module performs a two-layer combination three times on the feature map output of the lightweight segmentation module. The two-layer combination includes a 2×2 max pooling layer and a convolutional layer with a 5×5 convolutional kernel size. The number of channels is set to increase as the feature resolution decreases. The three convolutional layers are set to 8, 16, and 32 channels respectively. Then, global max pooling and global average pooling operations are performed to generate 64 output neurons. In addition, the results of global max pooling and global average pooling on the single-channel feature map output in the lightweight segmentation module are also connected into 2 output neurons respectively. Then, 66 neurons are output through a fully connected layer. After the integration process of the linear weight combination layer and with the use of the sigmoid activation function, the prediction result of whether the target truly belongs to a tear and its corresponding confidence level C are finally output. s 。
[0032] Optionally, the false detection region feature enhancement module obtains query vector Q, key vector K, and value vector V through linear transformation of the feature map output after the first use of the non-linear activation function ReLU. Then, through the axial global attention sub-module and the false detection region detail feature enhancement sub-module, the global features and false detection detail features are processed separately.
[0033] Among them, for the axial global attention sub-module to obtain globally connected context features, first, axial pooling operations are performed on the input feature map along the horizontal and vertical directions respectively to transform it into compact rows and columns. For the feature vector Q, where H and W are the height and width of the feature map respectively, and C is the number of channels, the horizontal axial pooling result Q of the H×W×C feature map h and the vertical axial pooling result Q v are expressed as follows:
[0034]
[0035] Among them, since the number of channels of vector Q and vector K is the same, they are uniformly represented as C QK , Q →(·) represents permutation of the dimension of Q, and Ⅰ W , Ⅰ H are both vectors with all elements equal to 1.
[0036] Then, horizontal compression and vertical compression are respectively performed on Q h and Q v to achieve dimension conversion, obtaining two axial feature vectors of H×1×C QK and 1×W×C QK .
[0037] Then, multi-channel and multi-angle attention is calculated for these two vectors to capture the dependencies between multi-angle axial feature vectors covering 360°, including:
[0038] The horizontal and vertical axial feature vectors used as input are concatenated into a new feature matrix using the Concat concatenation function. Multiple parallel channels are designed, and each channel uses a 3×3 directional convolution kernel with different directions. The parameter θ, 0°<θ<360°, is used to adjust the direction angle of the axial feature vector. The weight parameter values in the convolution kernel are all related to θ, so that the convolution kernel assigns a higher weight to the specified angle θ, realizing the extraction of multi-angle axial feature vectors. Then, a linear transformation is performed to obtain the Q, K, and V inputs corresponding to each channel, and the scaling dot product operation is performed to obtain the attention vector of each channel. This process is completed N times in the multi-channel multi-angle attention mechanism. The parameters of the linear transformation between the channels are not shared. The N attention vectors are then concatenated and fused by a linear transformation. The results of the multi-channel multi-angle attention mechanism are output. On this basis, a 1×1 convolution layer is used to regress the dimension of the original input feature map, and finally the global feature information enhanced by the axial global attention submodule is output to reduce the false detection phenomenon caused by the lack of global context feature information.
[0039] The submodule for enhancing the detail features of the misdetected area adopts a new convolution-based detail enhancement design based on the texture characteristics of the blurred boundaries of the interference objects in the image background that are easily misdetected as torn areas, which increases the clarity of the boundaries and facilitates the model to distinguish them from the torn areas, including:
[0040] Connect the channels of Q, K, and V. Since the channels of Q and K are the same, they are both C. QK , so the size is H×W×(2C QK +C V ) feature map, and then pass it to a block consisting of 3×3 depth-separable convolution and batch normalization to assist in aggregating local details from Q, K, V, and then use a 1×1 convolution layer with ReLU6 activation function and batch normalization for linear projection to convert (2C QK +C V ) dimension is compressed to the original dimension, and the false detection detail enhancement feature is output.
[0041] Optionally, the confidence fuser first converts C d and C s Splice into an input vector X = [C d ,C s ], and then input the vector X into a multi-layer perceptron neural network, whose output hi is expressed as:
[0042] hi=f(W i hi-1 +b i )
[0043] Among them, h 0 = X, f is the ReLU activation function, W i is the weight matrix of the i-th layer, b i is the bias vector, and the h of the last layer n generates the final fusion confidence through an output operation. The fusion formula is as follows:
[0044] C fusion = ɑ(W ou th n +b out )
[0045] Among them, ɑ is the Sigmoid function, which converts the output into a probability value between 0 and 1. W out is the weight matrix in the output operation, b out is the bias vector;
[0046] Finally, the dual-stream fusion network determines the final classification of the detected target according to the value of C fusion , and at the same time determines the detection anchor box position where the target is finally marked according to the predicted box position of the detected target of the lightweight target detection branch.
[0047] On the other hand, a belt tear detection system based on key feature fusion is provided. The system includes:
[0048] A collection preprocessing and partitioning module, which is used to collect and preprocess multiple belt surface images, label the multiple belt surface images for target detection tasks and segmentation tasks respectively, and partition the training set and the validation set;
[0049] A shared feature extraction module, which is used to input the images in the training set and train the shared feature layer of the dual-stream fusion network, extract general feature information, and perform feature fusion based on a bidirectional feature pyramid, and output a fused feature image for sharing by subsequent dual-branch tasks;
[0050] A target detection module, which is used to input the fused feature image and train the lightweight target detection branch of the dual-stream fusion network for the missed detection problem. Through the dual-branch perceptual attention mechanism, the global features are closely combined with the key features of the missed detection categories, and the sensitivity of the model to the feature information perception of small targets and small sample tear categories is enhanced specifically. The positioning detection anchor box for target detection and the category prediction result of the target are output, and at the same time the confidence C corresponding to the predicted category is output d ;
[0051] A segmentation decision module for inputting the fused feature image and training the lightweight segmentation decision branch of the dual-stream fusion network for the false detection problem. Through the false detection feature enhancement module, the recognition ability of the model for the texture features of interfering objects is specifically improved, and the interfering objects in the image background are accurately distinguished from the true tear area, and the prediction result of whether the target truly belongs to the tear and the corresponding confidence level C are output. s ;
[0052] A fusion module for inputting C d and C s into the confidence fusion device of the dual-stream fusion network for training and outputting the final classification of the detected target. At the same time, according to the predicted box position of the detected target of the lightweight target detection branch, the detection anchor box position where the target is finally marked is determined.
[0053] A detection module for using the trained dual-stream fusion network to perform belt tear detection on the surface image of the belt to be detected.
[0054] On the other hand, an electronic device is provided. The electronic device includes a processor and a memory. At least one instruction is stored in the memory, and the at least one instruction is loaded and executed by the processor to implement the above-mentioned belt tear detection method based on key feature fusion.
[0055] On the other hand, a computer-readable storage medium is provided. At least one instruction is stored in the storage medium, and the at least one instruction is loaded and executed by a processor to implement the above-mentioned belt tear detection method based on key feature fusion.
[0056] The beneficial effects brought by the technical solution provided by the present invention at least include:
[0057] The present invention proposes a method and system for detecting belt surface tears based on key feature fusion. First, shared feature information of image data is obtained through a shared feature extraction process. In the network structure of this process, a bidirectional feature pyramid fusion mechanism is added, which can effectively fuse the detailed information of the shallow layer and the high-level semantic information of the deep layer, thereby improving the model's processing ability for targets of different scales. Secondly, for the shared feature information, dual-task processing of target detection and segmentation decision is carried out respectively. At the same time, the lightweight design idea based on the depthwise separable network structure is introduced in the backbone network part of the dual-branch to ensure the high efficiency and practicality of the solution in industry. In the target detection branch, aiming at the tear category features that are easily missed, a dual-branch perception attention mechanism is designed, and combined with the category balance loss function, the missed detection rate of small target and small sample tear categories is more effectively reduced. At the same time, in the parallel segmentation decision branch, the image segmentation task is first executed. For the features of the easily misdetected area, a misdetected area feature enhancement module with the dual effects of axial global attention and misdetected area detail enhancement is set, and the more refined feature map is input to the decision module for the final judgment of whether the target is a tear, so as to reduce the situation of misdetecting the interference objects such as water stains and impurity particles on the belt surface as tears. Finally, through the confidence fusion process, the confidence output results of target detection and segmentation decision are integrated, and the tear type of the detected target and the position of the detection anchor box are selectively output in a weighted summation manner, so as to solve the problems of missed detection and misdetection of tears at the same time, and ensure the practicality and accuracy of the detection system. The advantages of the present invention are specifically as follows:
[0058] 1) Aiming at the problem that the existing single-stage detection scheme cannot synchronously improve the problems of missed detection and misdetection, a dual-stream fusion network is designed. Through the design of the shared feature extraction layer of the dual-branch of target detection and segmentation decision, the feature information of different levels and scales is efficiently fused, so that the detection accuracy of the dual-task is synchronously improved. At the same time, relying on the confidence fusion design, the detection results output by the dual-task are effectively combined, overcoming the technical difficulty that the missed detection problem and the misdetection problem restrict each other when optimized under the same model, and greatly improving the detection accuracy.
[0059] 2) Aiming at the missed detection problem caused by the limited perception ability of some tear category features, dual-branch perception attention is introduced in the target detection branch. The correlation between low-frequency features in different regions is strengthened in the global feature perception attention sub-module to realize the tight combination of global features. In the missed detection category feature perception attention sub-module, a perception enhancement design is added for the high-frequency features unique to the missed detection category to accurately capture the missed detection category features, and then combined with the category balance loss function to improve the detection ability of small samples and small targets of tears that are difficult to detect, and reduce the missed detection rate of the model.
[0060] 3) To address the problem of misdetection caused by the difficulty of existing models in distinguishing surface interference objects from true tears, a misdetection area feature enhancement module is added to the segmentation decision branch. In the axial global attention sub-module, a 360° axial decomposition of global features is achieved based on the multi-angle and multi-channel attention mechanism, enhancing the model's learning ability for global structures and overcoming the problem that the connection between the features of the misdetected area and the global features is not tight enough. At the same time, in cooperation with the misdetection area detail feature enhancement sub-module, the boundary recognition clarity of the misdetection area and the texture feature discrimination ability are improved, maximizing the recognition sensitivity of the key features in the misdetection area and reducing the model's misdetection rate.
[0061] 4) To address the problem of high computational complexity caused by the key feature enhancement process, a lightweight structure design is added to the dual-branch task at the same time. The convolutional structure in the backbone network is replaced with depthwise separable convolutions. In addition, a shortcut of single-channel mapping is introduced during the segmentation process. The above operations greatly reduce the number of model parameters and the computational amount, achieving model miniaturization and high efficiency, thus ensuring the practicality of the detection scheme in industry. Brief Description of the Drawings
[0062] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0063] Figure 1 It is a flowchart of a belt tear detection method based on key feature fusion provided by an embodiment of the present invention;
[0064] Figure 2 It is a general flowchart of the belt tear detection method provided by an embodiment of the present invention;
[0065] Figure 3 It is a schematic diagram of the shared feature layer provided by an embodiment of the present invention;
[0066] Figure 4 It is a schematic diagram of the lightweight object detection branch for the missed detection problem provided by an embodiment of the present invention;
[0067] Figure 5 It is a schematic diagram of the dual-branch perception attention mechanism module provided by an embodiment of the present invention;
[0068] Figure 6 It is a schematic diagram of the lightweight segmentation decision branch for the misdetection problem provided by an embodiment of the present invention;
[0069] Figure 7 It is a schematic diagram of the misdetection area feature enhancement module provided by an embodiment of the present invention;
[0070] Figure 8 It is a schematic diagram of a multi-channel and multi-angle attention mechanism provided by an embodiment of the present invention;
[0071] Figure 9 It is a schematic diagram of a confidence fusion device provided by an embodiment of the present invention;
[0072] Figure 10 It is a block diagram of a belt tear detection system based on key feature fusion provided by an embodiment of the present invention;
[0073] Figure 11 It is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. Detailed implementation manners
[0074] To make the technical problems, technical solutions and advantages to be solved by the present invention clearer, the following will be described in detail with reference to the accompanying drawings and specific embodiments.
[0075] An embodiment of the present invention provides a belt tear detection method based on key feature fusion. This method can be implemented by an electronic device, and the electronic device can be a terminal or a server. As Figure 1 shown in the flowchart of a belt tear detection method based on key feature fusion, the processing flow of this method can include the following steps:
[0076] S1. Collect and preprocess multiple belt surface images, perform labeling for the target detection task and the segmentation task on the multiple belt surface images respectively, and divide the training set and the validation set;
[0077] S2. Input the images in the training set into and train the shared feature layer of the dual-stream fusion network, extract general feature information, and perform feature fusion based on the bidirectional feature pyramid, and output the fused feature image for sharing by subsequent dual-branch tasks;
[0078] S3. Input the fused feature image into and train the lightweight target detection branch of the dual-stream fusion network for the missed detection problem. Through the dual-branch perception attention mechanism, closely combine the global feature with the key features of the missed detection categories, and specifically enhance the sensitivity of the model to the feature information perception of small targets and small sample tear categories, and output the localization detection anchor box for target detection and the category prediction result of the target, and at the same time output the confidence C corresponding to the predicted category d ;
[0079] S4. Input the fused feature image into and train the lightweight segmentation decision branch of the dual-stream fusion network for the false detection problem. Through the false detection feature enhancement module, specifically improve the recognition ability of the model for the texture features of interference objects, accurately distinguish the interference objects from the true tear area in the image background, and output the prediction result of whether the target truly belongs to a tear and the corresponding confidence Cs ;
[0080] S5. Input C d and C s into the confidence fusion unit of the two-stream fusion network for training, output the final classification of the detected target, and at the same time determine the detection anchor box position where the target is finally marked according to the predicted box position of the detected target in the lightweight target detection branch;
[0081] S6. Use the trained two-stream fusion network to perform belt tear detection on the surface image of the belt to be detected.
[0082] The embodiment of the present invention provides a belt surface tear detection method based on key feature fusion, designs a dual-task shared feature extraction layer and a confidence fusion layer, relies on the dual-branch perception attention set in the target detection task to enhance the detection ability of categories that are easily missed, and at the same time uses the misdetection feature enhancement module added in the segmentation task to specifically improve the misdetection problem of misdetecting the image background as a tear, realizing the dual role of parallel optimization of misdetection and missed detection problems. The overall network structure of the system is as Figure 2 shown: First, in order to obtain the most efficient general feature information such as edges, textures, and shapes shared by the dual tasks, a shared feature layer of the two-stream fusion network is constructed, and a bidirectional feature pyramid fusion mechanism is added to the shared feature extraction link therein to combine feature information of different depths, thereby improving the expression ability of the general feature information extracted in the early stage. Then, for the general feature information extracted by the shared feature layer, two prediction tasks of target detection and segmentation decision are respectively performed. In the backbone network of the target detection task, a lightweight design based on depthwise separable convolution is introduced, and the global pooling process is used to replace the fully connected layer to greatly reduce the number of calculation parameters. At the same time, a dual-branch perception attention mechanism for the feature fusion process and a class balance loss function for the post-processing process are designed to effectively handle the problem that small targets and small sample tear categories are easily missed. While the target detection task is optimized, the other segmentation decision task is divided into two steps: obtaining a more refined feature map in the segmentation process and integrating the feature information in the decision process to obtain the prediction result. Based on morphological features such as water stains and textures that are easily misdetected as tears, a misdetection feature enhancement mechanism is designed in the segmentation network. Relying on its strong pertinence to easily confused features, a faster and more accurate judgment is made on whether the target belongs to a tear. Then, the prediction results output by the dual tasks are input into the fusion network, and further judgments are made based on the class confidence of the detected target output by the target detection branch and the tear confidence of the detected target output by the segmentation decision network, so as to discard the actually misdetected part of the target and further avoid misdetecting interference objects such as water stains and impurity particles existing on the belt surface as defects, and finally achieve the ultimate goal of effectively reducing the model's missed detection rate and misdetection rate. The following is the complete technical solution:
[0083] S1. Collect and preprocess multiple images of the belt surface, perform annotations for the object detection task and the segmentation task on the multiple images of the belt surface respectively, and divide them into a training set and a validation set;
[0084] In the embodiment of the present invention, an industrial camera is used to periodically capture images (2D) of the belt surface during the operation of the belt in a complex industrial scenario. The pixel size of all images is 2048×2000. Due to the difficulty in obtaining data and the imbalance in the proportion of images containing torn areas, it is necessary to preprocess the data subsequently, including: performing appropriate artificial augmentation to achieve a balanced state of the number ratio of defective and non-defective images.
[0085] In the embodiment of the present invention, the labelme annotation tool is used to perform annotations for the object detection task and the segmentation task respectively: for the former, the torn area is selected by a rectangular tool, and the target box is divided into damage categories such as edge damage, belt stacking, patch scratch, large-area damage, and small-target tear according to the torn shape to be detected on-site, and a txt object detection annotation file is exported; for the latter, the boundary of the torn area at each pixel level is finely depicted by a polygon or brush tool, and all selected areas are uniformly classified as "belonging to the tear" category, and a json segmentation annotation file is exported.
[0086] In the embodiment of the present invention, all image-object detection label pairs are divided into a training set and a validation set according to a ratio of 4:1. The defective and non-defective images in all image-segmentation label pairs are divided into two groups of training sets and validation sets according to a ratio of 4:1 respectively. After merging the two training sets, a training set, a defective validation set, and a non-defective validation set are generated. Both branch tasks are trained on the training set, and then the model effect is tested on the validation set. The one with the best effect is taken as the final model.
[0087] S2. Input the images in the training set into and train the shared feature layer of the dual-stream fusion network, extract general feature information (edges, textures, shapes, etc.), and perform feature fusion based on a bidirectional feature pyramid to output a fused feature image for sharing by subsequent dual-branch tasks;
[0088] Optionally, as Figure 3As shown, after the shared feature layer performs general data augmentation (such as cropping and scaling, color adjustment, contrast adjustment, etc.) on the belt surface image, part of the image (randomly selected part of the image) is subjected to feature extraction through a convolutional layer with a 5×5 convolutional kernel for 9 layers. After each convolutional operation, a batch normalization process follows, normalizing each feature channel to a zero-mean distribution with unit variance, and then introducing a non-linear transformation through the ReLU activation function. Finally, the resolution of each layer is reduced by half to obtain general feature information of different scales; another part of the image passes through 3 max-pooling layers that reduce the image resolution, reducing the spatial dimension to reduce the model's computational complexity, while enhancing the model's generalization ability and capturing higher-level and more abstract global feature information;
[0089] Then, after merging the general feature information of different scales with the global feature information after pooling, through the bidirectional feature pyramid structure in the shared feature layer, efficient bidirectional cross-scale feature fusion is performed. Not only is the high-level abstract feature upsampled layer by layer along the top-down path to match the scale of the low-level feature and fuse with each other, but also a bottom-up feedback mechanism is introduced to gradually downsample the feature rich in details at the bottom and interact with the high-level feature. After that, a 1x1 convolutional layer outputs a fused feature image with the same dimension as the input feature.
[0090] S3. Input the fused feature image into the lightweight object detection branch of the dual-stream fusion network for training. Through the dual-branch perception attention mechanism, the global feature is closely combined with the key features of the missed detection category, specifically enhancing the model's sensitivity to the feature information perception of small targets and small sample tearing categories, and outputting the localization detection anchor box for object detection and the category prediction result of the object, and at the same time outputting the confidence level C corresponding to the predicted category d ;
[0091] Optionally, as Figure 4 shown, the lightweight object detection branch passes the fused feature image through an initial convolutional layer of 3×3 to extract preliminary feature information, then sends it to a batch normalization layer for standardization processing, and then performs non-linear conversion through the h-swish activation function to obtain the basic feature provided to the lightweight detection module;
[0092] The specific formula of the h-swish activation function is as follows:
[0093] f(x) = x·ReLU6(x + 3) / 6
[0094] where ReLU6(x + 3) = min(ReLU(x + 3), 6).
[0095] The lightweight detection module inputs the basic features into a depthwise separable convolutional layer with a stride of 2 and a size of 3×3 to extract depth convolution features, then performs pointwise convolution with a size of 1×1 to integrate the depth convolution features and reduce the number of feature channels. After that, the feature map output by the pointwise convolution is nonlinearly transformed through the h-swish activation function to enhance the expression ability of the model. The above steps are repeated 11 times (this process has an obvious effect on reducing the number of computational parameters of the model) to obtain 11 feature maps F with different resolutions i , where i = 1, 2, …, 11. A top-down method in a feature pyramid is used to gradually merge the feature maps of adjacent levels until a fused feature map F_fusion is formed;
[0096] The fused feature map F_fusion is input into the dual-branch perceptual attention mechanism module. The global feature perceptual attention sub-module of the dual-branch perceptual attention mechanism module outputs a global feature weighted vector X global , and the missed detection category feature perceptual attention sub-module of the dual-branch perceptual attention mechanism module outputs a missed detection category feature weighted vector X local . X global and X local are concatenated, and after the channel size is converted through a fully connected layer, a dual-branch perceptual attention output X output is obtained;
[0097] According to X output , the localization loss function, confidence loss function, and classification loss function are calculated for the object detection anchor box and object category prediction. At the same time, a class balance loss function is added to specifically address the problem of unbalanced data class ratios. The calculation formula of the class balance loss function is as follows:
[0098] F(p t ) = -(1 - p t ) γ log(p t )
[0099] where pt is the predicted probability of the model for a sample to be a positive class, and γ is a modulation parameter used to control the relative loss contributions of easy-to-classify (i.e., high-confidence) samples and difficult (i.e., low-confidence) samples. When γ > 0, difficult samples will obtain a larger loss weight, enabling the model to pay more attention to the sample categories that are easily misclassified. After this calculation process, the model outputs the localization detection anchor box for object detection and the category prediction result of the object, and at the same time outputs the confidence C corresponding to the predicted category d .
[0100] Optionally, as Figure 5As shown, the dual-branch perceptual attention mechanism module multiplies the input fused feature map \(F_{fusion}\) by different weight parameters to obtain the inputs of the three branches of Q, K, and V. Then, through three 1×1 convolutional layers Query_Conv, Key_Conv, and Value_Conv respectively, and at the same time, the view function in the Pytorch environment is used to perform dimensional transformation on the Query and Key feature vectors, and the permute function is used to invert the Query vector to obtain the output of three different-dimensional feature vectors of Q, K, and V.
[0101] The global feature perceptual attention sub-module of the dual-branch perceptual attention mechanism module passes through a fully connected layer to extract global information to obtain \(Q_{global}\), \(K_{global}\), and \(V_{global}\). Then, global pooling is used to downsample \(K_{global}\) and \(V_{global}\). Multiply \(Q_{global}\) by the downsampled \(K_{global}\) matrix and perform Softmax normalization. The resulting result is then multiplied by the downsampled \(V_{global}\) matrix. Finally, low-frequency global information is extracted to obtain the global feature weighted vector \(X_{global}\), and its formula is expressed as follows:
[0102] X global = Softmax(Q_global·Pool(K_global)·Pool(V_global)
[0103] The missed detection category feature perceptual attention sub-module of the dual-branch perceptual attention mechanism module specifically captures the high-frequency feature information corresponding to the missed detection category for perceptual enhancement. For the three different-dimensional feature vectors of Q, K, and V, three depthwise separable convolutions with a stride of 1×1 are respectively used to extract local information to obtain \(Q_{local}\), \(K_{local}\), and \(V_{local}\). At the same time, the weights of the depthwise separable convolutional layer are globally shared. Then, calculate the Hadamard product of \(Q_{local}\) and \(K_{local}\) to merge the two to generate the context-aware feature matrix M context , and then through a fully connected layer and the Swish activation function (the specific formula of the Swish activation function is as follows: where β is a custom training parameter), and then through another fully connected layer and the Tanh activation function, this result is subjected to feature integration and nonlinear transformation to obtain the context-aware feature W between -1 and 1 contex , and the relevant formula is expressed as follows:
[0104] M context = Q_local⊙k_local
[0105]
[0106] Where n represents the number of feature channels, the symbol ⊙ represents the Hadamard product, and FC represents the fully connected layer;
[0107] Then use W contex and V_local after integrating the 1×1 depthwise separable convolution feature information, and perform the Hadamard product again to obtain the feature weighted vector X after enhancing the perception of missed detection categories. local , X local = W context ⊙V_local;
[0108] Concatenate X global and X local , and after converting the channel size through the fully connected layer, obtain the dual-branch perception attention output X output , X output = FC(Contact(X global , X local ))), so that the model can effectively perceive high-frequency and low-frequency information simultaneously (solving the problem of missed detection of small targets and small sample torn categories caused by the neglect of high-frequency features).
[0109] S4. Input the fused feature image into the lightweight segmentation decision branch of the two-stream fusion network for the misdetection problem, and through the misdetection feature enhancement module, specifically improve the model's recognition ability of the texture features of interference objects, accurately distinguish the interference objects from the real torn area in the image background, and output the prediction result of whether the target truly belongs to the tear and the corresponding confidence level C s ;
[0110] Optionally, as Figure 6 shown, for the lightweight segmentation decision branch, first pass the fused feature image through the lightweight segmentation module and perform lightweight improvement using the depthwise separable convolution network, specifically including:
[0111] Use a depthwise separable convolution with a size of 15×15 and a stride of 1 to extract depth convolution features, then use a 1×1 pointwise convolution to integrate the depth convolution features and reduce the number of feature channels, first use the non-linear activation function ReLU to increase the information volume of the feature map, and then pass through a 1×1 convolution layer for reducing the number of output channels, and perform batch normalization and use the non-linear activation function ReLU again to obtain a single-channel feature map;
[0112] While obtaining the single-channel feature map, the feature map output after the first use of the non-linear activation function ReLU is input into the false detection area feature enhancement module. The axial global attention sub-module of the false detection area feature enhancement module outputs enhanced global feature information, and the false detection area detail feature enhancement sub-module of the false detection area feature enhancement module outputs false detection detail enhanced features. The enhanced global feature information and the false detection detail enhanced features are multiplied and fused to obtain a feature map with higher semantic information, which is concatenated with the single-channel feature map to obtain the feature map output of the lightweight segmentation module as the input of the decision-making module;
[0113] The decision-making module combines the feature map output of the lightweight segmentation module through a two-layer combination repeated 3 times. The two-layer combination includes a 2×2 max pooling layer and a convolutional layer with a convolutional kernel size of 5×5. The number of channels is set to increase as the feature resolution decreases. The three convolutional layers are set to 8, 16, and 32 channels respectively. Then, global max pooling and global average pooling operations are performed to generate 64 output neurons; in addition, the results of global max pooling and global average pooling of the single-channel feature map output in the lightweight segmentation module are also connected into 2 output neurons respectively (providing a shortcut in the case where the segmentation map has ensured detection accuracy, thus avoiding reducing the detection speed by using a large number of feature maps in unnecessary situations); and then 66 neurons are output through a fully connected layer. After the integration process of the linear weight combination layer and with the sigmoid activation function used, the final output is the prediction result of whether the target truly belongs to a tear and its corresponding confidence level C s 。
[0114] Optionally, as Figure 7 shown, the false detection area feature enhancement module obtains the query vector Q, the key vector K, and the value vector V by linear transformation of the feature map output after the first use of the non-linear activation function ReLU, and then respectively processes the global features and the false detection detail features through the axial global attention sub-module and the false detection area detail feature enhancement sub-module;
[0115] Among them, for the axial global attention sub-module to obtain the global features with close context connection, first, axial pooling operations are performed on the input feature map along the horizontal and vertical directions respectively to transform it into compact rows and columns. For the feature vector Q (taking the feature vector Q as an example, K and V are similar), H, W are the height and width of the feature map respectively, and C is the number of channels. The horizontal axial pooling result Q h and the vertical axial pooling result Q v are expressed as follows:
[0116]
[0117] Among them, since the number of channels of vector Q and vector K is the same, they are uniformly represented as C QK , Q →(·) represents arranging the dimensions of Q, and Ⅰ W , Ⅰ H are both vectors with all elements equal to 1;
[0118] Then, horizontal compression and vertical compression are respectively performed on Q h and Q v to achieve dimension conversion, obtaining two axial feature vectors of H×1×C QK and 1×W×C QK ;
[0119] After that, multi-channel and multi-angle attention is calculated for these two vectors respectively to capture the dependency relationship between the multi-angle axial feature vectors covering 360°, as Figure 8 shown, including:
[0120] Taking the horizontal and vertical axial feature vectors as inputs, using the Concat splicing function to splice them into a new feature matrix, designing multiple parallel channels and each channel using a 3×3 variable-direction convolutional kernel in a different direction, using the parameter θ, 0° < θ < 360°, to regulate the direction angle of the axial feature vector, and the weight parameter values in the convolutional kernel are all related to θ, so that the convolutional kernel assigns a higher weight at the specified angle θ to achieve the extraction of multi-angle axial feature vectors; then through a linear transformation, the corresponding Q, K, and V inputs for each channel are obtained, and the scaling dot product operation is respectively performed to obtain the Attention vector for each channel. This process is completed N times in the multi-channel and multi-angle attention mechanism, and the parameters of the linear transformation between channels are not shared. Then, the N Attention vectors are spliced, and after another linear transformation for fusion, the result obtained by the multi-channel and multi-angle attention mechanism is output;
[0121] On this basis, using a 1×1 convolutional layer to regress the dimension of the original input feature map, and finally outputting the global feature information enhanced by the axial global attention sub-module to reduce the misdetection phenomenon caused by the lack of combination of global context feature information;
[0122] The misdetection region detail feature enhancement sub-module, based on the texture characteristics of the interference objects (water stains, particles, etc.) in the image background that are easily misdetected as tears and have blurred boundaries, adopts a new detail enhancement design based on convolution to increase the boundary clarity for the model to distinguish them from the tear regions, as Figure 7 shown, including:
[0123] Performing channel splicing on Q, K, and V. Since the channels of Q and K are the same and both are C QK , so a size of H×W×(2C QK +CV ) feature map, and then pass it to a block consisting of 3×3 depth-separable convolution and batch normalization to assist in aggregating local details from Q, K, V, and then use a 1×1 convolution layer with ReLU6 activation function and batch normalization for linear projection to convert (2C QK +C V ) dimension is compressed to the original dimension, and the false detection detail enhancement feature is output.
[0124] S5. C d and C s Input and train the confidence fuser of the two-stream fusion network, output the final classification of the detected target, and determine the detection anchor frame position where the target is finally marked according to the predicted frame position of the detected target of the lightweight target detection branch;
[0125] Alternatively, if Figure 9 As shown, the confidence fusion device first converts C d and C s Splice into an input vector X = [C d ,C s ], and then input the vector X into a multi-layer perceptron neural network, whose output hi is expressed as:
[0126] hi=f(W i h i-1 +b i )
[0127] Among them, h 0 =X, f is the ReLU activation function, W i is the weight matrix of the i-th layer, b i is the bias vector, h of the last layer n The final fusion confidence is generated through an output operation. The fusion formula is as follows:
[0128] C fusion =ɑ(W out h n +b out )
[0129] Among them, ɑ is the Sigmoid function, which converts the output into a probability value between 0 and 1, W out is the weight matrix in the output operation, b out is the bias vector;
[0130] Finally, the two-stream fusion network is based on C fusion The value of is used to determine the final classification of the detected target, and at the same time, the detection anchor frame position where the target is finally marked is determined according to the predicted frame position of the detected target of the lightweight target detection branch.
[0131] S6. Use the trained dual-stream fusion network to perform belt tear detection on the belt surface image to be detected.
[0132] As Figure 10 shown, an embodiment of the present invention further provides a belt tear detection system based on key feature fusion. The system includes:
[0133] A collection, preprocessing, and partitioning module 1010, configured to collect and preprocess multiple belt surface images, perform annotation of object detection tasks and segmentation tasks on the multiple belt surface images respectively, and partition a training set and a validation set;
[0134] A shared feature extraction module 1020, configured to input the images in the training set and train the shared feature layer of the dual-stream fusion network, extract general feature information, and perform feature fusion based on a bidirectional feature pyramid, and output a fused feature image for sharing by subsequent dual-branch tasks;
[0135] An object detection module 1030, configured to input the fused feature image and train the lightweight object detection branch of the dual-stream fusion network for missed detection problems. Through a dual-branch perception attention mechanism, the global feature is closely combined with the key features of the missed detection category, and the sensitivity of the model to the feature information perception of small targets and small sample tear categories is enhanced specifically, and the localization detection anchor box of object detection and the category prediction result of the object are output. At the same time, the confidence C corresponding to the predicted category is output d ;
[0136] A segmentation decision module 1040, configured to input the fused feature image and train the lightweight segmentation decision branch of the dual-stream fusion network for false detection problems. Through a false detection feature enhancement module, the recognition ability of the model to the texture features of interference objects is improved specifically, and the interference objects in the image background are accurately distinguished from the true tear area, and the prediction result of whether the target truly belongs to a tear and the corresponding confidence C are output s ;
[0137] A fusion module 1050, configured to input C d and C s into the confidence fusion device of the dual-stream fusion network and train it, output the final classification of the detected target, and at the same time determine the position of the detection anchor box where the target is finally marked according to the position of the predicted box of the detected target of the lightweight object detection branch;
[0138] A detection module 1060, configured to use the trained dual-stream fusion network to perform belt tear detection on the belt surface image to be detected.
[0139] A belt tearing detection system based on key feature fusion provided by an embodiment of the present invention has a functional structure corresponding to a belt tearing detection method based on key feature fusion provided by an embodiment of the present invention, which will not be elaborated here.
[0140] Figure 11 FIG. 4 is a schematic structural diagram of an electronic device 1100 provided by an embodiment of the present invention. The electronic device 1100 may vary greatly due to different configurations or performances, and may include one or more central processing units (CPUs) 1101 and one or more memories 1102. Among them, at least one instruction is stored in the memory 302, and the at least one instruction is loaded and executed by the processor 1101 to implement the steps of the above-mentioned belt tearing detection method based on key feature fusion.
[0141] In an exemplary embodiment, a computer-readable storage medium is also provided, such as a memory including instructions. The above instructions can be executed by a processor in a terminal to complete the above-mentioned belt tearing detection method based on key feature fusion. For example, the computer-readable storage medium may be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.
[0142] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above embodiments can be completed by hardware, or can be completed by instructing relevant hardware through a program. The program can be stored in a computer-readable storage medium. The above-mentioned storage medium can be a read-only memory, a magnetic disk, or an optical disc, etc.
[0143] The above are only the preferred embodiments of the present invention, and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A belt tear detection method based on key feature fusion, characterized in that: The method comprises: S1. Collect and preprocess a plurality of belt surface images, annotate the plurality of belt surface images for target detection tasks and segmentation tasks, and divide the plurality of belt surface images into training sets and validation sets; S2, input the images in the training set and train the shared feature layer of the two-stream fusion network, extract common feature information, perform feature fusion based on the bidirectional feature pyramid, and output the fused feature image for sharing in subsequent two-branch tasks; The shared feature layer performs data enhancement for the belt surface image that is common to the dual-branch task, and then extracts features from part of the image through 9 convolution layers with a convolution kernel size of 5×5. Each convolution operation is followed by a batch normalization process to normalize each feature channel to a zero-mean distribution with unit variance, and then introduces a nonlinear transformation through a ReLU activation function, and finally reduces the resolution of each layer by half to obtain common feature information of different scales; another part of the image passes through 3 maximum pooling layers that will reduce the image resolution, reduces the spatial dimension and reduces the model calculation amount, while improving the generalization ability of the model and capturing higher-level and more abstract global feature information; After merging the common feature information of different scales with the pooled global feature information, efficient bidirectional cross-scale feature fusion is performed through the bidirectional feature pyramid structure in the shared feature layer. Not only are high-level abstract features upsampled layer by layer along the top-down path to match the scale of the bottom-level features and fuse with each other, but a bottom-up feedback mechanism is also introduced to gradually downsample the bottom-level features containing rich details and interact with the high-level features. After that, a fused feature image with the same dimension as the input feature is output through a 1x1 convolution layer. S3, input the fused feature image and train the lightweight target detection branch of the dual-stream fusion network for missed detection problems, and closely combine the global features with the key features of the missed detection category through the dual-branch perception attention mechanism, so as to enhance the model's sensitivity to the feature information of small targets and small sample tearing categories, output the positioning detection anchor frame of the target detection and the category prediction result of the target, and output the confidence C corresponding to the predicted category. d ; S4, input the fused feature image and train the lightweight segmentation decision branch of the dual-stream fusion network for the false detection problem, and improve the model's recognition ability of the texture features of the interference object through the false detection feature enhancement module, accurately distinguish the interference object in the image background from the real torn area, and output the prediction result of whether the target is really torn and the corresponding confidence C s ; S5. C d and C s Input and train the confidence fuser of the two-stream fusion network, output the final classification of the detected target, and determine the detection anchor frame position where the target is finally marked according to the predicted frame position of the detected target of the lightweight target detection branch; S6. Use the trained two-stream fusion network to perform belt tear detection on the belt surface image to be detected.
2. The method according to claim 1, characterized in that: The lightweight object detection branch extracts preliminary feature information from the fused feature image through a 3×3 initial convolution layer, and then sends it to a batch normalization layer for standardization, and then performs nonlinear transformation through an h-swish activation function, thereby obtaining basic features provided to the lightweight detection module; The lightweight detection module inputs the basic features into a 3×3 depth-separable convolution layer with a step size of 2 to extract the deep convolution features, and then integrates the deep convolution features through 1×1 point-by-point convolution and reduces the number of feature channels. The feature map output by the point-by-point convolution is then nonlinearly transformed through the h-swish activation function to enhance the expression ability of the model. The above steps are repeated 11 times to obtain 11 feature maps F with different resolutions. i , i = 1, 2, ..., 11, a top-down approach in a feature pyramid is used to gradually merge feature maps of adjacent levels until a fusion feature map F_fusion is formed; The fusion feature map F_fusion is input into the dual-branch perception attention mechanism module, and the global feature perception attention submodule of the dual-branch perception attention mechanism module outputs the global feature weighted vector X global , the missed detection category feature perception attention submodule of the dual-branch perception attention mechanism module outputs the missed detection category feature weighted vector X local , X global and X local The two are concatenated and the channel size is converted through the fully connected layer to obtain the dual-branch perception attention output X output ; According to X output , the positioning loss function, confidence loss function, and classification loss function are calculated for the target detection anchor box and target category prediction, and the category balance loss function is added to specifically solve the problem of unbalanced data category ratio. The calculation formula of the category balance loss function is as follows: F(p t )=-(1-p t ) γ log(p t ) Among them, p t is the model's predicted probability that a sample is a positive class. γ is used as a modulation parameter to control the relative loss contribution of easy-to-classify samples and difficult samples. When γ>0, difficult samples will obtain a larger loss weight, making the model pay more attention to sample categories that are easily misclassified. After this calculation process, the model outputs the positioning detection anchor box of the target detection and the target category prediction result, and also outputs the confidence C corresponding to the predicted category. d .
3. The method according to claim 2, characterized in that The dual-branch perception attention mechanism module multiplies the input fusion feature map F_fusion with different weight parameters to obtain the input of the three branches Q, K, and V, and then passes through three 1×1 convolutions Query_Conv, Key_Conv, and Value_Conv respectively, and uses the view function in the Pytorch environment to transform the dimensions of the Query and Key feature vectors, and uses the permute function to invert the Query vector to obtain feature vector outputs of three different dimensions Q, K, and V; The global feature perception attention submodule of the dual-branch perception attention mechanism module extracts global information through the fully connected layer to obtain Q_global, K_global and V_global, and then downsamples K_global and V_global through global pooling, multiplies Q_global with the downsampled K_global matrix and performs Softmax normalization, and then multiplies the result with the downsampled V_global matrix, finally extracts low-frequency global information, and obtains the global feature weighted vector X global , the formula is as follows: X global =Softmax(Q_global·Pool(K_global))·Pool(V_global) The missed category feature perception attention submodule of the dual-branch perception attention mechanism module specifically captures the high-frequency feature information corresponding to the missed category for perception enhancement. For feature vectors of three different dimensions, Q, K, and V, three depthwise separable convolutions with a step size of 1×1 are used to extract local information to obtain Q_local, K_local, and V_local. At the same time, the weights of the depthwise separable convolution layer are globally shared, and then the Hadamard product of Q_local and K_local is calculated, so that the two are combined to generate a context-aware feature matrix M. context , then through a fully connected layer and a Swish activation function, and then through a fully connected layer and a Tanh activation function, the result is feature integrated and nonlinearly transformed to obtain a context-aware feature W between -1 and 1 context , the relevant formula is expressed as follows: M context =Q_local⊙K_local Where n represents the number of feature channels, the symbol ⊙ represents the Hadamard product, and FC represents the fully connected layer; Then use W context The Hadamard product is performed again with V_local after the 1×1 depth-separable convolution feature information is integrated to obtain the feature weighted vector X after the feature perception of the missed category is enhanced. local , X local =W context ⊙V_local; X global and X local The two are concatenated and the channel size is converted through the fully connected layer to obtain the dual-branch perception attention output X output , X output =FC(Contact(X global ,X local )), so that the model can effectively perceive high and low frequency information at the same time.
4. The method according to claim 1, characterized in that: The lightweight segmentation decision branch first passes the fused feature image through a lightweight segmentation module and uses a deep separable convolutional network for lightweight improvement, specifically including: Use a depthwise separable convolution with a size of 15×15 and a stride of 1 to extract deep convolutional features, then use a 1×1 point-by-point convolution to integrate the deep convolutional features and reduce the number of feature channels. The nonlinear activation function ReLU is used for the first time to increase the amount of feature map information, and then a 1×1 convolution layer is used to reduce the number of output channels, and batch normalization is performed and the nonlinear activation function ReLU is used again to obtain a single-channel feature map; While obtaining the single-channel feature map, the feature map output after the first use of the nonlinear activation function ReLU is input into the false detection area feature enhancement module, the axial global attention submodule of the false detection area feature enhancement module outputs the enhanced global feature information, the false detection area detail feature enhancement submodule of the false detection area feature enhancement module outputs the false detection detail enhancement feature, the enhanced global feature information is multiplied and fused with the false detection detail enhancement feature to obtain a feature map with higher semantic information, which is spliced with the single-channel feature map to obtain the feature map output of the lightweight segmentation module as the input of the decision module; The decision module outputs the feature map of the lightweight segmentation module through a two-layer combination repeated three times, wherein the two-layer combination includes a 2×2 maximum pooling layer and a convolution layer with a convolution kernel size of 5×5, and the number of channels is set to increase as the feature resolution decreases. The three layers of convolution are set to 8, 16, and 32 channels respectively, and then global maximum pooling and global average pooling operations are performed to generate 64 output neurons; in addition, the single-channel feature map output in the lightweight segmentation module is also connected to 2 output neurons after the results of global maximum pooling and global average pooling respectively; and then 66 neurons are output through the fully connected layer, and after the integration process of the linear weight combination layer and the use of the sigmoid activation function, the prediction result of whether the target is truly torn and its corresponding confidence C are finally output. s .
5. The method according to claim 4, characterized in that The false detection area feature enhancement module obtains the query vector Q, the key vector K and the value vector V through linear transformation of the feature map output after the first use of the nonlinear activation function ReLU, and then processes the global features and the false detection detail features respectively through the axial global attention submodule and the false detection area detail feature enhancement submodule; Among them, the axial global attention submodule, in order to obtain global features closely related to the context, first performs axial pooling operations on the input feature map along the horizontal and vertical directions to transform it into compact rows and columns. For the feature vector Q, H and W are the height and width of the feature map respectively, C is the number of channels, and the horizontal axial pooling result Q of the H×W×C feature map is h And the vertical axis pooling result Q v The expression is as follows: Among them, since the number of channels of vector Q and vector K is the same, they are uniformly expressed as C QK , Q →(·) Indicates the arrangement of the dimensions of Q, I W Ⅰ H is a vector whose elements are all equal to 1; Then Q h and Q v Perform lateral and longitudinal compression to achieve dimensional conversion and obtain H×1×C QK and 1×W×C QK The two axial eigenvectors of ; Then, multi-channel and multi-angle attention is calculated for these two vectors to capture the dependencies between multi-angle axial feature vectors covering 360°, including: The horizontal and vertical axial feature vectors used as input are concatenated into a new feature matrix using the Concat concatenation function. Multiple parallel channels are designed, and each channel uses a 3×3 directional convolution kernel with different directions. The parameter θ, 0°<θ<360°, is used to adjust the direction angle of the axial feature vector. The weight parameter values in the convolution kernel are all related to θ, so that the convolution kernel assigns a higher weight to the specified angle θ, realizing the extraction of multi-angle axial feature vectors. Then, a linear transformation is performed to obtain the Q, K, and V inputs corresponding to each channel, and the scaling dot product operation is performed to obtain the attention vector of each channel. This process is completed N times in the multi-channel multi-angle attention mechanism. The parameters of the linear transformation between the channels are not shared. The N attention vectors are then concatenated and fused by a linear transformation. The results of the multi-channel multi-angle attention mechanism are output. On this basis, a 1×1 convolution layer is used to regress the dimension of the original input feature map, and finally the global feature information enhanced by the axial global attention submodule is output to reduce the false detection phenomenon caused by the lack of global context feature information. The submodule for enhancing the detail features of the misdetected area adopts a new convolution-based detail enhancement design based on the texture characteristics of the blurred boundaries of the interference objects in the image background that are easily misdetected as torn areas, which increases the clarity of the boundaries and facilitates the model to distinguish them from the torn areas, including: Connect the channels of Q, K, and V. Since the channels of Q and K are the same, they are both C. QK , so the size is H×W×(2C QK +C V ) feature map, and then pass it to a block consisting of 3×3 depth-separable convolution and batch normalization to assist in aggregating local details from Q, K, V, and then use a 1×1 convolution layer with ReLU6 activation function and batch normalization for linear projection to convert (2C QK +C V ) dimension is compressed to the original dimension, and the false detection detail enhancement feature is output.
6. The method according to claim 1, characterized in that The confidence fuser first converts C d and C s Splice into an input vector X = [C d ,C s ], and then input the vector X into a multi-layer perceptron neural network, whose output hi is expressed as: hi=f(W i h i-1 +b i ) Where h0=X, f is the ReLU activation function, W i is the weight matrix of the i-th layer, b i is the bias vector, h of the last layer n The final fusion confidence is generated through an output operation. The fusion formula is as follows: C fusion =ɑ(W out h n +b out ) Among them, ɑ is the Sigmoid function, which converts the output into a probability value between 0 and 1, W out is the weight matrix in the output operation, b out is the bias vector; Finally, the two-stream fusion network is based on C fusion The value of is used to determine the final classification of the detected target, and at the same time, the detection anchor frame position where the target is finally marked is determined according to the predicted frame position of the detected target of the lightweight target detection branch.
7. A belt tear detection system based on key feature fusion, characterized in that: The system comprises: A collection preprocessing and segmentation module is used to collect and preprocess multiple belt surface images, annotate the multiple belt surface images for target detection tasks and segmentation tasks, and divide the images into training sets and verification sets; The shared feature extraction module is used to input the images in the training set and train the shared feature layer of the two-stream fusion network, extract common feature information, perform feature fusion based on the bidirectional feature pyramid, and output the fused feature image for sharing by subsequent two-branch tasks; The shared feature layer performs data enhancement for the belt surface image that is common to the dual-branch task, and then extracts features from part of the image through 9 convolution layers with a convolution kernel size of 5×5. Each convolution operation is followed by a batch normalization process to normalize each feature channel to a zero-mean distribution with unit variance, and then introduces a nonlinear transformation through a ReLU activation function, and finally reduces the resolution of each layer by half to obtain common feature information of different scales; another part of the image passes through 3 maximum pooling layers that will reduce the image resolution, reduces the spatial dimension and reduces the model calculation amount, while improving the generalization ability of the model and capturing higher-level and more abstract global feature information; After merging the common feature information of different scales with the pooled global feature information, efficient bidirectional cross-scale feature fusion is performed through the bidirectional feature pyramid structure in the shared feature layer. Not only are high-level abstract features upsampled layer by layer along the top-down path to match the scale of the bottom-level features and fuse with each other, but a bottom-up feedback mechanism is also introduced to gradually downsample the bottom-level features containing rich details and interact with the high-level features. After that, a fused feature image with the same dimension as the input feature is output through a 1x1 convolution layer. The target detection module is used to input the fused feature image and train the lightweight target detection branch of the dual-stream fusion network for missed detection problems. Through the dual-branch perception attention mechanism, the global features are closely combined with the key features of the missed detection category, and the model is targeted to enhance the perception sensitivity of the feature information of small targets and small sample tearing categories. The positioning detection anchor frame of the target detection and the category prediction result of the target are output, and the confidence C corresponding to the predicted category is output at the same time. d ; The segmentation decision module is used to input the fusion feature image and train the lightweight segmentation decision branch of the dual-stream fusion network for the false detection problem. Through the false detection feature enhancement module, the model's ability to recognize the texture features of the interference object is improved in a targeted manner, and the interference object in the image background and the real torn area are accurately distinguished. The prediction result of whether the target is really torn and the corresponding confidence C are output. s ; Fusion module for C d and C s Input and train the confidence fuser of the two-stream fusion network, output the final classification of the detected target, and determine the detection anchor frame position where the target is finally marked according to the predicted frame position of the detected target of the lightweight target detection branch; The detection module is used to use the trained two-stream fusion network to perform belt tear detection on the belt surface image to be detected.
8. An electronic device, comprising a processor and a memory, wherein at least one instruction is stored in the memory, wherein: The at least one instruction is loaded and executed by the processor to implement the belt tear detection method based on key feature fusion as described in any one of claims 1-6.
9. A computer-readable storage medium, wherein at least one instruction is stored in the storage medium, characterized in that: The at least one instruction is loaded and executed by the processor to implement the belt tear detection method based on key feature fusion as described in any one of claims 1-6.
Citation Information
Patent Citations
An object detection method based on semantic segmentation enhancement
CN109214349A
Automatic gallstone recognition and segmentation system based on deep learning, computer equipment and storage medium
CN112233777A