Synthetic leather surface defect detection method based on space instance dissimilatory self-attention
By introducing spatial instance-specific self-attention bottleneck structure and self-attention mechanism into the object detection algorithm, the problem of single-sample feature specificity being ignored and the global weight weighting mechanism is missing, achieving higher detection accuracy and robustness.
Patent Information
- Application Number
- CN202510220947.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-02-27
AI Technical Summary
The existing object detection algorithm ignores the specificity of single-sample features in surface defect detection, lacks a global weight weighting mechanism, and fails to effectively integrate batch generalization features and single-sample features, resulting in unsatisfactory detection accuracy.
The synthetic leather surface defect detection method based on space instances is adopted. By constructing a space instance-specific self-attention bottleneck structure, a self-attention mechanism and a multi-level feature fusion strategy are introduced to replace the bottleneck structure in the backbone network of the YOLOv8 algorithm.
It improves the accuracy and robustness of object detection, enhances the perception of important areas in the image, effectively integrates batch generalization features and single-sample features, and improves detection accuracy and generalization capabilities.
Smart Images

Figure CN120070402A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of surface defect detection, and particularly relates to a synthetic leather surface defect detection method based on spatial instance-specific self-attention. Background Art
[0002] Synthetic leather surface defect detection is a very important application field in synthetic leather quality control. With the improvement of manufacturing automation, higher and higher requirements are put forward for surface defect detection on the production line. Traditional surface defect detection methods rely on manual inspection or rule-based machine vision algorithms, with low efficiency and prone to missed detection or false detection when facing complex environments or variable surface textures. In recent years, the successful application of deep learning in computer vision has made surface defect detection methods based on convolutional neural networks (CNNs) gradually become the mainstream. CNNs can automatically learn multi-level feature expressions of images, thus maintaining high detection accuracy in complex environments. Among them, the YOLO algorithm series is particularly suitable for surface defect detection scenarios in real-time monitoring and large-scale production lines due to its fast and efficient detection ability.
[0003] The YOLO (You Only Look Once) algorithm series is a very important technology in the field of object detection. Its core concept is to achieve object detection with only one forward pass on an image. Compared with traditional two-stage detectors (such as Faster R-CNN), YOLO regards the object detection problem as a single regression problem and directly predicts the category, location, and confidence score of the object from the input image. With the evolution of the YOLO series (from YOLOv1 to YOLOv7), its model has been continuously optimized, and significant improvements have been made in both detection speed and accuracy. Although using the YOLO algorithm series based on deep learning can solve the problem of scene diversity that cannot be solved by traditional vision and machine learning algorithms, there are also many problems or defects in the YOLO series algorithms themselves:
[0004] 1) Ignoring the specificity of single-sample features: In typical object detection algorithms, the extraction of image features mainly relies on the bottleneck structure in the backbone network. The bottleneck structure consists of multiple convolutional layers, BN layers, and activation functions. Among them, the BN layer calculates the mean and variance of batch images and adjusts the distribution of batch features using trainable parameters. This design tends to focus more on the generalization features of the entire batch of images and ignores the unique characteristics of each sample. This deficiency may lead to the algorithm being unable to fully capture single-sample features, thus affecting the overall detection accuracy.
[0005] 2) Lack of a global weight weighting mechanism for single-sample features: The bottleneck structures of existing object detection algorithms are mainly implemented through convolutional operations, lacking a global feature attention module. This design limits the network's ability to perceive global information in images, making it difficult to effectively extract the global weight features of single samples, which in turn has an adverse impact on the detection performance and reduces the detection accuracy.
[0006] 3) Failure to effectively fuse batch generalization features and single-sample features: The bottleneck structures of current object detection algorithms usually adopt a single convolutional method, with limited feature extraction ability, and do not combine multiple feature fusion strategies. This design is insufficient to achieve the effective cooperation of batch generalization features and single-sample features, further restricting the performance of the network in detection tasks and resulting in less than ideal detection accuracy. Summary of the Invention
[0007] To solve the above problems, the present invention provides a synthetic leather surface defect detection method based on spatial instance-specific self-attention. The specific scheme includes the following steps:
[0008] S1. Construct a spatial instance-specific self-attention bottleneck structure, which includes a main branch and an auxiliary branch; wherein, the output of the main branch and the output of the auxiliary branch are jointly fused with the input of the spatial instance-specific self-attention bottleneck structure, and the fusion result is used as the output of the spatial instance-specific self-attention bottleneck structure;
[0009] S2. Replace the bottleneck structure in the backbone network of the YOLOv8 algorithm with a spatial instance-specific self-attention bottleneck structure to obtain a surface defect detection model;
[0010] S3. Obtain a surface defect sample data set to train the surface defect detection model;
[0011] S4. Input the data to be detected into the trained surface defect detection model to obtain the detection result.
[0012] Further, the main branch includes two cascaded CBS modules, and each CBS module includes a 3×3 convolutional layer, a BN layer, and a SiLU activation function layer; the number of input channels of each CBS module is the same as the number of output channels.
[0013] Further, the auxiliary branch includes a first CIS module, an unfold layer, a self-attention mechanism module, a fold layer, and a second CIS module cascaded in sequence.
[0014] Further, the first CIS module includes a 1×1 convolutional layer, an IN layer, and a SiLU activation function layer; the second CIS module includes a 3×3 convolutional layer, an IN layer, and a SiLU activation function layer.
[0015] Furthermore, the strides of the first CIS module and the second CIS module are the same, denoted as s; the number of input channels of the spatially instance-specific self-attention bottleneck structure is denoted as c, the number of input channels of the first CIS module is c, and the number of output channels is (c / / s)×3; the number of input channels of the second CIS module is c / / s, and the number of output channels is c; where / / represents integer division operation, for example, 5 / / 2 = 2.
[0016] Furthermore, the calculation process of the self-attention mechanism module is expressed as
[0017]
[0018] Q = XW Q
[0019] K = XW K
[0020] V = XW V
[0021] where X represents the input of the self-attention mechanism module, and Output represents the output of the self-attention mechanism module; W Q 、W K 、W V represent weight matrices; Q represents the query vector matrix, K represents the key vector matrix, V represents the value vector matrix, Softmax() represents the softmax function, and d represents the dimension of the key vector.
[0022] Advantages of the present invention:
[0023] The present invention introduces a spatially instance-specific self-attention bottleneck structure, which helps to reduce the problems of false detection and missed detection caused by insufficient exploration of single-sample specificity, thereby improving the accuracy of object detection. The present invention models the specificity of single samples, better adapts to the differences between samples, and improves the performance of the overall network.
[0024] The introduction of the self-attention mechanism in the spatially instance-specific self-attention bottleneck structure enables the network to redistribute the weights of the feature map globally, enhancing the ability to perceive important regions in the image. Compared with traditional bottleneck structures that rely on convolutional operations for feature extraction, the self-attention mechanism of the present invention enhances the network's ability to capture global information, especially by weighting the relationship between space and channels, improving the accuracy and robustness of object detection.
[0025] In the spatial instance-specific self-attention bottleneck structure, three-branch fusion is performed, enabling the effective fusion of batch generalization features and single-sample specific features at the end of the network. Through multi-level feature fusion, the network can utilize features from different sources simultaneously, giving full play to the advantages of both and enhancing the detection ability for complex targets, while ensuring high detection accuracy in different detection scenarios. This optimization improves the generalization ability of the network, enabling it to perform well in various scenarios.
[0026] By adopting an improved bottleneck structure, the network can extract and fuse features more efficiently, thereby enhancing the model's expressive ability and non-linear characteristics. This enables the model to better adapt to diverse input data in complex object detection tasks and improves the detection effect.
[0027] Despite the introduction of the self-attention mechanism and more complex feature fusion strategies, the present invention significantly improves performance while maintaining the computational efficiency of the network. By reducing unnecessary floating-point operations and the number of parameters, the running speed of the model is increased, while ensuring low computational resource consumption, making it suitable for deployment in practical applications. Brief Description of the Drawings
[0028] Figure 1 It is a flowchart of the method of the present invention;
[0029] Figure 2 It is a comparison diagram of each bottleneck structure of the present invention;
[0030] Figure 3 It is a detailed schematic diagram of the spatial instance-specific self-attention bottleneck structure of the present invention. Detailed Embodiment
[0031] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0032] In the backbone network of the YOLO algorithm series, the most critical component is the bottleneck structure (Bottleneck), which is mainly used for the step-by-step extraction of image features and the fusion of features at different levels. The bottleneck structure helps the model maintain effective gradient propagation and improve the model's fitting ability and performance by reducing the network depth and avoiding the gradient vanishing problem. Specifically, the bottleneck structures adopted by YOLOv4 and its subsequent versions (such as YOLOv5, YOLOX) are as Figure 2As shown in (a), it is a CSPNet structure composed of 1×1 convolution and 3×3 convolution. This structure effectively alleviates the problem of gradient disappearance, accelerates network convergence, and enhances feature learning ability. However, the 1×1 convolution introduced to reduce network parameters leads to problems such as feature loss and a smaller receptive field.
[0033] To solve this problem, the bottleneck structure adopted in YOLOv7 is an improved structure based on the ELAN network, Figure 2 as shown in (b). The bottleneck structure adopted in YOLOv8 is further developed based on the ELAN network, Figure 2 as shown in (c). The bottleneck structure improved based on the ELAN network uses two layers of 3×3 convolution to replace the original combination of 1×1 convolution and 3×3 convolution, which can effectively improve the information flow of the network, reduce floating-point operations and the number of parameters, and at the same time enhance the expression ability and non-linear characteristics of the model, thereby improving the generalization ability of the model and optimizing the network performance. Although the bottleneck structure in the ELAN structure already has very good effects, there are still the following problems: 1) The specificity of single-sample features is ignored during the feature extraction process; 2) There is a lack of a global weight weighting mechanism for single-sample features; 3) It fails to effectively fuse batch generalization features and single-sample features.
[0034] For the above reasons, the present invention proposes an improved bottleneck structure and names it "Spatial Instance Specialized Self-Attention Bottleneck Module"; on the basis of this structure, a synthetic leather surface defect detection method based on spatial instance specialized self-attention is provided, as Figure 1 shown, including the following steps:
[0035] S1. Construct a spatial instance specialized self-attention bottleneck structure, which includes a main branch and an auxiliary branch; wherein, the output of the main branch and the output of the auxiliary branch are jointly fused with the input of the spatial instance specialized self-attention bottleneck structure, and the fusion result is used as the output of the spatial instance specialized self-attention bottleneck structure;
[0036] S2. Use the spatial instance specialized self-attention bottleneck structure to replace the bottleneck structure in the backbone network of the YOLOv8 algorithm, that is, replace the bottleneck layer Bottleneck in the feature pyramid layer C2f, to obtain a surface defect detection model;
[0037] S3. Obtain a surface defect sample data set to train the surface defect detection model;
[0038] S4. Input the data to be detected into the trained surface defect detection model to obtain the detection result.
[0039] Specifically, the spatial instance specialized self-attention bottleneck structure is as Figure 3As shown in the figure, on the basis of the existing network with a main branch plus a residual connection, an additional auxiliary branch is added. The auxiliary branch includes a first CIS module, an unfold layer, a self-attention mechanism module, a fold layer, and a second CIS module cascaded in sequence. The main branch includes two cascaded CBS modules. Each CBS module includes a 3×3 convolutional layer, a BN layer, and a SiLU activation function layer. The number of input channels of each CBS module is the same as the number of output channels. As Figure 3 shown, the number of input and output channels of the CBS module is both c, which is the same as the number of input image channels.
[0040] Specifically, the first CIS module includes a 1×1 convolutional layer, an Instance Normalization (IN) layer, and a SiLU activation function layer. The second CIS module includes a 3×3 convolutional layer, an IN layer, and a SiLU activation function layer. In the present invention, the instance normalization layer is introduced to independently normalize the features of each sample, and the feature distribution of each sample is learned through trainable parameters. This design effectively enhances the network's learning ability for the specific differences of single samples, thereby improving the performance of the overall network.
[0041] Specifically, the bottleneck structure in the existing object detection algorithms usually uses convolutional operations to extract features, but it cannot capture the global relationships in the image. In the present invention, a self-attention mechanism is introduced into the spatial instance specialization self-attention bottleneck structure to solve the problem of the lack of a global weight weighting mechanism in the feature extraction process. The self-attention mechanism divides the input feature map into spatial blocks and calculates the global self-attention according to the relationship between the channels and the positions of the spatial blocks, further capturing the interactions between the spatial blocks and enhancing the network's perception ability of spatial global information, effectively improving the feature extraction ability. Suppose the input feature map is where C is the number of channels of the feature map, H is the height of the feature map, and W is the width of the feature map. In the present invention, first, the input feature map X1 is unfolded into 4×4 image blocks through the unfold layer and converted into the channel dimension to obtain the feature map Z represents the dimension number of 4×4, and L is the number of the feature map unfolded into one dimension by image blocks. The specific calculation process is as follows:
[0042]
[0043] Q = XW Q
[0044] K = XW K
[0045] V = XW V
[0046] Among them, Output represents the output of the self-attention mechanism module; W Q , W K , W V represent weight matrices, and M represents the convolution kernel size; Q represents the query vector matrix, K represents the key vector matrix, V represents the value vector matrix, Softmax() represents the softmax function, d represents the dimension of the key vector, and T represents the matrix transpose operation.
[0047] Specifically, a multi-level feature fusion strategy is adopted at the end of the spatial instance specialization self-attention bottleneck structure, that is, the output of the main branch and the output of the auxiliary branch are effectively fused with the input of the spatial instance specialization self-attention bottleneck structure. Through this fusion mechanism, the collaborative work of batch generalization features and single-sample features can be achieved, thereby improving the overall detection accuracy of the network.
[0048] In the present invention, unless otherwise clearly defined and limited, terms such as "installation", "setting", "connection", "fixation", "rotation", etc. shall be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or integrated; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium. It can be the internal communication of two components or the interaction relationship between two components. Unless otherwise clearly defined, for those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0049] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A synthetic leather surface defect detection method based on spatial instance-specific self-attention, characterized in that: The following steps are involved: S1. construct a spatial instance-specific self-attention bottleneck structure, which includes a main branch and an auxiliary branch; wherein the output of the main branch and the output of the auxiliary branch are fused with the input of the spatial instance-specific self-attention bottleneck structure, and the fusion result is used as the output of the spatial instance-specific self-attention bottleneck structure; S2. The bottleneck structure in the YOLOv8 algorithm backbone network is replaced by the spatial instance-specific self-attention bottleneck structure to obtain a surface defect detection model; S3. Obtain a surface defect sample data set to train a surface defect detection model; S4. Input the data to be tested into the trained surface defect detection model to obtain the test results.
2. The method for synthetic leather surface defect detection based on spatial instance-specific self-attention according to claim 1, characterized in that: The main branch includes two cascaded CBS modules, each of which includes a 3×3 convolutional layer, a BN layer, and a SiLU activation function layer; the number of input channels of each CBS module is the same as the number of output channels.
3. The method for synthetic leather surface defect detection based on spatial instance-specific self-attention according to claim 1, characterized in that: The auxiliary branch includes a first CIS module, an unfold layer, a self-attention mechanism module, a fold layer and a second CIS module, which are cascaded in sequence.
4. The method for synthetic leather surface defect detection based on spatial instance-specific self-attention according to claim 3, characterized in that: The first CIS module includes a 1×1 convolution layer, an IN layer, and a SiLU activation function layer; the second CIS module includes a 3×3 convolution layer, an IN layer, and a SiLU activation function layer.
5. The method for synthetic leather surface defect detection based on spatial instance-specific self-attention according to claim 3, characterized in that: The stride of the first CIS module and the second CIS module is the same, denoted by s; the number of input channels of the spatial instance-specific self-attention bottleneck structure is c, the number of input channels of the first CIS module is c, and the number of output channels is (c / / s)×3; the number of input channels of the second CIS module is c / / s, and the number of output channels is c; where / / represents integer division operation.
6. The method for synthetic leather surface defect detection based on spatial instance-specific self-attention according to claim 3, characterized in that: The calculation process of the self-attention mechanism module is expressed as Q=XW Q K=XW K V=XW V Where X represents the input of the self-attention mechanism module, Output represents the output of the self-attention mechanism module; W Q , W K , W V represents the weight matrix; Q represents the query vector matrix, K represents the key vector matrix, V represents the value vector matrix, Softmax() represents the softmax function, and d represents the dimension of the key vector.
Citation Information
Patent Citations
Heterogeneous graph network node classification method based on graph self-attention mechanism
CN116821776A
Metal plate surface defect detection method and system
CN118212207A
Light-weight leather surface defect salient target detection method introducing RBU
CN118411514A
Fan surface defect detection method, device and equipment based on lightweight PC-EMA algorithm and storage medium
CN118691573A
Dam body defect identification method based on attention feature fusion enhancement network
CN119027795A