Surface defect detection method based on YOLOX multi-scale network reconstruction
By introducing multi-scale aggregation module and vision center module into the YOLOX algorithm, the shortcomings of the YOLO series algorithms in complex feature detection and multi-scene applications are solved, and more efficient surface defect detection performance and robustness are achieved.
Patent Information
- Application Number
- CN202510106590.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-05-23
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing surface defect detection method based on the YOLO series algorithm is not ideal in the detection of complex features, and is poorly robust to multi-scene applications, making it difficult to adapt to large-scale production processes.
By introducing multi-scale aggregation module and vision center module in the YOLOX algorithm, cross-layer feature fusion and attention to image center area are realized, and the feature extraction ability of the model and adaptability to defects at different scales are enhanced.
It significantly improves the performance of surface defect detection, improves the model's detection ability of small objects and dense scenes in complex environments, and enhances the robustness and adaptability of the model.
Smart Images

Figure CN120031828A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of deep learning and machine vision, and specifically relates to a surface defect detection method based on YOLOX multi-scale network reconstruction. Background Art
[0002] First, with the continuous development of deep learning, object detection methods based on convolutional neural networks are constantly being applied to surface defect detection scenarios. These methods can automatically extract features, reduce manual intervention, and improve detection efficiency and accuracy. However, traditional methods require manual creation of defect features. Faced with multi-scenario application environments, they have a large workload and poor robustness, and are not suitable for large-scale defect detection production processes.
[0003] Secondly, although deep learning algorithms have achieved detection speeds that traditional defect detection methods cannot achieve to a certain extent, they are still lacking in accuracy. For example, although the YOLO series of algorithms have fast detection speeds, their structures are relatively simple and the detection effect on complex features is not ideal. Therefore, researchers are committed to improving the YOLOX algorithm and improving detection accuracy through multi-scale feature fusion technology to meet the complex and diverse surface defect detection needs. Summary of the invention
[0004] In order to solve the technical problems existing in the prior art, the present invention provides a surface defect detection method based on YOLOX multi-scale network reconstruction. The YOLOX algorithm is designed to integrate a multi-scale aggregation module and a visual center module, which significantly improves the performance of surface defect detection. The scale aggregation module strengthens the information exchange between features of different scales through cross-layer feature fusion, so that the model can more effectively capture the features of defects of different sizes, which is crucial for surface defect detection because the size and shape of defects may vary greatly. At the same time, the visual center module improves the detection ability of small objects and objects in dense scenes by focusing on the visual center area of the image, which is particularly critical for accurate positioning and identification in surface defect detection, especially the identification of subtle defects in complex backgrounds. The fusion of these two modules not only enhances the feature extraction capability of the model, but also improves the adaptability and robustness of the model to defects of different scales and complex environments.
[0005] To achieve the above purpose, the technical solution adopted by the present invention is: a surface defect detection method based on YOLOX multi-scale network reconstruction, the specific steps are as follows:
[0006] Step S1, downloading a surface defect detection data set, and preprocessing the surface images in the data set;
[0007] Step S2: convert the preprocessed data set into the YOLO standard format, that is, convert the annotation file in XML format into a label file in TXT format;
[0008] Step S3, construct a backbone extraction network backbone based on YOLOX. The backbone extraction network backbone has five layers of feature extraction networks of different sizes. Each layer consists of a deep separable module and a multi-scale attention aggregation module, and uses a series connection method to perform deep information mining on the image;
[0009] Step S4: The feature extraction network uses a depth-separable module to extract features. The calculation is divided into two parts. One part is to apply a 3x3 depthwise convolution on a single input channel, and the other part is to use a 1x1 pointwise convolution. The pointwise convolution operation performs a weighted combination of the channel-by-channel convolution feature layers of the previous depthwise convolution in the depth direction.
[0010] Step S5: In the channel aggregation path of the multi-scale attention aggregation module, a channel attention map is generated through global average pooling, 1×1 convolution calculation and RELU activation function operation to achieve multi-scale information fusion in the channel dimension;
[0011] Step S6: In the spatial refinement path of the multi-scale attention aggregation module, the input features are first up-sampled by 1×1 convolution and divided into three aggregation nodes. 3×3, 5×5, and 7×7 convolution kernels are used to extract feature information of different depths, and multi-scale aggregation features are accumulated layer by layer.
[0012] Step S7, the features output by the two network branches after multi-scale aggregation processing are subjected to weight matrix calculation, and are concatenated with the feature input without any processing;
[0013] Step S8, constructing the YOLOX feature multi-scale fusion network Neck, before inputting the multi-scale fusion network, using the spatial explicit visual center mechanism and the lightweight MLP module to capture the global and local features of the outputs of multiple feature extraction layers, and retaining the local key area information in the input image;
[0014] Step S9, proposing a spatially explicit visual center scheme, including a lightweight MLP for capturing global long-range dependencies and a learnable visual center for aggregating local key areas;
[0015] Step S10, the learnable visual center mechanism aggregates the input image to average the operation, calculates the average value in the feature depth dimension direction and compresses it into a two-dimensional feature, defines a linear layer, sets the number of input neurons, the number of output neurons, and the bias of the parameters, so that the number of input neurons is equal to the number of output neurons, inputs the two-dimensional feature, and adopts a tiling operation to break up the two-dimensional feature into a one-dimensional vector form, and uses the broadcast mechanism and automatic reshaping function of the linear layer to achieve, and obtains the weight factor of the channel attention mechanism;
[0016] Step S11, adopting a lightweight MLP architecture to capture the global dependency of a specific feature size, performing averaging processing in the feature depth dimension direction, compressing it into a two-dimensional feature, and obtaining a set of weight factors for a channel attention mechanism through a learnable visual center mechanism to perform matrix calculation with the input image features;
[0017] Step S12, introduce the decoupling head of YOLOx to process the classification and positioning tasks respectively, separate the category prediction and the position prediction, and use two independent network branches for processing respectively. Each branch includes two 3×3 convolutional layers, which are used for classification and regression tasks respectively. The classification branch output is H×W×C, the regression branch output is H×W×4, and the IoU branch output is H×W×1;
[0018] Among them, H and W represent the size of the feature map, and C represents the number of feature channels;
[0019] Step S13, fine-tune the model parameters, set the learning rate to 0.0001, the number of iterations to 100 rounds, the input batch to 32, and the discard rate to 0.2, load the data set for model training to obtain the optimal weight, and load the trained model for defect detection.
[0020] In step S1, specific preprocessing steps include: performing signal enhancement through Fourier transform, removing noise through low-pass filtering, and enhancing edge high-frequency signals through high-pass filtering, thereby performing image enhancement.
[0021] In step S9, a set of scaling factors are sequentially mapped to corresponding position information, and the full image information about the kth codeword is calculated in the following manner:
[0022]
[0023] in, represents the i-th pixel, b k is the kth learnable visual codeword, s k is the kth scaling factor, x i -b krepresents the visual feature information, that is, the position of each pixel relative to the image, K is the total number of visual centers, N is the total number of feature layer channels, and e k is the information of the entire image of the kth codeword.
[0024] In step S12, the loss-based allocation strategy is used to transform the positive and negative sample allocation problem into an optimal transmission problem, and the positive and negative samples are allocated by minimizing the transmission cost:
[0025]
[0026] in, represents the confidence score of the reference box prediction, represents the bounding box predicted for the reference box, Represents the confidence score of the true box, represents the bounding box of the real box, L cls represents the cross entropy loss, L reg represents the intersection-over-union loss, α is the balance coefficient of the two losses, θ represents the probability of an object existing in the prediction box, and c ij represents the overall loss, i, j represent the feature position information.
[0027] Compared with the prior art, the present invention has the following beneficial effects:
[0028] 1. YOLOX adopts a lightweight model structure and decoupled head design, which enables the model to achieve fast inference speed while maintaining high detection accuracy. The multi-scale fusion module increases the redundancy and robustness of the model by introducing feature information of different scales. When feature information of certain scales is interfered by noise or occlusion, feature information of other scales can still provide effective support for the model. This feature makes YOLOX more reliable in dealing with complex and changing real-world environments.
[0029] Second, the multi-scale aggregation module fuses features of different scales through spatial and channel attention mechanisms to enhance feature expression capabilities. In the spatial refinement path, by summing convolutions of different kernel sizes and a series of spatial feature aggregation operations, the multi-scale fusion in the spatial refinement path can retain the features of small targets and achieve the fusion of multi-scale spatial information. In the channel aggregation path, the channel attention map is generated through operations such as global average pooling, convolution, and activation, and combined with the spatially refined map. The channel aggregation path can enhance the attention to channel features related to the target and achieve multi-scale information fusion in the channel dimension. In addition, its multi-scale fusion of space and channels can improve the effect of feature fusion, and can also enhance the ability to handle complex backgrounds and reduce background interference.
[0030] Third, the present invention proposes a visual center mechanism, where a lightweight MLP is used to capture global dependencies, while a parallel learnable visual center mechanism is used to capture local corner areas of the input image. Then, in a top-down manner, the present invention proposes a globally centralized feature pyramid, where the visual center information from the deepest layer is used to adjust the front-end shallow features, which can not only capture global long-distance dependencies, but also efficiently obtain comprehensive and discriminative feature representations. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1 This is a diagram of the multi-scale aggregation network structure based on YOLOx.
[0032] Figure 2 This is the structural diagram of the multi-scale aggregation and fusion module.
[0033] Figure 3 Schematic diagram of the structure of the YOLOx decoupling head. DETAILED DESCRIPTION
[0034] In order to make the technical problems, technical solutions and beneficial effects to be solved by the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0035] like Figure 1 As shown, a surface defect detection method based on YOLOX multi-scale network reconstruction, the specific steps are as follows:
[0036] Step S1: Download the official surface defect detection dataset (NEU-DET), and preprocess the surface images in the dataset, including image enhancement, noise removal and other steps, to improve the image quality.
[0037] Step S2: The enhanced dataset is converted into the YOLO standard format, which mainly involves converting the annotation file in XML format into the label file in TXT format.
[0038] Step S3: Build a backbone extraction network based on YOLOX and build a five-layer feature extraction network.
[0039] A five-layer feature extraction network is constructed, each layer consists of a deep separable module and a multi-scale attention aggregation module, and a series connection method is used to perform deep information mining on the image.
[0040] Step S4: The depthwise separable module performs feature extraction. The depthwise convolution processes each input channel independently, while the pointwise convolution is responsible for combining the features of these channels. Specifically:
[0041] Depthwise separable is based on conventional convolutional neural networks, and its calculation is divided into two parts: one is the depthwise convolution applied on a single input channel, and the other is the pointwise conventional convolution; this connection can effectively reduce the number of parameters and calculations, and is also beneficial to the feature extraction of the backbone network.
[0042] Depthwise convolution uses a convolution operation of size 3x3. One convolution kernel is responsible for one channel for channel-by-channel convolution. Pointwise convolution uses a convolution operation of size 1x1. The Pointwise convolution operation will perform a weighted combination of the feature layers of the previous Depthwise channel-by-channel convolution in the depth direction.
[0043] Step S5: The multi-scale attention aggregation module uses the multi-scale features of the aggregate encoder to improve the feature expression ability through dual aggregation of space and channels, and better support the multi-level information fusion of the decoder. Specifically:
[0044] The multi-scale attention aggregation module further refines the features and enhances the spatial and channel feature information through operations on both spatial and channel paths, making the output feature map higher quality in both spatial and channel dimensions.
[0045] Step S6: In the channel aggregation path, the channel attention map is generated through global average pooling, convolution and activation operations, and combined with the spatially refined map to achieve multi-scale information fusion in the channel dimension, specifically:
[0046] The input features are processed by adaptive average pooling and adaptive maximum pooling respectively, and the single-channel features are mapped to the corresponding average pooling feature values and maximum pooling feature values respectively.
[0047] A residual module is used to calculate the average pooled eigenvalue, so that the network can learn more complex feature representations. The residual module uses 1×1 convolution kernels successively and uses the RULE activation function between the two convolution kernels.
[0048] The same residual module is used to calculate the maximum pooling eigenvalue, and the output results of the two are added together to obtain a fusion feature that retains all background information and the most significant feature information.
[0049] By inputting into the Sigmoid function, the input features can be mapped to the range of 0 to 1, normalizing the feature data and enhancing the learning ability of the model.
[0050] Step S7: In the spatial refinement path, the fusion of multi-scale spatial information is achieved by summing the convolutions of different kernel sizes and a series of spatial feature aggregation operations, specifically:
[0051] The input features are first upsampled by 1×1 convolution and divided into three aggregation nodes. 3×3, 5×5, and 7×7 convolution kernels are used to extract feature information of different depths, and multi-scale aggregation features are accumulated layer by layer.
[0052] Step S8: The obtained multi-scale aggregated features are used as the input of the two network branches, one part is not processed, and the other part is processed by spatial refinement, specifically:
[0053] The specific method is to perform deep longitudinal compression on the output aggregated features, which is divided into maximum fusion compression and average fusion compression. The method is to perform maximum and average calculations on the pixel values at the same position on each layer of the feature layer to obtain single-channel features respectively.
[0054] The above two single-channel features are fused to obtain a feature layer with a depth of 2. A 7×7 convolution kernel is used to reduce the dimension of the feature layer to obtain a single-channel feature with feature information fusion.
[0055] By inputting into the Sigmoid function, the input features can be mapped to the range of 0 to 1, and the feature data can be normalized to enhance the learning ability of the model.
[0056] The feature weights of the outputs of the two network branches are reconstructed to obtain the features after spatial feature refinement. The downsampling operation is performed using a 1×1 convolution kernel to restore the original feature size.
[0057] Obtain the spatially refined features and the channel information compressed features, and multiply the two parts by weights to obtain the multi-scale attention aggregation features.
[0058] By fusing deformable convolutions, the receptive field of the defect feature extraction network is expanded to capture complete and comprehensive defect texture features.
[0059] Step S9: Construct the YOLOX feature multi-scale fusion network Neck. First, implement the explicit visual center mechanism on the top layer of the feature pyramid, and then use the obtained features containing the explicit visual center to simultaneously adjust all the previous shallow features.
[0060] By upsampling the deep features to the same spatial scale as the low-level features, then concatenating them along the channel dimension, and downsampling the concatenated features to 256 channels through 1×1 convolution, it is possible to explicitly increase the spatial weight of the global representation at each layer of the feature pyramid, thereby achieving a comprehensive and discriminative feature representation.
[0061] Step S10: A spatially explicit visual center scheme is proposed, including a lightweight MLP for capturing global dependencies and a learnable visual center for aggregating local key regions.
[0062] Step S11: The learnable visual center mechanism aggregates the local corner area of the input image to retain the local key area information, specifically:
[0063] Using a set of scaling factors to map and to the corresponding position information in turn, the full image information about the kth codeword can be calculated as follows:
[0064]
[0065] in, represents the i-th pixel, b k is the kth learnable visual codeword, s k is the kth scaling factor, x i -b k represents visual feature information, that is, represents the position of each pixel relative to the image, K is the total number of visual centers, e k is the information of the entire image of the kth codeword.
[0066] The number of visual centers is set to 64, a 1×1 convolution operation is performed on the input features, the number of feature dimensions is changed to 64 and encoded to obtain a feature vector.
[0067] Step S12: perform an averaging operation, perform averaging processing in the feature depth dimension direction, compress it into a two-dimensional feature, and obtain a set of weight factors for the channel attention mechanism through a learnable visual center mechanism, specifically:
[0068] Define a linear layer and set the parameters of the number of input neurons, the number of output neurons, and the bias so that the number of input neurons is equal to the number of output neurons.
[0069] Input two-dimensional features, use tiling operation to break up the two-dimensional features into a one-dimensional vector representation, and use the broadcast mechanism and automatic reshaping function of the linear layer to obtain the weight factor of the channel attention mechanism.
[0070] Step S13: A lightweight MLP architecture is adopted to capture the global dependency of specific feature sizes, and a learnable visual center mechanism is used to obtain weight factors to aggregate the input image features, specifically:
[0071] The lightweight MLP consists of two residual modules, a module based on depthwise separable convolution and a module based on channel MLP, where the input of the MLP module is the output of the depthwise separable convolution module. After the feature map is group-normalized, the depthwise separable convolution can improve the feature expression capability while reducing the computational cost:
[0072]
[0073] Among them, X i represents the output of the fourth layer of the backbone network, and GN represents normalization processing.
[0074] The input of the channel MLP module is the output of the depthwise separable convolution module, which is group normalized and then subjected to the channel MLP operation:
[0075]
[0076] in, represents the output of the fourth layer of the backbone network, and GN represents normalization processing.
[0077] Both modules undergo channel scaling and DropPath operation, which randomly deletes subpaths of the multi-branch structure to improve feature generalization and robustness.
[0078] Step S14: Introduce a decoupled head to process classification and positioning tasks respectively, separate category prediction and location prediction, and use two independent network branches to process them respectively, specifically:
[0079] For each feature fusion network output feature, a 1x1 convolution layer is first applied to reduce the number of channels to 256, and then it is divided into two parallel branches, each branch includes two 3×3 convolution layers, which are used for classification and regression tasks respectively. The output of the classification branch is H×W×C, the output of the regression branch is H×W×4, and the output of the IoU branch is H×W×1.
[0080] Step S15: The loss-based allocation strategy transforms the positive and negative sample allocation problem into an optimal transmission problem, and allocates positive and negative samples by minimizing the transmission cost, specifically:
[0081] Calculate the IoU loss and classification loss between each candidate box and GT.
[0082] For each predicted box, the IOU and category loss with the real box are calculated respectively, and then the overall loss is obtained by weighting. The calculation process is as follows:
[0083]
[0084] in, represents the confidence score of the reference box prediction, represents the bounding box predicted for the reference box, Represents the confidence score of the true box, represents the bounding box of the real box, L cls represents the cross entropy loss, L reg represents the intersection-over-union loss, α is the balance coefficient of the two losses, θ represents the probability of an object existing in the prediction box, and c ijrepresents the overall loss, i, j represent the feature position information.
[0085] Then sort the intersection-and-union ratios of each box and the true box, add up the intersection-and-union ratios of all boxes and take the integer to get the number of categories of positive samples.
[0086] Step S16: fine-tune the model parameters, load the data set for model training, obtain the optimal weight, and load the training model for defect detection, specifically:
[0087] Use the training set data to train the model and adjust the model parameters to minimize the loss function, including adjusting the hyperparameters and selecting the training strategy to obtain the defect detection model.
[0088] Use the loaded model to infer the data to be inspected and output defect detection results. The detection results may include defect probability maps, bounding boxes, etc., which are used to mark areas where defects may exist.
[0089] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principle of the present invention should be included in the scope of the present invention.
Claims
1. A surface defect detection method based on YOLOX multi-scale network reconstruction, characterized in that: The specific steps are as follows: Step S1, downloading a surface defect detection data set, and preprocessing the surface images in the data set; Step S2: convert the preprocessed data set into the YOLO standard format, that is, convert the annotation file in XML format into a label file in TXT format; Step S3, construct a backbone extraction network backbone based on YOLOX. The backbone extraction network backbone has five layers of feature extraction networks of different sizes. Each layer consists of a deep separable module and a multi-scale attention aggregation module, and uses a series connection method to perform deep information mining on the image; Step S4: The feature extraction network uses a depth-separable module to extract features. The calculation is divided into two parts. One part is to apply a 3x3 depthwise convolution on a single input channel, and the other part is to use a 1x1 pointwise convolution. The pointwise convolution operation performs a weighted combination of the channel-by-channel convolution feature layers of the previous depthwise convolution in the depth direction. Step S5: In the channel aggregation path of the multi-scale attention aggregation module, a channel attention map is generated through global average pooling, 1×1 convolution calculation and RELU activation function operation to achieve multi-scale information fusion in the channel dimension; Step S6: In the spatial refinement path of the multi-scale attention aggregation module, the input features are first up-sampled by 1×1 convolution and divided into three aggregation nodes. 3×3, 5×5, and 7×7 convolution kernels are used to extract feature information of different depths, and multi-scale aggregation features are accumulated layer by layer. Step S7, the features output by the two network branches after multi-scale aggregation processing are subjected to weight matrix calculation, and are concatenated with the feature input without any processing; Step S8, constructing the YOLOX feature multi-scale fusion network Neck, before inputting the multi-scale fusion network, using the spatial explicit visual center mechanism and the lightweight MLP module to capture the global and local features of the outputs of multiple feature extraction layers, and retaining the local key area information in the input image; Step S9, proposing a spatially explicit visual center scheme, including a lightweight MLP for capturing global long-range dependencies and a learnable visual center for aggregating local key areas; Step S10, the learnable visual center mechanism aggregates the input image to average the operation, calculates the average value in the feature depth dimension direction and compresses it into a two-dimensional feature, defines a linear layer, sets the number of input neurons, the number of output neurons, and the bias of the parameters, so that the number of input neurons is equal to the number of output neurons, inputs the two-dimensional feature, and adopts a tiling operation to break up the two-dimensional feature into a one-dimensional vector form, and uses the broadcast mechanism and automatic reshaping function of the linear layer to achieve, and obtains the weight factor of the channel attention mechanism; Step S11, adopting a lightweight MLP architecture to capture the global dependency of a specific feature size, performing averaging processing in the feature depth dimension direction, compressing it into a two-dimensional feature, and obtaining a set of weight factors for a channel attention mechanism through a learnable visual center mechanism to perform matrix calculation with the input image features; Step S12, introduce the decoupling head of YOLOx to process the classification and positioning tasks respectively, separate the category prediction and the position prediction, and use two independent network branches for processing respectively. Each branch includes two 3×3 convolutional layers, which are used for classification and regression tasks respectively. The classification branch output is H×W×C, the regression branch output is H×W×4, and the IoU branch output is H×W×1; Among them, H and W represent the size of the feature map, and C represents the number of feature channels; Step S13, fine-tune the model parameters, set the learning rate to 0.0001, the number of iterations to 100 rounds, the input batch to 32, and the discard rate to 0.2, load the data set for model training to obtain the optimal weight, and load the trained model for defect detection.
2. According to claim 1, a surface defect detection method based on YOLOX multi-scale network reconstruction is characterized in that: In step S1, specific preprocessing steps include: performing signal enhancement through Fourier transform, removing noise through low-pass filtering, and enhancing edge high-frequency signals through high-pass filtering, thereby performing image enhancement.
3. According to claim 2, a surface defect detection method based on YOLOX multi-scale network reconstruction is characterized in that: In step S9, a set of scaling factors are sequentially mapped to corresponding position information, and the full image information about the kth codeword is calculated in the following manner: in, represents the i-th pixel, b k is the kth learnable visual codeword, s k is the kth scaling factor, x i -b k represents the visual feature information, that is, the position of each pixel relative to the image, K is the total number of visual centers, N is the total number of feature layer channels, and e k is the information of the entire image of the kth codeword.
4. According to claim 3, a surface defect detection method based on YOLOX multi-scale network reconstruction is characterized in that: In step S12, the loss-based allocation strategy is used to transform the positive and negative sample allocation problem into an optimal transmission problem, and the positive and negative samples are allocated by minimizing the transmission cost: in, represents the confidence score of the reference box prediction, represents the bounding box predicted for the reference box, Represents the confidence score of the true box, represents the bounding box of the real box, L cls represents the cross entropy loss, L reg represents the intersection-over-union loss, α is the balance coefficient of the two losses, θ represents the probability of an object existing in the prediction box, and c ij represents the overall loss, i, j represent the feature position information.
Citation Information
Cited By
Silicon wafer surface defect detection method and device based on multi-scale feature fusion, and medium
CN120471918A
Multi-view target tracking detection method and device, terminal and medium
CN120707596A
Multi-mode sensing carbon fiber laying defect detection method, equipment and medium
CN120778751A
Power distribution network tower defect automatic identification and classification method and system based on deep learning
CN121564557A
A Deep Learning-Based Method and System for Automatic Identification and Classification of Defects in Power Distribution Network Towers
CN121564557B