Aluminum profile surface defect detection method based on DMSA-Swinin-Unet network model
Through the DMSA-Swin-Unet network model, combined with Swin Transformer and DMSANet modules, the problem of traditional methods weak surface defect detection capabilities for complex, small and diverse aluminum profiles is solved, and efficient and accurate defect detection is achieved, improving the robustness and detection stability of the model.
Patent Information
- Application Number
- CN202510666860.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-22
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-05-22
AI Technical Summary
Traditional methods have weak ability to detect surface defects of complex, small and diverse aluminum profiles and poor robustness.
Using the DMSA-Swin-Unet network model, combined with Swin Transformer and DMSANet modules, the window attention and shift window attention mechanism are introduced through multi-level local window attention and cross-region feature interaction, multi-scale feature extraction and semantic fusion are carried out, and the model is optimized using joint loss function and learning rate scheduler.
It significantly improves the accuracy of defect recognition, improves the positioning ability of small defects and blurred edges, enhances the generalization and detection stability of the model, and the edges of the image segmentation results are smoother and the structure is clearer.
Smart Images

Figure CN120339262A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of industrial inspection, and particularly to a method for detecting surface defects of aluminum profiles based on a DMSA-Swin-Unet network model. Background Art
[0002] In the process of industrial production, aluminum profiles, as important materials widely used in fields such as construction, transportation, and electronics, the surface quality directly affects the performance and aesthetics of the final products. However, during the production process, defects such as cracks, peeling, and scratches often appear on the surface of aluminum profiles. At present, the detection of surface defects of aluminum profiles still mainly relies on manual visual inspection, which has problems such as low efficiency, strong subjectivity, and high misjudgment rate, and it is difficult to control the efficiency of quality inspection. Therefore, researching efficient and accurate surface defect detection technology has become an important topic in industrial automation and intelligent manufacturing. Traditional image processing methods rely on preset rules and features, and it is difficult to adapt to complex and changeable defect types and actual environmental noise.
[0003] In recent years, deep learning has achieved rapid progress in fields such as image recognition. Researchers have proposed online visual detection methods based on deep learning, which use Convolutional Neural Network (CNN) to automatically extract image features, overcoming the dependence on feature design of traditional methods and improving the robustness and generality of detection. For example, the improved YOLOv3 algorithm introduces an attention mechanism and a multi-scale feature fusion structure, enhancing the detection ability for complex textures and multi-scale defects. However, although CNN performs excellently in local feature extraction, it has natural disadvantages in modeling global image information and long-distance dependence relationships, which easily leads to incomplete defect recognition.
[0004] Therefore, those skilled in the art are committed to developing a method for detecting surface defects of aluminum profiles based on a DMSA-Swin-Unet network model. Summary of the Invention
[0005] In view of the above-mentioned defects of the prior art, the technical problem to be solved by the present invention is that traditional methods have weak detection capabilities for complex, tiny, and diverse-shaped defect features and poor robustness.
[0006] To achieve the above object, the present invention provides a method for detecting surface defects of aluminum profiles based on a DMSA-Swin-Unet network model, and the method includes the following steps: S101: Preprocess the collected surface images of aluminum profiles, and the preprocessing includes image size normalization, color enhancement, and data standardization; S103: Use convolution operation with a predetermined step size on the preprocessed image, divide the image into image patches and perform linear projection to generate an embedded feature sequence; S105: Extract multi-level local and global features through the Swin Unet encoder module. In the encoder, the multi-layer nested Swin Transformer adopts alternating window attention and shifted window attention mechanisms; S107: In each layer of the decoding process, fuse the encoded features of the corresponding scale through skip connections; S109: Configure a DMSA module after the skip connection. The DMSA module is composed of a channel attention module and a spatial attention module in parallel; S111: Perform upsampling step by step through image patch expansion, and finally output a pixel-level defect segmentation map with the same size as the original image.
[0007] Further, in the step S105, the Swin Transformer module is a multi-layer nested structure. Each layer includes a window attention module and a shifted window attention module. The window attention module calculates self-attention in a local window, and the shifted window attention module performs cross-region feature interaction through window shifting.
[0008] Further, the step S105 includes the following sub-steps: S1051: The input feature map calculates self-attention through the window attention module:
[0009]
[0010] S1052: The shifted window attention module introduces cross-window feature dependency relationships to perform cross-region feature interaction:
[0011]
[0012] Among them, is the output of the W-MSA module of the th layer of the Swin Transformer module, is the output of the MLP layer of the l th layer of the SwinTransformer module, is the output of the SW-MSA module of the th + 1 layer of the Swin Transformer, is the l+1The output of the MLP layer of the Swin Transformer module, LN is layer normalization, and MLP is a multi-layer perceptron. is the window multi-head self-attention. is the shifted window multi-head attention mechanism. The self-attention mechanism is calculated according to the following formula:
[0013] where, Attention is the attention coefficient of the key-value pair. is the query matrix. is the key matrix. is the value matrix. is the transpose matrix of matrix K. is the normalization exponential function. is the dimension of the query or key. is the bias matrix.
[0014] Furthermore, the output of each layer in the Swin Transformer module is compressed in resolution and extended in channels through image patch merging to extract multi-scale features.
[0015] Furthermore, in the step S109, the channel attention module uses global average pooling and a multi-layer perceptron to weight the channel features to obtain the output features. where, the channel attention map is:
[0016]
[0017] The final output is:
[0018]
[0019] where, is the attention weight of channel c. is the global feature of channel c. is the input feature map of channel c. is the channel height. is the channel width. is the number of channels. and are the fully connected layer weights. represents the activation function. is the activation function. exp is the exponential function. is the channel identifier. is the output feature of the channel attention module. is the channel attention map, representing the influence of the i-th channel on the j-th channel, the original feature of the i-th channel, is the original feature of the j-th channel, is the scaling factor.
[0020] Furthermore, the spatial attention module uses a convolution operation to construct a spatial attention map for weighting the original features. The spatial attention map is:
[0021] The final output is:
[0022] Where, is the spatial attention map, representing the influence of the i-th channel on the j-th channel, is the output feature of the spatial attention module, is the local feature the new feature map generated by passing through the convolutional layer, is the channel identifier, is the number of channels, is the scaling factor, exp is the exponential function.
[0023] Furthermore, in the step S111, a 1×1 convolution is used at the output end of the DMSA-Swin-Unet network for class mapping, and a normalized exponential function is used to normalize the multi-class probabilities for each pixel.
[0024] Furthermore, the training of the DMSA-Swin-Unet network model uses a combined loss function. The combined loss function adopts a weighted combination of the cross-entropy loss function and the Dice loss function. The combined loss function is:
[0025]
[0026]
[0027] Where: is the combined loss function, is the cross-entropy loss function, is the Dice loss function, is the weighting coefficient, is the true label, is the predicted probability, A is the prediction result, and S represents the true label.
[0028] Furthermore, during the training process of the DMSA-Swin-Unet network model, the Adam optimizer is adopted, combined with a learning rate scheduler to dynamically adjust the learning rate, so as to improve the convergence speed and detection accuracy. The learning rate scheduling strategy used by the learning rate scheduler is as follows:
[0029] where is the learning rate of the th round, is the minimum learning rate, is the maximum learning rate, is the training round, is the total number of training rounds, and cos is the cosine function.
[0030] Furthermore, the method supports processing images of various complex surface defects, and the surface defects include convex powder, jet flow, dirty spots, scratches, pits, and exposed substrates.
[0031] In the preferred embodiment of the present invention, compared with the prior art, it has the following beneficial effects: 1. The present invention introduces the Swin-Unet structure to replace the traditional U-Net, realizes more effective feature extraction and semantic fusion. Using the Swin Transformer as the backbone network structure and adopting a multi-level local window attention structure, it improves the multi-scale feature expression ability; the skip connection ensures the retention of low-level semantic and spatial information. Compared with the traditional CNN structure, the defect recognition accuracy is improved by more than 5%; 2. The present invention introduces the window attention (W-MSA) and shifted window attention (SW-MSA) mechanisms to enhance the local and global modeling capabilities. W-MSA calculates self-attention in the local window, and SW-MSA realizes cross-region feature interaction through window shifting, enabling the model to simultaneously model local and context relationships, making the model better perceive the context and improving the positioning ability for irregular targets such as small defects and fuzzy edges; 3. The present invention integrates the DMSANet module in the decoder, introduces a "channel + space" dual attention mechanism. The DMSA module constructs channel attention (weighting the feature channels) and spatial attention (enhancing key positions) in parallel, enhances the skip connection features, improves the recognition ability of significant regions, improves the model's response ability to tiny defects and weak boundaries, avoids feature dilution, and enhances the model's generalization ability; 4. The present invention proposes a DMSA-Swin-Unet fusion structure to implement a multi-scale information extraction mechanism for semantic enhancement + spatial restoration. Through the joint optimization of the multi-scale encoder and the fusion attention module, the model has stronger category discrimination ability and regional edge expression ability; the upsampling and downsampling structure helps with spatial restoration, improves detection stability, and has stronger recognition ability for complex scenes and targets with blurred boundaries; the edge of the image segmentation result is smoother and the structure is clearer.
[0032] The following will further illustrate the concept, specific structure and technical effects of the present invention with reference to the accompanying drawings, so as to fully understand the purpose, features and effects of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 is a schematic flow chart of the detection method according to a preferred embodiment of the present invention; Figure 2 is a schematic diagram of the DMSA-Swin-Unet network structure according to a preferred embodiment of the present invention; Figure 3 is a schematic diagram of the overall architecture of the Swin Transformer according to a preferred embodiment of the present invention; Figure 4 is a schematic diagram of the Swin Transformer module according to a preferred embodiment of the present invention; Figure 5 is a schematic diagram of the DMSANet structure according to a preferred embodiment of the present invention; Figure 6 is a curve graph of the loss function according to a preferred embodiment of the present invention; Figure 7 is a schematic diagram of the semantic segmentation result of the convex powder defect according to a preferred embodiment of the present invention. Among them, (a) is the original image of the convex powder surface defect, (b) is the original label image of the convex powder, and (c) is the convex powder semantic segmentation map obtained by applying the defect detection method of the model of the present invention; Figure 8 is a schematic diagram of the semantic segmentation result of the dirt point defect according to a preferred embodiment of the present invention. Among them, (a) is the original image of the dirt point surface defect, (b) is the original label image of the dirt point, and (c) is the dirt point semantic segmentation map obtained by applying the defect detection method of the model of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0034] The following introduces multiple preferred embodiments of the present invention with reference to the accompanying drawings of the specification to make its technical content clearer and easier to understand. The present invention can be embodied in many different forms of embodiments, and the protection scope of the present invention is not limited to the embodiments mentioned in the text.
[0035] In the accompanying drawings, components with the same structure are denoted by the same numerical labels, and components with similar structures or functions everywhere are denoted by similar numerical labels. The dimensions and thicknesses of each component shown in the drawings are arbitrarily illustrated, and the present invention does not limit the dimensions and thicknesses of each component. To make the illustration clearer, the thicknesses of some components in the drawings are appropriately exaggerated.
[0036] Aiming at the problems of weak defect detection ability and poor robustness of traditional methods for complex, tiny, and diverse-shaped defects, the present invention proposes a Swin-Unet network model structure (DMSA-Swin-Unet) integrating a dual multi-scale attention mechanism (DMSA) for semantic segmentation of aluminum profile surface defects. This structure combines the powerful spatial context modeling ability of Swin Transformer and the enhanced ability of spatial attention and channel attention in DMSANet, can simultaneously focus on local details and global structures, significantly improve the accuracy of defect detection and the ability to restore details, and is particularly suitable for the segmentation task of multi-scale and multi-shaped defects.
[0037] As Figure 1 shown, an aluminum profile surface defect detection method based on the DMSA-Swin-Unet network model provided by an embodiment of the present invention includes the following steps: S101: Preprocess the collected aluminum profile surface image.
[0038] In this embodiment, the above preprocessing includes image size normalization, color enhancement, and data standardization.
[0039] Since the collected images have diverse sources and inconsistent sizes, directly inputting them into the model will cause waste of computing resources and unstable training. Therefore, it is necessary to uniformly crop the original images and use the mosaic data augmentation method to expand the training samples.
[0040] S103: Use convolutional operations with a predetermined stride on the preprocessed image, divide the image into image patches, and perform linear projection to generate an embedded feature sequence.
[0041] After the image is standardized, the model input part adopts a convolutional embedded image patch division method (Conv-Patch Embedding), divides it into non-overlapping image patches, and inputs them into the Swin Unet encoder. The specific method is as follows: Input the standardized image into a two-dimensional convolutional layer with a stride P and a convolutional kernel size of P P, where P is the image patch size, which is 4 here. The convolutional operation transforms the image and embeds the channel dimension in the image patch. The above process completes the division of the image patches and also realizes preliminary feature extraction and downsampling through convolution.
[0042] S105: Extract multi-level local and global features through the Swin Unet encoder module. The multi-layer nested Swin Transformer in the encoder adopts alternating window attention and shifted window attention mechanism.
[0043] In this embodiment, the Swin Transformer module is a multi-layer nested structure, each layer includes a window attention module and a shifted window attention module, the window attention module calculates self-attention in the local window, and the shifted window attention module realizes cross-region feature interaction through window shifting.
[0044] The specific steps include the following: S1051: Input feature map is passed through the window attention module to calculate self-attention:
[0045]
[0046] S1052: The shift window attention module introduces cross-window feature dependencies and performs cross-region feature interaction:
[0047]
[0048] in, For the The W-MSA module output of the Swin Transformer module of the layer, For the The output of the MLP layer of the SwinTransformer module, For the The output of the SW-MSA module of the Swin Transformer module of the layer, is the final output of the current Swin Transformer module, LN is layer normalization, MLP is multi-layer perceptron, is the window multi-head self-attention, It is a shift window multi-head attention mechanism; The self-attention mechanism is calculated as follows:
[0049] in, Attention is the attention coefficient of the key-value pair, is the query matrix, is the key matrix, is the value matrix, is the transposed matrix of matrix K, is the normalization exponential function, is the dimension of the query or key, is the bias matrix.
[0050] In this embodiment, the output of each layer in the Swin Transformer module is subjected to resolution compression and channel expansion through image patch merging to extract multi-scale features.
[0051] S107: During the decoding process of each layer, the encoded features of the corresponding scale are fused through skip connections.
[0052] S109: Configure the DMSA module after the skip connection. The DMSA module is composed of a channel attention module and a spatial attention module in parallel.
[0053] In this embodiment, the channel attention module uses global average pooling and a multi-layer perceptron to weight the channel features to obtain the output features; Among them, the channel attention weight is:
[0054]
[0055] The final output is:
[0056]
[0057] Among them, is the attention weight of channel c, is the global feature of channel c, is the input feature map of channel c, is the channel height, is the channel width, is the number of channels, and is the fully connected layer weight, represents the activation function, is the activation function, exp is the exponential function, is the channel identifier, is the output feature of the channel attention module, is the channel attention map, indicating the measurement of the influence of the i-th channel on the j-th channel, the original feature of the i-th channel, is the original feature of the j-th channel, is the scaling factor.
[0058] The spatial attention module uses convolution operations to construct a spatial attention map for weighting the original features. The spatial attention map is:
[0059] The final output is:
[0060] Among them, is the spatial attention map, representing the influence of the i-th channel on the j-th channel, is the output feature of the spatial attention module, is the local feature The new feature map generated by passing through the convolutional layer, is the channel identifier, is the number of channels, is the scaling factor, exp is the exponential function.
[0061] S111: Perform upsampling step by step through image patch expansion, and finally output a pixel-level defect segmentation map with the same size as the original image.
[0062] In this embodiment, at the output end of the DMSA-Swin-Unet network, 1×1 convolution is used for class mapping, and the Softmax function is used to normalize the multi-class probabilities of each pixel.
[0063] In the training of the DMSA-Swin-Unet network model, a joint loss function is used. The joint loss function adopts a weighted combination of the cross-entropy loss function and the Dice loss function. The joint loss function is:
[0064]
[0065]
[0066] Among them: is the joint loss function, is the cross-entropy loss function, is the Dice loss function, is the weighting coefficient, is the true label, is the predicted probability, A is the prediction result, and S represents the true label.
[0067] During the training process of the DMSA-Swin-Unet network model, the Adam optimizer is used, and the learning rate scheduler is combined to dynamically adjust the learning rate to improve the convergence speed and detection accuracy. The learning rate scheduling strategy used by the learning rate scheduler is:
[0068] Among them, is the round learning rate, is the minimum learning rate, is the maximum learning rate, is the number of training rounds, is the total number of training rounds, and cos is the cosine function.
[0069] A method for detecting surface defects of aluminum profiles based on the DMSA-Swin-Unet network model provided by an embodiment of the present invention supports processing multi-class complex surface defect images, including surface defects of aluminum profiles such as convex powder, spray flow, dirty points, scratching, pitting, and exposed bottom.
[0070] Compared with the prior art, a method for detecting surface defects of aluminum profiles based on the DMSA-Swin-Unet network model provided by an embodiment of the present invention has the following characteristics: 1. The present invention uses the Swin-Unet structure to replace the traditional U-Net, which can achieve more effective feature extraction and semantic fusion. By using the Swin Transformer as the backbone network structure and adopting a multi-level local window attention structure, the multi-scale feature expression ability is improved, and the skip connection ensures the retention of low-level semantic and spatial information. Compared with the traditional CNN structure, the defect recognition accuracy is improved by more than 5%; 2. The present invention introduces the window attention and shifted window attention mechanisms, which enhance the local and global modeling capabilities. W-MSA calculates self-attention in the local window, and SW-MSA realizes cross-region feature interaction through window shifting, enabling the model to simultaneously model local and context relationships, better perceive the context, and improve the positioning ability for small defects, fuzzy edges and other irregular targets; 3. The present invention integrates the DMSANet module in the decoder, introduces the "channel + space" dual attention mechanism DMSA module to parallelly construct channel attention and spatial attention, enhances the processing of skip connection features, improves the recognition rate of significant regions, and improves the response ability of the model to tiny defects and weak boundaries, avoiding feature dilution and enhancing the generalization ability of the model; 4. The DMSA-Swin-Unet fusion structure proposed by the present invention realizes a multi-scale information extraction mechanism for semantic enhancement and spatial restoration. Through the joint optimization of the multi-scale encoder and the fusion attention module, the model has stronger class discrimination ability and regional edge expression ability; the upsampling and downsampling structures contribute to spatial restoration, improve the detection stability, and have stronger recognition ability for complex scenes and targets with fuzzy boundaries; the edge of the image segmentation result is smoother and the structure is clearer.
[0071] The following combines the preferred embodiments of the present invention to detail the content of the present invention.
[0072] 1. Dataset Introduction The Tianchi aluminum profile surface defect dataset is provided by the Alibaba Cloud Tianchi platform and contains 10 types of typical defects (paint bubbles, spray streams, dirt spots, exposed substrate, scratches, orange peel, non-conductive, pitting, powder bumps, exposed substrate at corners). There are more than 10,000 monitoring image data of defective aluminum profiles from actual production, and each image contains one or more defects.
[0073] 2. Data Preprocessing Since the source of the collected images is diverse and the sizes are inconsistent, directly inputting them into the model will cause waste of computing resources and unstable training. Therefore, it is necessary to uniformly crop the original images to a fixed size and use the mosaic data augmentation method to expand the training samples.
[0074] 2-1. Image Size Standardization All images are scaled or center-cropped to a unified size of 512×512. Bilinear interpolation or nearest-neighbor interpolation is used to ensure that the image structure is not distorted during the scaling process, improve the batch training efficiency of the model, and maintain the matching between the convolutional kernel and the image features.
[0075] 2-2. Mosaic Image Augmentation To increase the generalization ability and robustness of the model and simulate the appearance states of defects at different angles on the actual production line, the mosaic augmentation method is used for data augmentation. Four pictures are randomly selected from the training dataset for processing.
[0076] Random Scaling and Cropping: Each picture is randomly scaled and cropped to adapt to the input size of the model.
[0077] Random Arrangement: The four pictures are randomly arranged into a rectangular area to form a new training sample. Data Expansion: The dataset is further expanded through operations such as flipping and rotation.
[0078] 3. Model Input After the images are standardized, the model input part adopts the convolutional embedded image patch division method (Conv-PatchEmbedding), divides the images into non-overlapping patches, and inputs them into the Swin Unet encoder. The specific method is as follows: The standardized images are input into a two-dimensional convolutional layer with a stride of P and a convolutional kernel size of P P, where P is the patch size, which is 4 here. The convolutional operation transforms the image into a shape of . Here, C is the number of output feature channels, that is, the patch embedding dimension. So far, the patch division of the image is completed, and preliminary feature extraction and downsampling are also achieved through convolution.
[0079] 4. Feature Extraction and Fusion The encoder extracts multi-level local and global features. The decoder gradually upsamples through the image patch expansion module and combines skip connections to restore the spatial structure. The DMSA module enhances the spatial and channel expression capabilities of the fused features.
[0080] 4-1. Swin Unet Encoder In the encoder, the C-dimensional input with a resolution of is fed into the multi-layer nested structure of the Swin Transformer for representation learning. Each layer includes a window attention (W-MSA) and a shifted window attention (SW-MSA) module. The feature dimension and resolution remain unchanged. At the same time, the image patch merging layer reduces the number of features (downsampling by 2 times) and increases the feature dimension to 2 times. Here .
[0081] First, for the l-th layer encoding, the input feature map undergoes multi-head self-attention calculation within the window:
[0082]
[0083] Then, the shifted window attention introduces cross-window feature dependencies, and the output is:
[0084]
[0085] Among them, is the output of the W-MSA module of the l-th layer Swin Transformer module, is the output of the MLP layer of the l-th layer Swin Transformer module, is the output of the SW-MSA module of the l-th layer Swin Transformer module, is the final output of the current Swin Transformer module. LN is layer normalization, MLP is a multi-layer perceptron, is the window multi-head self-attention, is the shifted window multi-head attention mechanism; The self-attention mechanism is calculated according to the following formula:
[0086] Among them, Attention is the attention coefficient of the key-value pair, , is the query matrix, is the key matrix, is the value matrix, is the transpose matrix of matrix K, is the normalized exponential function, is the dimension of the query or key, and B comes from the bias matrix .
[0087] Between each stage, the spatial dimension is reduced and the channel dimension is increased by image patch merging.
[0088] The input image patch is divided into 4 parts and connected together through the image patch merging layer. At this time, the feature resolution is downsampled by 2 times and the feature dimension is increased by 4 times. Therefore, a linear layer is applied to the connected features to unify the feature dimension to 2 times the original dimension:
[0089] 4-2, Swin Unet Bottleneck Region Since the Transformer is too deep to converge, only two consecutive Swin Transformer blocks are used to construct the bottleneck to learn deep feature representations. In the bottleneck region, the feature size and resolution remain unchanged.
[0090] 4-3, Swin Unet Decoder The decoder structure is symmetric to the encoder. In the decoder, the image patch expansion layer is used to upsample the extracted deep features. The image patch expansion layer reshapes the feature maps of adjacent dimensions into higher-resolution feature maps (2 times upsampling) and correspondingly reduces the feature dimension to half of the original dimension, gradually restoring the spatial structure. Taking the first image patch expansion layer as an example, before upsampling, a linear layer is applied to the input feature ( ) to increase the feature dimension to 2 times the original dimension ( ), and then the rearrangement operation is used to expand the resolution of the input feature to 2× the input resolution and reduce the feature dimension to one-fourth of the input dimension ( ).
[0091] For the k-th layer decoding:
[0092] where, is the skip connection feature from the same layer of the encoder, is the upsampling module, and the addition represents the feature fusion between the decoder and the encoder.
[0093] 4-4, Skip Connection Similar to U-Net, skip connections are used to fuse multi-scale features from the encoder with upsampled features. Shallow and deep features are concatenated together to reduce the loss of spatial information caused by downsampling. After the linear layer, the size of the concatenated features remains the same as that of the upsampled features.
[0094] 4-5, DMSANet After each level of skip connection in the decoder, the DMSA module is introduced, and the fused features are fed into the DMSA module.
[0095] The channel attention module is used to selectively weight the importance of each channel, thus generating the optimal output characteristics. Global average pooling or global max pooling is performed on the input feature map to obtain the global features of each channel:
[0096] The attention weight of the c-th channel:
[0097] Calculate the channel attention map from the original features and apply the softmax layer to obtain the channel attention map : :
[0098] Multiply the matrix multiplication result of it with a by the scaling parameter β and perform an element-wise summation operation on a to obtain the final output :
[0099] where represents the input feature map, and H, W, and C represent the height, width, and number of input channels respectively. and represent the weights of the fully connected (FC) layer. The symbol represents the activation function, where the Sigmoid function is usually used, exp being the exponential function. represents measuring the influence of the i-th channel on the j-th channel.
[0100] The spatial attention module is used to emphasize the important spatial regions in the image.
[0101] Feed the local features represented by A into the convolutional layer to generate two new feature maps and and apply the softmax layer to calculate the spatial attention map :
[0102] Input feature A into the convolutional layer to generate a new feature map . Multiply its matrix multiplication with S by the scaling parameter , and perform an element-wise summation operation on feature a to obtain the final output:
[0103] where represents the influence of the i-th channel on the j-th channel.
[0104] Finally, all sub-features are aggregated, and the output is:
[0105] where is the obtained multi-scale feature map.
[0106] 5. Defect segmentation output The attention module adaptively selects different spatial scales across channels under the guidance of the feature descriptor:
[0107] Multiply the recalibrated weights of the multi-scale channel attention by the corresponding scale of the feature map:
[0108] 6. Training strategy During the model training process, the following optimization configuration is adopted to improve the convergence speed and detection accuracy.
[0109] 6-1. Optimizer selection and parameter setting The Adam optimizer is selected, the learning rate is set to 1e-4, and it is dynamically adjusted in combination with the learning rate scheduler. The weight decay coefficient is 0.01 for regularization to prevent overfitting, and the momentum factor parameters are set to = 0.9, = 0.999.
[0110] 6-2. Learning rate scheduling strategy To improve the stability in the initial stage of training and the accuracy performance in the later stage, the Cosine Annealing learning rate adjustment strategy is used to gradually decrease the learning rate from the initial value to the lowest threshold:
[0111] where is the learning rate at the t-th round, T is the total number of training rounds, and are the minimum and maximum learning rates.
[0112] 6-3. Loss Function Design The cross-entropy loss (CrossEntropy Loss) and Dice Loss are jointly designed to effectively alleviate the problems of class imbalance and small target defect segmentation. The joint loss function is defined as follows:
[0113]
[0114]
[0115] Where: is the joint loss function, is the cross-entropy loss function, is the Dice loss function, is the weighting coefficient, is the true label, is the predicted probability, A is the prediction result, and S represents the true label.
[0116] 6-4. Overfitting Control Strategy The early stopping strategy (Early Stopping) on the validation set is used to terminate the training in advance when the validation performance has not improved for 10 consecutive rounds, avoiding overfitting. 6-5. Training Platform and Configuration Training platform: Implemented based on PyTorch, running on the NVIDIA RTX A5000 / GPU cluster; Batch size: 32, dynamically adjusted according to the video memory size; Total number of training rounds: 50; Supports resuming training from a breakpoint and saving the model.
[0117] 7. Evaluation Metrics The performance of the model is evaluated using metrics such as mIoU, F1-score, Precision, Recall, etc.:
[0118]
[0119]
[0120]
[0121] Among them, TP is the number of pixels predicted to be of this class and actually belonging to this class (True Positive), FP is the number of pixels predicted to be of this class but actually not belonging to this class (False Positive), FN is the number of pixels actually belonging to this class but predicted wrongly as other classes (False Negative), and N is the total number of classes.
[0122] As Figure 2 shown, it is the structure diagram of the Swin-Unet network of this embodiment. It is a U-Net variant based on the Swin Transformer. Replacements are made on the original convolutional UNet, where the convolutional module is replaced with the swin transformer module; the max pooling is replaced with patch merging, which groups adjacent minimum feature patches and concatenates them by depth, and it is a non-convolutional downsampling technique; at the same time, patch expansion is used to implement the upsampling of the feature map; the DMSANet dual multi-scale attention mechanism is introduced in the decoding stage of the Swin-Unet structure, paying attention to both spatial attention and channel attention simultaneously.
[0123] As Figure 3 shown, it is the overall architecture diagram of the Swin Transformer of this embodiment. It introduces hierarchical feature mapping and window attention transformation. The network cuts the picture into different patches through patch partitioning and then outputs them as embeddings. In each stage, non-convolutional downsampling is achieved through patch merging first, gradually increasing the receptive field. Each stage consists of patch merging and multiple Blocks. Entering the first stage, the number of channels is adjusted to c through linear embedding local linear transformation, and the number of channels doubles while the resolution is reduced subsequently.
[0124] As Figure 4 shown, it is the Swin Transformer module of this embodiment, mainly composed of LayerNorm, MLP, WindowAttention, and Shifted Window Attention. W-MSA and SW-MSA respectively represent window-based multi-head self-attention using regular and shifted window partitioning configurations. The shifted window partitioning method introduces connections between adjacent non-overlapping windows in the previous layer.
[0125] As Figure 5 shown, it is the structure diagram of the DMSANet of this embodiment. The dual multi-scale attention network consists of two parts: the first part separates the input image, extracts features of different scales and aggregates them, and the second part adaptively combines local features with their global dependencies in parallel using spatial and channel attention modules.
[0126] As Figure 6As shown, it is the loss function curve graph of this embodiment. As the number of training rounds increases, both the training loss and the validation loss show a downward trend overall, indicating that the performance of the model on the training set and the validation set is continuously optimized, and the data fitting ability is gradually enhanced.
[0127] As Figure 7 shown, it is the convex powder semantic segmentation result graph of this embodiment. Among them, Figure 7 (a) is the original image of the surface defect of the convex powder, Figure 7 (b) is the original label image of the convex powder, Figure 7 (c) is the convex powder semantic segmentation graph obtained by applying the model defect detection method of the present invention. Compared with Figure 7 the original state of (a) and Figure 7 (b) which is only used as a labeling reference, the Figure 7 (c) of the present invention can accurately segment defect areas such as convex powder, realize effective detection and positioning of defects, and improve the efficiency and accuracy of defect detection.
[0128] As Figure 8 shown, it is the dirt point semantic segmentation result graph of this embodiment. Among them, Figure 8 (a) is the original image of the surface defect of the dirt point, Figure 8 (b) is the original label image of the dirt point, Figure 8 (c) is the dirt point semantic segmentation graph obtained by applying the model defect detection method of the present invention. Compared with Figure 8 the original state of (a) and Figure 8 (b) which is only used as a labeling reference, the Figure 8 (c) of the present invention can accurately segment defect areas such as dirt points, realize effective detection and positioning of defects, and improve the efficiency and accuracy of defect detection.
[0129] Compared with the prior art, a method for detecting aluminum profile surface defects based on the DMSA-Swin-Unet network model provided by the present invention has extensive industrial practicability and promotion prospects.
[0130] 1. Significant technical advantages (1) It can adapt to various types of aluminum profile surface defects, such as linear scratches, surface peeling, dot-like dirt, regional wrinkles, edge breakage, etc.; (2) The model has strong robustness and fault tolerance, and still maintains high recognition accuracy under complex background interferences such as uneven illumination, complex surface texture, and weak defect contrast; (3) It adopts a structure combining Transformer and attention mechanism, effectively breaking through the problem of missed detection of tiny defects by traditional CNN methods.
[0131] 2. Flexible deployment (1)Support embedded device deployment, such as low-power platforms like Jetson NX, Raspberry Pi, and domestic industrial control motherboards; (2)Adapt to cloud server deployment and edge-side collaborative operation, and can achieve data closed-loop in combination with production line PLC and MES systems; (3)Provide an interface (API) between the inference module and the front-end interface, facilitating access and operation in different industrial control systems.
[0132] 3. Stable performance (1)The model has passed tests on samples collected from various industrial aluminum surfaces and is suitable for various processing techniques such as anodizing, spraying, wire drawing, and electrophoresis; (2)Verified in the actual environment, it can still stably output detection results under complex working conditions such as vibration, dust, noise, and electromagnetic interference; (3)The system supports online automatic update and model fine-tuning, and has the ability of continuous learning and self-adaptation.
[0133] 4. Strong scalability (1)The network structure has strong versatility and can be extended to the detection of other metal surfaces such as stainless steel, titanium alloy, and copper strip through lightweight parameter adjustment; (2)The modular design supports subsequent access to subsystems such as depth determination (such as defect severity assessment), traceability identification (QR code / bar code), and robot control; (3)It can cooperate with industrial camera arrays to achieve high-end detection requirements such as multi-angle detection and multi-channel fusion.
[0134] The preferred specific embodiments of the present invention have been described in detail above. It should be understood that those of ordinary skill in the art can make many modifications and variations based on the concept of the present invention without creative efforts. Therefore, all technical solutions that can be obtained by those skilled in the art in the technical field of the present invention based on the concept of the present invention through logical analysis, reasoning, or limited experiments on the basis of the prior art should be within the protection scope determined by the claims.
Claims
1. A method for detecting surface defects of aluminum profiles based on the DMSA-Swin-Unet network model, characterized in that, The method includes the following steps: S101: Preprocess the collected surface images of aluminum profiles, where the preprocessing includes image size normalization, color enhancement, and data standardization; S103: Use convolution operations with a predetermined step size on the preprocessed image, divide the image into image patches and perform linear projection to generate an embedded feature sequence; S105: Extract multi-level local and global features through the Swin Unet encoder module, where the multi-layer nested Swin Transformer in the encoder module adopts an alternating window attention and shifted window attention mechanism; S107: In each layer of the decoding process, fuse the encoded features of the corresponding scale through skip connections; S109: Configure a DMSA module after the skip connection, where the DMSA module is composed of a channel attention module and a spatial attention module in parallel; S111: Perform upsampling step by step through image patch expansion, and finally output a pixel-level defect segmentation map with the same size as the original image.
2. The method according to claim 1, wherein In the step S105, the Swin Transformer module has a multi-layer nested structure, and each layer includes a window attention module and a shifted window attention module. The window attention module calculates self-attention in a local window, and the shifted window attention module realizes cross-region feature interaction through window shifting.
3. The method according to claim 2, wherein The step S105 includes the following sub-steps: S1051: The input feature map calculates self-attention through the window attention module: S1052: The shifted window attention module introduces cross-window feature dependencies and performs cross-region feature interaction: Among them, is the output of the W-MSA module of the -th layer of the Swin Transformer module, is the output of the MLP layer of the l -th layer of the Swin Transformer module, is the output of the SW-MSA module of the +1-th layer of the Swin Transformer, is the output of the MLP layer of the l+1 -th layer of the Swin Transformer module, LN is layer normalization, and MLP is a multi-layer perceptron. is window multi-head self-attention, is the shifted window multi-head attention mechanism; The self-attention mechanism is calculated according to the following formula: Among them, Attention is the attention coefficient of the key-value pair, is the query matrix, is the key matrix, is the value matrix, is the transpose matrix of matrix K, is the normalization exponential function, is the dimension of the query or key, is the bias matrix.
4. The method according to claim 3, wherein The output of each layer in the Swin Transformer module is compressed in resolution and expanded in channels through image patch merging to extract multi-scale features.
5. The method according to claim 4, wherein In the step S109, the channel attention module weights the channel features using global average pooling and a multi-layer perceptron to obtain the output features; Among them, the channel attention map is: The final output is: Among them, is the attention weight of channel c, is the global feature of channel c, is the input feature map of channel c, is the channel height, is the channel width, is the number of channels, and are the fully connected layer weights, represents the activation function, is the activation function, exp is the exponential function, is the channel identifier, is the output feature of the channel attention module, is the channel attention map, denoted as the influence of the i th channel on the j th channel, is the original feature of the i-th channel, is the original feature of the j-th channel, is the scale factor.
6. The method according to claim 5, wherein The spatial attention module uses convolution operations to construct a spatial attention map for weighting the original features, and the spatial attention map is: The final output is: Among them, is the spatial attention map, representing the influence of the i-th channel on the j-th channel. is the output feature of the spatial attention module. is the local feature. is the new feature map generated by passing through the convolutional layer. is the channel identifier. is the number of channels. is the scale factor. exp is the exponential function.
7. The method according to claim 6, wherein In the step S111, at the output end of the DMSA-Swin-Unet network, 1×1 convolution is used for class mapping, and the normalized exponential function is used to normalize the multi-class probabilities of each pixel.
8. The method according to claim 7, wherein The training of the DMSA-Swin-Unet network model uses a joint loss function, and the joint loss function adopts a weighted combination of the cross-entropy loss function and the Dice loss function. The joint loss function is: Among them, is the combined loss function, is the cross-entropy loss function, is the Dice loss function, is the weighting coefficient, is the true label, is the predicted probability, A is the prediction result, and S represents the true label.
9. The method according to claim 8, wherein During the training process of the DMSA-Swin-Unet network model, the Adam optimizer is adopted, and the learning rate is dynamically adjusted in combination with a learning rate scheduler. The learning rate scheduling strategy used by the learning rate scheduler is: Among them, is the learning rate for the th round, is the minimum learning rate, is the maximum learning rate, is the number of training rounds, and cos is the cosine function.
10. The method according to claim 9, wherein The method supports processing multi-class complex surface defect images, and the surface defects include convex powder, spray flow, dirty points, scratches, pits, and exposed bottoms.
Citation Information
Patent Citations
H-shaped steel surface defect detection method and system
CN116167968A
Surface defect detection method and system based on Swin Transform
CN116703885A
Industrial defect detection method based on feature fusion and progressive supervision strategy
CN118967592A
Cited By
Power image retrieval method and system
CN121009203A
A method and system for retrieving power images
CN121009203B
Gravity inversion imaging method based on DT-UNet
CN121033323A
Steel plate surface defect detection method based on Swin-Transform network structure and electronic equipment
CN121353763A