Aluminum product surface damage detection method and system
By introducing the skip connection and focus mechanism of the encoder and decoder in the surface damage detection of aluminum products, the problems of low detection accuracy and poor robustness in the existing technology are solved, and higher detection accuracy and faster model convergence are achieved.
Patent Information
- Application Number
- CN202510992542.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-18
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2045-07-18
AI Technical Summary
Existing surface defect detection methods for aluminum products have poor robustness and low detection accuracy when faced with complex lighting conditions and changes in defect morphology. Traditional methods are also prone to blurred and incomplete boundaries.
The skip connection between the encoder and decoder is adopted, combined with the transfer learning unit and the focus mechanism introduced in the Dice loss. The model is trained on historical aluminum product image data, and the detail information and semantic information are utilized to improve the detection accuracy.
The accuracy and robustness of surface damage detection for aluminum products are improved, the ability to detect damaged areas is enhanced, the model training time is shortened, and data imbalance is adapted.
Smart Images

Figure CN120510453B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of fault detection, and more particularly to a method and system for detecting surface damage of aluminum products. Background Art
[0002] In the production process of aluminum products, surface defect detection is a key link in ensuring product quality. In early industrial production, manual visual inspection relied on the experience of quality inspectors and naked eye observation to intuitively identify obvious surface defects, but its efficiency was low and subjectivity was strong, making it difficult to adapt to high-speed assembly line operations, and the missed detection rate of minor defects was high. Traditional methods based on image processing extract defect features through algorithms such as threshold segmentation and edge detection, which improved detection efficiency to a certain extent. However, they have poor robustness when faced with complex lighting conditions and changes in defect morphology. In recent years, with the development of deep learning, methods based on convolutional neural networks have been widely used in surface defect detection of aluminum products. However, the segmentation results output by existing models are prone to blurred and incomplete boundaries in damaged areas, resulting in low detection accuracy. Therefore, there are deficiencies in existing technologies. Summary of the Invention
[0003] In response to the shortcomings of the existing technology, the purpose of the present invention is to provide a method and system for detecting surface damage of aluminum products. Through the jump connection between the encoder and the decoder, the model can simultaneously utilize detail information and semantic information when performing damage detection, thereby improving the accuracy of detection. In addition, by introducing a focus mechanism into the traditional Dice loss, the problem of imbalance between the damaged area and the background is better handled, thereby improving the detection accuracy.
[0004] To achieve the above object, the present invention provides the following technical solutions:
[0005] The present invention provides a method for detecting surface damage of aluminum products, comprising:
[0006] Acquire a surface image of the aluminum product to be inspected;
[0007] Based on the surface image and the preset model, a first damage image is obtained, wherein the value of each pixel in the first damage image represents the probability that damage exists at the position corresponding to the pixel. The preset model is obtained by training historical aluminum product image data. The preset model includes an encoder, a decoder and a transfer learning unit, and the encoder and the decoder are jump-connected.
[0008] As a further improvement of the present invention, the preset model is obtained by training historical aluminum product image data and includes:
[0009] Dividing the historical aluminum product image data into a training set and a test set;
[0010] Performing staged training on the initial model according to the training set to obtain a training model, wherein the staged training includes three stages, wherein the first stage is used to train the decoder, the second stage is used to train some layers of the encoder, and the third stage is used to train all layers of the encoder;
[0011] The preset model is obtained according to the test set and the training model.
[0012] As a further improvement of the present invention, the training set includes multiple batches, and the initial model is trained in stages according to the training set to obtain a training model, including:
[0013] For each stage, a first iterative operation is performed, which includes: obtaining a current training model, preprocessing the current batch to obtain a normalized image and a multi-scale edge prior feature map, obtaining a semantic feature map based on the normalized image, the multi-scale edge prior feature map and the encoder, inputting the semantic feature map into the transfer learning unit to obtain an enhanced feature map, obtaining a damage segmentation probability map based on the enhanced feature map, the semantic feature map and the decoder, calculating a loss function based on the damage segmentation probability map, and updating the parameters of the current training model according to the loss function until a preset termination condition is reached to obtain the training model.
[0014] As a further improvement of the present invention, the calculating of the loss function according to the damage segmentation probability map includes:
[0015] obtaining a second damage image according to the damage segmentation probability map;
[0016] For each pixel in the second damaged image, obtaining a mean and a variance within a preset neighborhood;
[0017] Obtaining a focus parameter corresponding to each pixel point according to the value, variance, and hyperparameter corresponding to each pixel point, wherein the value of the hyperparameter is determined according to the mean and variance;
[0018] Acquiring a real image corresponding to the second damaged image;
[0019] Obtaining a dice loss according to a value corresponding to each pixel in the real image and the focus parameter;
[0020] Obtaining a binary cross entropy loss and a perceptual loss according to the second damaged image and the true image;
[0021] The loss function is calculated according to the dice loss, the binary cross entropy loss, and the perceptual loss.
[0022] As a further improvement of the present invention, the preprocessing of the current batch to obtain a normalized image and a multi-scale edge prior feature map includes:
[0023] Performing image enhancement processing on the current batch to obtain an enhanced image;
[0024] performing normalization processing on the enhanced image to obtain the normalized image;
[0025] Performing convolution processing on the normalized image to obtain a blurred image;
[0026] The multi-scale edge prior feature map is obtained according to the blurred image and the edge detection algorithm.
[0027] As a further improvement of the present invention, obtaining a semantic feature map based on the normalized image, the multi-scale edge prior feature map and the encoder includes:
[0028] Concatenating the normalized image and the multi-scale edge prior feature map in the channel dimension to form a tensor;
[0029] The tensor is input into the encoder to obtain the semantic feature map, wherein the semantic feature map has a one-to-one correspondence with the convolutional layer in the encoder.
[0030] As a further improvement of the present invention, the transfer learning unit includes a boundary prior transfer module and a boundary attention mechanism module, and the semantic feature map is input into the transfer learning unit to obtain an enhanced feature map, including:
[0031] Obtaining a contour prior map according to the semantic feature map and the boundary prior migration module;
[0032] The semantic feature map and the contour prior map are input into the boundary attention mechanism module to obtain the enhanced feature map.
[0033] As a further improvement of the present invention, obtaining a damage segmentation probability map according to the enhanced feature map, the semantic feature map, and the decoder includes:
[0034] Obtaining a feature map according to the enhanced feature map and the encoder;
[0035] For each deconvolution layer in the decoder, a second iterative operation is performed, which includes: obtaining a current feature map, performing deconvolution processing on the current feature map, obtaining the semantic feature map corresponding to the deconvolution layer, splicing the deconvolution-processed image with the semantic feature map and performing convolution processing until the convolved image reaches a preset size, and outputting the damage segmentation probability map.
[0036] As a further improvement of the present invention, updating the parameters of the current training model according to the loss function includes:
[0037] Calculate the gradient of each layer parameter in the current training model according to the loss function and the back propagation algorithm;
[0038] Update the parameters of the current training model based on the gradient and the optimizer.
[0039] The present invention provides an aluminum product surface damage detection system, comprising:
[0040] An acquisition module, used to obtain a surface image of the aluminum product to be inspected;
[0041] A calculation module is used to obtain a first damage image based on the surface image and a preset model, wherein the value of each pixel in the first damage image represents the probability that damage exists at the position corresponding to the pixel, and the preset model is obtained by training historical aluminum product image data. The preset model includes an encoder, a decoder and a transfer learning unit, and the encoder and the decoder are jump-connected.
[0042] The beneficial effects of the present invention are as follows: through the jump connection between the encoder and the decoder, features of different scales are integrated into the corresponding layers of the decoder, so that the model can simultaneously utilize detail information and semantic information when performing damage detection, thereby improving the accuracy of detection. Moreover, since the jump connection provides an additional information transmission path, the model can find the optimal solution more quickly during the training process, thereby accelerating the convergence of the model. At the same time, the present invention introduces a focus mechanism into the Dice loss, and through an adaptive weight adjustment strategy, reduces the weight of the background and increases the weight of the damaged area, so that the model can better adapt to the imbalance of the data and improve the detection capability of the damaged area. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 is a flow chart of the method steps of the present invention;
[0044] Figure 2 It is a structural diagram of the encoder;
[0045] Figure 3 It is a structural diagram of the transfer learning unit;
[0046] Figure 4 A schematic diagram of the decoder structure. DETAILED DESCRIPTION
[0047] The technical solution of the present invention is described in detail below through the accompanying drawings and specific embodiments. It should be understood that the embodiments of the present invention and the specific features in the embodiments are detailed descriptions of the technical solution of the present invention, rather than limitations of the technical solution of the present invention.
[0048] The term "and / or" in the following text simply describes an association between related objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A exists alone, A and B exist simultaneously, or B exists alone. Additionally, the character " / " generally indicates an "or" relationship between the related objects.
[0049] like Figure 1 As shown, a method for detecting surface damage of aluminum products provided by the present invention comprises:
[0050] Acquire a surface image of the aluminum product to be inspected;
[0051] Based on the surface image and the preset model, a first damage image is obtained. The value of each pixel in the first damage image represents the probability that damage exists at the position corresponding to the pixel. The preset model is obtained by training historical aluminum product image data. The preset model includes an encoder, a decoder and a transfer learning unit. The encoder and the decoder are jump-connected.
[0052] This embodiment uses the U-Net model as its basic architecture. U-Net is a symmetrical encoder-decoder structure. The encoder is responsible for extracting image features, and the decoder restores the features to the original image size for damage prediction. On this basis, a transfer learning unit is introduced to enhance the boundary features of the damaged area. Specifically, the encoder is composed of multiple convolutional layers and pooling layers stacked alternately. The convolutional layers are used to extract features, and the pooling layers are used to downsample the images output by the convolutional layers. As the network level deepens, the number of channels in the convolutional layers gradually increases. The decoder is composed of multiple deconvolutional layers (also called upsampling layers) stacked alternately with convolutional layers. The upsampling layers are used to restore the image size, and the convolutional layers are used to extract features based on the restored image.
[0053] The initial parameters in the encoder are obtained according to the ResNet34 model. Specifically, the ResNet34 model can be trained based on large-scale image datasets such as ImageNet. The ImageNet dataset contains a large number of images of different categories and scenes. The ResNet34 model has been fully trained on this dataset and has learned a wealth of general image features. These features have strong generalization capabilities and can adapt to a variety of image-related tasks. For the task of detecting damage to the surface of aluminum products, although the image content is specific to the surface of aluminum products, these general features can still be used to quickly extract basic visual information in the image, such as the texture and contour of the surface of the aluminum product, laying the foundation for more accurate extraction of damage features in the future. However, this embodiment is not limited to this, and those skilled in the art can choose other datasets according to actual conditions.
[0054] This embodiment then directly loads the weight parameters of the ResNet34 model into the encoder as the initial parameters of the encoder network. Specifically, this embodiment removes the classification head of the ResNet34 model (i.e., the final fully connected layer used for image classification). Because the original classification task differs from the aluminum product surface damage detection task, the parameters of the classification head are no longer applicable to the new task. Therefore, this embodiment retains only the convolutional layer parameters of the pre-trained model and uses them as the initial parameters of the encoder. These convolutional layer parameters incorporate the learned general image feature extraction capabilities and continue to be effective in the new task, performing preliminary feature extraction on the input aluminum product surface image. The initial parameters in the decoder and transfer learning unit are generated through random initialization.
[0055] Among them, skip connections, also known as shortcut connections or cross-layer connections, allow the network to directly transfer information between different layers, thereby directly transferring features learned by shallow networks to deeper networks. In this embodiment, semantic feature maps corresponding to different layers in the encoder are directly transferred to the corresponding layers of the decoder via skip connections. These semantic feature maps contain information at different scales and levels. In the decoder, these semantic feature maps from the encoder are spliced with the decoder's upsampled image. This allows the decoder to utilize richer and more comprehensive feature information when restoring the image size and performing damage prediction, thereby improving the accuracy of damage prediction.
[0056] This embodiment uses skip connections between the encoder and decoder to fuse features of different scales into the corresponding layers of the decoder. This allows the model to utilize both detailed and semantic information during damage detection, improving detection accuracy. Furthermore, because skip connections provide additional information transmission paths, the model can find the optimal solution more quickly during training, accelerating model convergence.
[0057] Furthermore, this embodiment provides a step of obtaining a preset model by training historical aluminum product image data, including:
[0058] The historical aluminum product image data is divided into a training set and a test set;
[0059] The initial model is trained in stages according to the training set to obtain a training model. The staged training includes three stages: the first stage is used to train the decoder, the second stage is used to train some layers of the encoder, and the third stage is used to train all layers of the encoder;
[0060] Get the preset model based on the test set and training model.
[0061] Specifically, the data ratio of the training set and the test set is 3:7. In the first stage, because the encoder's initial parameters use the parameters of the convolutional layers of the ResNet34 model, the ResNet34 model has learned a wealth of common image features on large-scale image datasets. These features are of certain fundamental value for aluminum surface damage detection. Therefore, this embodiment fixes all encoder layers so that their parameters are not updated during this stage to avoid excessive parameter modification and information loss. In this stage, only the decoder and boundary attention mechanism module are trained for 5 epochs with a learning rate of 0.001, so that these two components are initially adapted to the aluminum surface damage detection task. The boundary attention mechanism module is located in the transfer learning unit.
[0062] After the first stage of training, the model enters the second stage. In this stage, the last two layers of the encoder are unfrozen, allowing the parameters of these two layers to participate in training. Training continues for three epochs. During these three epochs, the decoder and boundary attention mechanism modules are updated at a learning rate of 0.001, and the parameters of the last two layers of the encoder are updated at a learning rate of 0.0001. At this time, the last two layers of the encoder are fine-tuned for the surface damage data of aluminum products based on the existing general features, and work together with the decoder and boundary attention mechanism modules to improve the overall performance of the model.
[0063] After the second phase, the third phase begins. During this phase, all encoder layers are unfrozen, allowing all model parameters to participate in training and updating. Training lasts for two epochs. To avoid excessive degradation of pre-trained features, the learning rate for the encoder layer is set to 0.00001, while the learning rate for the decoder and boundary attention mechanism modules remains at 0.001. This third phase of training allows the model to better adapt to the aluminum surface damage detection task without degrading pre-trained features. The parameters are gradually adjusted to their optimal state, improving the model's ability to extract and detect aluminum surface damage features.
[0064] The learning rate is a hyperparameter that determines the step size for parameter updates during model training. If the learning rate is set too high, the optimal solution may be skipped during parameter updates, resulting in model convergence failure and even a continuous increase in the loss function value. If the learning rate is set too low, parameter updates will be very slow, requiring more training time and computing resources to achieve good performance. The learning rate is fixed in each stage of this embodiment. In the first stage, a relatively large learning rate of 0.001 is set to enable the decoder and boundary attention mechanism modules to quickly adapt to the damage detection task, allowing for rapid parameter adjustment of these two components. In the second stage, the last two layers of the encoder are unfrozen. To prevent over-adjustment of their parameters, their learning rates are set to 0.0001. In the third stage, all encoder layers are unfrozen. To fine-tune the model parameters and prevent over-modification, the encoder learning rate is further reduced to 0.00001. The learning rate settings in this embodiment facilitate efficient and stable model training at different stages. At the same time, this embodiment sets different numbers of epochs for different stages to control the degree of parameter adjustment in each stage and prevent excessive parameter adjustment. However, the specific values of the learning rate and the number of epochs in this embodiment are only examples. Those skilled in the art can make adjustments based on the values provided in this embodiment, and this embodiment does not impose any restrictions on this.
[0065] After obtaining the training model, the test set is input into the training model, and the damaged image corresponding to the test set is output. Then, the real image corresponding to the test set is obtained, and a series of evaluation indicators such as accuracy, precision, recall, F1-score, intersection over union (IoU), etc. are calculated based on the real image and the damaged image. If the accuracy, precision, recall, F1-score and IoU are high, it means that the model performs well on the test set and can more accurately detect the damaged area on the surface of the aluminum product. At this time, the training model is used as the preset model; if some indicators are low, such as low recall, which means that the model has many missed detections, and low precision, which means that the model has many false detections, it is necessary to further analyze the reasons and retrain to obtain the preset model.
[0066] This embodiment uses staged training to avoid large-scale parameter adjustments of the entire model at the beginning, allowing the model to learn more robustly during training and avoiding overfitting of the model to the training data due to premature or excessive adjustment of all parameters. It also reduces the complexity and instability of training. By gradually adapting to the training data, the model can better generalize to new, unseen images of surface damage on aluminum products, thereby improving the robustness and reliability of the model.
[0067] Furthermore, this embodiment provides a step of performing stage-by-stage training on the initial model according to the training set to obtain a training model, including:
[0068] For each stage, a first iterative operation is performed, which includes: obtaining the current training model, preprocessing the current batch to obtain a normalized image and a multi-scale edge prior feature map, obtaining a semantic feature map based on the normalized image, the multi-scale edge prior feature map and the encoder, inputting the semantic feature map into the transfer learning unit to obtain an enhanced feature map, obtaining a damage segmentation probability map based on the enhanced feature map, the semantic feature map and the decoder, calculating the loss function based on the damage segmentation probability map, and updating the parameters of the current training model according to the loss function until the preset termination condition is reached to obtain the training model.
[0069] The preset termination condition is that the preset number of iterations is reached. The number of iterations is determined by the number of batches contained in each epoch. An epoch represents the process of the model performing a complete training of the entire training set. That is, when the model inputs all images in the training set and completes the parameter update, an epoch is completed.
[0070] Specifically, each epoch processes multiple batches of images, preferably 32 images per batch. Therefore, before training, the training set needs to be divided into multiple batches. Assuming the number of batches is n, the steps from preprocessing to parameter updating are performed sequentially for each batch. For example, for the first stage, the number of epochs is 5 and the number of batches is n, corresponding to a number of iterations of 5n. Within each epoch, the steps from preprocessing to parameter updating are performed sequentially for each batch.
[0071] For the first stage, the current training model is the initialization model, and the parameters in the model have not been updated. For the second stage, the current training model is the model obtained after the first stage training. For the third stage, the current training model is the model obtained after the second stage training, and the model obtained after the third stage training is the training model, which needs to be tested with the test set to obtain the preset model.
[0072] In this embodiment, the normalized image and the multi-scale edge prior feature map are spliced and input into the encoder of the U-Net model, combining the original image information with the edge prior features to broaden the feature dimension. The encoder then quickly extracts rich features from the bottom-level edges and textures to the high-level semantics of the aluminum product surface image to obtain a semantic feature map. A boundary prior migration module is introduced to generate a contour prior map. The contour prior map and the semantic feature map are fused using the boundary attention mechanism module and the encoder to obtain an enhanced feature map. The above steps can highlight the features of the damaged area, significantly enhancing the model's perception and extraction capabilities of the damage boundary features, enabling the model to more accurately identify the damaged area and its boundaries.
[0073] Furthermore, this embodiment provides a step of preprocessing the current batch to obtain a normalized image and a multi-scale edge prior feature map, including:
[0074] Perform image enhancement processing on the current batch to obtain an enhanced image;
[0075] Normalizing the enhanced image to obtain a normalized image;
[0076] Perform convolution on the normalized image to obtain a blurred image;
[0077] According to the blurred image and edge detection algorithm, a multi-scale edge prior feature map is obtained.
[0078] Specifically, each image in the current batch is first enhanced by performing image enhancement. This involves horizontally flipping each image, rotating it 15° clockwise, and boosting its brightness by 20%. This results in four enhanced images for each image. Each enhanced image is then normalized, mapping pixel values from 0-255 to the range [0,1]. For example, a pixel with an original RGB value of (100, 120, 150) becomes (0.392, 0.471, 0.588) after normalization. Each enhanced image now corresponds to a normalized image. Each normalized image is then convolved with Gaussian kernels with standard deviations of 1, 2, and 3, respectively. This results in three blurred images for each normalized image. Gaussian kernels with different standard deviations can extract details at different scales. A kernel with a small standard deviation extracts fine edges, while a kernel with a large standard deviation extracts coarse edges. Then, for the three blurred images corresponding to each normalized image, the Canny edge detection algorithm is applied respectively to generate three single-channel edge maps. Each blurred image corresponds to a single-channel edge map. Finally, the three single-channel edge maps are spliced in the channel dimension to obtain a multi-scale edge prior feature map. At this time, each normalized image corresponds to a multi-scale edge prior feature map, that is, each image in the current batch corresponds to 4 multi-scale edge prior feature maps. The Canny edge detection algorithm is an existing technology and is not described in detail in this embodiment. The values of the image rotation angle, brightness enhancement value, and standard deviation of the Gaussian kernel are only examples. Those skilled in the art can select other appropriate values for image enhancement according to actual conditions, and this embodiment does not limit this.
[0079] This embodiment uses image enhancement to simulate images under different viewing angles and lighting conditions, and normalizes the image data to the same magnitude. Finally, Gaussian kernel convolution and the Canny edge detection algorithm are used to generate a multi-scale edge prior feature map. Through these preprocessing steps, the edge information of aluminum product surface damage at different scales can be accurately obtained, thereby enabling the model to learn richer damage features, improve the model's generalization ability, and enhance the accuracy of damage detection.
[0080] Furthermore, this embodiment provides a step of obtaining a semantic feature map based on a normalized image, a multi-scale edge prior feature map, and an encoder, including:
[0081] Concatenate the normalized image and the multi-scale edge prior feature map in the channel dimension to form a tensor;
[0082] The tensor is input into the encoder to obtain a semantic feature map, where the semantic feature map corresponds one-to-one to the convolutional layer in the encoder.
[0083] Specifically, the normalized image and the multi-scale edge prior feature map are spliced in the channel dimension to form a tensor. For example, assuming that the sizes of a normalized image and its corresponding multi-scale edge prior feature map are both 2048×1536×3, a tensor with a shape of 2048×1536×6 is formed after splicing. In this embodiment, all images corresponding to the current batch need to be input into the encoder at the same time, and convolution and pooling operations are performed on these images at the same time. Therefore, the shape of the tensor obtained in this embodiment is a×b×c×d, where a represents the number of normalized images corresponding to the current batch, that is, the number of multi-scale edge prior feature maps, b and c represent the length and width of the normalized image, that is, the length and width of the multi-scale edge prior feature map, and d represents the number of RGB channels. The tensor includes information of all images in the current batch. Based on this, Figure 2 Each semantic feature map and enhanced feature map also includes information about all images in the current batch and needs to be represented in the form of a four-dimensional tensor.
[0084] The tensor is then fed into the encoder, where the convolution and pooling steps are performed as follows: Figure 2 As shown, it is assumed that the encoder includes three convolutional layers and three pooling layers. For each convolutional layer, the corresponding semantic feature map will be output based on the corresponding input of the layer, and the semantic feature map will be input into the transfer learning unit to obtain the enhanced feature map corresponding to the semantic feature map, and the enhanced feature map will be used as the input of the next adjacent pooling layer. The pooling layer downsamples based on the input enhanced feature map and uses the downsampled image as the input of the next adjacent convolutional layer. Downsampling is also called downsampling. Its core purpose is to reduce the data dimension and the amount of computation, while enhancing the robustness of the features to a certain extent. Finally, the feature map is output by the last pooling layer as the input of the decoder, and the encoder passes each semantic feature map to the corresponding layer in the decoder, where the feature map is also a four-dimensional tensor.
[0085] The encoder provided in this embodiment is initialized according to the ResNet34 model and receives the preprocessed image. After that, it extracts features from the bottom edge and texture to the high-level semantics through multi-layer convolution and pooling operations, while reducing the image size, reducing the amount of calculation, and enhancing the robustness of features. The semantic features output by the convolution layer are Figure 1 On the one hand, the jump connection provides semantic information to the decoder, helping the decoder to accurately identify the damage area by combining the underlying details and high-level semantics. On the other hand, it is input into the transfer learning unit to highlight the damage boundary features and obtain an enhanced feature map, thereby improving the accuracy of damage detection.
[0086] Furthermore, this embodiment provides a step of inputting the semantic feature map into the transfer learning unit to obtain an enhanced feature map, including:
[0087] According to the semantic feature map and the boundary prior migration module, the contour prior map is obtained;
[0088] The semantic feature map and contour prior map are input into the boundary attention mechanism module to obtain the enhanced feature map.
[0089] The transfer learning unit includes a boundary prior transfer module and a boundary attention mechanism module. For example, Figure 3 As shown, first obtain all the multi-scale edge prior feature maps corresponding to the batch of images to form a four-dimensional tensor, and then resize the four-dimensional tensor. Specifically, the interpolation method can be used to resize it so that the resized four-dimensional tensor is consistent with the semantic feature. Figure 1 The size of the resized four-dimensional tensor is then input into the boundary prior migration module. The function of the boundary prior migration module is to use the Canny edge detection algorithm to obtain the contour prior map corresponding to the four-dimensional tensor, that is, the semantic feature, through Gaussian filtering, non-maximum suppression, and hysteresis threshold processing on the basis of the resized four-dimensional tensor. Figure 1 The corresponding contour prior map.
[0090] Then the generated contour prior map and semantic features Figure 1 The boundary attention mechanism module is input together. In the boundary attention mechanism module, the contour prior map and semantic features are first calculated. Figure 1 The similarity of the two layers is obtained by measuring their similarity matrix. This step can be achieved through a convolution layer. Then the similarity matrix is softmax-operated to obtain the attention weight matrix. The semantic features are analyzed according to the attention weight matrix. Figure 1 Perform weighted operations to enhance the features related to the damage boundary, while the weights of other non-boundary features are relatively reduced, thereby realizing the contour prior map and semantic features. Figure 1 The effective fusion of the enhanced feature map is obtained, and the size of the enhanced feature map is consistent with the semantic features. Figure 1 The same is true for all other semantic feature maps. Repeat the above steps for other semantic feature maps to obtain all enhanced feature maps.
[0091] This embodiment effectively integrates multi-scale edge prior information and semantic feature information through the transfer learning unit, enhances the model's ability to perceive and extract damage boundaries on the surface of aluminum products, and thus improves the accuracy and reliability of damage detection.
[0092] Furthermore, this embodiment provides a step of obtaining a damage segmentation probability map based on the enhanced feature map, the semantic feature map, and the decoder, including:
[0093] Obtain a feature map according to the enhanced feature map and the encoder;
[0094] For each deconvolution layer in the decoder, a second iterative operation is performed. The second iterative operation includes: obtaining the current feature map, performing deconvolution processing on the current feature map, obtaining the semantic feature map corresponding to the deconvolution layer, splicing the deconvolution image with the semantic feature map and performing convolution processing until the convolved image reaches a preset size, and outputting a damage segmentation probability map.
[0095] For example, follow Figure 2 In the example, it is assumed that the decoder includes three deconvolution layers and three convolution layers, such as Figure 4 As shown, each deconvolution layer needs to obtain the current feature map and perform deconvolution processing on the current feature map, that is, upsampling processing, and then obtain the semantic feature map corresponding to the deconvolution layer, splice the deconvolution image with the semantic feature map, and finally use the spliced result as the input of the adjacent next convolution layer. For the first deconvolution layer, the corresponding current feature map is the feature map output by the encoder. For the subsequent deconvolution layers, the corresponding current feature map is the image output by the adjacent previous convolution layer. In addition, in order to ensure that the two images to be spliced are of the same size and the spliced image has sufficient feature information, the semantic feature map corresponding to the first deconvolution layer is the semantic feature map output by the last convolution layer in the encoder. The preset size is the original size of the image, that is, the size before preprocessing. Finally, the last convolution layer outputs the damage segmentation probability map, which is still a four-dimensional tensor.
[0096] Since the encoder loses some spatial detail information while continuously reducing the feature map size through convolution and pooling operations to extract high-level semantic features, this embodiment sets a decoder to gradually increase the feature map size through upsampling operations to restore the spatial resolution of the image. In addition, during the restoration process, the features are refined through convolution operations so that the model can more accurately locate and identify the damaged area. At the same time, after upsampling each layer, the decoder will splice it with the semantic feature map output by the corresponding layer of the encoder to introduce feature information of different scales in the encoder into the decoder. When restoring the image, the decoder can not only use high-level semantic information to determine the category of the damage, but also combine the underlying detail information to accurately depict the boundary of the damage, thereby improving the accuracy of damage detection.
[0097] Furthermore, this embodiment provides a step of calculating a loss function based on a damage segmentation probability map, including:
[0098] obtaining a second damage image according to the damage segmentation probability map;
[0099] For each pixel in the second damaged image, obtaining a mean and a variance within a preset neighborhood;
[0100] According to the value, variance and hyperparameter corresponding to each pixel, the focus parameter corresponding to each pixel is obtained. The value of the hyperparameter is determined according to the mean and variance;
[0101] Acquire a real image corresponding to the second damaged image;
[0102] According to the numerical value and focus parameters corresponding to each pixel in the real image, the dice loss is obtained;
[0103] According to the second damaged image and the real image, the binary cross entropy loss and the perceptual loss are obtained;
[0104] The loss function is calculated based on dice loss, binary cross entropy loss and perceptual loss.
[0105] Among them, first, an activation function such as the Sigmoid function is used to convert each pixel value in the damage segmentation probability map into a probability value between 0 and 1. This probability value indicates the possibility that the position corresponding to the pixel belongs to the damaged area. At this time, when the converted damage segmentation probability map is represented as a four-dimensional tensor, the value corresponding to the last dimension is 1. Since the damage segmentation probability map contains the features of all images in the current batch, it is then sliced to split the four-dimensional tensor into multiple three-dimensional tensors. At this time, the value corresponding to the last dimension in the three-dimensional tensor is 1. Each three-dimensional tensor serves as a second damage image, and each second damage image corresponds to a normalized image. The value of each pixel point in the second damage image indicates the probability that there is damage at the position corresponding to the pixel point.
[0106] Specifically, taking one of the second loss images as an example, calculate its corresponding Dice loss, binary cross entropy (BCE) loss and perceptual (SSIM) loss, and express the second loss image as , the height of the image is , with a width of , Indicates the first Rank For each pixel value, get the other pixel values in the preset neighborhood centered on it. The size of the preset neighborhood is , then calculate the mean of all pixel values in the neighborhood and variance They are:
[0107]
[0108]
[0109] in, Represents the pixel value within the preset neighborhood, and relative to and By changing the offset and You can traverse the pixels at different positions in the preset neighborhood and then calculate the focus parameters :
[0110]
[0111] in and is a hyperparameter, a hyperparameter It can be determined based on the variance of each pixel value in all loss images in the batch. Specifically, the variance of each pixel value in all loss images within a preset neighborhood can be calculated, and then the mean and standard deviation of these variances can be calculated. The coefficient of variation is calculated based on the mean and standard deviation. If the coefficient of variation is greater than the preset value, it means that the variance fluctuates greatly. In order to make the focus parameter pay more attention to the difference in variance, at this time It is the inverse of the mean of these variances. If the coefficient of variation is less than the preset value, it can be appropriately increased. , to enhance the influence of variance on focal parameters, then It is twice the inverse of the mean of these variances. Preferably, the default value is 0.5. , can be achieved through The distribution of Indicates the loss of any pixel value in the image. If most When the value is concentrated in a small range and close to 0, it means that the model has a high degree of confidence in predicting most pixel values. In this case, you can increase The value of, for example =3, otherwise, if If the data is more dispersed, set a smaller value to avoid over-weighting, for example =1, determining whether a group of data is centralized or decentralized is a prior art, and this embodiment will not be described in detail here.
[0112] Then, the aluminum product image corresponding to the second damage image is obtained from the historical aluminum product image data, and the real image is obtained based on the aluminum product image. The real image is a binary image, that is, if there is damage at the position corresponding to a pixel point, the pixel value corresponding to the pixel point is 1, otherwise it is 0. The size of the real image is the same as the second damage image. The real image is recorded as No. Rank The pixel value of the column is recorded as .
[0113] Then calculate the sum of the product of the weighted pixel value and the pixel value of the real image :
[0114]
[0115] Then calculate the weighted sum of the squares of the pixel values :
[0116]
[0117] Then calculate the sum of the squares of the pixel values of the weighted real image :
[0118]
[0119] Finally, the Dice loss corresponding to the second damaged image is:
[0120]
[0121] Then calculate the binary cross entropy (BCE) loss of the second loss image :
[0122]
[0123] in, is the set of boundary pixels, Indicates the number of boundary pixels. Boundary pixels are pixels with a pixel value of 1 in the real image. Indicates the index of the boundary pixel point, Indicates that the index in the real image is The pixel value of the pixel point, Indicates that the index in the second loss image is The pixel value of the pixel.
[0124] Then calculate the perceptual (SSIM) loss of the second loss image Specifically, first, the second loss image is divided into multiple sub-regions, and correspondingly, the real image is also divided into the same multiple sub-regions. The corresponding SSIM value of the sub-region is:
[0125]
[0126] in, and represent the second loss image and the first sub-regions, and Respectively and The mean of the inner pixel values, and Respectively and The standard deviation of the pixel values inside, express and The covariance of the pixel values within the and is a constant to avoid the denominator being zero.
[0127] Finally got ,in is the number of sub-regions, Indicates the The weight of a sub-region can be determined according to the density of boundary pixels in the sub-region.
[0128] Assume that the batch has The above steps need to be repeated for each second loss image to obtain the Dice loss, binary cross entropy (BCE) loss and perceptual (SSIM) loss corresponding to each second loss image, and then the loss function is obtained. :
[0129]
[0130] in, 、 and Respectively represent The Dice loss, Binary Cross Entropy (BCE) loss, and Perceptual (SSIM) loss corresponding to the second loss image, 、 and Specifically, this embodiment does not limit the specific value of the weight, and the weight can be determined according to the damage detection requirements. For example, if in the surface damage detection of aluminum products, more attention is paid to the accurate identification of the damaged area, the weight can be appropriately increased. If you want the model to identify the boundary of the damaged area more accurately, you can increase If the structural authenticity of the damaged area is required to be high, the value can be increased. For example, when inspecting high-precision aviation aluminum parts, the integrity and boundary accuracy of the damaged area identification are extremely high. In this case, the weight can be set to , , .
[0131] In existing technologies, the damaged areas in the model output results are prone to blurred boundaries and incompleteness, resulting in low detection accuracy. However, this embodiment discovered that the traditional Dice loss used in existing technologies assigns equal weight to all samples during calculation. However, the historical aluminum product image data currently used exhibits a significant imbalance between the damaged area and the background. For example, in historical aluminum product image data, the normal background area occupies the majority of the image, while the damaged area is only a small portion, even accounting for only a few percent of the total image pixels. For example, in a 1000×1000 pixel image, the damaged area only has 100 pixels, while the background area has 999,900 pixels. This significant disparity in numbers reflects a significant imbalance between the damaged area and the background.
[0132] This embodiment takes into account that in models such as convolutional neural networks, feature extraction is achieved through multi-layer convolution and pooling operations. If the damaged area appears infrequently in the training data, the model will not be able to learn enough detailed features about the texture, shape, etc. of the damaged area, resulting in the inability to accurately outline the contour of the damaged area during prediction, causing blurred surface defects. In addition, during the training process, the model will adjust its own parameters based on the input sample data to minimize the loss function. However, due to the large number of background areas, the model will tend to predict most samples as background to reduce the overall loss. However, this will also lead to insufficient learning of the damaged area by the model, and it is easy to ignore the characteristics of the damaged area, thereby affecting the model's ability to detect and identify the damaged area.
[0133] To address the severe imbalance between damage areas and background, this embodiment introduces a focus mechanism into the traditional Dice loss. This mechanism uses a modulation factor to reduce the weight of background areas and increase the weight of damage areas. This allows the model to focus more on learning the features of minority samples, such as damage areas, thereby improving damage detection accuracy. For example, the model might originally ignore or misclassify subtle, less distinct damage areas. However, with the focus mechanism, it will place greater emphasis on learning and classifying these areas. Furthermore, the focus mechanism adaptively adjusts weights based on the difficulty of classifying image samples. The more difficult the area, the greater the weight, while the weight gradually decreases for easier areas. The difficulty of classification is determined by the complexity of the damage. For example, a clear, linear scratch has distinct features, making it easy for the model to extract its features. However, complex defects, such as irregularly shaped corrosion areas with textures similar to the background, have blurred edges and variable textures, making it difficult for the model to extract effective features. Therefore, the focus mechanism dynamically adjusts the weights of pixels in the loss function based on the actual conditions of each image, allowing the model to better adapt to data imbalance and comprehensively improve damage detection capabilities. Furthermore, this embodiment combines the other two losses to enable the model to enhance its perception of boundary details while ensuring the accuracy of feature extraction.
[0134] Furthermore, this embodiment provides a step of updating the parameters of the current training model according to the loss function, including:
[0135] Calculate the gradient of each layer parameter in the current training model according to the loss function and backpropagation algorithm;
[0136] Update the parameters of the currently trained model based on the gradient and the optimizer.
[0137] Specifically, after obtaining the loss function, the gradient needs to be calculated using the backpropagation algorithm. During backpropagation, the gradient of the loss function with respect to the output is propagated back through each layer of the network using the chain rule. During this process, the gradient of the loss function with respect to each parameter (such as the weights and biases of the convolutional layers in the encoder and decoder) is calculated. The gradient value reflects the degree and direction of the parameter's influence on the loss function. Calculating the gradient of the loss function with respect to the output is to calculate the vector consisting of the partial derivatives of the loss function with respect to the model output (i.e., each pixel value in the lesion segmentation probability map).
[0138] Then, an optimizer is used to adjust the parameters based on the calculated gradients, updating the parameters in a direction that reduces the loss function value, thereby optimizing the model. Specific optimizers can be Adam or SGD.
[0139] Finally, after obtaining the preset model, the surface image of the aluminum product to be inspected is input into the preset model, wherein the surface image of the aluminum product to be inspected has been preprocessed, and then the model outputs a damage image, wherein the value of each pixel in the damage image represents the probability of damage existing at the position corresponding to the pixel. Then, each pixel in the damage image is traversed, and for each pixel, its corresponding probability value is compared with the set threshold. If the probability value of the pixel is greater than or equal to the threshold, the value of the pixel is set to 1, indicating that the point belongs to the damaged area; if the probability value of the pixel is less than the threshold, the value of the pixel is set to 0, indicating that the point belongs to the normal area. By performing the above judgment and conversion operations on all pixels, a binary image is finally obtained to clearly present the distribution of damage and normal areas on the surface of the aluminum product. The preferred threshold is equal to 0.5.
[0140] Furthermore, each pixel in the binary image can be traversed. When a point with a pixel value of "1" is encountered, its coordinate information is recorded to determine the boundary range of the damaged area and the positioning coordinates of the damaged area. Specifically, for a continuous damaged area, the point with the smallest horizontal coordinate and the smallest vertical coordinate among all the points with a pixel value of "1" is found as the coordinate of the upper left corner; the point with the largest horizontal coordinate and the largest vertical coordinate is found as the coordinate of the lower right corner. For example, during the traversal process, it is found that the minimum horizontal coordinate value of the pixel points in the damaged area is 500, the minimum vertical coordinate value is 600, the maximum horizontal coordinate value is 1200, and the maximum vertical coordinate value is 700. Then the coordinates of the upper left corner of the damaged area are determined to be (500, 600) and the coordinates of the lower right corner are determined to be (1200, 700).
[0141] An embodiment of the present application provides a method for detecting surface damage of aluminum products. By using a jump connection between an encoder and a decoder, features of different scales are integrated into the corresponding levels of the decoder. This allows the model to utilize both detail information and semantic information when performing damage detection, thereby improving detection accuracy. Furthermore, because the jump connection provides an additional information transmission path, the model can find the optimal solution more quickly during training, accelerating the convergence of the model. Furthermore, the present invention introduces a focus mechanism into the Dice loss. By using an adaptive weight adjustment strategy, the weight of the background is reduced and the weight of the damaged area is increased. This enables the model to better adapt to data imbalance and improve the ability to detect damaged areas.
[0142] Furthermore, an embodiment of the present application provides an aluminum product surface damage detection system, comprising:
[0143] An acquisition module, used to obtain a surface image of the aluminum product to be inspected;
[0144] A calculation module is used to obtain a first damage image based on the surface image and a preset model. The value of each pixel in the first damage image represents the probability that damage exists at the position corresponding to the pixel. The preset model is obtained by training historical aluminum product image data. The preset model includes an encoder, a decoder and a transfer learning unit. The encoder and the decoder are jump-connected.
[0145] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0146] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0147] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0148] The above description is merely a preferred embodiment of the present invention. The scope of protection of the present invention is not limited to the above embodiment. All technical solutions based on the concept of the present invention are within the scope of protection of the present invention. It should be noted that for those skilled in the art, various improvements and modifications that do not depart from the principles of the present invention should also be considered within the scope of protection of the present invention.
Claims
1. A method for detecting surface damage of aluminum products, characterized in that: include: Acquire a surface image of the aluminum product to be inspected; A first damage image is obtained based on the surface image and a preset model, wherein the value of each pixel in the first damage image represents a probability that damage exists at the position corresponding to the pixel, and the preset model is trained using historical aluminum product image data. The preset model includes an encoder, a decoder, and a transfer learning unit, and the encoder and the decoder are jump-connected; The preset model is trained by historical aluminum product image data and includes: Dividing the historical aluminum product image data into a training set and a test set; Performing staged training on the initial model according to the training set to obtain a training model, wherein the staged training includes three stages, wherein the first stage is used to train the decoder, the second stage is used to train some layers of the encoder, and the third stage is used to train all layers of the encoder; Obtaining the preset model according to the test set and the training model; The training set includes multiple batches, and the initial model is trained in stages according to the training set to obtain a training model, including: For each stage, a first iterative operation is performed, which includes: obtaining a current training model, preprocessing the current batch to obtain a normalized image and a multi-scale edge prior feature map, obtaining a semantic feature map based on the normalized image, the multi-scale edge prior feature map and the encoder, inputting the semantic feature map into the transfer learning unit to obtain an enhanced feature map, obtaining a damage segmentation probability map based on the enhanced feature map, the semantic feature map and the decoder, calculating a loss function based on the damage segmentation probability map, and updating the parameters of the current training model according to the loss function until a preset termination condition is reached to obtain the training model.
2. The method for detecting surface damage of aluminum products according to claim 1, characterized in that: The calculating of the loss function according to the damage segmentation probability map includes: obtaining a second damage image according to the damage segmentation probability map; For each pixel in the second damaged image, obtaining a mean and a variance within a preset neighborhood; Obtaining a focus parameter corresponding to each pixel point according to the value, variance, and hyperparameter corresponding to each pixel point, wherein the value of the hyperparameter is determined according to the mean and variance; Acquiring a real image corresponding to the second damaged image; Obtaining a dice loss according to a value corresponding to each pixel point in the real image and the focus parameter; Obtaining a binary cross entropy loss and a perceptual loss according to the second damaged image and the true image; The loss function is calculated according to the dice loss, the binary cross entropy loss, and the perceptual loss.
3. The method for detecting surface damage of aluminum products according to claim 1, characterized in that: The preprocessing of the current batch to obtain a normalized image and a multi-scale edge prior feature map includes: Performing image enhancement processing on the current batch to obtain an enhanced image; performing normalization processing on the enhanced image to obtain the normalized image; Performing convolution processing on the normalized image to obtain a blurred image; The multi-scale edge prior feature map is obtained according to the blurred image and the edge detection algorithm.
4. The method for detecting surface damage of aluminum products according to claim 1, characterized in that: The step of obtaining a semantic feature map according to the normalized image, the multi-scale edge prior feature map, and the encoder includes: Concatenating the normalized image and the multi-scale edge prior feature map in the channel dimension to form a tensor; The tensor is input into the encoder to obtain the semantic feature map, wherein the semantic feature map has a one-to-one correspondence with the convolutional layer in the encoder.
5. The method for detecting surface damage of aluminum products according to claim 1, characterized in that: The transfer learning unit includes a boundary prior transfer module and a boundary attention mechanism module. The inputting the semantic feature map into the transfer learning unit to obtain an enhanced feature map includes: Obtaining a contour prior map according to the semantic feature map and the boundary prior migration module; The semantic feature map and the contour prior map are input into the boundary attention mechanism module to obtain the enhanced feature map.
6. The method for detecting surface damage of aluminum products according to claim 1, characterized in that: The obtaining of a damage segmentation probability map according to the enhanced feature map, the semantic feature map, and the decoder includes: Obtaining a feature map according to the enhanced feature map and the encoder; For each deconvolution layer in the decoder, a second iterative operation is performed, which includes: obtaining a current feature map, performing deconvolution processing on the current feature map, obtaining the semantic feature map corresponding to the deconvolution layer, splicing the deconvolution-processed image with the semantic feature map and performing convolution processing until the convolved image reaches a preset size, and outputting the damage segmentation probability map.
7. The method for detecting surface damage of aluminum products according to claim 1, characterized in that: The updating of the parameters of the current training model according to the loss function includes: Calculate the gradient of each layer parameter in the current training model according to the loss function and the back-propagation algorithm; Update the parameters of the current training model based on the gradient and the optimizer.
8. A system for detecting surface damage of aluminum products, used to implement the method for detecting surface damage of aluminum products according to any one of claims 1 to 7, characterized in that: include: An acquisition module, used to obtain a surface image of the aluminum product to be inspected; a calculation module, configured to obtain a first damage image based on the surface image and a preset model, wherein the value of each pixel in the first damage image represents a probability of damage existing at the location corresponding to the pixel, the preset model being trained using historical aluminum product image data, the preset model comprising an encoder, a decoder, and a transfer learning unit, the encoder and the decoder being jump-connected; The preset model is trained by historical aluminum product image data and includes: Dividing the historical aluminum product image data into a training set and a test set; Performing staged training on the initial model according to the training set to obtain a training model, wherein the staged training includes three stages, wherein the first stage is used to train the decoder, the second stage is used to train some layers of the encoder, and the third stage is used to train all layers of the encoder; Obtaining the preset model according to the test set and the training model; The training set includes multiple batches, and the initial model is trained in stages according to the training set to obtain a training model, including: For each stage, a first iterative operation is performed, which includes: obtaining a current training model, preprocessing the current batch to obtain a normalized image and a multi-scale edge prior feature map, obtaining a semantic feature map based on the normalized image, the multi-scale edge prior feature map and the encoder, inputting the semantic feature map into the transfer learning unit to obtain an enhanced feature map, obtaining a damage segmentation probability map based on the enhanced feature map, the semantic feature map and the decoder, calculating a loss function based on the damage segmentation probability map, and updating the parameters of the current training model according to the loss function until a preset termination condition is reached to obtain the training model.
Citation Information
Patent Citations
Skin disease image segmentation method and system based on joint attention convolutional neural network
CN115457021A