Image segmentation method based on feature fusion and attention in industrial flaw detection
By using the image segmentation method of feature fusion and attention in industrial defect detection, the problem of insufficient application of multi-scale feature fusion and attention mechanism in the prior art is solved, and high-precision and high-efficiency industrial defect detection is achieved to adapt to complex industrial environments.
Patent Information
- Application Number
- CN202510204504.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2025-06-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing industrial defect detection methods are difficult to effectively integrate multi-scale features, use attention mechanisms to highlight defect areas, and adapt to complex industrial environments, resulting in low detection accuracy and efficiency.
The image segmentation method based on feature fusion and attention is adopted, and multi-scale features are extracted through convolutional neural networks, feature fusion is performed, and the features of defective areas are highlighted by combining channel and spatial attention mechanisms. Finally, the segmentation network of U-Net structure is used for image segmentation.
It improves the detection ability of defects of different sizes, enhances the adaptability to complex backgrounds and diverse products, realizes high-precision and high-efficiency industrial defect detection, and meets the needs of industrial production for real-time inspection.
Smart Images

Figure CN120147242A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and particularly relates to an image segmentation method based on feature fusion and attention in industrial defect detection. Background Art
[0002] In the field of industrial production, product quality inspection is of crucial importance, directly affecting the production efficiency and market competitiveness of enterprises. Among them, industrial defect detection is a key link to ensure product quality. Traditional industrial defect detection methods mainly rely on manual inspection, where inspectors check the product surface with the naked eye and experience. However, this method has many limitations. The efficiency of manual inspection is low. With the continuous expansion of industrial production scale, the output of products has increased significantly, and it is difficult for manual inspection to meet the needs of large-scale production, resulting in the detection speed being unable to keep up with the production rhythm and affecting production efficiency. Moreover, manual inspection is easily affected by subjective factors, fatigue levels, and work experience of inspectors, making it difficult to guarantee the accuracy and consistency of detection results. Different inspectors may have different judgment criteria for the same defect, and even the same inspector may have fluctuations in detection results under different working conditions. With the development of computer vision technology, industrial defect detection methods based on image processing have gradually been applied. Early methods mainly relied on simple image processing algorithms, such as threshold segmentation, edge detection, etc. Although these methods can detect some obvious defects, they have poor detection effects for defects in complex backgrounds, small defects, or irregularly shaped defects. Because they often can only extract single features of the image, cannot make full use of the rich information in the image, and have poor robustness to factors such as illumination changes and noise interference. In recent years, deep learning technology has achieved great success in the field of images, and some industrial defect detection methods based on convolutional neural networks (CNNs) have emerged. These methods can automatically learn features from images and improve the accuracy and efficiency of detection to a certain extent. However, existing CNN-based methods still have some problems. On the one hand, they usually only focus on features at a single scale, ignoring the complementarity between different scale features, resulting in unbalanced detection capabilities for defects of different sizes. For example, they can detect large-sized defects well, but are prone to missing small-sized defects. On the other hand, existing methods lack an effective attention mechanism and cannot automatically focus on the defect areas in the image, making the model vulnerable to background information interference when processing complex images and reducing the detection accuracy. In addition, on the industrial production line, the environmental conditions are complex and variable, and the image acquisition process may be affected by factors such as illumination, vibration, and dust, resulting in unstable image quality. Moreover, different types of industrial products have different surface textures, colors, and shapes, which further increases the difficulty of defect detection. Existing industrial defect detection methods are difficult to adapt to these complex environments and diverse product characteristics and cannot meet the requirements of industrial production for high-precision and high-efficiency defect detection. Therefore, it is of great practical significance to develop an image segmentation method that can effectively fuse multi-scale features, use the attention mechanism to highlight the defect areas, and adapt to complex industrial environments. The image segmentation method based on feature fusion and attention in industrial defect detection proposed in this patent is designed to solve the above problems. Summary of the Invention
[0003] An image segmentation method based on feature fusion and attention in industrial defect detection proposed by the present invention is used to solve the problems mentioned in the above-mentioned prior art.
[0004] To achieve the above object, the present invention adopts the following technical solutions: An image segmentation method based on feature fusion and attention in industrial defect detection, comprising the following steps:
[0005] Data preprocessing step: Obtain the original image of the industrial product, and the image is sourced from cameras and industrial camera devices on the industrial production line; perform grayscale conversion, normalization, and filtering on the image to remove noise interference; the grayscale conversion adopts the weighted average method, and the formula is Gray = 0.299R + 0.587G + 0.114B. After normalization, the pixel values are scaled to the interval [0, 1], specifically through the formula where p is the original pixel value, p min and p max are respectively the minimum and maximum values of the image pixel values. Gaussian filtering is used for filtering, and the size of the Gaussian filter kernel is adjusted according to the resolution and noise conditions of the image;
[0006] Feature extraction step: Use the convolutional neural network CNN to extract features from the preprocessed image, adopt different sizes of convolutional kernels and different numbers of convolutional layers to obtain feature maps of different scales. The size of the convolutional kernel is selected based on experience and experimental results. Different numbers of convolutional layers construct feature hierarchies of different depths. Shallow convolutional layers extract the edge and texture underlying features of the image, and deeper convolutional layers extract more abstract semantic features;
[0007] Feature fusion step: Fuse the feature maps of different scales obtained in the feature extraction step, and adopt the methods of feature splicing and weighted summation; first, splice the feature maps of different scales in the channel dimension, and the weights can be learned. Then, calculate the fused feature map F according to the formula fused where F i is the i-th feature map, and w i is the corresponding weight;
[0008] Attention mechanism application step: Apply the attention mechanism to the fused feature map, and adopt the combination of channel attention and spatial attention; channel attention compresses the feature map in the spatial dimension through global average pooling to obtain the global feature information of each channel, and then calculates the channel weights through a fully connected layer and an activation function. Spatial attention processes the feature map through convolutional operations to calculate the spatial weights and highlight the important spatial regions in the feature map; according to the formula A final = A channel × Aspace Calculate the final attention map A final , multiply the results of channel attention and spatial attention to highlight the features of the defect areas;
[0009] Image segmentation step: Use a segmentation network to segment the feature map processed by the attention mechanism and output the segmentation results of the defect areas; The segmentation network adopts a U-Net structure and realizes image segmentation through an encoder-decoder structure; The encoder part gradually reduces the size of the feature map through convolution and pooling operations to extract the high-level semantic features of the image; The decoder part gradually restores the size of the feature map through deconvolution and upsampling operations, and combines skip connections to splice the feature maps of the same size in the encoder, and finally outputs a segmentation result with the same size as the original image, and each pixel is classified as a defect area or a normal area.
[0010] Furthermore, it also includes the following steps:
[0011] Post-processing step: Post-process the segmentation results, using morphological operations and connected component analysis; The dilation operation expands the pixels around the defect area to connect adjacent small defect areas; The erosion operation shrinks the defect area, and the size of the structural element used in the dilation and erosion operations is adjusted according to the size and shape of the defect; Connected component analysis is to label each connected component in the segmentation result, and judge whether it is a noise area according to the area threshold S of the connected component th of the area less than S th is removed, and for adjacent defect areas, if the distance between them is less than the set threshold d th , then they are merged into one defect area.
[0012] Model training step: Use the labeled industrial defect image dataset to train the entire segmentation model, adopt the cross-entropy loss function and the Adam optimizer, and calculate the loss L according to the formula where y i is the true label, and p i is the predicted probability. During the training process, the dataset is divided into a training set and a validation set, and the model parameters are continuously adjusted.
[0013] Performance evaluation step: Use the test dataset to evaluate the performance of the trained model. The test dataset should be independent of the training dataset. Adopt accuracy, recall rate, and F1-score metrics, and calculate according to the formula where TP is predicted as a defect and actually is a defect, TN is predicted as normal and actually is normal, FP is predicted as a defect but actually is normal, and FN is predicted as normal but actually is a defect.
[0014] Real-time detection step: The trained model is deployed to the industrial production line to detect and segment defects in real-time collected industrial product images. Images are obtained in real-time through image acquisition devices, and the model outputs the segmentation results. GPU acceleration is used for computing to improve the inference speed of the model.
[0015] Dynamic adjustment step: According to the real-time detection results and performance evaluation metrics, dynamically adjust the parameters of the model and the weights of the attention mechanism. When the accuracy rate is lower than the set threshold, adjust the convolution kernel size and learning rate parameters. According to the characteristics of different types of defects and real-time images, adjust the weights of the attention mechanism.
[0016] Data augmentation step: During the model training process, perform data augmentation on the training dataset, using random flipping, rotation, and scaling operations. Random flipping can be horizontal flipping or vertical flipping. The rotation operation randomly rotates the image within an angular range to simulate industrial products under different perspectives. The scaling operation enlarges or reduces the image to change the size and proportion of the image. Data transformation is performed according to the formula I' = T(I), where I is the original image, I′ is the transformed image, and T is the transformation operation.
[0017] 8. An image segmentation system based on feature fusion and attention in the industrial defect detection described above, comprising:
[0018] Data preprocessing module: Perform grayscale conversion, normalization, and filtering on the original images of industrial products. Grayscale conversion uses the weighted average method with the formula Gray = 0.299R + 0.587G + 0.114B, and the calculation of this formula is implemented through a hardware acceleration circuit or software algorithm. The normalization operation scales the pixel values according to the formula and is processed using a normalization circuit or software function. Filtering uses Gaussian filtering, and the Gaussian filter kernel size is selected according to the image situation. The image is filtered through convolution operations, and parallel computing technology is used to improve the filtering speed.
[0019] Multi-scale feature extraction module: Use a convolutional neural network (CNN) to extract features from the preprocessed images. In hardware implementation, GPU or FPGA is used to accelerate the convolution operation.
[0020] Feature fusion module: Fuse the feature maps of different scales obtained by the feature extraction module, using the methods of feature splicing and weighted summation. First, splice the feature maps in the channel dimension, and then calculate the fused feature map F according to the formula ; The module learns and normalizes the weights through a fully connected layer and a softmax function. fused
[0021] Attention mechanism module: Apply the attention mechanism to the fused feature map, adopting a combination of channel attention and spatial attention; for the channel attention part, calculate the channel weights through global average pooling and fully connected layers, and for the spatial attention part, calculate the spatial weights through convolution operations; calculate the final attention map A final = A channel × A space to highlight the features of the defect area; final
[0022] Image segmentation module: Use a segmentation network to segment the feature map processed by the attention mechanism and output the segmentation result of the defect area; the module runs on the GPU and completes the segmentation task through parallel computing capabilities, outputting a segmentation result with the same size as the original image.
[0023] Furthermore, it also includes the following modules:
[0024] Post-processing module: Post-process the segmentation result, adopting morphological operations and connected component analysis. The morphological operations are implemented through a hardware morphological processing unit or a software algorithm, and the size of the structuring element for dilation and erosion operations is adjusted according to the size and shape of the defects; the connected component analysis marks each connected component and determines whether it is a noise area according to the area threshold S th of the connected component, and remove the area smaller than S th ; for adjacent defect areas, if the distance between them is less than the set threshold d th , merge them into one defect area;
[0025] Model training module: Use the labeled industrial defect image dataset to train the entire segmentation model, adopting the cross-entropy loss function and the Adam optimizer; the module divides the dataset into a training set and a validation set, and continuously adjusts the model parameters to minimize the loss function Use distributed training technology to accelerate the training process, and at the same time adopt an early stopping strategy to prevent overfitting.
[0026] Performance evaluation module: Use the test dataset to evaluate the performance of the trained model, adopting accuracy, recall, and F1-score metrics; this module calculates and evaluates the segmentation effect of the model according to the formula
[0027] ;
[0028] Real-time detection module: Deploy the trained model to the industrial production line to perform defect detection and segmentation on the industrially produced product images collected in real time; the module obtains images in real time through an image acquisition device, uses the GPU to accelerate the inference process of the model, and quickly outputs the segmentation result;
[0029] Dynamic adjustment module: Dynamically adjust the parameters of the model and the weights of the attention mechanism according to the real-time detection results and performance evaluation metrics; when the accuracy rate is lower than the set threshold, the module adjusts the convolution kernel size and learning rate parameters, and adjusts the weights of the attention mechanism according to different types of defects and the characteristics of real-time images.
[0030] Data augmentation module: During the model training process, perform data augmentation on the training dataset, using random flipping, rotation, and scaling operations; the module performs data transformation according to the formula I′ = T(I), where I is the original image, I′ is the transformed image, and T is the transformation operation.
[0031] Compared with the existing technologies, the beneficial effects of the present invention are as follows:
[0032] In terms of detection accuracy, through multi-scale feature extraction and fusion, the complementary nature between different-scale features is fully utilized. Small-scale features can capture the detailed information of defects, while large-scale features can provide more global semantic information, enabling the model to have good detection capabilities for defects of different sizes, reducing the situations of missed detection and false detection. At the same time, the introduction of the attention mechanism enables the model to automatically focus on the defect areas in the image, suppressing the interference of background information, and further improving the accuracy and precision of segmentation, and being able to accurately identify tiny defects and irregularly shaped defects under complex backgrounds. In terms of efficiency, the method and system can quickly process the images of industrial products to achieve real-time defect detection. By using GPU-accelerated computing and an optimized algorithm structure, the image segmentation time is greatly shortened, meeting the large-scale and high-efficiency detection requirements in industrial production lines. Compared with traditional manual detection methods, the detection speed has been significantly improved, being able to promptly discover the defects in products, preventing defective products from flowing into subsequent production links, and improving production efficiency and product quality. In terms of adaptability, the data preprocessing step performs grayscale conversion, normalization, and filtering on the images, enhancing the clarity and feature expression of the images, and improving the robustness of the model to factors such as illumination changes and noise interference. At the same time, the dynamic adjustment step can dynamically adjust the parameters of the model and the weights of the attention mechanism according to the real-time detection results and performance evaluation metrics, enabling the model to adapt to different types of industrial products and complex and changeable industrial environments, and having stronger versatility and adaptability. In terms of model training, the data augmentation step increases the diversity of training data, improves the generalization ability of the model, and reduces the risk of overfitting. The post-processing step optimizes the segmentation results, removes small noise regions, and merges adjacent defect regions, further improving the accuracy and integrity of the segmentation results. In summary, the method and system of this patent can provide a high-precision, high-efficiency, and strongly adaptable defect detection solution for industrial production, contribute to improving the quality and production efficiency of industrial products, and have important application value and market prospects. Description of the Drawings
[0033] Figure 1 This is a schematic block diagram of an image segmentation system based on feature fusion and attention in industrial defect detection proposed by the present invention;
[0034] Figure 2 This is a schematic block diagram of an image segmentation method based on feature fusion and attention in industrial defect detection proposed by the present invention. Specific implementation manners
[0035] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0036] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by the terms "center", "longitudinal", "transverse", "length", "width", "thickness", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", "clockwise", "counterclockwise", etc. is based on the orientation or positional relationship shown in the accompanying drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus cannot be understood as a limitation to the present invention.
[0037] In addition, the terms "first" and "second" are only used for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of the described features. In the description of the present invention, "a plurality" means two or more, unless otherwise specifically defined. In addition, the terms "installation", "connection", and "connection" should be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or an integral connection; it may be a mechanical connection or an electrical connection; it may be directly connected, or indirectly connected through an intermediate medium, and it may be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations. The present invention will be further described in detail below with reference to the accompanying drawings.
[0038] Refer to Figure 1-2 : An image segmentation method based on feature fusion and attention in industrial defect detection, including the following steps:
[0039] Data preprocessing steps: Get the original images of industrial products, which may come from high-definition cameras, industrial cameras and other equipment on industrial production lines. Grayscale, normalize and filter the images to remove noise interference. Grayscale uses the weighted average method with the formula Gray = 0.299R + 0.587G + 0.114B to convert color images into grayscale images, reducing the amount of calculation while retaining key information. Normalization scales the pixel values to the [0,1] interval, specifically through the formula Implementation, where p is the original pixel value, p min and p max They are the minimum and maximum values of the image pixel values, respectively, so that the pixel values of different images are in a unified range, which is conducive to subsequent processing. Gaussian filtering is used for filtering. The size of the Gaussian filter kernel can be adjusted according to the resolution and noise of the image. Generally, a 3×3, 5×5 or 7×7 kernel is selected to enhance the clarity and feature expression of the image and suppress high-frequency noise in the image.
[0040] Multi-scale feature extraction step: Use convolutional neural network (CNN) to extract multi-scale features from the preprocessed image, using convolution kernels of different sizes and different numbers of convolution layers to obtain feature maps of different scales, such as 4×4, 8×8, and 16×16. The size of the convolution kernel can be selected based on experience and experimental results. For example, a small convolution kernel (such as 3×3) can capture local detail features of the image, and a large convolution kernel (such as 7×7) can extract more global feature information. Different numbers of convolution layers can construct feature hierarchies of different depths. Shallower convolution layers extract underlying features such as edges and textures of the image, while deeper convolution layers extract more abstract semantic features. This multi-scale feature extraction method provides rich feature information for subsequent feature fusion.
[0041] Feature fusion step: The feature maps of different scales obtained in the multi-scale feature extraction step are fused by feature concatenation and weighted summation. First, the feature maps of different scales are concatenated in the channel dimension to increase the dimension and richness of the features. Then, according to the formula Calculate the fused feature map F fused , where F i is the i-th feature map, w i is the corresponding weight. The weight can be obtained through learning, for example, using a fully connected layer and a softmax function to normalize the weight to make full use of feature information of different scales so that the fused feature map can retain local details and contain global semantic information.
[0042] Steps of applying the attention mechanism: Apply the attention mechanism to the fused feature map, adopting a combination of channel attention and spatial attention. Channel attention compresses the feature map in the spatial dimension through global average pooling to obtain the global feature information of each channel, and then calculates the channel weights through a fully connected layer and an activation function to emphasize important channel features. Spatial attention processes the feature map through convolution operations to calculate spatial weights and highlight important spatial regions in the feature map. According to the formula A final = A channel × A space Calculate the final attention map A final , multiply the results of channel attention and spatial attention to further highlight the features of the defect region and suppress the interference of irrelevant regions.
[0043] Steps of image segmentation: Use a segmentation network to segment the feature map processed by the attention mechanism and output the segmentation results of the defect region. The segmentation network adopts a U-Net structure and realizes image segmentation through an encoder-decoder structure. The encoder part gradually reduces the size of the feature map through convolution and pooling operations to extract high-level semantic features of the image; the decoder part gradually restores the size of the feature map through deconvolution and upsampling operations, and combines skip connections to splice the feature maps of the same size in the encoder to retain detailed information. Finally, the segmentation results of the same size as the original image are output, and each pixel point is classified as a defect region or a normal region.
[0044] In the present invention, the following steps are further included:
[0045] Post-processing steps: Post-process the segmentation results, adopting morphological operations (such as dilation and erosion) and connected component analysis. The dilation operation expands pixels around the defect region to connect adjacent small defect regions; the erosion operation shrinks the defect region to remove some small noises. The size of the structuring element used in the dilation and erosion operations can be adjusted according to the size and shape of the defect. Connected component analysis is to label each connected component in the segmentation results, and judge whether it is a noise region according to the area threshold S th of the connected component, and remove the region with an area smaller than S th . At the same time, for adjacent defect regions, if the distance between them is less than the set threshold d th , they are merged into one defect region to improve the accuracy and integrity of the segmentation results.
[0046] Model training steps: Use a labeled industrial defect image dataset to train the entire segmentation model. The dataset should contain industrial defect images of different types and degrees, and the labels should be accurate and clear. Adopt the cross-entropy loss function and the Adam optimizer, and calculate the loss L according to the formula where yi is the true label, p i is the predicted probability. During the training process, the dataset is divided into a training set and a validation set, usually in a ratio of 8:2. By continuously adjusting the model parameters, such as the weights and biases of the convolutional kernels, the loss function is minimized to improve the segmentation performance of the model. At the same time, an early stopping strategy can be adopted. When the loss on the validation set no longer decreases, the training is stopped to prevent overfitting.
[0047] Performance evaluation steps: Use the test dataset to evaluate the performance of the trained model. The test dataset should be independent of the training dataset and have a similar distribution. Adopt metrics such as accuracy, recall, and F1-score, and calculate according to the formula
[0048] where TP is the true positive (predicted as a defect and actually a defect), TN is the true negative (predicted as normal and actually normal), FP is the false positive (predicted as a defect but actually normal), and FN is the false negative (predicted as normal but actually a defect). Through these metrics, the segmentation effect of the model is comprehensively evaluated. For example, accuracy reflects the overall classification correctness of the model, recall reflects the detection ability of the model for defect regions, and the F1-score comprehensively considers accuracy and recall.
[0049] Real-time detection steps: Deploy the trained model to the industrial production line to perform defect detection and segmentation on the industrially produced product images collected in real time. The images are obtained in real time through an image acquisition device, which should have a high frame rate and high resolution to meet the real-time and accuracy requirements of industrial production. The model is used to quickly output the segmentation results, and GPU acceleration can be used to improve the inference speed of the model. At the same time, the segmentation results are displayed and recorded in real time to facilitate the staff to promptly discover and process defective products, realizing real-time defect detection of industrial products.
[0050] Dynamic adjustment steps: According to the real-time detection results and performance evaluation metrics, dynamically adjust the model parameters and the weights of the attention mechanism. When the accuracy is lower than the set threshold, adjust parameters such as the convolutional kernel size and learning rate. For example, reducing the convolutional kernel size may help capture more subtle defect features; adjusting the learning rate can control the step size of model parameter updates to avoid the model falling into a local optimum. At the same time, according to the characteristics of different types of defects and real-time images, adjust the weights of the attention mechanism to make the model pay more attention to the defect regions that are easily missed or misjudged, optimizing the performance of the model.
[0051] Data augmentation steps: During the model training process, data augmentation is performed on the training dataset, using operations such as random flipping, rotation, and scaling. Random flipping can be horizontal flipping or vertical flipping to increase the diversity of images; the rotation operation can randomly rotate the image within a certain angle range (such as -45° to 45°) to simulate industrial products under different perspectives; the scaling operation can enlarge or reduce the image to change the size and proportion of the image. Data transformation is performed according to the formula I′ = T(I), where I is the original image, I′ is the transformed image, and T is the transformation operation. Through data augmentation, the diversity of data is increased, enabling the model to encounter more images in different forms during the training process, improving the generalization ability of the model, and reducing the risk of overfitting.
[0052] The present invention also discloses an image segmentation system based on feature fusion and attention in industrial defect detection, including:
[0053] Data preprocessing module: Perform grayscale conversion, normalization, and filtering on the original image of the industrial product. Grayscale conversion uses the weighted average method, and the formula is Gray = 0.299R + 0.587G + 0.114B. The calculation of this formula is implemented through a hardware acceleration circuit or a software algorithm to quickly convert the color image into a grayscale image. The normalization operation scales the pixel values according to the formula and is processed using a dedicated normalization circuit or software function. Gaussian filtering is used for filtering. The appropriate Gaussian filter kernel size is selected according to the specific situation of the image, and the image is filtered through convolution operations. Parallel computing technology can be used to improve the filtering speed to enhance the clarity and feature expression of the image and provide high-quality image data for subsequent processing.
[0054] Multi-scale feature extraction module: Use a convolutional neural network (CNN) to perform multi-scale feature extraction on the preprocessed image. This module contains convolutional layers and pooling layers with different sizes of convolutional kernels, and the parameters of the convolutional kernels are obtained through model training. In hardware implementation, GPU or FPGA can be used to accelerate the convolutional operation to improve the efficiency of feature extraction. Through different numbers of convolutional layers and sizes of convolutional kernels, feature maps of different scales are obtained, such as 4×4, 8×8, 16×16 scales, providing rich feature information for feature fusion.
[0055] Feature fusion module: Fuse the feature maps of different scales obtained by the multi-scale feature extraction module, using the methods of feature splicing and weighted summation. First, splice the feature maps in the channel dimension, and then calculate the fused feature map F according to the formula fusedThis module learns and normalizes the weights through a fully connected layer and a softmax function, and this function can be implemented using deep learning frameworks (such as TensorFlow, PyTorch) to make full use of feature information at different scales and improve the feature expression ability.
[0056] Attention mechanism module: Apply the attention mechanism to the fused feature map, adopting a combination of channel attention and spatial attention. For the channel attention part, the channel weights are calculated through global average pooling and a fully connected layer, and for the spatial attention part, the spatial weights are calculated through convolution operations. According to the formula A final = A channel × A space calculate the final attention map A final , highlighting the features of the defect area. This module can use a hardware acceleration unit (such as a GPU) to accelerate the calculation process and improve the calculation efficiency of the attention mechanism.
[0057] Image segmentation module: Use a segmentation network to segment the feature map processed by the attention mechanism and output the segmentation result of the defect area. The segmentation network adopts a U-Net structure and realizes image segmentation through an encoder-decoder structure, and combines skip connections to retain detailed information. This module can run on a GPU, utilize its powerful parallel computing ability to quickly complete the segmentation task, and output a segmentation result with the same size as the original image, providing accurate segmentation information for industrial defect detection.
[0058] In the present invention, the following modules are further included:
[0059] Post-processing module: Post-process the segmentation result, adopting morphological operations (such as dilation, erosion) and connected component analysis. The morphological operations are implemented through a hardware morphological processing unit or a software algorithm, and the size of the structural element for dilation and erosion operations is adjusted according to the size and shape of the defect. Connected component analysis marks each connected component and determines whether it is a noise area according to the area threshold S th of the connected component, and the area less than S th is removed. At the same time, for adjacent defect areas, if the distance between them is less than the set threshold d th , they are merged into one defect area. This module can improve the accuracy and integrity of the segmentation result and reduce the influence of noise and misjudgment.
[0060] Model training module: Use a labeled industrial defect image dataset to train the entire segmentation model, adopting a cross-entropy loss function and an Adam optimizer. This module divides the dataset into a training set and a validation set, and by continuously adjusting model parameters, such as the weights and biases of the convolution kernels, minimizes the loss function Distributed training techniques (such as data parallelism and model parallelism) can be used to accelerate the training process. At the same time, an early stopping strategy is adopted to prevent overfitting and improve the segmentation performance of the model.
[0061] Performance evaluation module: Use the test data set to evaluate the performance of the trained model, and adopt indicators such as accuracy, recall rate, and F1-score. This module calculates according to the formula
[0062] to comprehensively evaluate the segmentation effect of the model. Based on the performance evaluation results, provide a basis for optimizing and adjusting the model, such as judging whether it is necessary to adjust the model structure, parameters, or the weights of the attention mechanism.
[0063] Real-time detection module: Deploy the trained model to the industrial production line to detect and segment defects in real-time collected industrial product images. This module obtains images in real-time through image acquisition devices, uses GPU to accelerate the inference process of the model, and quickly outputs the segmentation results. At the same time, the segmentation results are displayed and recorded in real-time, facilitating the staff to promptly discover and handle defective products, and realizing real-time defect detection of industrial products.
[0064] Dynamic adjustment module: Dynamically adjust the parameters of the model and the weights of the attention mechanism according to the real-time detection results and performance evaluation indicators. When the accuracy is lower than the set threshold, this module adjusts parameters such as the convolution kernel size and learning rate, and at the same time adjusts the weights of the attention mechanism according to the characteristics of different types of defects and real-time images. The adjustment of parameters and weights can be automatically realized through software algorithms to optimize the performance of the model and make it adapt to different industrial production environments and defect types.
[0065] Data augmentation module: During the model training process, perform data augmentation on the training data set, and adopt operations such as random flipping, rotation, and scaling. This module performs data transformation according to the formula I′ = T(I), where I is the original image, I′ is the transformed image, and T is the transformation operation. The data augmentation operation is realized through a hardware acceleration circuit or a software algorithm to increase the diversity of data, improve the generalization ability of the model, and reduce the risk of overfitting.
[0066] The above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, makes equivalent substitutions or changes, and should be covered by the protection scope of the present invention.
Claims
1. An image segmentation method based on feature fusion and attention in industrial defect detection, characterized in that: The following steps are involved: Data preprocessing steps: Get the original image of the industrial product, which comes from the camera and industrial camera equipment on the industrial production line; grayscale, normalize and filter the image to remove noise interference; grayscale uses the weighted average method, the formula is Gray = 0.299R + 0.587G + 0.114B, and the pixel value is scaled to the [0,1] interval after normalization, specifically through the formula Implementation, where p is the original pixel value, p min and p max are the minimum and maximum values of the image pixel values, respectively. Gaussian filtering is used for filtering, and the size of the Gaussian filter kernel is adjusted according to the resolution and noise of the image; Feature extraction step: Use convolutional neural network (CNN) to extract features from the preprocessed image. Use convolution kernels of different sizes and different numbers of convolution layers to obtain feature maps of different scales. The size of the convolution kernel is selected based on experience and experimental results. Different numbers of convolution layers construct feature hierarchies of different depths. Shallow convolution layers extract edge and texture bottom features of the image, while deeper convolution layers extract more abstract semantic features. Feature fusion steps :The feature maps of different scales obtained in the feature extraction step are fused by feature splicing and weighted summation. First, the feature maps of different scales are spliced in the channel dimension. The weights can be obtained through learning and then according to the formula Calculate the fused feature map F fused , where F i is the i-th feature map, w i is the corresponding weight; Attention mechanism application steps: Apply the attention mechanism to the fused feature map, using a combination of channel attention and spatial attention; channel attention compresses the feature map in the spatial dimension through global average pooling to obtain the global feature information of each channel, and then calculates the channel weight through the fully connected layer and activation function. Spatial attention processes the feature map through convolution operations, calculates the spatial weight, and highlights the important spatial areas in the feature map; According to formula A final =A channel ×A space Calculate the final attention map A final , the results of channel attention and spatial attention are multiplied to highlight the characteristics of the defect area; Image segmentation steps: Use the segmentation network to segment the feature map processed by the attention mechanism and output the segmentation result of the defect area; the segmentation network adopts the U-Net structure and realizes image segmentation through the encoder-decoder structure; the encoder part gradually reduces the size of the feature map through convolution and pooling operations to extract high-level semantic features of the image; The decoder part gradually restores the size of the feature map through deconvolution and upsampling operations, and combines jump connections to splice the feature maps of the same size in the encoder, and finally outputs a segmentation result of the same size as the original image, and each pixel is classified as a defect area or a normal area.
2. The image segmentation method based on feature fusion and attention in industrial defect detection according to claim 1 is characterized in that: Also includes: Post-processing steps: The segmentation results are post-processed using morphological operations and connected region analysis. The dilation operation expands pixels around the defect area to connect adjacent small defect areas. The erosion operation shrinks the defect area. The size of the structural element used in the dilation and erosion operations is adjusted according to the size and shape of the defect. The connected region analysis is to mark each connected region in the segmentation result and calculate the area threshold S of the connected region. th Determine whether it is a noise area, the area is smaller than S th For adjacent defective areas, if the distance between them is less than the set threshold d th , merged into one defect area.
3. The image segmentation method based on feature fusion and attention in industrial defect detection according to claim 1 is characterized in that: Also includes: Model training steps: Use the labeled industrial defect image dataset to train the entire segmentation model, using the cross entropy loss function and Adam optimizer, according to the formula Calculate the loss L, where y i is the true label, p i It is the prediction probability. During the training process, the data set is divided into a training set and a validation set, and the model parameters are continuously adjusted.
4. The image segmentation method based on feature fusion and attention in industrial defect detection according to claim 1 is characterized in that: Also includes: Performance evaluation steps: Use the test data set to evaluate the performance of the trained model. The test data set should be independent of the training data set. Use accuracy, recall, and F1-score indicators according to the formula Calculation is performed, where TP is predicted to be a defect and is actually a defect, TN is predicted to be normal and is actually normal, FP is predicted to be a defect but is actually normal, and FN is predicted to be normal but is actually a defect.
5. The image segmentation method based on feature fusion and attention in industrial defect detection according to claim 1 is characterized in that: Also includes: Real-time detection step: The trained model is deployed on the industrial production line to detect and segment defects on the industrial product images collected in real time; Images are acquired in real time through image acquisition devices, and the model outputs segmentation results. GPU is used to accelerate calculations to improve the reasoning speed of the model.
6. The image segmentation method based on feature fusion and attention in industrial defect detection according to claim 1 is characterized in that: Also includes: Dynamic adjustment steps: According to the real-time detection results and performance evaluation indicators, dynamically adjust the model parameters and the weight of the attention mechanism. If the accuracy is lower than the set threshold, adjust the convolution kernel size and learning rate parameters; adjust the weight of the attention mechanism according to the characteristics of different types of defects and real-time images.
7. The image segmentation method based on feature fusion and attention in industrial defect detection according to claim 1 is characterized in that: Also includes: Data enhancement step: During the model training process, the training data set is enhanced by random flipping, rotation, and scaling operations; Random flipping is horizontal flipping or vertical flipping; the rotation operation is to randomly rotate the image within an angle range to simulate industrial products under different viewing angles; the scaling operation enlarges or reduces the image, changing the size and proportion of the image; data transformation is performed according to the formula I′=T(I), where I is the original image, I′ is the transformed image, and T is the transformation operation.
8. An image segmentation system based on feature fusion and attention for realizing industrial defect detection according to any one of claims 1 to 7, characterized in that: include: Data preprocessing module: grayscale, normalize and filter the original image of industrial products. The grayscale is processed by weighted average method, and the formula is Gray = 0.299R + 0.587G + 0.114B. The calculation of this formula is realized by hardware acceleration circuit or software algorithm. The normalization operation is based on the formula Pixel values are scaled and processed using normalization circuits or software functions; Gaussian filtering is used for filtering, and the Gaussian filter kernel size is selected according to the image situation. The image is filtered through convolution operations, and parallel computing technology is used to improve the filtering speed; Multi-scale feature extraction module: Use convolutional neural network (CNN) to extract features from preprocessed images; in hardware implementation, GPU or FPGA is used to accelerate convolution operations; Feature fusion module: The feature maps of different scales obtained by the feature extraction module are fused by feature splicing and weighted summation. First, the feature maps are spliced in the channel dimension, and then according to the formula Calculate the fused feature map F fused ; The module learns and normalizes the weights through the fully connected layer and the softmax function; Attention mechanism module: The attention mechanism is applied to the fused feature map, combining channel attention and spatial attention. The channel attention part calculates the channel weight through global average pooling and fully connected layers, and the spatial attention part calculates the spatial weight through convolution operations. According to formula A final =A channel ×A space Calculate the final attention map A final , highlight the characteristics of the defective area; Image segmentation module: Use the segmentation network to segment the feature map processed by the attention mechanism and output the segmentation result of the defect area; the module runs on the GPU to complete the segmentation task through parallel computing power and outputs the segmentation result of the same size as the original image.
9. The image segmentation system based on feature fusion and attention in industrial defect detection according to claim 8, characterized in that: Also includes: Post-processing module: post-process the segmentation results using morphological operations and connected region analysis. The morphological operation is implemented by hardware morphological processing unit or software algorithm, and the size of the structural element of the expansion and corrosion operation is adjusted according to the size and shape of the defect; the connected region analysis is performed by marking each connected region and calculating the area threshold S of the connected region. th Determine whether it is a noise area, the area is smaller than S th For adjacent defective areas, if the distance between them is less than the set threshold d th , merged into a defect area; Model training module: Use the labeled industrial defect image dataset to train the entire segmentation model, using the cross entropy loss function and Adam optimizer; the module divides the dataset into a training set and a validation set, and minimizes the loss function by continuously adjusting the model parameters. Distributed training techniques are used to speed up the training process, while early stopping strategies are adopted to prevent overfitting.
10. The image segmentation system based on feature fusion and attention in industrial defect detection according to claim 8, characterized in that: Also includes: Performance evaluation module: Use the test data set to evaluate the performance of the trained model, using accuracy, recall, and F1-score indicators; This module is based on the formula Perform calculations to evaluate the segmentation effect of the model; Real-time detection module: The trained model is deployed on the industrial production line to detect and segment defects through real-time collected industrial product images. The module acquires images in real time through image acquisition equipment, uses GPU to accelerate the model's reasoning process, and quickly outputs segmentation results. Dynamic adjustment module: dynamically adjusts the model parameters and the weights of the attention mechanism based on real-time detection results and performance evaluation indicators; If the accuracy is lower than the set threshold, the module adjusts the convolution kernel size and learning rate parameters, and adjusts the weight of the attention mechanism according to the characteristics of different types of defects and real-time images; Data enhancement module: During the model training process, the training data set is enhanced by random flipping, rotation, and scaling operations; the module transforms the data according to the formula I′=T(I), where I is the original image, I′ is the transformed image, and T is the transformation operation.
Citation Information
Cited By
Industrial machine tool circuit wiring detection method and equipment suitable for industrial machine tool circuit wiring detection method
CN120823609A
Industrial image detection method based on DINOv3
CN121414723A