Lightweight HAN-YOLO defect detection method based on machine vision

By improving the YOLOv8 model, designing lightweight adaptive feature bottlenecks and introducing a variety of convolution and attention mechanisms, the problem of high computing resources of traditional detection methods is solved, and efficient and accurate defect detection is achieved, which is suitable for equipment with limited resources.

CN120236136APending Publication Date: 2025-07-01CHINA THREE GORGES UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510351607.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

The traditional arc additive manufacturing melt layer defect detection method is difficult to deploy on terminal equipment with limited resources due to high demand for computing resources and storage space, resulting in low detection efficiency and accuracy.

Method used

By improving the YOLOv8 model, the lightweight adaptive feature bottleneck (NLAIB) is designed and the introduction of AKConv convolution, HS-FPN feature fusion pyramid network, ConvTranspose2d transposed convolution, Conv2d two-dimensional convolution and CA lightweight attention mechanisms are optimized to optimize the lightweight performance and defect detection performance of the model.

Benefits of technology

It reduces the number of parameters and calculation amount of the model, improves detection efficiency and accuracy, and is suitable for deployment in embedded equipment with limited resources, and meets the needs of arc additive manufacturing defect detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120236136A_ABST
    Figure CN120236136A_ABST
Patent Text Reader

Abstract

The invention provides a machine vision-based lightweight HAN-YOLO defect detection method, which belongs to the field of machine vision target detection, and comprises the following steps of: adding lightweight changeable kernel convolution into a backbone network Backbone to solve the problem of fixed sampling shape of common convolution, and dynamically adjusting input characteristics to improve the detection performance of the network. Secondly, a feature fusion pyramid network is added into the Neck neck network, the model is optimized, and the feature selection and fusion capability of the model is enhanced; and finally, a novel lightweight adaptive bottleneck module NLAIB is designed to reduce the parameter quantity and complexity of the model, improve the reasoning speed and detection efficiency of the model, and realize relatively high precision detection. According to the method, the detection performance and the hardware adaptability of the model are remarkably improved, and the method can be deployed on terminal equipment with limited computing resources and storage space. According to the method, the defects can be identified and detected in a more efficient and lightweight manner.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image recognition and processing, relates to defect detection technology, and particularly relates to a defect detection method based on machine vision lightweight HAN-YOLO. Background Art

[0002] Wire and arc additive manufacturing (WAAM) is a wire-fed directed energy deposition technology. An electric arc is used as a heat source to melt metal materials and stack them layer by layer to form parts. Compared with other technologies, WAAM has high material utilization rate, low manufacturing cost, reasonable accuracy and higher efficiency due to its characteristics. In addition, WAAM can also produce large and medium-sized, high-strength, and complex parts. However, due to the need for a large amount of energy input and metal raw materials, bottleneck problems such as cracks, pores, and deformation frequently occur, seriously affecting the surface quality such as surface tolerance and forming accuracy of WAAM parts. Therefore, defect detection plays a key role in ensuring the surface quality of parts.

[0003] The methods for detecting defects in deposited layers mainly include visual inspection and machine vision inspection. Due to the fact that the quantity, size, and position distribution of surface defects on wire and arc additive manufacturing parts have no significant rules and are affected by manufacturing processes, with high hardware requirements, complex model structures, large parameter quantities and large amounts of calculations, the detection efficiency and detection accuracy are relatively low. As a result, traditional detection methods require a large amount of storage space and computing resources and are difficult to be deployed to resource-constrained and computationally-limited terminal devices. Therefore, they cannot adapt to the actual detection of surface defects in the deposited layers of WAMM.

[0004] With the rapid development of artificial intelligence and the increasing maturity of machine vision, convolutional neural networks have better performance in object detection, such as networks like R-CNN and YOLO. The continuous development of machine vision technology provides new ideas and methods for detecting defects in deposited layers. The object detection method based on machine vision can automatically select and extract the features of defects according to surface defect pictures and has advantages such as good robustness. These advantages improve the performance of defect detection in the deposited layers of wire and arc additive manufacturing and provide a solution with low cost, high efficiency, and detection accuracy for defect detection in wire and arc additive manufacturing. Summary of the Invention

[0005] To solve the above problems, the present invention provides a defect detection method based on machine vision lightweight HAN-YOLO, which optimizes and improves the original YOLOv8 model. When used for detecting defects in wire and arc additive manufacturing, it can improve its lightweight performance. At the same time, compared with the original YOLOv8 model, the number of parameters and the amount of calculation of the model are significantly reduced, and the accuracy of the model reaches 79.7%, which is suitable for deployment to resource-limited embedded devices and can effectively solve the above problems.

[0006] To achieve the above technical features, the object of the present invention is achieved as follows: A defect detection method for lightweight HAN-YOLO based on machine vision, comprising the following steps: Step 1, obtain pictures of surface defects of the fused deposition layer, expand the fused deposition layer defect dataset through preprocessing and enhancement methods, annotate the defect samples and convert them into label files in corresponding formats, and divide them into training, validation, and test experimental datasets according to the corresponding ratio; Step 2, construct an improved YOLOv8 defect detection model, which includes AKConv variable kernel convolution, HS-FPN feature fusion pyramid network, ConvTranspose2d transposed convolution, Conv2d two-dimensional convolution, and CA lightweight attention mechanism embedded. A lightweight adaptive bottleneck module NLAIB is adopted in the backbone and neck networks to enhance the detection and recognition performance of the model for fused deposition layer defects; Step 3, use the fused deposition layer defect training set to extract features and make predictions for the improved YOLOv8 network model, and use the divided validation set to evaluate the performance of the improved YOLOv8 detection model to obtain the optimal lightweight defect detection model; Step 4, input the fused deposition layer defect test set into the optimal lightweight defect detection model, stop detection when the set epoch is reached, and output the defect test results.

[0007] Preferably, the specific operation steps for establishing the fused deposition layer defect dataset in Step 1 are as follows: Use the roboflow public dataset and combine it with the laboratory arc additive manufacturing equipment to make defect image samples, and use preprocessing and data augmentation to expand the fused deposition layer defect dataset. Use the LabelImg annotation software to accurately annotate the defect picture samples, and classify the defect categories into surface pores, surface crack, spatter, overlap, workpiece and weld; and save them in the dataset and label formats of YOLO. Then, randomly divide the fused deposition layer defect dataset into a training set, a test set, and a validation set according to a ratio of 7:2:1; the training set is used for the performance training of the model, the test set is used for the performance testing of the model, and the validation set is used for the performance verification of the model; At the same time, methods such as gray-scale transformation and contrast enhancement are used to expand the experimental dataset.

[0008] Preferably, in the step 2, the specific operation steps of the improved YOLOv8 defect detection model include: adopting AKConv lightweight convolution in the YOLOv8 network structure; introducing HS-FPN screening feature fusion pyramid network to use selective feature fusion to optimize the model performance and structure; introducing ConvTranspose2d transposed convolution, Conv2d two-dimensional convolution, and CA lightweight channel attention mechanism to optimize the feature extraction and expression of the neck network; in the Backbone backbone network and Neck neck network parts, an NLAIB adaptive feature bottleneck is obtained by combining the IB bottleneck of MobileNetV2, the variant convolution of ConvNext, the ExtraDW extended depthwise separable convolution, the FFN feed-forward neural network, and the elements of the ViT vision transform block, and the original C2f layer is replaced.

[0009] Preferably, in the step 2, the specific process of the neural network model for improving YOLOv8 is as follows: replacing the original ordinary Conv convolution with AKConv convolution; replacing the Conv convolution of the original neck network with Conv2d; adding a ConvTranspose2d transposed convolution module to the neck network to optimize the original upsampling; embedding a CA attention mechanism to optimize the model's understanding and extraction of defect features; replacing the original C2f module with an NLAIB feature bottleneck in the backbone and neck networks.

[0010] Preferably, the neck network is composed of a CA module, an HS-FPN module, a Conv2d module, a ConvTranspose2d module, and an NLAIB module.

[0011] Preferably, the CA coordinate attention mechanism combines the channel correlation and spatial information of the input feature map in the neck network; The Conv2d two-dimensional convolution has parameter sharing, translational invariance, and can integrate features of different scales in object detection, so as to extract richer effective features.

[0012] Preferably, the ConvTranspose2d transposed convolution layer improves the Neck part on the basis of the original neck network to enhance the resolution of the feature map.

[0013] Preferably, the NLAIB feature bottleneck module is mainly composed of a Depthwise Conv initial optional depthwise separable convolution module, a Conv channel expansion convolution module, a Depthwise Conv intermediate optional depthwise separable convolution module, a Conv channel compression convolution module, etc.

[0014] Preferably, in step 3, feature extraction and training are performed on the improved YOLOv8 model. The specific steps include: inputting the fused deposition layer defect training set and validation set into the improved YOLOv8 defect detection model. The experimental environment configuration includes: Intel Core (TM) i5-12490F processor, NVIDIA GeForce RTX 3060 Graphics card, Windows 10 computer operating system, Pycharm software development platform, Anaconda2023 development environment tool, Python 3.8 programming language, CUDA 11.7 GPU computing platform; Training hyperparameter settings: batch size is set to 64, number of iterations is set to 200, image data size is set to 640 × 640, initial learning rate is 0.0001, momentum parameter is 0.937, and weight decay coefficient is 0.0005.

[0015] Preferably, in step 4, the fused deposition layer defect dataset test set is loaded into the improved YOLOv8 defect detection model for performance testing, and the target detection results are output.

[0016] The present invention has the following beneficial effects: 1. By improving the YOLOv8 model, the present invention first designs a lightweight adaptive feature bottleneck (NLAIB) to improve the lightweight performance of the network and defect detection performance, and optimize the inference speed and computational efficiency of the network. Secondly, AKConv convolution and Conv2d convolution are used to replace Conv convolution to improve the detection accuracy of the model, while reducing the number of model parameters and model complexity. In order to improve the feature upsampling of the YOLOv8 model, ConvTranspose2d convolution is embedded in the Neck network part for feature extraction. Then, the HS-FPN feature fusion pyramid network and CA attention mechanism are used to achieve multi-scale object detection, enhancing the feature selection and fusion ability of the model.

[0017] 2. The present invention reduces the model parameters, memory occupancy and computational amount of YOLOv8, reduces the model size and structural complexity, and optimizes the detection efficiency and detection accuracy of the model.

[0018] 3. The present invention can realize real-time detection of defects in the fused deposition layer of arc additive manufacturing, has better comprehensive performance, can effectively identify and detect the specific categories of defects, is suitable for deployment on terminal devices with limited computing resources and space, and meets the requirements of arc additive manufacturing defect detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The present invention will be further described below with reference to the drawings and embodiments.

[0020] Figure 1This is the network structure diagram of HAN-YOLO in the embodiments of the present invention.

[0021] Figure 2 This is the network structure diagram of NLAIB in the embodiments of the present invention.

[0022] Figure 3 This is the schematic diagram of dataset augmentation of HAN-YOLO in the embodiments of the present invention.

[0023] Figure 4 This is the F1 schematic diagram of network training of HAN-YOLO in the embodiments of the present invention.

[0024] Figure 5 This is the detection effect diagram of HAN-YOLO in the embodiments of the present invention. Detailed implementation manners

[0025] The following will detail the technical solutions in the embodiments of the present invention. The described embodiments are only a part of the embodiments of the present invention, rather than all of them. Without departing from the design concept of the present invention, various improvements made by those of ordinary skill in the art to the technical solutions of the present invention shall fall within the protection scope of the present invention.

[0026] Embodiment 1: Step 1: Obtain pictures of the surface defects of the fused deposition layer, expand the fused deposition layer defect dataset through preprocessing and enhancement methods, label the defect samples and convert them into label files in the corresponding format, and divide them into training, validation, and test experimental datasets according to the corresponding ratio; Step 2: Construct an improved YOLOv8 defect detection model, which includes AKConv variable kernel convolution, HS-FPN feature fusion pyramid network, ConvTranspose2d transposed convolution, Conv2d two-dimensional convolution, and CA lightweight attention mechanism embedded. The lightweight adaptive bottleneck module NLAIB is adopted in the backbone and neck networks to enhance the detection and recognition performance of the model for the fused deposition layer defects; Step 3: Use the fused deposition layer defect training set to extract features and make predictions for the improved YOLOv8 network model, and use the divided validation set to evaluate the performance of the improved YOLOv8 detection model to obtain the optimal lightweight defect detection model; Step 4: Input the fused deposition layer defect test set into the optimal lightweight defect detection model, stop detection and output the defect test results when the set epoch is reached.

[0027] Embodiment 2: As Figure 1 shown, a defect detection method based on machine vision for lightweight HAN-YOLO includes the following steps: Step 1: Use the Roboflow open-source dataset and the laboratory arc additive manufacturing equipment, and use a high-speed industrial camera to capture images of the surface defects of the deposited layer. Expand the deposited layer defect dataset through preprocessing and data augmentation, and use the LabelImg professional annotation software to annotate 6 different types of defects. Divide the deposited layer defect experimental dataset into a training set, a test set, and a validation set according to a ratio; Step 2: Build an improved YOLOv8 deposited layer defect detection model. This model includes an embedded CA attention mechanism, an HS-FPN feature fusion pyramid network, a ConvTranspose2d transposed convolution module, and a Conv2d two-dimensional convolution. The AKConv convolution is used in the backbone network to replace the original ordinary Conv convolution, and the lightweight adaptive feature bottleneck NLAIB is used in the head and neck networks to enhance the model's recognition and detection performance for defects; Among them, the CA attention mechanism, also known as the coordinate attention mechanism, includes parts such as X Avg Pool global average pooling, Y AvgPool global average pooling, Concat feature concatenation layer, Conv2d two-dimensional convolution, BatchNorm batch normalization layer, Non-linear non-linear activation layer, Split feature segmentation module, Sigmoid activation function, Re-weight re-weighting layer, etc. First, it generates feature maps with shapes of C×1×W and C×H×1 by using X-direction average pooling (X Avg Pool) and Y-direction average pooling (Y Avg Pool), retaining the spatial information in different directions. Secondly, the Concat + Conv2 module is used to splice and fuse the feature maps obtained by average pooling to obtain a feature map with a shape of C / r×1×(W + H), where C is the number of feature channels, W is the feature width, H is the feature height, and r is a reduction ratio used to reduce the number of channels and lower the computational complexity. Then, the BatchNorm + Non-linear module is used to perform batch normalization (BatchNorm) on the convolved features and apply a non-linear activation function to enhance the feature representation. In combination with the split + Conv2d module, the feature map is segmented and the channels are restored. Finally, the Sigmoid activation function and Re-weight are used to obtain the feature weight map and weight the output again to obtain the final effective features. It can effectively learn the weights at different positions and enhance the features of the target regions of interest to improve the feature representation and detection performance of the model. The HS-FPN feature fusion pyramid network mainly includes two parts. The first part is the feature selection module, and the second part is the feature fusion module. First, the feature selection module is used to perform size matching and efficient selection on the input feature maps of different scales. Subsequently, the high-level and low-level information contained in these feature maps will be integrated through the feature selection process. In this selection process, the feature maps will undergo two pooling operations: global average pooling and global max pooling. The features obtained through these poolings are then combined to eliminate redundant data and minimize information loss. And the sigmoid activation function is used to determine the weight value of each channel. Finally, the weight of each channel is achieved. Multiplying the weight information by the feature maps corresponding to each scale can generate the output feature map. In the process of feature fusion, the upsampled high-level features and low-level features are combined together by pixel summation to enhance the semantic information of each layer. This module adopts a feature fusion method of Semantic Feature Filtering, using the high-level features as weight coefficients to selectively filter the important semantic information in the low-level features.Upsample or downsample the high-level features through bilinear interpolation to unify features at different levels. Apply the channel attention mechanism to transform the high-level features into corresponding attention weights, and then filter the low-level features to ensure that the dimensions of the features remain consistent. Finally, the filtered low-level features are fused with the high-level features, thereby enhancing the feature representation of the model and improving the model detection performance. The Conv2d two-dimensional convolution module is mainly composed of an input feature map, a convolution kernel, padding, an output feature map, an activation function, and a bias term. By sliding the convolution kernel on the input feature map and interacting with the local area of the input feature map, it can capture target features under the condition of feature map position change to achieve sparse interaction and translational invariance. The ConvTranspose2d transposed convolution module is mainly composed of an input feature map, a convolution kernel, output padding, an activation function, and a bias term, and can complete upsampling while performing feature learning and fusion.

[0028] The adaptive feature bottleneck NLAIB is the bottleneck part of the model. The structural principle of the YOLOv8 model consists of main components such as the input layer (Input), initial convolution (Initial Conv), split module 1 (Split), split module 2 (Split), feature map concatenation module (Concatenate), final convolution (Final Conv), output layer (Output), optional intermediate depthwise convolution (Depthwise Conv), convolution for channel expansion (Conv_expand), convolution for channel compression (Conv_projection), etc. First, the input feature map passes through an initial convolution layer (Initial Conv) to extract basic features, and then branches into two parts. One part undergoes the first split operation (Split 1) to split the feature map after the initial convolution into multiple parts for fusion on the feature map concatenation module (Concatenate). The feature map can be split by channels or by spatial dimensions, which enhances the model's feature extraction ability and flexibility and improves the model's computational efficiency. Secondly, one part undergoes the second split operation (Split2) to divide the feature map for parallel processing. Applying different convolutions on different branches can enhance the model's ability to express features and better improve the model's performance. Thirdly, multiple feature extraction and enhancement operations will be performed according to the actual detection needs, decomposing the feature detection task into multiple subtasks. Through the initial Depthwise Conv (start, optional) layer, spatial features are further extracted, and a non-linear activation function is used to learn effective model features. And Conv(expand) is used to increase the number of model channels, and then combined with the intermediate DepthwiseConv(middle, optional) to extract richer spatial features. At this time, the number of channels is compressed in Conv (project) to achieve the non-linear transformation of the model and feature expression between different channels; then the input features are fused on the feature map concatenation module (Concatenate).Fourth, the feature maps with reduced number of channels sequentially enter the feature extraction and enhancement loop of the Depthwise Conv (start, optional) layer, Conv(expand) layer, Depthwise Conv (middle, optional) layer, and Conv (projection) layer, and are fused on the Concatenate module each time after the output from the Conv (projection) layer. The number of Block loops is adjusted according to different object detection tasks and actual requirements, and dynamic convolution kernel sizes and numbers of channels are adopted to adapt to multi-scale feature extraction. An additional depth convolution layer is introduced, which can expand the network depth and receptive field of feature information of the detection model at a low computational cost, thereby achieving the lightweight of the network and efficient detection performance. Finally, the Final Conv layer further adjusts the number of channels and feature representation of the concatenated feature maps and outputs the final detection features to ensure effective feature extraction and enhancement. The IB bottleneck includes a channel expansion layer, a channel compression layer, a linear bottleneck, and a residual connection. It uses a low-dimensional feature map as input, expands it to a high-dimensional one, and then returns to the low-dimensional one using depth convolution and linear convolution for feature extraction and model calculation. The ConvNext variant convolution mainly consists of depth convolution, layer normalization, 1x1 convolution, GELU activation function, and residual connection, etc. It can perform spatial mixing with deeper convolutions of larger sizes and use more efficient non-linear activation functions to enhance the model's learning of complex features, effectively capturing global features. At the same time, the convolution kernel size and number of channels are adjusted at different stages to meet the requirements of multi-scale feature extraction. The ExtraDW extended depthwise separable convolution mainly includes depth convolution, which can expand the network depth and receptive field of feature information of the model at a low computational cost, thereby maximizing the computational utilization rate. The FFN feed-forward neural network includes an input layer, a hidden layer, and an output layer, which is used for channel mixing, enhancing the model's non-linear transformation and feature representation ability between different channels, and helping to improve the computational efficiency of the model. The ViT vision transformer block mainly consists of structures such as multi-layer perceptron, layer normalization, and residual connection, etc. It can effectively capture the global relationships in the image and improve the model's performance in image recognition tasks. The Depthwise Conv convolution includes a per-channel convolution kernel, a batch normalization layer, an activation function, etc., which are used to perform convolution operations on each channel of the input feature map respectively, complete the extraction of the spatial information of the feature map, reduce the number of model parameters and computational amount to improve the efficiency of the model.The AKConv convolution mainly includes structures such as coordinate offset, feature resampling, depth convolution, and batch normalization. It generates offsets for adjusting the sampling shape through p_conv, combines the initial sampling position and the offsets to obtain new sampling coordinates, and according to the sampling coordinates, bilinearly interpolates to obtain feature values from the input feature map. Then, it resamples the sampled features and applies depth convolution operations, and further processes the convolution results through batch normalization and activation functions to obtain the final output feature map. The AKConv convolution provides a richer selection of convolution kernels, uses efficient convolution kernels with arbitrary numbers of parameters and convolution shapes to extract richer features, reduces the number of model parameters and computational load, makes up for the deficiencies of conventional convolutions, can adapt to defective targets of different shapes, and at the same time supports linearly increasing or decreasing convolution parameters to lightweight the model.

[0029] Step 3: Use the defective training set to extract features and make predictions on the improved YOLOv8 network model, and use the divided validation set to evaluate the performance of the improved YOLOv8 detection model during training to obtain the optimal lightweight defective detection model. After improving the model using the above three main methods, not only the number of model parameters and the structure are optimized, the model detection efficiency is improved, but also the detection accuracy is relatively high, thus improving the overall detection performance of the model. Step 4: Input the fused deposition layer defect test set into the optimal lightweight defective detection model, stop detection and output the defect test results when the set epoch is reached.

[0030] As a preferred embodiment of the present invention, in Step 2, the specific process of improving the YOLOv8 network model is as follows: In the YOLOv8 network, use AKConv convolution to replace the original ordinary Conv convolution to adapt to defective targets of different shapes to extract richer features; introduce Conv2d to replace the Conv convolution in the original neck network to optimize feature extraction. By combining elements of the IB bottleneck of MobileNetV2, the ConvNext variant convolution, the ExtraDW extended depthwise separable convolution, the FFN feedforward neural network, and the ViT vision transformer block, obtain the NLAIB adaptive feature bottleneck, replace the original C2f module to improve the model's feature extraction and flexibility, reduce the number of model parameters and complexity, and improve the hardware adaptability of the model, making the model have better performance. In the Neck part of the neck network, optimize the feature selection and fusion of the model neck by integrating the CA attention mechanism, the HS-FPN feature pyramid network, the ConvTranspose2d transposed convolution, the Conv2d convolution, and the NLAIB adaptive feature bottleneck module, and optimize the structural complexity and detection performance of the model.

[0031] In Step 2: The HAN-YOLOv8 network optimizes the detection performance and lightweight structure of the model through the AKConv module, NLAIB module, CA module, Conv2d module, ConvTranspose2d module, and HS-FPN module, and improves the detection efficiency and accuracy of the model.

[0032] It should be noted that this model uses AKConv convolution to replace the original Conv convolution in the backbone network, providing convolutional kernels of arbitrary sampling shapes and sizes, and using efficient convolutional kernels with arbitrary numbers of parameters and convolutional shapes to extract richer features. The Conv2d two-dimensional convolution replaces the ordinary Conv convolution in the neck network. By sliding the convolutional kernel on the input data to extract effective features and achieve parameter sharing, it can extract rich features and enhance feature learning. In the neck network, ConvTranspose2d transposed convolution is used to generate a larger-sized output feature map by expanding the input feature map and applying a transposed convolutional kernel, enhancing the upsampling operation of the model.

[0033] Furthermore, in the improved neck network, the HS-FPN feature fusion pyramid network is introduced. It uses a feature selection and feature fusion structure to achieve feature enhancement and fusion, screens features through channel attention and dimension matching, and uses selective feature fusion to enhance the model's performance and achieve multi-scale object detection. The CA attention processes the input features through average pooling in the X and Y directions to reduce the loss of effective information in the feature space; at the same time, it obtains the fused features on the Concat layer and Conv2d convolution, and flexibly changes the number of channels to re-weight the features. In addition, the CA attention mechanism embeds position information into channel attention, enabling the model to capture the dependencies between features more comprehensively and improving the performance and expressive ability of the model.

[0034] Furthermore, compared with the original Bottleneck block, the NLAIB bottleneck network integrates elements of the IB bottleneck, ConvNext variant convolution, ExtraDW extended depthwise separable convolution, FFN feed-forward neural network, and ViT layer, including parts such as the Depthwise Conv (start, optional) convolution layer, Conv(expand) convolution layer, Depthwise Conv (middle, optional) layer convolution, Conv (projection) convolution layer, etc. Specifically, the input feature map passes through the Initial Conv module, and then the input feature fusion is achieved through splicing by the Split module and the Concatenate module. And the Depthwise Conv (start, optional) module, Conv(expand) module, Depthwise Conv (middle, optional) module, and Conv (projection) module are used to perform multiple feature extraction and enhancement operations, decomposing the feature detection task into multiple subtasks, adjusting the Block loop times according to different object detection tasks and actual requirements. Finally, the Concatenate module and the Final Conv layer are used to further adjust the number of channels and feature representation of the spliced feature map and output the final detection features to ensure effective extraction, fusion, and enhancement of features, capable of adapting to different optimization objectives, thereby achieving the lightweight and efficient detection performance of the network.

[0035] Through the above operations, the improved NLAIB lightweight adaptive feature bottleneck can enhance the model's ability to extract and enhance weld defect features, reduce the model parameters and model complexity, improve the model's computational efficiency and detection accuracy, and be applicable to terminal devices with limited memory space and computing resources, improving the model's hardware adaptability and lightweight characteristics.

[0036] In summary, the introduction of the AKConv module, NLAIB module, CA module, Conv2d module, ConvTranspose2d module, and HS-FPN module optimizes the model's detection performance and lightweight network structure, reduces the number of model parameters and computational volume, and improves the model's detection efficiency and detection accuracy.

[0037] In step 4, the designed HAN-YOLOv8 neural network model is trained using the PyTorch network object detection framework. The set hyperparameters are: the number of iterations is 200, the batch size is 64, the model learning rate is 0.0001, the momentum parameter is 0.937, the weight decay coefficient is 0.0005, and the learning rate is adjusted using linear decay.

[0038] Performance analysis and comparison were carried out for different improved networks of the present invention, as shown in Table 1: Table 1 Performance Comparison of Different Modules

[0039] As can be seen from Table 1, the P, R, F1, and mAP@0.5:.95 of the YOLOv8n+AHN network structure of the present invention are 78.0%, 77.5%, and 77.7% respectively, while the number of model parameters Params and the computational GFLOPs of the model are 1.685M and 5.0G respectively. The comprehensive performance is significantly better than other model networks, and a high detection accuracy is maintained. Therefore, the detection performance of the HAN-YOLO network of the present invention is better.

[0040] Analysis and comparison were carried out for the trained model and other detection models, as shown in Table 2: Table 2 Comparison of Different Models

[0041] As can be seen from Table 2, the HAN-YOLO model of the present invention, with the number of model parameters of 1.685M and the computational cost of 5G GFLOPs, achieved an mAP@0.5 of 51.1%. In the comparison of several current mainstream algorithms, its overall performance in terms of detection accuracy, detection speed, and model size is better.

[0042] As Figure 5 shown, the detection results of the present invention in the arc additive manufacturing deposition layer defect dataset are given. It can be seen from the figure that the model can accurately identify 6 types of label type defects, namely surface pores, surface crack, spatter, overlap, workpiece, and weld. The experimental results show that the present invention reduces computational and storage resources, optimizes the structural complexity and computational complexity of the model, improves the detection speed and detection efficiency of the model, and has a high detection accuracy, achieving lightweight defect detection with fewer model parameters and computational amounts.

[0043] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than limiting the present invention; although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that any simple modification, equivalent transformation, and non-substantial modification do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the inventive content of each embodiment of the present application, and should all be included in the protection scope of the present invention.

Claims

1. A defect detection method based on lightweight HAN-YOLO of machine vision, characterized in that: The following steps are involved: Step 1: Obtain images of surface defects of the deposited layer, expand the deposited layer defect dataset through preprocessing and enhancement methods, annotate the defect samples and convert them into label files of corresponding formats, and divide them into training, verification and test experimental datasets according to corresponding proportions; Step 2: Build an improved YOLOv8 defect detection model, which includes embedded AKConv variable kernel convolution, HS-FPN feature fusion pyramid network, ConvTranspose2d transposed convolution, Conv2d two-dimensional convolution, CA lightweight attention mechanism, and adopt lightweight adaptive bottleneck module NLAIB in the trunk and neck network to enhance the model's detection and recognition performance of deposited layer defects; Step 3, using the deposited layer defect training set to perform feature extraction and prediction on the improved YOLOv8 network model, and using the divided validation set to evaluate the performance of the improved YOLOv8 detection model to obtain the optimal lightweight defect detection model; Step 4: Input the deposited layer defect test set into the optimal lightweight defect detection model, stop the detection when the set epoch is reached, and output the defect test results.

2. According to claim 1, a defect detection method based on lightweight HAN-YOLO of machine vision is characterized in that: The specific operation steps of establishing the fused layer defect dataset in step 1 are: using the roboflow public dataset and combining with the laboratory arc additive manufacturing equipment to make defect image samples, and using preprocessing and data enhancement to expand the fused layer defect dataset, and using LabelImg annotation software to accurately annotate the defect image samples, and classify the defect categories into surface pores, surface crack, spatter, overlap, workpiece and weld; and save it as the dataset and label format of YOLO, and then randomly divide the fused layer defect dataset into a training set, a test set and a validation set in a ratio of 7:2:1; the training set is used for performance training of the model, the test set is used for performance testing of the model, and the validation set is used for performance verification of the model; At the same time, grayscale transformation and contrast enhancement methods are used to expand the experimental data set.

3. According to the defect detection method of the lightweight HAN-YOLO based on machine vision in claim 1, it is characterized in that: In the step 2, the specific operation steps of the improved YOLOv8 defect detection model include: using AKConv lightweight convolution in the YOLOv8 network structure; introducing the HS-FPN screening feature fusion pyramid network to use selective feature fusion to optimize the model performance and structure; introducing ConvTranspose2d transposed convolution, Conv2d two-dimensional convolution, and CA lightweight channel attention mechanism to optimize the neck network feature extraction and expression; in the Backbone backbone network and Neck neck network parts, by combining the IB bottleneck of MobileNetV2, ConvNext variant convolution, ExtraDW extended depth separation convolution, FFN feedforward neural network and ViT visual conversion block elements, the NLAIB adaptive feature bottleneck is obtained to replace the original C2f layer.

4. According to claim 3, a defect detection method based on lightweight HAN-YOLO of machine vision is characterized in that: In step 2, the specific process of improving the neural network model of YOLOv8 is: replacing the original ordinary Conv convolution with AKConv convolution; replacing the Conv convolution of the original neck network with Conv2d; adding the ConvTranspose2d transposed convolution module to the neck network to optimize the original upsampling; The CA attention mechanism is embedded to optimize the model's understanding and extraction of defect features; the original C2f module is replaced with the NLAIB feature bottleneck in the backbone and neck networks.

5. According to claim 4, a defect detection method based on lightweight HAN-YOLO of machine vision is characterized in that: The neck network consists of a CA module, a HS-FPN module, a Conv2d module, a ConvTranspose2d module, and a NLAIB module.

6. According to claim 5, a defect detection method based on lightweight HAN-YOLO of machine vision is characterized in that: The CA coordinate attention mechanism combines the correlation between channels and spatial information of the input feature map in the neck network; The Conv2d two-dimensional convolution has parameter sharing, translation invariance and the ability to integrate features of different scales in target detection, thereby extracting richer effective features.

7. The defect detection method based on lightweight HAN-YOLO of machine vision according to claim 6 is characterized in that: The ConvTranspose2d transposed convolution layer improves the Neck part based on the original neck network and enhances the resolution of the feature map.

8. The defect detection method based on lightweight HAN-YOLO of machine vision according to claim 7 is characterized in that: The NLAIB feature bottleneck module is composed of main structures such as Depthwise Conv initial optional depth separable convolution module, Conv channel expansion convolution module, Depthwise Conv intermediate optional depth separable convolution module, and Conv channel compression convolution module.

9. The defect detection method based on lightweight HAN-YOLO of machine vision according to claim 1, characterized in that: The step 3 performs feature extraction and training on the improved YOLOv8 model, and the specific steps include: inputting the deposition layer defect training set and the verification set into the improved YOLOv8 defect detection model, and the experimental environment configuration includes: Intel Core (TM) i5-12490F processor, NVIDIA GeForce RTX 3060 Graphics card, Windows 10 computer operating system, Pycharm software development platform, Anaconda2023 development environment tool, Python3.8 programming language, CUDA11.7GPU computing platform; training hyperparameter settings: batch size is set to 64, number of iterations is set to 200, image data size is set to 640 × 640, initial learning rate is 0.0001, momentum parameter is 0.937, and weight attenuation coefficient is 0.0005.

10. The defect detection method based on lightweight HAN-YOLO of machine vision according to claim 1, characterized in that: In step 4, the test set of the deposited layer defect data set is loaded into the improved YOLOv8 defect detection model for performance testing, and the target detection result is output.

Citation Information

Cited By

  • Visual analysis-based crane crane operation risk detection method and system

    CN121121646A