Aluminum profile surface defect detection method and system based on CDA-YOLOv8s

By improving the YOLOv8 model, combined with CG Block, C2f_DWR and ASFP2 modules, the problem of insufficient recognition accuracy of small and medium-sized targets for surface defect detection of aluminum profiles is solved, efficient and accurate defect detection is achieved, and labor costs are reduced.

CN119991614APending Publication Date: 2025-05-13YAZHOU BAY INNOVATION RESEARCH INSTITUTE HAINAN TROPICAL OCEAN UNIVERSITY +1
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510081265.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-19
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

In the detection of surface defects of aluminum profiles, the traditional method is limited by the subjective judgment of the detector and the high labor cost, and the existing object detection algorithms have shortcomings in accuracy and speed, making it difficult to effectively identify small object defects in the image.

Method used

The surface defect detection method of aluminum profile based on CDA-YOLOv8s is adopted. By improving the YOLOv8 model, feature extraction of the target global context is enhanced, and combined with CG Block, C2f_DWR and ASFP2 modules, the extraction capability and detection accuracy of small target features are improved.

Benefits of technology

It significantly improves the identification accuracy of surface defect detection of aluminum profiles, reduces the rate of missed detection and error detection, improves detection efficiency and reliability, and reduces manual intervention and costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991614A_ABST
    Figure CN119991614A_ABST
Patent Text Reader

Abstract

The invention provides an aluminum profile surface defect detection method and system based on CDA-YOLOv8s, and relates to the technical field of aluminum profile defect detection. The method comprises the following steps: firstly, obtaining an aluminum profile surface image, preprocessing the aluminum profile surface image, constructing a defect data set, updating an initial network structure of YOLOv8s, obtaining an aluminum profile defect detection model based on CDA-YOLOv8s, then training the model by using the defect data set, and finally detecting the aluminum profile surface image by using the trained model to obtain a detection result. According to the method, through the improved CDA-YOLOv8s model, the recognition capability of the small defects on the surface of the aluminum profile is improved, and the problem that the precision is low in the existing aluminum profile defect detection process is solved; according to the method, the utilization of global context information of the target is enhanced, and the capability of extracting small target features in the image is remarkably improved, so that more accurate defect identification is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of aluminum profile defect detection, and in particular to a method and system for detecting surface defects of aluminum profiles based on CDA-YOLOv8s. Background Art

[0002] Aluminum profiles have become the core raw materials for many industrial products and daily necessities due to their good corrosion resistance and easy processing. With the rapid advancement of my country's industrialization, the market demand for aluminum profiles has become increasingly diversified, and the requirements for their quality have also increased. However, during the production process, due to multiple factors such as raw material quality, manufacturing process and casting environment, aluminum profiles often have defects such as scratches, non-conductivity, and coating cracks on the surface. These defects not only damage the appearance of the product and reduce material performance, but may also cause unstable performance and pose a safety hazard. Therefore, implementing efficient and accurate surface defect detection is crucial to ensuring product quality.

[0003] Traditional manual inspection methods are limited by the subjective judgment, fatigue and experience level of the inspector, resulting in low detection accuracy, low efficiency and high labor costs. With the development of computer vision technology, target detection algorithms have provided new solutions for aluminum profile surface defect detection. For example, algorithms such as Faster RCNN, Mask-RCNN, SSD and YOLO have shown great potential in the field of metal surface defect detection. Among them, two-stage algorithms such as Faster RCNN and Mask-RCNN have excellent performance in accuracy, but the processing speed is limited, while one-stage algorithms such as SSD and YOLO take into account both speed and accuracy and are more suitable for real-time detection needs. Due to the high efficiency of the YOLO algorithm, researchers have also made many improvements to the YOLO algorithm to improve its performance in aluminum profile surface defect detection, but it still faces challenges such as unclear defect features, unbalanced samples and insufficient detection accuracy.

[0004] Therefore, in the process of aluminum profile defect detection, there is an urgent need for a method that can enhance the ability to extract small target features in images and effectively improve the recognition accuracy of aluminum profile surface defects, so as to reduce the missed detection and false detection rates and promote the development of this field to a higher level. Summary of the invention

[0005] In view of this, the embodiment of the present application provides an aluminum profile surface defect detection method based on CDA-YOLOv8s, which improves the YOLOv8 model to enhance the feature extraction of the target global context and significantly improves the feature extraction capability of small targets in the image. The embodiment of the present application provides the following technical solutions:

[0006] On the one hand, an embodiment of the present application provides an aluminum profile surface defect detection method based on CDA-YOLOv8s, comprising the following steps:

[0007] Acquire aluminum profile surface image data, preprocess the aluminum profile surface image, and construct a defect data set;

[0008] The initial network structure of YOLOv8 is updated to obtain the aluminum profile defect detection model based on CDA-YOLOv8;

[0009] Training the aluminum profile defect detection model according to the defect data set;

[0010] A detection result is obtained based on the trained aluminum profile defect detection model and the aluminum profile surface image.

[0011] Furthermore, the detection steps of the aluminum profile defect detection model include:

[0012] Use the improved backbone network to perform multi-level feature extraction on the input image to obtain basic features;

[0013] Using the neck network to perform multi-scale feature fusion on the basic features to obtain enhanced features;

[0014] The enhanced features are classified and bounding box regressed using a decoupled detection head to generate results.

[0015] Furthermore, the backbone network uses CGBlock as the downsampling of the model, and the steps are as follows:

[0016] S1. Use f loc and f sur To learn local features and corresponding surrounding context, and output intermediate feature maps;

[0017] S2. Use f joi Obtain joint features from the intermediate feature maps, and perform batch normalization and parameterized ReLU activation;

[0018] S3. Use f glo The global context information is used as a weight vector and applied to the processing of channel joint features to output the final feature map.

[0019] Furthermore, the backbone network and the neck network use the C2f module combined with the DWR module as feature extraction components, the Bottleneck unit in the C2f module is improved using the DWR module in the ultra-real-time semantic segmentation model DWR-Seg, and the improved C2f module is named C2f_DWR.

[0020] Furthermore, the steps of C2f_DWR are as follows:

[0021] S1. Based on the original input features, we introduce deep separation dilated convolutions at different rates to generate multiple feature maps with different receptive fields.

[0022] S2. splicing the feature maps along the channel dimension to generate a fused feature map containing multi-scale information;

[0023] S3. Pass the fused feature map through a 1x1 convolution layer to generate an intermediate feature map;

[0024] S4. Add the original input feature to the intermediate feature map through the DWR module to form a residual connection.

[0025] Furthermore, the neck network integrates an ASFP2 detection layer, and the ASFP2 detection layer includes a scale sequence feature fusion module and a small target detection layer module.

[0026] Furthermore, the scale sequence feature fusion module fuses the underlying features with the high-level features, and uses an upsampling operation to generate a high-resolution feature map; the small target detection layer uses the nearest neighbor interpolation method to align the resolution of the high-resolution feature map, and fuses the shallow feature map containing rich target position information and higher resolution with the deep feature map, and further extracts the fused scale sequence feature information through a 3D convolution operation.

[0027] On the other hand, an embodiment of the present application provides an aluminum profile surface defect detection system based on CDA-YOLOv8s, the system comprising:

[0028] A data acquisition module is used to obtain aluminum profile surface image data, pre-process the aluminum profile surface image, and construct a defect data set;

[0029] The model update module is used to update the initial network structure of YOLOv8 to obtain an aluminum profile defect detection model based on CDA-YOLOv8;

[0030] A model training module, connected to the data acquisition module and the model updating module, and used to train the aluminum profile defect detection model according to the defect data set;

[0031] The defect detection module is connected to the data acquisition module and the model training module, and obtains the detection result according to the trained aluminum profile defect detection model and the aluminum profile surface image.

[0032] Furthermore, the detection steps of the aluminum profile defect detection model in the defect detection module include:

[0033] Use the improved backbone network unit to perform multi-level feature extraction on the input image to obtain basic features;

[0034] Using the neck network unit to perform multi-scale feature fusion on the basic features to obtain enhanced features;

[0035] The enhanced features are classified and bounding box regressed using a decoupled detection head unit to generate results.

[0036] In a third aspect, an embodiment of the present application provides a computer-readable storage medium having machine-executable instructions stored thereon. When the machine-executable instructions are executed by at least one processor, the at least one processor implements the aluminum profile surface defect detection method based on CDA-YOLOv8s according to any of the above items.

[0037] Compared with the prior art, the at least one technical solution adopted in the embodiment of the present application can achieve the following beneficial effects:

[0038] 1. The present invention adopts CG Block to improve the downsampling operation of the YOLOv8 model, and makes the context information saved in the extraction process more complete by supplementing the information loss of various modules in the downsampling process, thereby improving the accuracy of the model in detecting small defect targets of aluminum profiles; the DWR module in the ultra-real-time semantic segmentation model DWR-Seg is used to improve the Bottleneck module in the C2f module, improve the feature representation, and enhance the model's ability to extract high-level detail features of the network. C2f_DWR can expand the receptive field by expanding the convolution without increasing the number of convolution kernel parameters, thereby improving the multi-scale feature extraction capability of aluminum profile defect types, and improving the detection accuracy of small target defect types on the surface of aluminum profiles while maintaining efficient calculation.

[0039] 2. In the inference process, the present invention effectively manages the storage of intermediate feature maps through dynamic memory scheduling strategies, avoiding the problem of excessive memory usage caused by high-resolution feature maps. By integrating the ASFP2 small target detection layer and adopting a lightweight design, the additional calculation and memory overhead are reduced, while the small target detection performance is improved.

[0040] 3. Through the optimized DWR module and ASFP2 layer, CDA-YOLOv8 reduces the dependencies in the convolution operation, so that more convolution operations can be performed in parallel. The decoupled detection head processes the classification task and the regression task through independent branches, reducing the mutual interference between tasks and allowing the two branches to be calculated in parallel. This design fully utilizes the parallel capabilities of multi-core processors or GPUs.

[0041] 4. The present invention introduces the DFL (Distribution Focal Loss) loss function to make the boundary prediction of the target box more accurate. This optimization not only improves the detection accuracy, but also indirectly improves the computational efficiency. The accuracy of boundary prediction reduces the number of redundant frames in the subsequent non-maximum suppression (NMS) operation, thereby reducing the computational overhead in the post-processing stage. Higher prediction accuracy makes the detection results more reliable, reduces the dependence on subsequent steps (such as manual review or automatic repair), and reduces the computational overhead of the overall system. The CDA-YOLOv8 algorithm enhances the network's detection capability for small targets and improves the detection accuracy of aluminum profile surface defects by introducing new modules. Traditional aluminum surface defect detection methods often rely on manual visual inspection or simple machine vision technology, which is easily affected by human factors and lighting conditions. The CDA-YOLOv8 algorithm uses deep learning technology to accurately identify aluminum surface defects and improve the reliability and stability of detection. The CDA-YOLOv8 algorithm helps to reduce manual intervention, reduce labor costs, and improve production efficiency and product quality. Illustrations

[0042] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0043] Figure 1 It is a flow chart of a method for detecting surface defects of aluminum profiles provided by one embodiment of the present invention;

[0044] Figure 2 It is a flow chart of an aluminum profile surface defect detection model provided by an embodiment of the present invention;

[0045] Figure 3 It is a flowchart of the steps of the CG Block module of the aluminum profile surface defect detection model provided by one embodiment of the present invention;

[0046] Figure 4 It is a flow chart of the steps of a C2f_DWR module of an aluminum profile surface defect detection model provided by an embodiment of the present invention;

[0047] Figure 5 It is an architecture diagram of an aluminum profile surface defect detection system provided by an embodiment of the present invention;

[0048] Figure 6 It is a YOLOv8 model network structure diagram provided by an embodiment of the present invention;

[0049] Figure 7It is a network structure diagram of a CDA-YOLOv8 model provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0050] The embodiments of the present invention are described in detail below with reference to the accompanying drawings.

[0051] The following describes the embodiments of the present invention through specific examples, and those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that, in the absence of conflict, the following embodiments and the features in the embodiments can be combined with each other. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work belong to the scope of protection of the present invention. It should be noted that the terms "including" and "having" in the specification, claims and drawings of the present invention and any of their variations are intended to cover non-exclusive inclusions.

[0052] On the one hand, the embodiment of the present application provides a method for detecting surface defects of aluminum profiles based on CDA-YOLOv8s, such as Figure 1 As shown, the following steps are included:

[0053] Acquire aluminum profile surface image data, preprocess the aluminum profile surface image, and construct a defect data set;

[0054] Specifically, in the process of aluminum profile surface defect detection, it is first necessary to obtain high-quality surface image data: use an industrial camera or a high-resolution camera to scan the surface of the aluminum profile and collect surface images. The hardware configuration is an industrial camera, a light source (such as a ring light source, a linear light source), and a mobile platform. Images can be obtained by using a fixed camera + a mobile aluminum profile, or a fixed aluminum profile + a mobile camera. It is necessary to ensure uniform lighting during image acquisition to avoid overexposure or shadow interference. The collected raw image data is usually a high-resolution image, which is convenient for subsequent processing and defect detection.

[0055] In order to improve the detection accuracy and efficiency, the collected image data is preprocessed and the image pixel values ​​are normalized to the range of [0, 1] or [-1, 1] to ensure the uniformity of the model input. The original high-resolution image is scaled to a suitable size for the CDA-YOLOv8 model input (such as 640×640) while retaining the key features of the defect. The training data is augmented to improve the generalization ability of the model. This includes: random rotation, translation, cropping; adjusting brightness and contrast; adding noise to simulate possible interference during the detection process. The preprocessed data is divided into training set, validation set, and test set for model training and evaluation.

[0056] The initial network structure of YOLOv8 is updated to obtain the aluminum profile defect detection model based on CDA-YOLOv8;

[0057] Specifically, the Context Guided Block (CG Block) introduced in the backbone network is used to downsample the defect image data; in the backbone network and the neck network, the C2f module is combined with the DWR module for feature extraction. The DWR module is a three-branch receptive field expansion module; the ASFP2 small target detection layer is integrated in the neck network to improve the detection capability of small target defects; the dynamic label matching strategy and DFL loss are used to optimize boundary prediction.

[0058] According to the defect data set, the aluminum profile defect detection model is trained; according to the trained aluminum profile defect detection model and the aluminum profile surface image, the detection results are obtained. The detected defect box is marked on the original image to display the defect type, location and confidence for manual verification or display. Generate a test report, which contains the following contents: defect type and quantity statistics; defect location coordinates and size information; overall test results (such as whether it is qualified). The test results are transmitted to the subsequent system (such as the production line control system) for automatic sorting or alarm. In addition, according to the test results and actual needs, the system is continuously optimized. If the false alarm or missed alarm rate is high, the model performance can be optimized by marking more data or adjusting the training parameters (such as learning rate, loss weight). In the actual production line, if the detection speed does not meet the requirements, the real-time performance can be improved by adjusting the model lightweight module (such as reducing the network depth or resolution). For different hardware platforms (such as GPU, embedded devices), adjust the model deployment parameters to improve the operation efficiency.

[0059] Furthermore, if Figure 2 As shown, the detection steps of the aluminum profile defect detection model include:

[0060] Use the improved backbone network to perform multi-level feature extraction on the input image to obtain basic features;

[0061] Use the neck network to perform multi-scale feature fusion on basic features to obtain enhanced features;

[0062] The enhanced features are used to perform classification and bounding box regression using a disentangled detection head to generate the results.

[0063] The classification branch processes the predicted defect types, such as scratches, cracks, bubbles, etc., and the regression branch processes the predicted defect bounding box position and size. It can also optimize the prediction results of the target bounding box through the distributed focus loss function (DFL) to improve the accuracy of bounding box positioning.

[0064] Furthermore, the backbone network uses CGBlock as the downsampling of the model, such as Figure 3 As shown, the steps are as follows:

[0065] S1. Use f loc and f sur To learn local features and the corresponding surrounding environment context, output the intermediate feature map; use convolution operation to extract local features and capture the fine-grained content of the input features. Use a larger receptive field to extract coarse-grained surrounding environment context information. Use local features and context features as the fused intermediate feature map to provide original information support for the next step of joint feature generation.

[0066] S2. Use f joi Get the joint feature from the intermediate feature map and convert f joi Designed as a link layer with batch normalization and parameterized ReLU activation; local features and context features are fused by concatenation or element-by-element addition, features are normalized, the training process is stabilized, and the convergence of the model is improved. Use the ReLU activation function with parameters to enhance the nonlinear fitting ability of the model.

[0067] S3. Use f glo The global context information is used as a weight vector and applied to the processing of the channel joint features. The final feature map is output, and the global context information is generated using a global pooling operation to compress the spatial dimension of the feature map into a global representation of the channel dimension. A set of fully connected layers or convolutions are used to generate a channel weight vector, the weight vector is normalized, and the generated channel weight vector is multiplied by the joint features channel by channel to weight the joint features, highlighting important channels and weakening irrelevant channels.

[0068] In general, the improved backbone network is used to perform multi-level feature extraction on the input image, and CG Block is added in the feature extraction process to enhance context information and extract local and global features of the aluminum profile surface. Information is effectively captured at different scales, and the distinguishability and expression of features are improved through the guidance of the global context.

[0069] Furthermore, the C2f module combined with the DWR module is used as a feature extraction component in the backbone network and the neck network, and the Bottleneck unit in the C2f module is improved with the DWR module in the ultra-real-time semantic segmentation model DWR-Seg, and the improved C2f module is named C2f_DWR. The three-branch DWR module replaces the traditional C2f module to improve the expansion capability of the receptive field and make the defect features more prominent.

[0070] Furthermore, if Figure 4 As shown, the steps of C2f_DWR are as follows:

[0071] S1. Based on the original input features, we introduce deep separable dilated convolutions at different rates to generate multiple feature maps with different receptive fields; we introduce deep separable dilated convolutions with different dilation rates to capture multi-scale context information. By setting different dilation coefficients, we generate multiple feature maps with different receptive fields, thereby achieving efficient modeling of multi-scale features. Through multi-rate deep separable dilated convolutions, we can achieve feature extraction with different receptive fields.

[0072] S2. Concatenate the feature maps along the channel dimension to generate a fused feature map containing multi-scale information; concatenate the multi-scale feature maps generated above along the channel dimension to integrate feature information from different receptive fields to form a fused feature representation containing multi-scale contextual information.

[0073] S3. Pass the fused feature map through a 1x1 convolution layer to generate an intermediate feature map; apply a 1x1 convolution operation to the spliced ​​fused feature map to adjust the channel dimension of the feature map, and further fuse and compress feature information from different scales.

[0074] S4. The original input features are added to the intermediate feature maps through the DWR module to form a residual connection. The original input features are added element by element to the features after the above convolution processing through the dynamic weighted residual (DWR) module to form a residual connection. This operation not only retains the original feature information, but also enhances the expression ability of multi-scale fusion features.

[0075] Furthermore, the neck network integrates the ASFP2 detection layer, which includes a scale sequence feature fusion (SSFF) module and a small target detection layer module. By adding the ASFP2 layer to the neck network, the detection performance of small target defects (such as scratches, bubbles, pits, etc.) is specifically improved. Through the resolution alignment and 3D convolution extraction of the scale sequence feature fusion (SSFF) and the small target detection layer, effective detection of targets of different scales can be achieved, especially in the small target detection scenario.

[0076] Furthermore, the scale sequence feature fusion module fuses the underlying features with the high-level features and uses an upsampling operation to generate a high-resolution feature map, which is used in the small target detection layer, thereby improving the detection capability of small targets. The SSFF module fuses the underlying features with the high-level features and uses an upsampling operation to generate a high-resolution feature map for use in the small target detection layer. The principle is shown in formulas (1) and (2):

[0077] F σ (i,j)=∑ u ∑ v f(iu,iv)×G σ (u,v) (1)

[0078]

[0079] In the small target detection layer, the nearest neighbor interpolation method is used to align all feature maps to the same resolution as P3. The shallow feature maps with high resolution and rich target location information are effectively fused with the deep feature maps. Their scale sequence features are further extracted through 3D convolution to improve the accuracy and robustness of small target detection.

[0080] On the other hand, Figure 5 As shown, the embodiment of the present application provides an aluminum profile surface defect detection system based on CDA-YOLOv8s, and the system includes:

[0081] The data acquisition module is used to obtain the surface image data of the aluminum profile, pre-process the surface image of the aluminum profile, and construct a defect data set;

[0082] The model update module is used to update the initial network structure of YOLOv8 to obtain an aluminum profile defect detection model based on CDA-YOLOv8;

[0083] A model training module, connected to the data acquisition module and the model updating module, is used to train the aluminum profile defect detection model according to the defect data set;

[0084] The defect detection module is connected with the data acquisition module and the model training module, and obtains the detection result according to the trained aluminum profile defect detection model and the aluminum profile surface image.

[0085] The system runs on the internal system of the computer, which includes: a central processing unit (CPU) and / or a graphics processing unit (GPU) for performing inference calculations of the CDA-YOLOv8 model; a memory for storing input image data, model parameters and inference results; a communication interface for receiving defect image data from an external data acquisition device and outputting the detection results to an external display device or control system. Specifically, the convolution acceleration module runs on the GPU and is used to accelerate the calculation of the C2f_DWR module in the CDA-YOLOv8 model; the dynamic memory scheduling unit is used to optimize the use of memory and reduce the space occupied by feature map storage during model inference; the boundary optimization submodule runs on the CPU and is used to introduce DFL loss in the inference stage, thereby improving the boundary prediction accuracy of defect detection.

[0086] Furthermore, the detection steps of the aluminum profile defect detection model in the defect detection module include:

[0087] Use the improved backbone network unit to perform multi-level feature extraction on the input image to obtain basic features;

[0088] Use the neck network unit to perform multi-scale feature fusion on the basic features to obtain enhanced features;

[0089] The enhanced features are used for classification and bounding box regression to generate results using a disentangled detection head unit.

[0090] The backbone network unit uses CG Block as the downsampling of the model, and the backbone network unit and the neck network unit use the C2f module combined with the DWR module as feature extraction components. The Bottleneck unit in the C2f module is improved with the DWR module in the ultra-real-time semantic segmentation model DWR-Seg, and the improved C2f module is named C2f_DWR. Furthermore, the neck network unit integrates the ASFP2 detection layer, which includes a scale sequence feature fusion module and a small target detection layer module.

[0091] In a third aspect, an embodiment of the present application provides a computer-readable storage medium having machine-executable instructions stored thereon. When the machine-executable instructions are executed by at least one processor, the at least one processor implements any of the above-mentioned methods for detecting surface defects of aluminum profiles based on CDA-YOLOv8s.

[0092] Those skilled in the art know that, in addition to realizing the system and its various devices, modules, and units provided by the present invention in a purely computer-readable program code, it is entirely possible to realize the same functions in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers by logically programming the method steps. Therefore, the system and its various devices, modules, and units provided by the present invention can be considered as a hardware component, and the devices, modules, and units included therein for realizing various functions can also be regarded as structures within the hardware component; the devices, modules, and units for realizing various functions can also be regarded as both software modules for realizing the method and structures within the hardware component.

[0093] like Figure 6 As shown in the figure, the YOLOv8 model improves the C3 module of the traditional YOLO series into a C2f module, providing the network with rich feature information, enhancing the feature extraction capability, and improving the computational efficiency of the network. The PANet structure is used in the neck, which adds a bottom-up fusion architecture compared to FPN, aggregating feature maps from different scales, enhancing multi-scale detection capabilities, and enabling the model to simultaneously detect targets of different sizes. The detection head is improved from the coupling head of the traditional YOLO series to a decoupling head. The classification task and regression task are processed through independent branches to avoid mutual interference. YOLOv8 adopts a dynamic label matching strategy, adapts to different data characteristics, enhances the flexibility of positive sample selection, and improves processing speed. YOLOv8 also introduces DFL loss, which uses the idea of ​​cross entropy to fit the boundary of the target box to make boundary prediction more accurate. For the detection of surface defects of aluminum profiles, such as Figure 7 As shown in the network structure diagram of the CDA-YOLOv8 model, the addition of the three modules of CG Block, C2f_DWR and ASFP2 can better enhance the detection capability of subtle defects, reduce the false detection rate in complex backgrounds, improve the model reasoning speed, and enhance the robustness of detection of defects of different scales.

[0094] The following sampling comparison experiment uses a 64-bit Windows 11 operating system, an Intel Core i7 13700K CPU, an NVIDIA GeForce RTX4090 graphics card, and 22GB of video memory. The deep learning framework is Pytorch 1.11.0, the Python version is 3.8, and the CUDA version is 11.3. The parameters used in the implementation process are: Optimizer: SGD; Batch_Size: 16; Epochs: 300; Momentum: 0.937; Irf: 0.01; weight decay: 0.0005.

[0095] Experiment 1:

[0096] Experimental group 1: Based on the YOLOv8 model, the network backbone uses the convolution module Conv as the downsampling of the model. In the backbone network and the neck, the C2f module is combined with the DWR module to replace the C2f in the original network. The newly proposed ASFP2 small target detection layer is used and integrated into the neck of YOLOv8.

[0097] Control group 1.1: The difference from experimental group 1 is that the backbone of the network uses CG Block as the downsampling of the model, and the rest remains the same.

[0098] Control group 1.2: The difference from experimental group 1 is that the backbone of the network uses SPDConv as the downsampling of the model, and the rest remains the same.

[0099] Control group 1.3: The difference from experimental group 1 is that the backbone of the network uses ADown as the downsampling of the model, and the rest remains the same.

[0100] Control group 1.4: The difference from experimental group 1 is that the backbone of the network uses LDConv as the downsampling of the model, and the rest remains the same.

[0101] Control group 1.5: The difference from experimental group 1 is that the backbone of the network uses WaveletPool as the downsampling of the model, and the rest remains the same.

[0102] Table 1 Downsampling comparison experimental results

[0103] Case Model Parameter quantity / M Calculation amount / G Detection accuracy / % Experimental Group 1 Conv(Baseline) 11.2 28.5 83.7 Control group 1.1 CGBlock 10.6 27.1 84.8 Control group 1.2 SPDConv 15.8 39.8 83.7 Control group 1.3 ADown 10.0 25.7 84.5 Control group 1.4 LDConv 10.6 27.7 81.6 Control group 1.5 WaveletPool 9.9 25.7 84.1

[0104] It can be seen from Table 1 that the detection effects of CG Block, Adown, and WaveletPool improved downsampling have been improved, proving that the improved downsampling module can more effectively capture and retain the key information of the input image. Among them, the effect of using the CG Block module to improve downsampling is the best. With the reduction of both the number of parameters and the amount of calculation, the detection accuracy of the model is 1.1% higher than that of the baseline model.

[0105] Experiment 2:

[0106] Experimental Group 2: Based on the YOLOv8 model, the network backbone uses CG Block as the downsampling of the model, uses C2f modules in the backbone network and the neck, and adopts the newly proposed ASFP2 small target detection layer, which is integrated into the neck of YOLOv8.

[0107] Control group 2.1: The difference from experimental group 2 is that the attention mechanism DWR module is used in the backbone network and the neck to improve C2f. The obtained new module replaces all C2f modules in the model, and the rest remain the same.

[0108] Control group 2.2: The difference from experimental group 2 is that the convolution module DySnakeConv is used in the backbone network and the neck to improve C2f, and the obtained new module replaces all C2f modules in the model, and the rest remain the same.

[0109] Control group 2.3: The difference from experimental group 2 is that the convolution module RFCAConv is used in the backbone network and the neck to improve C2f, and the obtained new module replaces all C2f modules in the model, and the rest remain the same.

[0110] Control group 2.4: The difference from experimental group 2 is that the attention mechanism MLCA module is used in the backbone network and the neck to improve C2f. The obtained new module replaces all C2f modules in the model, and the rest remain the same.

[0111] Control group 2.5: The difference from experimental group 2 is that the attention mechanism PPA module is used in the backbone network and the neck to improve C2f. The obtained new module replaces all C2f modules in the model, and the rest remains the same.

[0112] Table 2 Comparative experimental results of C2f modules

[0113] Case Model Parameter quantity / M Calculation amount / G Detection accuracy / % Experimental Group 2 C2f(Baseline) 11.2 28.5 83.7 Control group 2.1 C2f_DWR 10.6 27.1 85.1 Control group 2.2 C2f_DySnakeConv 14.4 33.8 83.9 Control group 2.3 C2f_RFCAConv 11.2 29.3 84.9 Control group 2.4 C2f_MLCA 11.1 28.5 84.3 Control group 2.5 C2f_PPA 14.3 34.9 84.2

[0114] As can be seen from Table 2, the convolution modules DySnakeConv and RFCAConv, the attention mechanisms MLCA, PPA and DWR modules are used to improve C2f. The new modules replace all C2f modules in the model. All improvements are effective, proving the feasibility of improving the C2f module. Among them, C2f_DWR has the best effect. Compared with the baseline model, the detection accuracy of the model is improved by 1.4%, and the number of model parameters and calculation amount are reduced, indicating that C2f_DWR is lightweight while taking into account performance.

[0115] Experiment 3:

[0116] Experimental Group 3: Based on the YOLOv8 model, the network backbone uses CG Block as the downsampling of the model, uses C2f modules in the backbone network and the neck, and adopts the newly proposed ASFP2 small target detection layer, which is integrated into the neck of YOLOv8.

[0117] Control group 3.1: The difference from experimental group 3 is that a new ASFP2 small target detection layer is proposed and integrated into the neck of YOLOv8 to improve the detection ability of small targets, and the rest remain the same.

[0118] Control group 3.2: The difference from experimental group 3 is that a new P2 small target detection layer is proposed and integrated into the neck of YOLOv8, and the rest remains the same.

[0119] Table 3 Comparative experimental results of small target detection layer

[0120] Case Model P2 ASFP2 Parameter quantity / M Calculation amount / G Detection accuracy / % Experimental Group 3 CGBlock+C2f_DWR 11.3 29.0 86.0 Control group 3.1 CGBlock+C2f_DWR √ 10.8 36.9 86.0 Control group 3.2 CGBlock+C2f_DWR √ 9.2 34.6 88.1

[0121] As can be seen from Table 3, adding a small target detection layer to the original model can improve the detection accuracy. The ASFP2 designed by the present invention alone has a lower parameter amount, calculation amount and accuracy than P2, but the ASFP2 is combined with the algorithm of CG Block and C2f_DWR, and the detection accuracy of the model is improved by 2.1% compared with the model with P2 added. The effectiveness of the ASFP2 structure designed by the present invention is proved.

[0122] Experiment 4:

[0123] Experimental Group 4: YOLOv8 has made structural optimizations based on YOLOv5, and improved the C3 module to the C2f module. The PANet structure is used in the neck. PANet adds a bottom-up fusion architecture compared to FPN, aggregating feature maps from different scales. The detection head is improved from the coupling head of the traditional YOLO series to the decoupling head. The classification task and regression task are processed through independent branches to avoid mutual interference. YOLOv8 adopts a dynamic label matching strategy to adapt to different data characteristics. YOLOv8 also introduces DFL loss.

[0124] Control group 4.1: The difference from experimental group 4 is that the network backbone is different. Specifically, the network backbone uses CG Block as the downsampling of the model and is added to YOLOv8 to form a new network structure, and the rest remains the same.

[0125] Control group 4.2: The difference from experimental group 4 is that the network backbone and neck are different. Specifically, in the backbone network and the neck, the C2f module is combined with the DWR module to replace the C2f in the original network and add it to YOLOv8 to form a new network structure, and the rest remains the same.

[0126] Control group 4.3: The difference from experimental group 4 is that the neck network part is different. Specifically, a new ASFP2 small target detection layer is proposed and integrated into the neck of YOLOv8, and the rest remains the same.

[0127] Control group 4.4: The difference from experimental group 4 is that the network backbone and neck are different. Specifically, the network backbone uses CG Block as the downsampling of the model. In the backbone network and the neck, the C2f module is combined with the DWR module to replace the C2f in the original network and add it to YOLOv8 to form a new network structure. The rest remains the same.

[0128] Control group 4.5: The difference from experimental group 4 is that the network backbone and neck parts are different. Specifically, the network backbone uses CG Block as the downsampling of the model, proposes a new ASFP2 small target detection layer, and integrates it into the neck of YOLOv8 to form a new network structure. The rest remains consistent.

[0129] Control group 4.6: The difference from experimental group 4 is that the network backbone and neck parts are different. Specifically, the C2f module is combined with the DWR module in the backbone network and the neck to replace the C2f in the original network; a new ASFP2 small target detection layer is proposed and integrated into the neck of YOLOv8 to form a new network structure, and the rest remains consistent.

[0130] Control group 4.7: The difference from experimental group 4 is that the backbone of the network uses CG Block as the downsampling of the model; the C2f module is combined with the DWR module in the backbone network and the neck to replace the C2f in the original network; a new ASFP2 small target detection layer is proposed and integrated into the neck of YOLOv8 to form a new network structure, and the rest remains the same.

[0131] Table 4 Ablation experiment results

[0132]

[0133]

[0134] The experimental group and the control group of the present invention added CG Block, C2f_DWR, and ASFP2 to YOLOv8 respectively to form a new network structure. By conducting ablation experiments, we tested the proposed improvements one by one to confirm their respective effectiveness. The experimental results are shown in Table 4. Through comparative analysis, it can be obtained that the precision P, recall R, and detection accuracy of the model of the improved algorithm in this paper are improved by 5.1%, 2.4%, and 4.4%, respectively, which verifies that the improved algorithm improves the effectiveness of the model for surface defects of aluminum profiles.

[0135] In this specification, the same or similar parts between the various embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the embodiments described later, the description is relatively simple, and the relevant parts can be referred to the partial description of the previous embodiments.

[0136] The above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by a person skilled in the art within the technical scope disclosed by the present invention should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.

Claims

1. A method for detecting surface defects of aluminum profiles based on CDA-YOLOv8s, characterized in that: The following steps are involved: Acquire aluminum profile surface image data, preprocess the aluminum profile surface image, and construct a defect data set; The initial network structure of YOLOv8 is updated to obtain the aluminum profile defect detection model based on CDA-YOLOv8; Training the aluminum profile defect detection model according to the defect data set; A detection result is obtained based on the trained aluminum profile defect detection model and the aluminum profile surface image.

2. The aluminum profile surface defect detection method based on CDA-YOLOv8s according to claim 1, characterized in that: The detection steps of the aluminum profile defect detection model include: Use the improved backbone network to perform multi-level feature extraction on the input image to obtain basic features; Using the neck network to perform multi-scale feature fusion on the basic features to obtain enhanced features; The enhanced features are classified and bounding box regressed using a decoupled detection head to generate results.

3. The aluminum profile surface defect detection method based on CDA-YOLOv8s according to claim 2 is characterized in that: The backbone network uses CGBlock as the downsampling of the model, and the steps are as follows: S1. Use f loc and f sur To learn local features and corresponding surrounding environment context, and output intermediate feature maps; S2. Use f joi Obtain joint features from the intermediate feature maps, and perform batch normalization and parameterized ReLU activation; S3. Use f glo The global context information is used as a weight vector and applied to the processing of channel joint features to output the final feature map.

4. The aluminum profile surface defect detection method based on CDA-YOLOv8s according to claim 2 is characterized in that: The backbone network and the neck network use the C2f module combined with the DWR module as feature extraction components, the Bottleneck unit in the C2f module is improved using the DWR module in the ultra-real-time semantic segmentation model DWR-Seg, and the improved C2f module is named C2f_DWR.

5. The aluminum profile surface defect detection method based on CDA-YOLOv8s according to claim 4 is characterized in that: The steps of C2f_DWR are as follows: S1. Based on the original input features, we introduce deep separation dilated convolutions at different rates to generate multiple feature maps with different receptive fields. S2. splicing the feature maps along the channel dimension to generate a fused feature map containing multi-scale information; S3. Pass the fused feature map through a 1x1 convolution layer to generate an intermediate feature map; S4. Add the original input feature to the intermediate feature map through the DWR module to form a residual connection.

6. The aluminum profile surface defect detection method based on CDA-YOLOv8s according to claim 2, characterized in that: The neck network integrates an ASFP2 detection layer, and the ASFP2 detection layer includes a scale sequence feature fusion module and a small target detection layer module.

7. The aluminum profile surface defect detection method based on CDA-YOLOv8s according to claim 6, characterized in that: The scale sequence feature fusion module fuses the underlying features with the high-level features, and uses an upsampling operation to generate a high-resolution feature map; the small target detection layer uses the nearest neighbor interpolation method to align the resolution of the high-resolution feature map, fuses the shallow feature map containing rich target position information and high resolution with the deep feature map, and further extracts the fused scale sequence feature information through a 3D convolution operation.

8. An aluminum profile surface defect detection system based on CDA-YOLOv8s, characterized in that: The system comprises: A data acquisition module is used to obtain aluminum profile surface image data, pre-process the aluminum profile surface image, and construct a defect data set; The model update module is used to update the initial network structure of YOLOv8 to obtain an aluminum profile defect detection model based on CDA-YOLOv8; A model training module, connected to the data acquisition module and the model updating module, and used to train the aluminum profile defect detection model according to the defect data set; The defect detection module is connected to the data acquisition module and the model training module, and obtains the detection result according to the trained aluminum profile defect detection model and the aluminum profile surface image.

9. The aluminum profile surface defect detection system based on CDA-YOLOv8s according to claim 8, characterized in that: The detection steps of the aluminum profile defect detection model in the defect detection module include: Use the improved backbone network unit to perform multi-level feature extraction on the input image to obtain basic features; Using the neck network unit to perform multi-scale feature fusion on the basic features to obtain enhanced features; The enhanced features are classified and bounding box regressed using a decoupled detection head unit to generate results.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores machine executable instructions, and when the machine executable instructions are executed by at least one processor, the at least one processor implements the aluminum profile surface defect detection method based on CDA-YOLOv8s according to any one of claims 1-7.

Citation Information

Cited By

  • Lightweight traffic target detection method and device based on LM-YOLO and medium

    CN120689676A

  • Low-altitude unmanned aerial vehicle detection method based on super-resolution and multi-dimensional attention fusion

    CN120783029A

  • Road surface crack evaluation method and system based on improved deep learning model

    CN121392483A