Fan metal surface defect detection method and device based on lightweight YOLO11 and medium

By building the SED-YOLO model, combining the lightweight YOLO11 network and multi-scale feature fusion technology, the problems of model deployment, feature extraction and positioning accuracy in metal surface defect detection are solved, and efficient and accurate defect detection is achieved, which is suitable for intelligent operation and maintenance of infrastructure such as fans.

CN120580205APending Publication Date: 2025-09-02GUILIN UNIV OF ELECTRONIC TECH
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202510676960.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-26
Publication Date
2025-09-02

AI Technical Summary

Technical Problem

The existing metal surface defect detection technology has problems such as difficult to deploy models to edge devices, insufficient feature extraction capabilities for lightweight networks, loss of multi-scale feature fusion details, and low positioning accuracy of traditional detection heads, which makes it difficult to resolve the contradiction between detection efficiency and accuracy.

Method used

The SED-YOLO model is built, and the lightweight YOLO11 network is adopted, the StarNet backbone network with a four-level hierarchical architecture and a star operation feature fusion mechanism is combined with the bottleneck module CSP and the CSP_MSCB module of the multi-scale convolution block MSCB, and the EMEFPN neck network of the efficient upsampling module is used, and the DELSCD detection head of the convolution DEConv and the group normalized GN strategy is enhanced by the DELSCD detection head of the multi-scale feature fusion and efficient detection.

Benefits of technology

It realizes high-precision identification of metal weld defects while ensuring detection efficiency, and is suitable for high-real-time and high-precision requirements in industrial quality inspection scenarios, and has strong engineering implementation value.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120580205A_ABST
    Figure CN120580205A_ABST
Patent Text Reader

Abstract

The invention discloses a fan metal surface defect detection method and device based on lightweight YOLO11 and a medium, and relates to the technical field of computer vision and industrial detection. The method comprises the following steps: acquiring fan metal surface defect image data, and preprocessing to obtain a training data set; yOLO11n is used as a basic model, and a lightweight StarNet adopting a four-level layered architecture and a star operation feature fusion mechanism is used as a model backbone network; combining a bottleneck module with a multi-scale convolution block, and constructing a neck network by applying a global heterogeneous kernel selection mechanism and an efficient up-sampling module; a heavy parameterized detail enhanced convolution and group normalization GN strategy is used to construct a detail enhanced lightweight shared convolution detection head; and sequentially connecting the backbone network, the neck network and the output layer of the detection head to form the lightweight metal surface defect detection model. On the premise that the detection efficiency is guaranteed, high-precision identification of metal weld defects can be achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision and industrial inspection technology, and more specifically, to a method, device, and medium for detecting metal surface defects of a fan based on lightweight YOLO11. Background Art

[0002] With the continuous expansion and significant increase in the complexity of metal structures, metal surface quality has become a key factor affecting building safety, durability, and maintenance costs. Weld defects, including cracks, pores, and lack of fusion, can lead to safety hazards such as structural stress concentration and fatigue fracture, and even major accidents if not detected promptly.

[0003] Traditional inspection methods rely primarily on manual visual inspection, ultrasonic testing, or X-ray testing. However, manual inspection is inefficient and susceptible to subjective factors. Ultrasonic and X-ray techniques are sensitive to interference from metal surface reflections, rust, and other factors, making them difficult to adapt to the high-precision inspection requirements in complex environments. Furthermore, the large number and complex distribution of welds make it difficult for traditional methods to achieve full coverage and real-time monitoring, necessitating an urgent need for intelligent, automated defect detection technology.

[0004] In recent years, deep learning-based defect detection methods have gradually become mainstream, but they still face many challenges. On the one hand, existing models such as YOLO and Faster R-CNN rely on complex backbone networks such as ResNet and CSPDarkNet, which have large parameters and high computational costs, making them difficult to deploy on edge computing devices. Lightweight networks such as MobileNet and ShuffleNet can reduce computational complexity, but their feature extraction capabilities are insufficient and they tend to overlook minor defects. On the other hand, multi-scale feature fusion modules such as FPN and PANet suffer from detail loss during cross-scale information exchange, resulting in a high rate of missed detection of small targets. The high reflectivity, complex textures, and environmental noise of metal surfaces further reduce the robustness of the model. In addition, traditional detection heads rely on fixed convolution kernels, which are difficult to adapt to the dynamic changes of irregular defect shapes and have limited positioning accuracy. These problems seriously restrict the practicality and large-scale application of metal surface defect detection. Summary of the Invention

[0005] In order to solve the above technical problems, the present invention provides a method, device and medium for detecting metal surface defects of fans based on lightweight YOLO11, so as to solve the key problems existing in metal surface defect detection - complex models are difficult to deploy to edge devices, lightweight network feature extraction capabilities are insufficient, multi-scale feature fusion details are lost, and traditional detection head positioning accuracy is low. By constructing the SED-YOLO model, high-precision identification of metal weld defects is achieved while ensuring detection efficiency, effectively solving the contradiction between detection efficiency and accuracy, and providing reliable technical support for the intelligent operation and maintenance of infrastructure such as fans.

[0006] In a first aspect, the present invention provides a method for detecting metal surface defects of a fan based on a lightweight YOLO11, the method comprising:

[0007] Acquire fan metal surface defect image data, preprocess the fan metal surface defect image data, and divide the data set into a training set, a test set, and a validation set in proportion;

[0008] Based on the YOLO11n model, the lightweight StarNet with a four-level hierarchical architecture and star operation feature fusion mechanism is used as the model backbone network;

[0009] The bottleneck module CSP and the multi-scale convolution block MSCB are combined to form the CSP_MSCB module, and the global heterogeneous kernel selection mechanism and the efficient upsampling module are applied to construct the neck network EMEFPN;

[0010] Using the re-parameterized detail enhancement convolution DEConv and group normalization GN strategy, we construct the detail enhancement lightweight shared convolution detection head DELSCD;

[0011] The output layers of the backbone network StarNet, the neck network EMEFPN, and the detection head DELSCD are connected in sequence to form the lightweight metal surface defect detection model SED-YOLO;

[0012] Based on the training set, the test set and the validation set, the lightweight metal surface defect detection model SED-YOLO is trained to obtain a trained lightweight metal surface defect detection model SED-YOLO.

[0013] Furthermore, in the neck network EMEFPN, the CSP_MSCB module combines the bottleneck module CSP and the multi-scale convolution block MSCB, and extracts features of shallow image information together with the Conv convolution layer. In the CSP_MSCB module, the incremental convolution kernel design that balances model performance and speed is extended to the global heterogeneous kernel selection mechanism GHKS, and the concept of heterogeneous large convolution kernels is applied to the CSP_MSCB module to adapt to the needs of different resolutions, thereby gradually obtaining multi-scale perception field information. Based on the weighted fusion of multi-scale features in the bidirectional feature pyramid network BiFPN, Concat is replaced by Add to reduce the number of parameters and calculations, and then the extracted features of different scales are adaptively selected and weighted fused through the Fusion module. The output of the CSP_MSCB module is subjected to the depthwise separable upsampling convolution module EUCB to increase the spatial resolution of the features to the resolution of the target output.

[0014] Furthermore, in the detail enhanced lightweight shared convolution detection head DELSCD, group normalization GN is used instead of normalized BN, and a re-parameterized detail enhancement convolution DEConv is introduced as a shared convolution. While adjusting the number of feature map channels and feature fusion, the spatial size of the feature map is not changed, and the Scale layer is introduced. The detection head receives feature map inputs of three different scales: P3, P4, and P5. After 1x1 convolution adjusts the number of channels and 3x3 convolution extracts local features, the feature maps of each scale enter the corresponding Conv_Reg layer for position regression prediction.

[0015] Furthermore, the neck network EMEFPN includes ten Conv modules, six CSP_MSCB modules, six Fusion modules and two EUCB modules.

[0016] Furthermore, the detail enhanced lightweight shared convolutional detection head DELSCD includes three Conv_GN modules, two DEConv modules, six Conv modules and three Scale modules.

[0017] Furthermore, in the lightweight metal surface defect detection model SED-YOLO, the neck network module EMEFPN takes the three scale outputs of the backbone network module StarNet as input, and performs top-down and bottom-up multi-scale feature extraction, fusion, and sampling operations on the input features. The neck network outputs the optimized features of three different scales to the detail enhanced lightweight shared convolution detection head DELSCD. The features of the three scales are first group-normalized and convolved, and then the group-normalized outputs are merged and passed through the shared convolution DEConv. The results are then distributed to multiple branches for different convolution operations and scaling. The detection head DELSCD outputs three target detection result parameters: large-scale, medium-scale, and small-scale target object positioning boxes Box, detection confidence Conf, and output category Class, respectively, to realize lightweight metal surface defect detection and recognition functions.

[0018] Furthermore, when the lightweight metal surface defect detection model SED-YOLO is trained based on the training set, the test set and the validation set, the loss functions used include a bounding box loss function and a classification loss function.

[0019] Furthermore, the bounding box loss function and the classification loss function are respectively expressed as:

[0020] L bbox =SmoothL1Loss(y pred ,y true )

[0021] L cls =CrossEntropyLoss(y pred ,y true )

[0022] Among them, L bbox is the bounding box loss, L cls is the classification loss, y pred is the predicted value, y true is the true value, SmoothL1Loss is a hybrid loss function that combines the mean square error and the mean absolute error, and CrossEntropyLoss is a loss function that evaluates model performance by measuring the difference between the probability distribution of the model output and the true label.

[0023] In a second aspect, the present invention provides a device for detecting metal surface defects of a fan based on a lightweight YOLO11, the device comprising:

[0024] a data acquisition unit configured to acquire fan metal surface defect image data, preprocess the fan metal surface defect image data, and divide the data set into a training set, a test set, and a validation set in proportion;

[0025] The backbone grid construction unit is configured to use the YOLO11n model as the base model, using the lightweight StarNet with a four-level hierarchical architecture and star operation feature fusion mechanism as the model backbone network;

[0026] The neck network construction unit is configured to combine the bottleneck module CSP and the multi-scale convolution block MSCB to form a CSP_MSCB module, and apply the global heterogeneous kernel selection mechanism and the efficient upsampling module to construct the neck network EMEFPN;

[0027] The detection head construction unit is configured to use the re-parameterized detail enhancement convolution DEConv and the group normalization GN strategy to construct the detail enhancement lightweight shared convolution detection head DELSCD;

[0028] The model building unit is configured to sequentially connect the output layers of the backbone network StarNet, the neck network EMEFPN, and the detection head DELSCD to form a lightweight metal surface defect detection model SED-YOLO;

[0029] The model training unit is configured to train the lightweight metal surface defect detection model SED-YOLO based on the training set, the test set and the validation set to obtain a trained lightweight metal surface defect detection model SED-YOLO.

[0030] In a third aspect, the present invention provides a readable storage medium, wherein the readable storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the method as described above.

[0031] The present invention has at least the following beneficial effects:

[0032] In the task of detecting metal surface defects on fans, this invention achieves a balance between accuracy, speed, and generalization through lightweight architecture innovation and detail enhancement technology. It is suitable for the high real-time and high-precision requirements in industrial quality inspection scenarios and has strong engineering implementation value.

[0033] Specifically, the four-level hierarchical architecture and star operation feature fusion mechanism of the StarNet backbone network can significantly compress the number of model parameters. Combined with the cross-stage local connection and global heterogeneous kernel selection mechanism of the CSP_MSCB module, it further optimizes computing efficiency and greatly reduces the model size.

[0034] Through re-parameterized DEConv, efficient upsampling modules and shared detection head design, the model's inference speed on GPU / edge devices is improved to meet industrial real-time detection needs.

[0035] The EMEFPN neck network combines the heterogeneous convolution kernels of the MSCB module (such as 3×3 and 5×5 in parallel) and the global dynamic kernel selection mechanism to significantly improve the robustness to small-size defects (cracks, scratches) and complex background interference. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 A flowchart of a method for detecting metal surface defects of a fan based on lightweight YOLO11 according to an embodiment of the present invention is shown;

[0037] Figure 2 FIG2 shows an architecture diagram of a detail enhanced lightweight shared convolutional detection head DELSCD according to an embodiment of the present invention;

[0038] Figure 3 The following is an architecture diagram of a lightweight metal surface defect detection model SED-YOLO according to an embodiment of the present invention;

[0039] Figure 4 The figure shows a structural diagram of a fan metal surface defect detection device based on lightweight YOLO11 according to an embodiment of the present invention. DETAILED DESCRIPTION

[0040] In order to enable those skilled in the art to better understand the technical solution of the present invention, the present invention is described in detail below with reference to the accompanying drawings and specific embodiments. The embodiments of the present invention are further described in detail below with reference to the accompanying drawings and specific embodiments, but are not intended to limit the present invention. For the various steps described herein, if there is no necessity for a contextual relationship between each other, the order in which they are described as examples herein should not be regarded as limiting, and those skilled in the art should know that they can be adjusted in order as long as the logic between them is not destroyed, resulting in the inability to implement the entire process.

[0041] The embodiment of the present invention provides a method for detecting metal surface defects of a fan based on lightweight YOLO11. Figure 1 As shown, the fan metal surface defect detection method based on lightweight YOLO11 includes the following steps S100-S500.

[0042] S100, obtain fan metal surface defect image data, preprocess the fan metal surface defect image data, and divide the data set into a training set, a test set, and a validation set in proportion.

[0043] Exemplarily, step S100 takes into account issues such as the image quality of the

[65] dataset, preprocesses the images to meet the input model requirements, and randomly divides the dataset into training set, validation set, and test set in a ratio of 7:2:1.

[0044] S200, based on YOLO11n model, uses the lightweight StarNet with a four-level hierarchical architecture and star operation feature fusion mechanism as the model backbone network.

[0045] In this embodiment, the YOLO11n network model is used as the basic model, and the basic model structure is improved to make the model detection more accurate and lightweight.

[0046] S300, combine the bottleneck module CSP and the multi-scale convolution block MSCB to form a CSP_MSCB module, and apply the global heterogeneous kernel selection mechanism and the efficient upsampling module to construct the neck network EMEFPN.

[0047] In some embodiments, in the neck network EMEFPN, the CSP_MSCB module combines the bottleneck module CSP and the multi-scale convolution block MSCB, and performs feature extraction on shallow image information with the Conv convolution layer. In the CSP_MSCB module, the incremental convolution kernel design that balances model performance and speed is extended to the global heterogeneous kernel selection mechanism GHKS, and the concept of heterogeneous large convolution kernels is applied to the CSP_MSCB module to adapt to the needs of different resolutions, thereby gradually obtaining multi-scale perception field information. Based on the weighted fusion of multi-scale features in the bidirectional feature pyramid network BiFPN, Concat is replaced by Add to reduce the number of parameters and calculations, and then the extracted features of different scales are adaptively selected and weighted fused through the Fusion module. The output of the CSP_MSCB module is subjected to the depthwise separable upsampling convolution module EUCB to increase the spatial resolution of the features to the resolution of the target output.

[0048] Exemplarily, the neck network EMEFPN includes ten Conv modules, six CSP_MSCB modules, six Fusion modules and two EUCB modules.

[0049] S400. Use the re-parameterized detail enhancement convolution DEConv and group normalization GN strategy to construct a detail enhancement lightweight shared convolution detection head DELSCD.

[0050] In some embodiments, as Figure 2As shown in the figure, this is the architecture diagram of the detail enhanced lightweight shared convolution detection head DELSCD. In the detail enhanced lightweight shared convolution detection head DELSCD, group normalization GN is used instead of normalization BN, and a re-parameterized detail enhancement convolution DEConv is introduced as a shared convolution. While adjusting the number of feature map channels and feature fusion, the spatial size of the feature map is not changed, and the Scale layer is introduced. The detection head receives feature map inputs of three different scales: P3, P4, and P5. After 1x1 convolution adjusts the number of channels and 3x3 convolution extracts local features, the feature maps of each scale enter the corresponding Conv_Reg layer for position regression prediction.

[0051] Exemplarily, the detail enhanced lightweight shared convolutional detection head DELSCD includes three Conv_GN modules, two DEConv modules, six Conv modules and three Scale modules.

[0052] S500, connect the output layers of the backbone network StarNet, the neck network EMEFPN, and the detection head DELSCD in sequence to form a lightweight metal surface defect detection model SED-YOLO;

[0053] In some embodiments, as Figure 3 As shown in the figure, it is an architecture diagram of the lightweight metal surface defect detection model SED-YOLO. In the lightweight metal surface defect detection model SED-YOLO, the neck network module EMEFPN takes the three scale outputs of the backbone network module StarNet as input, and performs top-down and bottom-up multi-scale feature extraction, fusion, and sampling operations on the input features. The neck network outputs the optimized features of three different scales to the detail enhanced lightweight shared convolution detection head DELSCD. The features of the three scales are first group-normalized and convolved, and then the group-normalized outputs are merged and passed through the shared convolution DEConv. The results are then distributed to multiple branches for different convolution operations and scaling. The detection head DELSCD outputs three target detection result parameters: large-scale, medium-scale, and small-scale target object positioning boxes Box, detection confidence Conf, and output category Class, respectively, to realize lightweight metal surface defect detection and recognition functions.

[0054] S600 : Based on the training set, the test set, and the validation set, the lightweight metal surface defect detection model SED-YOLO is trained to obtain a trained lightweight metal surface defect detection model SED-YOLO.

[0055] In step S600, the specific process of training the lightweight metal surface defect detection model SED-YOLO based on the training set, test set and validation set is as follows:

[0056] The 640x640 resolution image is input into the model and passes through a convolution layer Conv with 64 channels, a 3x3 convolution kernel, and a stride of 2. The convolution layer expression is:

[0057] y=σ(W*x+b)

[0058] Among them, y is the output, W is the convolution kernel, x is the input, b is the bias, σ is the activation function, and * represents the convolution operation.

[0059] The feature maps output by the four star blocks StarBlocks are 160×160×128, 80×80×256, 40×40×512, and 20×20×1024 respectively.

[0060] After a series of convolution and star block processing, the feature map is reduced in dimension by global average pooling GAP and fully connected layer FC, and then multi-scale features are further extracted by spatial pyramid pooling SPPF. The expression of global average pooling layer is:

[0061]

[0062] Among them, y i is the output of the i-th channel, x ij is the jth element of the i-th channel, and N is the total number of elements in each channel.

[0063] The expression of the spatial pyramid pooling layer is:

[0064]

[0065] Among them, y i is the output of the i-th pooling region, x ij is the jth element of the i-th pooling region, and N is the total number of elements in each pooling region.

[0066] The feature map passes through multiple maximum pooling layers MaxPool2d with different pooling kernel sizes. Each pooling layer captures features at multiple scales.

[0067] The feature maps are concatenated together to form a feature map containing multi-scale information. Then a convolutional layer Conv is added to integrate all the extracted features through the feature fusion and enhancement module C2PSA. The concatenation expression is:

[0068] y=x1+x2+…+x n

[0069] Among them, y is the fused feature, x1, x2, ..., x n are the features to be fused.

[0070] After the input integrated feature map enters the network structure based on EMEFPN, the extracted feature map undergoes convolution processing in the seventh, eighth, and ninth layers respectively to achieve feature coverage of multi-scale targets in the image.

[0071] The number of channels in this part of the convolutional layer dynamically generates network structure parameters according to the configuration.

[0072] The feature maps output by the seventh and eighth layers pass through a convolutional layer again, which are the tenth and fourteenth layers respectively.

[0073] The convolution parameters of this layer are 256 channels, a 3x3 kernel, and a stride of 2 to capture the preliminary feature information of the image.

[0074] The feature maps output by the ninth and tenth convolutional layers are fused through the eleventh layer to merge multi-scale feature information.

[0075] This is followed by three stacked CSB_MSCB modules.

[0076] The module combines the bottleneck module CSP and multi-scale convolution MSCB to extract features and reduce computational complexity.

[0077] In this module, the input feature map first passes through the CSP layer for preliminary feature extraction, and then the data is sent to three different MSCBs for feature extraction.

[0078] Perform residual connection to further extract image feature information.

[0079] Finally, the feature maps of different MSCB modules are spliced ​​together through Concat.

[0080] The spliced ​​feature map is processed again through the CSP layer.

[0081] Then, the thirteenth upsampling convolution block EUCB is responsible for merging feature information from different processing paths to achieve feature fusion, which is expressed as:

[0082] y=EUCB(x,W,b,scale f actor)

[0083] Among them, y is the output, x is the input, W is the convolution kernel, b is the bias, scale f actor is the upsampling factor.

[0084] The output information of the tenth, thirteenth, and fourteenth layers is fused through the features of the fifteenth layer and then passes through the stacked CSB_MSCB module of the sixteenth layer.

[0085] The convolution kernels used in the MSCB feature extraction of this layer are 3x3, 5x5, and 7x7 respectively to further refine the features.

[0086] The refined feature map is passed through the EUCB module to organize the feature information.

[0087] The features are extracted by the eighteenth layer of convolution kernel with parameters of 3x3 and step size of 2.

[0088] It is then handed over to the 19th layer for feature fusion, and again passes through the 20th layer stacked CSB_MSCB module, the MSCB module of this layer.

[0089] The convolution kernels used are 1×1, 3×3, and 5×5 respectively.

[0090] The feature information extracted from the twenty layers and the feature information fused from the seventeen layers are further fused through the twenty-first layer of fusion.

[0091] Twenty-two layers of stacked CSB_MSCB modules.

[0092] The convolution kernel parameters of MSCB are 1×1, 3×3, and 5×5 respectively.

[0093] After processing the feature information, 24 convolution layers are used to extract deeper feature information, and then the 13th layer is combined with it through fusion for feature fusion processing.

[0094] The Fusion module provides a variety of different fusion methods, including weighted fusion bifp, adaptive fusion adaptive, concat and spatial feature fusion SDI.

[0095] Then, it passes through three CSB_MSCB modules in the 26th layer, and the convolution kernels with MSCB parameters of 3x3, 5x5, and 7x7 are used to further refine the deep features.

[0096] After processing, it passes through the twenty-eighth convolution layer and the output information of the tenth layer through the twenty-seventh convolution layer passes through the twenty-ninth layer for feature fusion.

[0097] Then pass through three CSB_MSCB modules.

[0098] MSCB uses convolution kernels of 5×5, 7×7, and 9×9, respectively, which are used to further extract and refine features, helping to accelerate training and improve the generalization ability of the model.

[0099] The processed feature information will enter the DELSCD lightweight detection head decoder.

[0100] The feature map is processed through a series of convolutional layers Conv_GN to extract finer features.

[0101] The resolution of the feature map is increased by upsampling Scale, and then two parallel deconvolution layers DEConv2d are used to achieve multi-scale feature fusion to improve the detection ability of targets of different sizes.

[0102] Finally, these feature maps are used to calculate the bounding box loss Bbox_loss and classification loss Cls.Loss. The expressions of Bbox_loss and Cls_loss are:

[0103] L bbox =SmoothL1Loss(y pred ,y true )

[0104] L cls =CrossEntropyLoss(y pred ,y true )

[0105] Among them, L bbox is the bounding box loss, L cls is the classification loss, y pred is the predicted value, y true is the true value, SmoothL1Loss is a hybrid loss function that combines the mean square error and the mean absolute error, and CrossEntropyLoss is a loss function that evaluates model performance by measuring the difference between the probability distribution of the model output and the true label.

[0106] These two loss functions optimize the target location and category prediction respectively to generate accurate target detection results.

[0107] Based on the trained lightweight metal surface defect detection model SED-YOLO obtained above, the following experiments will demonstrate the effectiveness of the model SED-YOLO in applying it to fan metal surface defect detection.

[0108] The experimental environment and settings of the present invention are as follows: The experimental equipment system of the present invention is Windows 10, equipped with an NVIDIA GeForce RTX 3060 graphics card, and runs under the Ultralytics 8.3.9+Python-3.10.16torch-2.2.1+cu121 deep learning framework. The training hyperparameters are as follows: the optimizer is stochastic gradient descent SGD, using a linear decay learning rate adjustment strategy, the batch is 32, and the epochs are 300 rounds. The experimental dataset is a fan metal surface defect detection dataset, of which the training set is 3543 pictures, the test set is 1013 pictures, and the validation set is 506 pictures, with a total of 4 categories, namely crack, damage, dirt, and erosion. Table 1 gives the comparative results of the lightweight model ablation experiment proposed in the present invention.

[0109] Table 1 Comparative test results of lightweight target detection models

[0110]

[0111] From the experimental results in Table 1, it can be seen that the SED-YOLO model proposed in the present invention has improved mAP0.5 and accuracy by 5.26% and 7.28% respectively compared with the traditional YOLO11n model on the fan metal surface defect detection dataset, and the amount of calculation and parameters has been reduced by 23.81% and 50.39% respectively, and the model memory has been reduced by about 38.3%. The model of the present invention has achieved a balance in terms of parameter amount, calculation amount and model detection accuracy, so that the model can be deployed on edge devices for real-time processing (FPS>35) while still maintaining strong robustness and effective recognition capabilities, making the model of the present invention more suitable for application in mobile terminal embedded environments. Ultimately, the lightweight target positioning and recognition function of the model is achieved. This method makes the feature expression ability of the feature map of target detection better and the accuracy of target detection high.

[0112] The embodiment of the present invention also provides a fan metal surface defect detection device based on lightweight YOLO11, such as Figure 4 As shown, the device includes:

[0113] The data acquisition unit 401 is configured to acquire fan metal surface defect image data, pre-process the fan metal surface defect image data, and divide the data set into a training set, a test set, and a validation set in proportion;

[0114] The backbone grid construction unit 402 is configured to use the YOLO11n model as the basic model and a lightweight StarNet with a four-level hierarchical architecture and a star operation feature fusion mechanism as the model backbone network;

[0115] The neck network construction unit 403 is configured to combine the bottleneck module CSP and the multi-scale convolution block MSCB to form a CSP_MSCB module, and apply the global heterogeneous kernel selection mechanism and the efficient upsampling module to construct the neck network EMEFPN;

[0116] The detection head construction unit 404 is configured to use the re-parameterized detail enhancement convolution DEConv and the group normalization GN strategy to construct a detail enhancement lightweight shared convolution detection head DELSCD;

[0117] The model building unit 405 is configured to sequentially connect the output layers of the backbone network StarNet, the neck network EMEFPN, and the detection head DELSCD to form a lightweight metal surface defect detection model SED-YOLO;

[0118] The model training unit 406 is configured to train the lightweight metal surface defect detection model SED-YOLO based on the training set, the test set and the validation set to obtain a trained lightweight metal surface defect detection model SED-YOLO.

[0119] In some embodiments, in the neck network EMEFPN, the CSP_MSCB module combines the bottleneck module CSP and the multi-scale convolution block MSCB, and performs feature extraction on shallow image information with the Conv convolution layer. In the CSP_MSCB module, the incremental convolution kernel design that balances model performance and speed is extended to the global heterogeneous kernel selection mechanism GHKS, and the concept of heterogeneous large convolution kernels is applied to the CSP_MSCB module to adapt to the needs of different resolutions, thereby gradually obtaining multi-scale perception field information. Based on the weighted fusion of multi-scale features in the bidirectional feature pyramid network BiFPN, Concat is replaced by Add to reduce the number of parameters and calculations, and then the extracted features of different scales are adaptively selected and weighted fused through the Fusion module. The output of the CSP_MSCB module is subjected to the depthwise separable upsampling convolution module EUCB to increase the spatial resolution of the features to the resolution of the target output.

[0120] In some embodiments, in the detail enhanced lightweight shared convolution detection head DELSCD, group normalization GN is used instead of normalized BN, and a re-parameterized detail enhancement convolution DEConv is introduced as a shared convolution. While adjusting the number of feature map channels and feature fusion, the spatial size of the feature map is not changed, and a Scale layer is introduced. The detection head receives feature map inputs of three different scales: P3, P4, and P5. After 1x1 convolution adjusts the number of channels and 3x3 convolution extracts local features, the feature maps of each scale enter the corresponding Conv_Reg layer for position regression prediction.

[0121] In some embodiments, the neck network EMEFPN includes ten Conv modules, six CSP_MSCB modules, six Fusion modules and two EUCB modules.

[0122] In some embodiments, the detail enhanced lightweight shared convolutional detection head DELSCD includes three Conv_GN modules, two DEConv modules, six Conv modules and three Scale modules.

[0123] In some embodiments, in the lightweight metal surface defect detection model SED-YOLO, the neck network module EMEFPN takes the three scale outputs of the backbone network module StarNet as input, and performs top-down and bottom-up multi-scale feature extraction, fusion, and sampling operations on the input features. The neck network outputs optimized features of three different scales to the detail enhanced lightweight shared convolution detection head DELSCD. The features of the three scales are first group-normalized and convolved, and the group-normalized outputs are merged and passed through the shared convolution DEConv. The results are then distributed to multiple branches for different convolution operations and scaling. The detection head DELSCD outputs three target detection result parameters: large-scale, medium-scale, and small-scale target object positioning boxes Box, detection confidence Conf, and output category Class, respectively, to realize lightweight metal surface defect detection and recognition functions.

[0124] In some embodiments, when the lightweight metal surface defect detection model SED-YOLO is trained based on the training set, the test set, and the validation set, the loss functions used include a bounding box loss function and a classification loss function.

[0125] In some embodiments, the bounding box loss function and the classification loss function are respectively expressed as:

[0126] L bbox =SmoothL1Loss(y pred ,y true )

[0127] L cls =CrossEntropyLoss(y pred ,y true )

[0128] Among them, L bbox is the bounding box loss, L cls is the classification loss, y pred is the predicted value, y trueis the true value, SmoothL1Loss is a hybrid loss function that combines the mean square error and the mean absolute error, and CrossEntropyLoss is a loss function that evaluates model performance by measuring the difference between the probability distribution of the model output and the true label.

[0129] It should be noted that the structures of the various fan metal surface defect detection devices based on lightweight YOLO11 described in this embodiment belong to the same technical concept as the fan metal surface defect detection method based on lightweight YOLO11 described previously, and achieve the same beneficial effects through the same principles, which will not be repeated here.

[0130] An embodiment of the present invention further provides a readable storage medium, which stores one or more programs. The one or more programs can be executed by one or more processors to implement the method described in any of the above embodiments.

[0131] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions or improvements made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A method for detecting metal surface defects of a fan based on lightweight YOLO11, characterized in that: The method comprises: Acquire fan metal surface defect image data, preprocess the fan metal surface defect image data, and divide the data set into a training set, a test set, and a validation set in proportion; Based on the YOLO11n model, the lightweight StarNet with a four-level hierarchical architecture and star operation feature fusion mechanism is used as the model backbone network; The bottleneck module CSP and the multi-scale convolution block MSCB are combined to form the CSP_MSCB module, and the global heterogeneous kernel selection mechanism and the efficient upsampling module are applied to construct the neck network EMEFPN; Using the re-parameterized detail enhancement convolution DEConv and group normalization GN strategy, we construct the detail enhancement lightweight shared convolution detection head DELSCD; The output layers of the backbone network StarNet, the neck network EMEFPN, and the detection head DELSCD are connected in sequence to form the lightweight metal surface defect detection model SED-YOLO; Based on the training set, the test set and the validation set, the lightweight metal surface defect detection model SED-YOLO is trained to obtain a trained lightweight metal surface defect detection model SED-YOLO.

2. The method for detecting metal surface defects of a fan based on lightweight YOLO11 according to claim 1 is characterized in that: In the neck network EMEFPN, the CSP_MSCB module combines the bottleneck module CSP and the multi-scale convolution block MSCB, and extracts features of shallow image information together with the Conv convolution layer. In the CSP_MSCB module, the incremental convolution kernel design that balances model performance and speed is extended to the global heterogeneous kernel selection mechanism GHKS, and the concept of heterogeneous large convolution kernels is applied to the CSP_MSCB module to adapt to the needs of different resolutions, thereby gradually obtaining multi-scale perception field information. Based on the weighted fusion of multi-scale features in the bidirectional feature pyramid network BiFPN, Concat is replaced by Add to reduce the number of parameters and calculations. The extracted features of different scales are then adaptively selected and weighted fused through the Fusion module. The spatial resolution of the features of the output of the CSP_MSCB module is improved to the resolution of the target output through the depthwise separable upsampling convolution module EUCB.

3. The method for detecting metal surface defects of a fan based on lightweight YOLO11 according to claim 1 is characterized in that: In the detail enhancement lightweight shared convolution detection head DELSCD, group normalization GN is used instead of normalized BN, and a re-parameterized detail enhancement convolution DEConv is introduced as a shared convolution. The number of feature map channels and feature fusion are adjusted while the spatial size of the feature map is not changed. The Scale layer is introduced, and the detection head receives feature map inputs of three different scales: P3, P4, and P5. After 1x1 convolution adjusts the number of channels and 3x3 convolution extracts local features, the feature maps of each scale enter the corresponding Conv_Reg layer for position regression prediction.

4. The method for detecting metal surface defects of a fan based on lightweight YOLO11 according to claim 1 is characterized in that: The neck network EMEFPN includes ten Conv modules, six CSP_MSCB modules, six Fusion modules and two EUCB modules.

5. The method for detecting metal surface defects of a fan based on lightweight YOLO11 according to claim 1 is characterized in that: The detail enhanced lightweight shared convolutional detection head DELSCD includes three Conv_GN modules, two DEConv modules, six Conv modules and three Scale modules.

6. The method for detecting metal surface defects of a fan based on lightweight YOLO11 according to claim 1 is characterized in that: In the lightweight metal surface defect detection model SED-YOLO, the neck network module EMEFPN takes the three-scale outputs of the backbone network module StarNet as input, and performs top-down and bottom-up multi-scale feature extraction, fusion, and sampling operations on the input features. The neck network outputs optimized features of three different scales to the detail-enhanced lightweight shared convolution detection head DELSCD. The features of the three scales are first group-normalized and convolved, and the group-normalized outputs are merged and passed through the shared convolution DEConv. The results are then distributed to multiple branches for different convolution operations and scaling. The detection head DELSCD outputs three target detection result parameters: large-scale, medium-scale, and small-scale target object positioning boxes Box, detection confidence Conf, and output category Class, respectively, to realize lightweight metal surface defect detection and recognition functions.

7. The method for detecting metal surface defects of a fan based on lightweight YOLO11 according to claim 1 is characterized in that: When the lightweight metal surface defect detection model SED-YOLO is trained based on the training set, the test set, and the validation set, the loss functions used include a bounding box loss function and a classification loss function.

8. The method for detecting metal surface defects of a fan based on lightweight YOLO11 according to claim 8 is characterized in that: The bounding box loss function and classification loss function are respectively expressed as: L bbox =SmoothL1Loss(y pred ,y true ) L cls =CrossEntropyLoss(y pred ,y true ) Among them, L bbox is the bounding box loss, L cls is the classification loss, y pred is the predicted value, y true is the true value, SmoothL1Loss is a hybrid loss function that combines the mean square error and the mean absolute error, and CrossEntropyLoss is a loss function that evaluates model performance by measuring the difference between the probability distribution of the model output and the true label.

9. A fan metal surface defect detection device based on lightweight YOLO11, characterized in that: The device comprises: a data acquisition unit configured to acquire fan metal surface defect image data, preprocess the fan metal surface defect image data, and divide the data set into a training set, a test set, and a validation set in proportion; The backbone grid construction unit is configured to use the YOLO11n model as the base model, using the lightweight StarNet with a four-level hierarchical architecture and star operation feature fusion mechanism as the model backbone network; The neck network construction unit is configured to combine the bottleneck module CSP and the multi-scale convolution block MSCB to form a CSP_MSCB module, and apply the global heterogeneous kernel selection mechanism and the efficient upsampling module to construct the neck network EMEFPN; The detection head construction unit is configured to use the re-parameterized detail enhancement convolution DEConv and the group normalization GN strategy to construct the detail enhancement lightweight shared convolution detection head DELSCD; The model building unit is configured to sequentially connect the output layers of the backbone network StarNet, the neck network EMEFPN, and the detection head DELSCD to form a lightweight metal surface defect detection model SED-YOLO; The model training unit is configured to train the lightweight metal surface defect detection model SED-YOLO based on the training set, the test set and the validation set to obtain a trained lightweight metal surface defect detection model SED-YOLO. 10 . A non-transitory computer-readable storage medium storing instructions, which, when executed by a processor, perform the method according to claim 1 .

Citation Information

Cited By

  • Unmanned aerial vehicle target detection method based on improved YOLOv11n

    CN121236643A

  • Lightweight target detection edge deployment method based on high-frequency detail enhancement

    CN121746687A

  • Lightweight model-based forklift tray tracking method, system and equipment and medium

    CN121982072A

  • Continuous casting slab running state monitoring method and system based on lightweight YOLOv8 and continuous learning

    CN122244802A

  • Method and system for navel orange defect detection

    CN122435604A