Defect detection method and device, storage medium and computer program product
By using a lightweight feature extraction network and an adaptive cross-level feature pyramid to optimize feature fusion, the problem of missed detection and false detection in small-sized defect detection under complex backgrounds is solved, and efficient defect detection is achieved.
Patent Information
- Application Number
- CN202511104751.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-07
- Publication Date
- 2025-11-18
AI Technical Summary
Existing technologies suffer from missed detections and false detections in detecting small-sized defects in complex backgrounds, and high-performance models have high computational overhead, making them difficult to deploy on mobile devices with limited hardware resources.
We designed a lightweight feature extraction network, combined with an adaptive cross-level feature pyramid and lightweight convolution operations, and optimized the feature fusion network to improve the model's ability to learn global information and its computational efficiency.
It improves the accuracy of identifying and locating small-sized defects, reduces the number of model parameters, and increases detection speed and efficiency.
Smart Images

Figure CN120976160A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of defect detection technology, specifically to a defect detection method, a defect detection device, a storage medium, and a computer program product. Background Technology
[0002] As an indispensable key tool in the field of industrial quality inspection, industrial defect detection is widely used in production and operation scenarios such as unmanned quality inspection, intelligent inspection, and quality control. It helps enterprises understand the distribution of product quality, identify weak points, reduce product quality fluctuations, and form a closed-loop control for production and quality improvement. Industrial defect detection aims to inspect the surface defects of industrial products, such as scratches, stains, and dents. It is a crucial link in ensuring product quality and maintaining stable production, and is widely used in various industrial scenarios such as aerospace, machinery manufacturing, and semiconductors. Traditional defect detection relies on human experience and subjective factors, resulting in high costs, low efficiency, and high false positive and false negative rates, making it difficult to cover large-scale quality inspection needs. With the rapid development of emerging technologies in fields such as industrial imaging and computer vision, machine vision-based industrial defect detection technology has become an effective solution for product surface quality inspection due to its advantages in high precision, high efficiency, low cost, and non-destructive nature.
[0003] In recent years, with the superior performance of deep learning in processing complex image data, vision-based industrial defect detection has evolved from manually designing defect features and using image processing or machine learning-based detection methods to deep learning methods based on convolutional neural networks. Compared with traditional technologies, deep learning has the advantage of autonomously extracting and learning defect features from industrial images, greatly reducing the cost of manually designing features, while also possessing good detection accuracy and generalization ability.
[0004] However, when dealing with the detection of minute defects such as cracks and scratches, the small and irregular area occupied by these defects, their low signal-to-noise ratio, and tendency to cluster and obscure them, coupled with the difficulty in extracting feature information, make them highly susceptible to interference from complex textured backgrounds and noise introduced during industrial imaging. This results in general-purpose target detection models struggling to accurately locate and identify small-sized defects, leading to missed detections and false detections. Furthermore, to achieve high-precision detection, high-performance defect detection models typically construct complex network structures, incurring high computational costs and making them difficult to deploy on mobile devices with limited hardware resources. This hinders their practical application in industrial scenarios requiring real-time image data processing, such as production line quality inspection and intelligent inspection.
[0005] It should be noted that the information disclosed in the background section above is only used to enhance the understanding of the background of this disclosure, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention
[0006] This disclosure provides a defect detection method, a defect detection device, a storage medium, a computer program product, and an electronic device, which can improve the accuracy of identifying and locating small-sized defects in complex backgrounds and effectively overcome the defects existing in the prior art.
[0007] Other features and advantages of this disclosure will become apparent from the following detailed description, or may be learned in part from practice of this disclosure.
[0008] According to a first aspect of this disclosure, a defect detection method is provided, the method comprising:
[0009] Acquire the target detection image and input the target detection model into the lightweight detection model;
[0010] The target detection image is processed by a feature extraction network to extract features, and the target feature map is obtained based on the feature map output by the target layer backbone network.
[0011] A feature fusion network is used to perform feature fusion processing on target feature maps at different levels, combined with the fusion space weight parameters corresponding to target feature maps at each scale obtained by adaptive learning, in order to obtain fused feature information.
[0012] The fused feature information is detected using the detection heads of the feature detection network to obtain defect detection results.
[0013] In some exemplary embodiments, the feature extraction network includes a MobileViTv3-based backbone network;
[0014] The method further includes:
[0015] The feature maps extracted from the third, sixth, eighth, and tenth layers of the backbone network are used as the target feature maps.
[0016] In some exemplary embodiments, the step of using a feature extraction network to perform feature extraction processing on the target detection image, and obtaining a target feature map based on the feature map output by the target layer backbone network, includes:
[0017] Local feature extraction is performed on the target detection image to obtain a first feature map embedded with local feature information;
[0018] The first feature map is processed by vector encoding so that the pixels can acquire global information, and a second feature map with global representation is obtained.
[0019] The first feature map and the second feature map are concatenated along the channel dimension and then fused through a 1×1 convolution to obtain a fused feature map.
[0020] The fused feature map is combined with the target detection image to obtain the target feature map.
[0021] In some exemplary embodiments, the step of extracting features from the target detection image to obtain a first feature map embedding local feature information includes:
[0022] A 3×3 depth-separable convolutional layer is used to perform convolution processing on the target detection image to obtain the first feature information;
[0023] A 1×1 convolutional layer is used to perform channel dimensionality reduction on the first feature information to generate a first feature map that embeds local feature information.
[0024] In some exemplary embodiments, the step of performing vector encoding processing on the first feature map to enable pixels to acquire global information and obtain a second feature map with global representation includes:
[0025] The first feature map, which embeds local feature information, is subjected to data transformation processing and mapped to the corresponding first vector group.
[0026] The first vector group is encoded using a Linear Transformer network, and a second feature map with global representation is extracted.
[0027] In some exemplary embodiments, a feature fusion network is used to perform feature fusion processing on target feature maps at different levels, combined with the fusion space weight parameters corresponding to the target feature maps at each scale obtained through adaptive learning, to obtain fused feature information, including:
[0028] Based on the Adaptive Cross-Level Feature Pyramid (ACFPN), the C2f module is used to perform 2×2 transposed convolution upsampling and 2×2 standard convolution downsampling to unify the size of target feature maps at different levels.
[0029] The fusion spatial weight parameters of feature maps at various scales are adaptively learned through 1×1 convolution and Softmax activation function.
[0030] The weighted feature map is obtained by multiplying the feature maps of different levels with their corresponding spatial weights element by element. The fused feature information after adaptive spatial fusion is obtained by summing the weighted feature maps element by element.
[0031] In some exemplary embodiments,
[0032] The C2f module includes a first standard convolutional layer, multiple Bottleneck layers, and a second standard convolutional layer arranged sequentially.
[0033] The Bottleneck layer includes a Faster module; the Faster module includes a partial convolution PConv computation module and a pointwise convolution PWConv computation module; the partial convolution PConv computation module is used to extract spatial features by performing regular convolution operations on some channels, while the remaining channels are mapped identically; the pointwise convolution PWConv computation module is used to perform pointwise convolution operations on all channels and fuse all channels to extract feature information.
[0034] According to a second aspect of this disclosure, a defect detection apparatus is provided, comprising:
[0035] The image acquisition module is used to acquire the target detection image and input the target detection model into the lightweight detection model;
[0036] The feature extraction module is used to perform feature extraction processing on the target detection image using a feature extraction network, and to obtain the target feature map based on the feature map output by the target layer backbone network;
[0037] The feature fusion module is used to perform feature fusion processing on target feature maps of different levels using the feature fusion network, combined with the fusion space weight parameters corresponding to the target feature maps of each scale obtained by adaptive learning, in order to obtain fused feature information.
[0038] The detection result output module is used to detect the fused feature information using the detection heads of the feature detection network to obtain defect detection results.
[0039] According to a third aspect of this disclosure, a storage medium is provided that stores a computer program thereon, which, when executed by a processor, implements the defect detection method described above.
[0040] According to a fourth aspect of this disclosure, an electronic device is provided, comprising:
[0041] Processor; and
[0042] Memory for storing the executable instructions of the processor;
[0043] The processor is configured to implement the aforementioned defect detection method by executing the executable instructions.
[0044] According to a fifth aspect of this disclosure, a computer program product is provided, on which a computer program is stored, and when the computer program is executed by a processor, it implements the above-described defect detection method.
[0045] The defect detection method provided in this disclosure enhances the model's ability to learn global information by designing a lightweight feature extraction network. It also constructs a direct interaction between features at non-adjacent levels through an adaptive cross-level feature pyramid optimization feature fusion network, compensating for the feature information loss and degradation caused by the reduced feature extraction capability of the lightweight backbone network. Simultaneously, it optimizes convolution operations while maintaining model detection performance, improving the algorithm's computational efficiency. The improved model effectively reduces the number of parameters and increases detection speed while ensuring defect detection accuracy.
[0046] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description
[0047] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure. It is obvious that the drawings described below are merely some embodiments of this disclosure, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort.
[0048] Figure 1 This illustration schematically depicts an application scenario architecture in an exemplary embodiment of the present disclosure;
[0049] Figure 2 The schematic diagram illustrates a defect detection method according to an exemplary embodiment of the present disclosure;
[0050] Figure 3 This illustration schematically shows a schematic diagram of the MVAF-Net model in an exemplary embodiment of the present disclosure;
[0051] Figure 4 This schematic diagram illustrates the backbone feature extraction network of the MVAF-Net model in an exemplary embodiment of the present disclosure.
[0052] Figure 5 This schematic diagram illustrates the structure of the MobileViTv3 Block in an exemplary embodiment of the present disclosure.
[0053] Figure 6 This schematic diagram illustrates a feature fusion network of MVAF-Net in an exemplary embodiment of the present disclosure.
[0054] Figure 7 This illustration schematically shows the Faster_C2f structure and its application in exemplary embodiments of this disclosure;
[0055] Figure 8This illustration shows a comparison of lightweight strategies for model training convolution operations in exemplary embodiments of the present disclosure.
[0056] Figure 9 This schematic diagram illustrates the composition of a defect detection apparatus according to an exemplary embodiment of the present disclosure;
[0057] Figure 10 This schematic diagram illustrates the composition of an electronic device according to an exemplary embodiment of the present disclosure. Detailed Implementation
[0058] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, they are provided so that this disclosure will be more comprehensive and complete, and will fully convey the concept of the exemplary embodiments to those skilled in the art. The described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.
[0059] Furthermore, the accompanying drawings are merely illustrative of this disclosure and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and therefore repeated descriptions of them will be omitted. Some block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.
[0060] To address the shortcomings and deficiencies of existing technologies, this example embodiment provides a defect detection method applicable to detecting minute defects such as cracks and scratches in industrial products. Specifically, this method can be applied to scenarios such as... Figure 1In the application environment shown, terminal 102 communicates with server 104 via network 101. Data storage system 103 stores the data that server 104 needs to process. Data storage system 103 can be integrated onto server 104, or it can be located in the cloud or on other network servers. Terminal 102 can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. Portable wearable devices can be smartwatches, smart bracelets, head-mounted devices, etc. Head-mounted devices can be virtual reality (VR) devices, augmented reality (AR) devices, smart glasses, etc. Server 104 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. The user information collected by terminal 102 and the user information processed by server 104 can be information and data authorized by the user or fully authorized by all parties. The collection, storage, use, processing, transmission, provision, disclosure and application of related data shall comply with the relevant laws, regulations and standards of relevant countries and regions, take necessary confidentiality measures, not violate public order and good morals, and provide corresponding operation entry points for users to choose to authorize or refuse.
[0061] For example, a user can initiate an image detection request on terminal 102. Server 104 can receive and execute the request from the terminal, thereby obtaining the detection results and saving the detection results data to data storage system 103. Then, the detection results data is sent to the terminal for display.
[0062] In one exemplary embodiment, such as Figure 2 As shown, a defect detection method is provided, which can be applied to... Figure 1 The following steps, S11 to S14, are used as an example of servers and / or terminals in the system. Wherein:
[0063] In step S11, the target detection image is acquired, and the target detection model is input into the lightweight detection model.
[0064] In this example implementation, the target detection image can be a pre-acquired or real-time acquired image of the device surface, such as a surface image of a gear device. A corresponding image detection task can be created for the target detection image, and this task can be executed, inputting the target detection image into a trained lightweight detection model. This lightweight detection model can be a lightweight small target defect detection model based on YOLOv8. Based on the original model, this invention designs a defect detection algorithm EM-YOLO based on hierarchical downsampling and feature comparison. Addressing the problems of complex EM-YOLO network model structure, large number of parameters, and difficulty in practical deployment and application, this invention improves the EM-YOLO defect detection algorithm and proposes a lightweight defect detection model MVAF-Net based on separable self-attention and adaptive cross-level feature fusion. The model structure is as follows: Figure 3 As shown, this enables rapid identification and precise localization of minute defects in complex backgrounds. The entire MVAF-Net model can be divided into three parts: 1) Single-view feature extraction (SVFE), 2) Multi-view feature fusion (MVFF), and 3) Fusion feature detection (FFD).
[0065] In step S12, the target detection image is processed by a feature extraction network to extract features, and the target feature map is obtained based on the feature map output by the target layer backbone network.
[0066] In this example implementation, the feature extraction network includes a backbone network based on MobileViTv3; the method further includes using the feature maps extracted from the third, sixth, eighth, and tenth layers of the backbone network as the target feature maps.
[0067] In this example implementation, the step of using a feature extraction network to perform feature extraction processing on the target detection image, and obtaining a target feature map based on the feature map output by the target layer backbone network, includes:
[0068] Local feature extraction is performed on the target detection image to obtain a first feature map embedded with local feature information;
[0069] The first feature map is processed by vector encoding so that the pixels can acquire global information, and a second feature map with global representation is obtained.
[0070] The first feature map and the second feature map are concatenated along the channel dimension and then fused through a 1×1 convolution to obtain a fused feature map.
[0071] The fused feature map is combined with the target detection image to obtain the target feature map.
[0072] For example, this invention constructs the backbone network based on the lightweight hybrid architecture model MobileViTv3 of CNN-ViT; the backbone feature extraction network of the model is as follows: Figure 4 As shown, based on the local information perception capability of CNN and the global information modeling capability of ViT, the model effectively combines the advantages of spatial inductive bias of convolutional neural networks and visual self-attention models in capturing long-distance dependencies. This allows the model to extract richer surface texture features while extracting local feature information, improving the detection rate of small-sized defects in complex industrial images, and reducing the number of parameters and computational complexity. Furthermore, to achieve multi-scale defect detection, feature maps C2, C3, C4, and C5 extracted from the third, sixth, eighth, and tenth layers of the backbone network are used as inputs to the feature fusion network.
[0073] In this example implementation, the step of extracting features from the target detection image to obtain a first feature map embedding local feature information includes:
[0074] A 3×3 depth-separable convolutional layer is used to perform convolution processing on the target detection image to obtain the first feature information;
[0075] A 1×1 convolutional layer is used to perform channel dimensionality reduction on the first feature information to generate a first feature map that embeds local feature information.
[0076] Specifically, the structure of the MobileViTv3 Block is as follows: Figure 5 As shown, the specific process of the MobileViTv3 Block can be summarized as follows: The feature map of size H×W×C is first subjected to a 3×3 depthwise separable convolution operation to extract features, and then subjected to a 1×1 convolution to reduce the dimensionality of the channels, generating a feature map embedding local information; then, it is transformed into a vector group through data transformation and input into the Linear Transformer module, so that the pixels can obtain global information, thereby obtaining a feature map with global representation; the local representation feature map and the global representation feature map are concatenated along the channel dimension, and then subjected to a 1×1 convolution to fuse features; finally, the fused feature map is added to the original input feature map to obtain the final output feature map.
[0077] For example, Depthwise Separable Convolution is a convolution operation that combines Depthwise and Pointwise Convolution. It first performs independent spatial convolution on each channel of the input using Depthwise Convolution, and then combines the features between channels using Pointwise Convolution. By separating these two steps, Depthwise Separable Convolution can significantly reduce the number of parameters, thereby reducing computational and storage costs. Depthwise Convolution is a convolution operation that performs convolution operations only on the input channels. It uses a convolution kernel (filter) with the same depth as the input, and each kernel operates on only one channel of the input data. This means that for an input with C input channels, C different convolution kernels will be used for the convolution operation, with each kernel associated with only one channel. This allows for convolution operations on the spatial information of the input while reducing the number of parameters. Pointwise Convolution is a 1x1 convolution operation that uses only a single 1x1 convolution kernel. Pointwise convolution applies a linear combination of channels and a non-linear activation function to the input without changing its spatial resolution. It allows control over the number of parameters while altering the channel dimension.
[0078] In this example implementation, the step of performing vector encoding processing on the first feature map to enable pixels to acquire global information and obtain a second feature map with global representation includes:
[0079] The first feature map, which embeds local feature information, is subjected to data transformation processing and mapped to the corresponding first vector group.
[0080] The first vector group is encoded using a Linear Transformer network, and a second feature map with global representation is extracted.
[0081] Specifically, the Linear Transformer network structure includes three branches: I, K, and V. Branch I linearly maps each latent vector sequence, transforming the input vector set X into a k-dimensional vector. After softmax normalization, the context scores are obtained. The context scores are then multiplied by the output broadcast mapped by branch K, and summed element-wise to obtain the context vector representing the weights. The context information encoded in this vector is shared by all vector sequences in the input vector set X. The context vector is then multiplied by the output broadcast mapped by branch V, and after a linear transformation, the final output Y is obtained.
[0082] In step S13, the feature fusion network is used to perform feature fusion processing on the target feature maps at different levels, combined with the fusion space weight parameters corresponding to the target feature maps at each scale obtained by adaptive learning, in order to obtain fused feature information.
[0083] In this example embodiment, step S13 described above may include:
[0084] Based on the Adaptive Cross-Level Feature Pyramid (ACFPN), the C2f module is used to perform 2×2 transposed convolution upsampling and 2×2 standard convolution downsampling to unify the size of target feature maps at different levels.
[0085] The fusion spatial weight parameters of feature maps at various scales are adaptively learned through 1×1 convolution and Softmax activation function.
[0086] The weighted feature map is obtained by multiplying the feature maps of different levels with their corresponding spatial weights element by element. The fused feature information after adaptive spatial fusion is obtained by summing the weighted feature maps element by element.
[0087] Specifically, the structure of the MVAF-Net feature fusion network is as follows: Figure 6 As shown, this invention designs an adaptive cross-level feature pyramid (ACFPN), which improves the loss and degradation of feature information in the original EM-YOLO model and enhances the model's localization accuracy by supporting direct interaction between non-adjacent levels.
[0088] ACFPN preserves the resolution and semantic information of the current layer's feature map through lateral connections, fuses original features, and alleviates semantic imbalance between different layers. Based on cross-level feature propagation, it concatenates contextual feature information with current feature information, enhancing the flow of feature information across different levels. This allows for more comprehensive and richer information transfer and feature fusion between levels, improving the model's understanding and abstraction of input data, better representing the complexity and diversity of input data, and thus improving model accuracy.
[0089] An adaptive spatial fusion approach is employed, assigning an adaptive fusion factor to each path to achieve feature adjustment and feature selection. Feature adjustment optimizes the fusion result by adjusting feature weights, giving the model better representational capabilities. Feature selection dynamically chooses useful features based on their importance and confidence level.
[0090] In the adaptive spatial fusion process, identical scaling is performed first. Since feature maps at different levels have different scales and number of channels, to facilitate feature fusion, a 2×2 transposed convolution with stride=2 is used for upsampling, and a 2×2 standard convolution with stride=2 is used for downsampling to unify the size of features at different levels. Using transposed convolution instead of ordinary upsampling effectively reduces feature loss for small targets caused by manual feature engineering. Then, the fusion spatial weight parameters for feature maps at each scale are adaptively learned through 1×1 convolution and the Softmax activation function; their values determine whether feature points are enhanced or suppressed. Finally, feature maps at different levels are multiplied element-wise with their corresponding spatial weights to obtain three weighted feature maps. The final output after adaptive spatial fusion is obtained by summing the feature maps element-wise.
[0091] In this example implementation, the C2f module includes a first standard convolutional layer, multiple Bottleneck layers, and a second standard convolutional layer arranged sequentially.
[0092] The Bottleneck layer includes a Faster module; the Faster module includes a partial convolution PConv computation module and a pointwise convolution PWConv computation module; the partial convolution PConv computation module is used to extract spatial features by performing regular convolution operations on some channels, while the remaining channels are mapped identically; the pointwise convolution PWConv computation module is used to perform pointwise convolution operations on all channels and fuse all channels to extract feature information.
[0093] Specifically, the feature fusion network of the model mainly consists of C2f modules, as referenced. Figure 7 As shown, the C2f module contains two standard convolutions and multiple stacked Bottleneck modules. In industrial product surface defect detection, complex convolutional structures are not conducive to deployment in resource-constrained industrial scenarios. To achieve a lightweight feature fusion network structure, this invention introduces FasterBlock into the C2f design to create the Faster_C2f module, and uses it to replace the adaptive spatial fusion C2f module. This lightweight convolutional structure balances the number of parameters and the running speed, meeting the requirements for rapid defect detection.
[0094] FasterNet achieves faster detection speeds than other mainstream lightweight networks on various devices while maintaining accuracy. The Faster module consists of partial convolution (PConv) and pointwise convolution (PWConv). PConv, based on the redundancy of feature channels, only uses regular convolution operations to extract spatial features from a subset of channels, while maintaining identity mapping for the remaining channels. Then, PWConv operations are applied to all channels to fuse them, thereby extracting feature information.
[0095] Step S14: Use the detection heads of the feature detection network to detect the fused feature information to obtain defect detection results.
[0096] In some exemplary embodiments, after acquiring the fused feature information, a feature detection network can be used to extract and identify features from the fused feature information, thereby obtaining detection results for minor defects. For example, the detection head may include: classification (focal loss), bounding box regression (SmoothL1 loss), and orientation classification (softMax loss).
[0097] In some exemplary embodiments, a model training method is provided for providing a lightweight model for small target defect detection based on YOLOv8. The model training method may include:
[0098] Step 1: Select the dataset for model training and collect the raw data needed for defect detection.
[0099] Step 2: Construct the EM-YOLO defect detection algorithm based on hierarchical downsampling and feature comparison.
[0100] This invention, based on the YOLOv8 network, improves the feature extraction network by designing a hierarchical downsampling module. This addresses the information loss issue during feature extraction due to minor defects, enhancing the model's ability to learn defect features. The feature fusion network is optimized by adding a shallow detection head as a branch for small-sized defect detection, improving the model's ability to classify and locate minor defects. A multi-scale feature comparison module is designed to suppress interference from complex backgrounds, enhancing the model's noise resistance. The bounding box loss function is optimized to better regress minor defects to their ground truth bounding boxes, improving the issues of missed and false detections.
[0101] Finally, an EM-YOLO defect detection algorithm based on hierarchical downsampling and feature comparison was constructed to improve the accuracy of identifying and locating small-sized defects in complex backgrounds.
[0102] In some exemplary embodiments, the training strategy of the model can be pre-configured, and the defect detection model described above can be trained based on the training strategy.
[0103] Specifically, the training strategy can be to preset the main parameters of the defect detection model proposed in the previous section based on the gear surface defect detection dataset: image input size is 640×640, momentum parameter is 0.937, weight decay coefficient is 0.0005, learning rate is 0.01, batch size is set to 16, and Mosaic data augmentation is enabled and disabled in the last 10 epochs.
[0104] In some exemplary embodiments, based on a gear surface defect detection dataset, multiple sets of experiments are set up to evaluate the detection performance of each algorithm on the test set, and the training is divided into three stages:
[0105] a) The first stage mainly analyzes the impact of different lightweight backbone feature extraction networks on model performance. Taking YOLOV8s as the benchmark model, FasterNet, EfficientNet, MobileNext, GhostNetv2 and MobileViTv3 are selected to improve the architecture of the backbone part of the model. The model is trained and tested on the gear defect detection training set and test set. The average class detection accuracy, model parameter count and computational cost are comprehensively compared to evaluate the detection effect of each algorithm on the test set.
[0106] b) The second stage is to verify the effectiveness of the adaptive cross-level feature fusion structure. Using EM-YOLO as the benchmark model, the backbone feature extraction network is improved based on the evaluation results of the first stage experiment. A comparative experiment is designed based on the gear defect detection dataset to explore the impact of the feature fusion network optimization strategy on the performance of the defect detection algorithm.
[0107] c) The third stage is to verify the effectiveness of the lightweight convolution operation improvement strategy and the rationality of the improvement location, referring to... Figure 8 As shown, based on the analysis and comparison of the detection results of the first and second stages, the feature extraction network and feature fusion network of EM-YOLO are optimized. Using this as the baseline model, C2f at different positions of the feature fusion network is replaced with Faster_C2f. Then, the impact of the improvement before and after and different improvement strategies on the detection effect of the model is compared and analyzed on the test set to explore the algorithm performance and evaluate the improvement strategy.
[0108] The defect detection method provided in this example implementation focuses on three aspects in its model structure design: lightweight feature extraction, enhanced feature fusion, and optimized convolutional operations. The model uses MobileViTv3 as its network backbone, lightweighting the feature extraction network and enhancing its ability to learn global information. An adaptive cross-level feature pyramid is designed to optimize the feature fusion network, constructing interactions between features at non-adjacent levels to compensate for the loss and degradation of feature information caused by the reduced feature extraction capability due to the lightweight backbone network. Optimized convolutional operations and a lightweight network structure improve computational efficiency, addressing the issues of large parameter count and slow inference speed in defect detection networks while maintaining accuracy. It should be noted that the above figures are merely illustrative of the processes included in the method according to an exemplary embodiment of the present invention and are not intended to be limiting. It is readily understood that the processes shown in the above figures do not indicate or limit the temporal order of these processes. Furthermore, it is readily understood that these processes may be executed synchronously or asynchronously in multiple modules, for example.
[0109] Further reference Figure 9 As shown, this example embodiment also provides a defect detection device 90, the device comprising:
[0110] The image acquisition module 401 is used to acquire the target detection image and input the target detection model into the lightweight detection model;
[0111] Feature extraction module 402 is used to perform feature extraction processing on the target detection image using a feature extraction network, and obtain a target feature map based on the feature map output by the target layer backbone network;
[0112] The feature fusion module 403 is used to perform feature fusion processing on target feature maps of different levels by combining the fusion space weight parameters corresponding to target feature maps of each scale obtained by adaptive learning, so as to obtain fused feature information.
[0113] The detection result output module 404 is used to detect the fused feature information using each detection head of the feature detection network to obtain defect detection results.
[0114] The specific details of each module in the aforementioned defect detection device have been described in detail in the corresponding defect detection methods, so they will not be repeated here.
[0115] It should be noted that although several modules or units for the device used to perform actions have been mentioned in the detailed description above, this division is not mandatory. In fact, according to embodiments of this disclosure, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.
[0116] Figure 10 A schematic diagram of an electronic device suitable for implementing embodiments of the present invention is shown.
[0117] It should be noted that, Figure 10 The illustrated electronic device 1000 is merely an example and should not be construed as limiting the functionality and scope of use of the embodiments disclosed herein.
[0118] like Figure 10As shown, the electronic device 1000 includes a Central Processing Unit (CPU) 1001, which can perform various appropriate actions and processes based on programs stored in Read-Only Memory (ROM) 1002 or programs loaded from Storage Unit 1008 into Random Access Memory (RAM) 1003. The RAM 1003 also stores various programs and data required for system operation. The CPU 1001, ROM 1002, and RAM 1003 are interconnected via a bus 1004. An Input / Output (I / O) interface 1005 is also connected to the bus 1004.
[0119] The following components are connected to I / O interface 1005: an input section 1006 including a keyboard, mouse, etc.; an output section 1007 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and speakers, etc.; a storage section 1008 including a hard disk, etc.; and a communication section 1009 including a network interface card such as a LAN (Local Area Network) card, modem, etc. The communication section 1009 performs communication processing via a network such as the Internet. A drive 1010 is also connected to I / O interface 1005 as needed. Removable media 1011, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., are installed on drive 1010 as needed so that computer programs read from them can be installed into storage section 1008 as needed.
[0120] In particular, according to embodiments of the present invention, the processes described below with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of the present invention include a computer program product comprising a computer program carried on a storage medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 1009, and / or installed from removable medium 1011. When the computer program is executed by central processing unit (CPU) 1001, it performs various functions defined in the system of this application.
[0121] Specifically, the aforementioned electronic devices can be smart mobile electronic devices such as mobile phones, tablets, or laptops. Alternatively, the aforementioned electronic devices can also be smart electronic devices such as desktop computers.
[0122] It should be noted that the storage medium shown in the embodiments of the present invention can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical fiber, portable compact disc read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In the present invention, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present invention, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, wherein computer-readable program code is carried. Such transmitted data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The computer-readable signal medium can also be any storage medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the storage medium can be transmitted using any suitable medium, including but not limited to wireless, wired, etc., or any suitable combination thereof.
[0123] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0124] The units described in the embodiments of the present invention can be implemented in software or hardware, and the described units can also be located in a processor. The names of these units do not necessarily limit the specific unit itself.
[0125] It should be noted that, as another aspect, this application also provides a storage medium, which may be included in an electronic device or may exist independently without being assembled into the electronic device. The aforementioned storage medium carries one or more programs, which, when executed by an electronic device, cause the electronic device to perform the methods described in the following embodiments. For example, the electronic device may perform... Figure 2 The steps shown.
[0126] Furthermore, the above figures are merely illustrative of the processes included in the method according to exemplary embodiments of the present invention, and are not intended to be limiting. It is readily understood that the processes shown in the above figures do not indicate or limit the temporal order of these processes. Additionally, it is readily understood that these processes may be executed synchronously or asynchronously, for example, in multiple modules.
[0127] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.
[0128] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and embodiments are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the claims.
[0129] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.
Claims
1. A defect detection method, characterized in that, The method includes: Acquire the target detection image and input the target detection image into the lightweight defect detection model; A lightweight feature extraction network is used to extract features from the target detection image, and the resulting target feature map is output. By using the improved feature fusion network and combining the fusion space weight parameters corresponding to the target feature maps at each scale obtained through adaptive learning, feature fusion processing is performed on the target feature maps output by the backbone network to obtain fused feature information. The fused feature information is detected using the detection heads of the feature detection network to obtain defect detection results.
2. The method according to claim 1, characterized in that, The feature extraction network is designed with a lightweight approach based on MobileViTv3, effectively combining the advantages of spatial inductive bias of convolutional neural networks and visual self-attention models in capturing long-distance dependencies. Meanwhile, in order to achieve multi-scale target detection, the feature maps extracted from the third, sixth, eighth and tenth layers of the backbone network are used as the target feature maps.
3. The method according to claim 1 or 2, characterized in that, The process of using a lightweight feature extraction network to extract features from the target detection image, obtaining local and global feature maps of the target based on the outputs of the backbone networks of different target layers, and achieving multi-scale feature extraction includes: InvertedResidual module: Extracts local feature information and performs downsampling. MobileViTv3 Block module: Captures and processes global information in images. The InvertedResidual module consists of an inverted residual structure and a linear bottleneck structure. It decomposes the standard convolution of the third layer information of the backbone network into depthwise convolution and 1×1 convolution to obtain the local feature map of the target. The process of the MobileViTv3 Block module is as follows: Local feature extraction is performed on the target detection image to obtain a first feature map embedded with local feature information; The first feature map is processed by vector encoding so that the pixels can acquire global information, and a second feature map with global representation is obtained. The first feature map and the second feature map are concatenated along the channel dimension and then fused through a 1×1 convolution to obtain a fused feature map. The fused feature map is combined with the target detection image to obtain a global feature map of the target.
4. The method according to claim 3, characterized in that, The step of extracting features from the target detection image to obtain a first feature map embedding local feature information includes: A 3×3 depth-separable convolutional layer is used to perform convolution processing on the target detection image to obtain the first feature information; A 1×1 convolutional layer is used to perform channel dimensionality reduction on the first feature information to generate a first feature map that embeds local feature information.
5. The method according to claim 3, characterized in that, The step of performing vector encoding processing on the first feature map to enable pixels to acquire global information and obtain a second feature map with global representation includes: The first feature map, which embeds local feature information, is subjected to data transformation processing and mapped to the corresponding first vector group. The first vector group is encoded using a Linear Transformer network, and a second feature map with global representation is extracted.
6. The method according to claim 1, characterized in that, Using a feature fusion network, and combining the fusion space weight parameters corresponding to the target feature maps at different scales obtained through adaptive learning, feature fusion processing is performed on target feature maps at different levels to obtain fused feature information, including: Based on the Adaptive Cross-Level Feature Pyramid (ACFPN), the Faster_C2f module is used to perform 2×2 transposed convolution upsampling and 2×2 standard convolution downsampling to unify the size of target feature maps at different levels. The fusion spatial weight parameters of feature maps at various scales are adaptively learned through 1×1 convolution and Softmax activation function. The weighted feature map is obtained by multiplying the feature maps of different levels with their corresponding spatial weights element by element. The fused feature information after adaptive spatial fusion is obtained by summing the weighted feature maps element by element.
7. The method according to claim 6, characterized in that, The Faster_C2f module has a lighter structure compared to the traditional C2f module; Specifically, the Faster Block from the Faster module is introduced into the C2f design to create the Faster_C2f module, which replaces the adaptive spatial fusion C2f module. The Faster module includes a partial convolution PConv computation module and a pointwise convolution PWConv computation module. The partial convolution PConv computation module is used to extract spatial features from a subset of channels using regular convolution operations, while the remaining channels are mapped identically. The pointwise convolution PWConv computation module performs pointwise convolution operations on all channels and fuses all channels to extract feature information.
8. A defect detection device, characterized in that, The device includes: The image acquisition module is used to acquire the target detection image and input the target detection model into the lightweight detection model; The feature extraction module is used to perform feature extraction processing on the target detection image using a feature extraction network, and to obtain the target feature map based on the feature map output by the target layer backbone network; The feature fusion module is used to perform feature fusion processing on target feature maps of different levels using the feature fusion network, combined with the fusion space weight parameters corresponding to the target feature maps of each scale obtained by adaptive learning, in order to obtain fused feature information. The detection result output module is used to detect the fused feature information using the detection heads of the feature detection network to obtain defect detection results.
9. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the defect detection method as described in any one of claims 1 to 7.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the defect detection method according to any one of claims 1 to 7.
Citation Information
Cited By
Image defect segmentation method, device and equipment and computer readable storage medium
CN121661065A