An Intelligent Defect Detection Method for the Appearance of Capacitors Based on the Fusion of YOLO-NAS and Transformer

By adopting the detection method of YOLO-NAS and Transformer in the intelligent defect detection of capacitor appearance, combining lightweight backbone network and multi-scale feature fusion, the problem of small-objective miss detection and multiple defect distinction is solved, and the calculation complexity of the model is reduced, achieving high-precision and real-time detection effects.

CN119762889BActive Publication Date: 2025-06-24CENT SOUTH UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411971122.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2025-06-24
Estimated Expiration
2044-12-30

AI Technical Summary

Technical Problem

The prior art has problems such as small-objective miss detection, difficulty in distinguishing multiple defects, and high model calculation complexity in the intelligent defect detection of capacitor appearance.

Method used

The detection method based on the fusion of YOLO-NAS and Transformer is adopted, combining lightweight backbone network, Transformer module, multi-scale feature fusion and improved loss function to solve the problems of small-objective miss detection, multiple defect distinction difficulties and high model calculation complexity.

Benefits of technology

It improves the accuracy and robustness of capacitor appearance defect detection, reduces the computational complexity, and adapts to real-time industrial scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119762889B_ABST
    Figure CN119762889B_ABST
Patent Text Reader

Abstract

The present invention discloses an intelligent defect detection method for the appearance of capacitors based on the fusion of YOLO-NAS and Transformer, belonging to the fields of artificial intelligence and industrial automation. It uses a lightweight backbone network combined with MobileViT, optimizes the feature extraction architecture through NAS to enhance the global feature expression ability, introduces Transformer-FPN to achieve multi-scale feature fusion, utilizes the high-order features of the teacher model to guide the learning of the student model through feature distillation, adopts an improved localization loss function to optimize the detection accuracy of small targets and irregular defects, combines Soft-NMS and DIoU-NMS, and reduces the computational complexity through model pruning and mixed quantization to achieve real-time detection and deployment on embedded devices. The present invention adopts the above-mentioned intelligent defect detection method for the appearance of capacitors based on the fusion of YOLO-NAS and Transformer, and solves the problems of missed detection of small targets, difficulty in distinguishing multiple defects, and high computational complexity of the model by combining a lightweight backbone network, a Transformer module, multi-scale feature fusion, and an improved loss function.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of artificial intelligence and industrial automation, and in particular to an intelligent defect detection method for the appearance of capacitors based on the fusion of YOLO-NAS and Transformer. Background Art

[0002] As a core component in electronic devices, the appearance quality of capacitors directly affects the performance and lifespan of the devices. Traditional manual inspection methods are difficult to meet the requirements of industrial production due to low efficiency, strong subjectivity, and susceptibility to fatigue. In recent years, object detection methods based on deep learning (such as the YOLO series) have demonstrated powerful performance in object detection tasks, but still face challenges in the following aspects:

[0003] (1) Limitations in small object detection: Defects such as cracks and surface marks are small in size and are easily missed.

[0004] (2) Similarity of multiple defects: Different types of defects are easily confused in complex backgrounds, such as cracks and surface marks.

[0005] (3) Requirement for model lightweight: Traditional YOLO models have a large computational volume and are difficult to meet the real-time detection requirements on embedded devices.

[0006] Therefore, a defect detection method that combines lightweight networks and Transformer characteristics is needed to improve detection accuracy, reduce computational complexity, and adapt to industrial real-time scenarios. Summary of the Invention

[0007] The purpose of the present invention is to provide an intelligent defect detection method for the appearance of capacitors based on the fusion of YOLO-NAS and Transformer, which solves the problems of missed detection of small objects, difficulty in distinguishing multiple defects, and high computational complexity of the model by combining a lightweight backbone network, a Transformer module, multi-scale feature fusion, and an improved loss function.

[0008] To achieve the above purpose, the present invention provides an intelligent defect detection method for the appearance of capacitors based on the fusion of YOLO-NAS and Transformer, including the following steps:

[0009] S1. Produce a dataset, including collecting capacitor appearance images and annotating the defects of the capacitors, and classifying and organizing the collected image data to form a dataset for model training and testing;

[0010] S2. Construct a network model based on YOLO-NAS and Transformer, where the network model based on YOLO-NAS and Transformer integrates a lightweight backbone network and a Transformer module for multi-scale feature fusion;

[0011] S3. Train the network model based on YOLO-NAS and Transformer, using the collected and labeled dataset, and accurately detect capacitor defects by iteratively optimizing the model parameters;

[0012] S4. Test the network model based on YOLO-NAS and Transformer, and verify the detection effect and practical application value of the model through an independent test dataset.

[0013] Preferably, a high-resolution industrial camera is used for data collection in S1 to collect capacitor defect images including cracks, surface marks, and deformations.

[0014] Preferably, in S1, professional personnel use labeling tools to accurately label capacitor defects to ensure that each defect type is correctly identified and classified.

[0015] Preferably, the defect types include four states: crack, surface mark, deformation, and no defect, and the image data in each state is classified and labeled.

[0016] Preferably, the network model based on YOLO-NAS and Transformer in S2 includes a lightweight Backbone, a Transformer-FPN module for multi-scale feature fusion, and a YOLO detection head.

[0017] Preferably, the lightweight Backbone adopts the GhostNet or ShuffleNetV2 structure and combines MobileViT blocks to enhance the feature extraction efficiency and the acquisition of global information.

[0018] Preferably, the Transformer-FPN module uses a multi-head attention mechanism to fuse features of different scales, improving the robustness of the model to defects of variable sizes.

[0019] Preferably, the multi-head attention mechanism includes applying linear projection and positional encoding to the feature maps of each scale to generate query Q, key K, and value V, and enhancing the feature fusion effect by calculating the cross-attention of different feature maps.

[0020] Preferably, the YOLO detection head is responsible for detecting each target, including anchor box localization, target category prediction, and confidence score.

[0021] Preferably, the training process of S3 includes the following steps:

[0022] S3.1. Optimize object detection using an improved loss function, including EIoU - type and SIoU - type losses;

[0023] S3.2. Automatically adjust the learning rate and momentum using an adaptive hyperparameter optimization tool;

[0024] S3.3. Apply feature distillation technology to guide the student model using the high - order features of the teacher model, improving the classification and localization capabilities of the lightweight model.

[0025] Therefore, the present invention adopts the above - mentioned intelligent defect detection method for the appearance of capacitors based on the fusion of YOLO - NAS and Transformer. By combining a lightweight backbone network, a Transformer module, multi - scale feature fusion, and an improved loss function, it solves the problems of missed detection of small targets, difficulty in distinguishing multiple defects, and high model computational complexity.

[0026] The technical solution of the present invention will be further described in detail below with reference to the drawings and embodiments. Description of the Drawings

[0027] Figure 1 is the overall system architecture flowchart of an embodiment of the intelligent defect detection method for the appearance of capacitors based on the fusion of YOLO - NAS and Transformer of the present invention;

[0028] Figure 2 is the structural schematic diagram of the MobileViT module of an embodiment of the intelligent defect detection method for the appearance of capacitors based on the fusion of YOLO - NAS and Transformer of the present invention;

[0029] Figure 3 is the multi - scale feature fusion schematic diagram of the Transformer - FPN module of an embodiment of the intelligent defect detection method for the appearance of capacitors based on the fusion of YOLO - NAS and Transformer of the present invention;

[0030] Figure 4 is the bar chart of the defect classification results of an embodiment of the intelligent defect detection method for the appearance of capacitors based on the fusion of YOLO - NAS and Transformer of the present invention;

[0031] Figure 5 is the trend chart of the change in accuracy and recall rate of an embodiment of the intelligent defect detection method for the appearance of capacitors based on the fusion of YOLO - NAS and Transformer of the present invention. Detailed Embodiments

[0032] The technical solution of the present invention will be further described below with reference to the accompanying drawings and embodiments.

[0033] Unless otherwise defined, the technical terms or scientific terms used in the present invention shall have the ordinary meanings understood by those of ordinary skill in the field to which the present invention pertains.

[0034] Embodiment 1

[0035] Based on the CSD-YOLO model, the present invention introduces a lightweight Backbone and a Transformer structure. Adopting the idea of automated architecture search (Neural Architecture Search) of YOLO-NAS, a lightweight Backbone (such as GhostNet or ShuffleNetV2) is selected, and MobileViT blocks are introduced into the Transformer structure to enhance the globality and efficiency of feature extraction. It mainly includes:

[0036] (1) Lightweight Backbone design: Use a lightweight Backbone (GhostNet or ShuffleNetV2) and MobileViT blocks to construct an efficient feature extraction backbone; use NAS automated search to determine the optimal structure.

[0037] (2) Multi-scale feature fusion: Replace the traditional FPN / PAN module with Transformer-FPN to enhance the expression ability of multi-scale features; use the multi-head attention mechanism in the Neck to efficiently fuse features of different scales, and improve the robustness to variable-size defects.

[0038] (3) Feature distillation: Use the high-order features in the teacher model to distill the student model, and enhance the feature identification ability of the lightweight model. Especially focus on distilling the feature dimensions of Crack and Surface Mark.

[0039] (4) Improve the loss function: Adopt losses such as EIoU, SIoU or alpha-DIoU to improve the localization accuracy.

[0040] (5) Adaptive hyperparameter optimization: Use automated tools to optimize hyperparameters such as the learning rate and momentum, and use Bayesian optimization or Optuna tools to automatically tune hyperparameters such as the learning rate, weight decay, and momentum.

[0041] (6) Post-processing improvement: Use Soft-NMS and DIoU-NMS strategies to optimize the detection results, and perform domain adaptation on data under different lighting and background conditions to improve generalization and robustness.

[0042] (7) Structural pruning and quantization: By pruning and quantizing the model, the computational load is reduced and the performance of embedded deployment is enhanced.

[0043] (8) Data augmentation and domain adaptation: Focus on data augmentation for Crack and Surface Mark class samples (such as morphology-based sample augmentation or synthetic mixed defect datasets), and weight the loss of these two types of samples during training to improve the model's sensitivity to them.

[0044] The following is the idea of automated architecture search (Neural Architecture Search, NAS) using YOLO-NAS, combined with a lightweight Backbone (such as GhostNet or ShuffleNetV2) and the model block diagram process design of introducing MobileViT blocks into the Transformer structure. The YOLO-NAS architecture flow diagram is as Figure 1 shown and explained as follows:

[0045] (1) Input data:

[0046] Input image size: For example, 640×640×3;

[0047] Data preprocessing: Includes image resizing, normalization, data augmentation, etc.

[0048] (2) Automated architecture search (NAS):

[0049] Search space: Define different candidate model architectures (including different Backbones, Transformer modules, etc.).

[0050] Architecture search: Search for the optimal network architecture through NAS algorithms (such as reinforcement learning, evolutionary algorithms, etc.).

[0051] (3) Backbone (lightweight network):

[0052] GhostNet / ShuffleNetV2: Select a lightweight convolutional neural network as the backbone network of the network. It is mainly used in the initial stage of feature extraction and has a low computational complexity.

[0053] GhostNet: Through efficient Ghost modules, rich features are extracted using a small number of convolutional operations.

[0054] ShuffleNetV2: By channel shuffling and depthwise convolution, the computational load is reduced while maintaining a high feature extraction ability.

[0055] (4) Introduction of MobileViT blocks:

[0056] Transformer Module: Introduce the Transformer structure on the basis of the YOLO model to capture long-range dependencies and global features.

[0057] MobileViT Block: MobileViT is a module that integrates Convolutional Neural Network (CNN) and Vision Transformer (ViT), mainly used to enhance the network's extraction of local features and global feature modeling. The MobileViT block combines the advantages of convolution and Transformer, which can not only maintain high efficiency but also improve the model's global feature modeling ability. In YOLO-NAS, the MobileViT module helps to further improve the model's feature extraction efficiency and detection accuracy.

[0058] (5)YOLO Head:

[0059] After feature extraction, each target is detected through the detection head of YOLO.

[0060] Anchor-based: Use anchor boxes for object detection.

[0061] Bounding Box: Regress the position and size of the target box.

[0062] Class Prediction: Predict the object class corresponding to each box.

[0063] Confidence Score: Predict the confidence level for each box.

[0064] (6)Loss Calculation:

[0065] Objective Function: Adopt the YOLO loss function, including coordinate loss, class loss, and confidence loss.

[0066] (7)Optimization and Training:

[0067] Optimizer: Use the Adam optimizer or other optimization algorithms to train the network.

[0068] Learning Rate Adjustment: Use a learning rate scheduler (such as CosineAnnealingLR, etc.) to dynamically adjust the learning rate.

[0069] (8)Output Results:

[0070] Predicted Boxes: Return the position, class, and confidence of each detection box.

[0071] Post-processing: Remove redundant boxes through NMS (Non-Maximum Suppression) to obtain the final detection results.

[0072] Among them, the role of automated architecture search (NAS) is to search for the optimal network architecture. In this process, YOLO-NAS will select a suitable Backbone according to different objectives and conditions and introduce the MobileViT block. The lightweight Backbone is responsible for extracting features from images, and both GhostNet and ShuffleNetV2 have low computational complexity and high accuracy. The role of the MobileViT block is to enhance the model's ability to capture global information through the Transformer structure and improve the model's performance in complex environments. The YOLO Head uses the YOLO architecture for object detection to complete localization, classification, and confidence prediction.

[0073] (1) MobileViT module

[0074] MobileViT combines local convolutional features (Conv) and global attention (Transformer). Its structural block diagram is as Figure 2 shown. Let the input feature map be x ∈ R H×W×C , and first obtain the feature X conv through depthwise separable convolution (Depthwise Separable Conv):

[0075] X conv = DWConv(X) (1);

[0076] where DWConv() represents depthwise separable convolution.

[0077] Then flatten the feature into a sequence form X seq and input it into the lightweight Transformer:

[0078] X seq = Reshape(X conv ) ∈ R N×D , N = H × W, D = C (2);

[0079] where the Reshape() function can be used to change the shape of a multi-dimensional array. N represents the length of the feature sequence, that is, the height of the feature map, the product of H and W. D represents the dimension of each sequence element, that is, the number of channels C of the feature map, H represents the height of the feature map, W represents the width of the feature map, and C represents the number of channels of the feature map.

[0080] In the Transformer, multi-head self-attention (MHSA) is used:

[0081] MHSA(X seq) = Concat(head1, … head m )W 0 (3);

[0082] Among them, MHSA(X seq ) represents processing the information in the X seq sequence, that is, the output after processing by the multi-head self-attention mechanism. Concat() represents concatenating the outputs of multiple heads. head1 represents the output of the i-th attention head, m represents the number of attention heads, and W 0 represents the weight matrix for linear transformation.

[0083]

[0084] Among them, Q, K, and V are the query, key, and value matrices obtained through linear transformation. Q i represents the query matrix of the i-th attention head, K i represents the key matrix of the i-th attention head, V i represents the value matrix of the i-th attention head. These three parameters are all obtained by their respective linear transformations. Generating d k represents the dimension of the key matrix, and i represents the serial number of the attention head. Softmax() can convert the output of the model into a probability distribution. Finally, transform the Transformer output back to the feature map size X out :

[0085] X out = Reshape(MHSA(X seq )) + X conv (5);

[0086] Through the above module, MobileViT obtains a global receptive field and enhances feature expression under lightweight conditions.

[0087] Replace the Backbone through NAS search: Select a suitable MobileViT module configuration (number of layers, number of channels, dilation rate, etc.) and GhostNet or ShuffleNetV2 structure through NAS search, and finally obtain a lightweight and high-performance Backbone.

[0088] This architecture makes full use of the automatic search feature of the NAS algorithm, combines a lightweight Backbone (GhostNet / ShuffleNetV2) and MobileViT blocks to improve the performance of the YOLO model in object detection, especially in global feature extraction and computational efficiency in complex scenarios.

[0089] (2) Transformer-FPN module for multi-scale feature fusion

[0090] Replace the existing PAN and FPN structures with an efficient Transformer-FPN in the Neck structure. By combining multi-scale features and the self-attention mechanism of the Transformer, improve the robustness of the model to various sizes and deformation defects. The structural block diagram is as shown in Figure 3 Figure []. Description of the Transformer-FPN algorithm: Given the multi-scale feature maps {F3, F4, F5} output by the Backbone (taking C3, C4, C5 as examples), use Transformer-FPN for fusion. For each scale feature, first obtain the query Q, key K, and value V through linear projection and positional encoding:

[0091] Q = W Q F i , K = W K F j , V = W V F j (6);

[0092] where W Q is the linear transformation matrix for generating the query matrix Q, F i is the feature map of the i-th scale, W K is the linear transformation matrix for generating the key matrix K, F j is the feature map of the j-th scale, and W V is the linear transformation matrix for generating the value matrix V.

[0093] During multi-scale feature interaction, selectively extract local information from high-resolution features and obtain global context from low-resolution features through Cross-Attention:

[0094]

[0095] where the Attention(Q, K, V) function represents the output of the attention mechanism, and d represents the dimension of the query and key matrices.

[0096] Repeat this process at multiple scales to achieve multi-level Transformer fusion, thereby obtaining enhanced features P3, P4, and P5. The final Transformer-FPN structure fuses these multi-scale features for subsequent detection heads.

[0097] (3) Structural Pruning and Channel Hybrid Quantization

[0098] Structural Pruning: Use L1 norm or weight pruning based on importance score to reduce redundant channels:

[0099] ImportanceScore(w) = ∥w∥1 (8);

[0100] where w is the weight parameter in the network, and ∥ ∥1 represents the L1 norm of the weights, which is used to measure the importance of the weights.

[0101] Channels with relatively small absolute values of weights are pruned to reduce the model complexity.

[0102] Channel Mixing Quantization: Quantize the activation values and weights using mixed precision (such as 8-bit or lower):

[0103]

[0104] where Q(x) is the quantized value, s is the quantization step (scale), x is the activation value or weight to be quantized, and the round function rounds the result to a given number of digits. The appropriate s is found by minimizing the quantization error. Perceptual quantization or optimizing quantization parameters can be used to maximize the accuracy while maintaining a low bit width.

[0105] (4) Improved localization loss function

[0106] Adopt improved IoU loss functions such as EIoU (Enhanced IoU) or SIoU (Smooth IoU) to optimize the localization of small objects and irregular objects. Use the Feature Distillation technique to distill the high-order features of the teacher model into the lightweight student model to improve the classification ability. In the IoU-based losses, introduce shape-sensitive constraints such as EIoU, SIoU, or alpha-DIoU to improve the localization accuracy of irregular shapes and small object defects. Taking EIoU as an example, its definition:

[0107]

[0108] where L EIoU represents the enhanced IoU loss function, IoU represents the Intersection over Union, ρ represents the distance between the center points of the predicted box b and the ground truth box b gt , b represents the current predicted box, c represents the diagonal length of the smallest enclosing box containing the predicted box and the ground truth box, α represents the weight coefficient used to balance the angular error, h represents the height of the predicted box, gt represents the correction coefficient for the predicted box, w represents the width of the predicted box, w gt represents the width of the ground truth box, h gt represents the height of the ground truth box.

[0109] Similar to other loss functions such as SIoU, it optimizes the positioning by jointly considering multiple factors such as the aspect ratio of the target box, the distance between the center points, and the overlap degree. alpha-DIoU introduces an exponential transformation on the basis of DIoU to accelerate the convergence in the high IoU range:

[0110]

[0111] Among them, L α-DIoU represents the improved loss function based on DIoU, which introduces an exponential transformation to accelerate the convergence in the high IoU range.

[0112] 5. Adaptive Hyperparameter Optimization (HPO)

[0113] Automated hyperparameter search is used to optimize the learning rate, momentum, weight decay, etc. Tools such as BayesianOptimization, Hyperband, or Optuna can be used to automatically find the optimal hyperparameters. Learning rate scheduling: CosineAnnealing or OneCycle is adopted:

[0114]

[0115] Among them, η t is the learning rate at step t, η max and η min are the upper and lower limits of the learning rate, T cur is the current step, and T max is the total number of steps.

[0116] 6. Multimodal and Feature Distillation

[0117] Multilevel and multimodal feature distillation is adopted. Channel-level distillation is performed on the feature maps F T and F S of the teacher model T and the student model S:

[0118]

[0119] Among them, L distill is the feature distillation loss, and P T (c) and P S (c) are the activation distributions of the teacher model and the student model for channel c, respectively. By minimizing the KL divergence, the feature distribution of the student model is made close to that of the teacher model, so as to maintain high accuracy under lightweight.

[0120] 7. Data Augmentation and Domain Adaptation

[0121] Data augmentation: Mixup, Cutmix, Mosaic-9, random brightness and contrast adjustment, and affine transformation are used to enhance data diversity.

[0122] Domain Adaptation: Use adversarial training or adaptive normalization layers on target domain images to enable the model to maintain stable detection under domain biases such as different lighting and backgrounds.

[0123] 8. Improved NMS Post-Processing

[0124] Replace the standard NMS with Soft-NMS and DIoU-NMS techniques to improve the suppression strategy of target boxes, reduce the probability of small targets being suppressed, increase the correct detection rate of adjacent targets, and improve both accuracy and recall. For example, in DIoU-NMS, the suppression weight considers the distance between the centers of the predicted box and the candidate box:

[0125]

[0126] where w is the suppression weight, b is the current predicted box, b other is the candidate box, c is the diagonal length of the smallest enclosing box containing the current predicted box and the candidate box, and d is the distance between the centers of the predicted box and the candidate box.

[0127] Perform weighted attenuation on the confidence instead of directly deleting.

[0128] 9. Multi-Scale Training and Testing Techniques

[0129] Multi-Scale Training: Randomly change the input resolution during training, randomly select from {640, 672, 704...} to improve the robustness of the model to defects of different sizes.

[0130] Multi-Scale Strategy during Testing: Try different scaling ratios during the testing phase and fuse the results to improve the mAP with a negligible increase in computational cost.

[0131] 10. Practical Deployment and Online Learning

[0132] Deploy to embedded platforms (such as Jetson Nano, Raspberry Pi 4), and measure the actual inference time and memory occupancy of the model. Further optimize according to the feedback. Online learning refers to performing lightweight updates using continuously collected data at the edge to adapt to the long-term changing defect types and distributions on the production line.

[0133] After using the above improvements, the FLOPs of the model are reduced by more than 50% compared to the original YOLOv7-Tiny, the inference latency is significantly reduced, and the mAP@0.5 on the CSD dataset reaches 98.9%. By adding channel-level feature distillation, the lightweight model can ensure a small number of parameters while achieving an accuracy close to or exceeding that of the teacher model. Multi-scale training and improved NMS further enhance the small object detection performance. Domain adaptation and data augmentation strategies improve the robustness of the model under different lighting and background conditions.

[0134] The above details the construction process of the appearance intelligent defect detection system from aspects such as algorithms, mathematical expressions, training strategies, and pseudocode implementation. By introducing the YOLO-NAS idea, the multi-scale fusion structure of MobileViT and Transformer-FPN, structure pruning and channel quantization, optimizing the localization loss function, adaptive hyperparameter search, multi-modal feature distillation, data augmentation and domain adaptation, improving the NMS post-processing, as well as multi-scale training and testing strategies, a defect detection model with high accuracy, lightweight, and high real-time performance is finally achieved, providing an effective solution for real-time defect detection deployment on embedded devices and industrial production lines. Through the present invention, high-precision detection of capacitor cracks, surface marks, deformations, and defect-free states can be realized. In industrial practical applications, the running time of the model on embedded devices is less than 30ms, and it shows extremely strong robustness under complex background and lighting conditions.

[0135] After 1000 training epochs, the training accuracy and recall rate of the model have reached 98.9%. The bar chart of the capacitor defect classification results is as Figure 4 shown, demonstrating the classification quantity distribution of cracks, surface marks, deformations, and no defects, with a significant improvement in the classification effect.

[0136] As Figure 5 shown, where Figure (a) represents the accuracy change trend during the training process, and Figure (b) represents the recall rate change trend during the training process; the accuracy and recall rate trend charts during training show that after expanding the training epochs, the accuracy and recall rate gradually approach 98.9%, indicating excellent performance.

[0137] In the actual measurement session of the present invention, the upgraded CSD-YOLO is deployed and tested on embedded devices such as Jetson Nano and Raspberry Pi 4, and further optimized according to the running time and memory occupancy. Combining the online learning mechanism on edge devices, lightweight updates are made according to the continuously collected defect data in the actual measurement scenario, enabling the model to maintain high accuracy, adaptability, and real-time performance during long-term operation.

[0138] Therefore, the present invention adopts the above-mentioned intelligent defect detection method for the appearance of capacitors based on the fusion of YOLO-NAS and Transformer. By utilizing a more advanced model structure, a multi-dimensional feature distillation strategy, a more efficient lightweight and hardware acceleration method, as well as a better loss function and post-processing strategy, it can achieve higher detection accuracy, faster response, and better online real-time performance on the basis of CSD-YOLO, providing a more reliable and efficient solution for the immediate automated detection of capacitor side defects in industrial production lines.

[0139] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that they can still modify or equivalently replace the technical solutions of the present invention, and these modifications or equivalent replacements do not cause the modified technical solutions to deviate from the spirit and scope of the technical solutions of the present invention.

Claims

1. A capacitor appearance intelligent defect detection method based on YOLO-NAS and Transformer fusion, characterized in that: The following steps are involved: S1. Create a data set, including collecting capacitor appearance images and marking capacitor defects, and classifying and arranging the collected image data to form a data set for model training and testing; S2. Construct a network model based on YOLO-NAS and Transformer, wherein the network model based on YOLO-NAS and Transformer integrates a lightweight backbone network and a Transformer module with multi-scale feature fusion; the network model based on YOLO-NAS and Transformer includes a lightweight Backbone, a Transformer-FPN module with multi-scale feature fusion, and a YOLO detection head; the lightweight Backbone adopts a GhostNet or ShuffleNetV2 structure, and is combined with a MobileViT block to enhance feature extraction efficiency and acquisition of global information; S3, training the network model based on YOLO-NAS and Transformer, using the collected and annotated data set, and iteratively optimizing the model parameters to accurately detect capacitor defects; S4. Test the network model based on YOLO-NAS and Transformer, and verify the detection effect and practical application value of the model through an independent test data set.

2. According to claim 1, a capacitor appearance intelligent defect detection method based on YOLO-NAS and Transformer fusion is characterized in that: The data acquisition in S1 uses a high-resolution industrial camera to collect capacitor defect images including cracks, surface marks, and deformations.

3. According to claim 1, the method for intelligent defect detection of capacitor appearance based on YOLO-NAS and Transformer fusion is characterized in that: The data annotation in S1 is performed by professionals using annotation tools to accurately annotate the capacitor defects, ensuring that each defect type is correctly identified and classified.

4. The method for intelligent defect detection of capacitor appearance based on YOLO-NAS and Transformer fusion according to claim 3 is characterized in that: The defect types include four states: crack, surface mark, deformation and no defect, and the image data in each state are classified and labeled.

5. The method for intelligent defect detection of capacitor appearance based on YOLO-NAS and Transformer fusion according to claim 1 is characterized in that: The Transformer-FPN module uses a multi-head attention mechanism to fuse features of different scales, thereby improving the robustness of the model to defects of variable sizes.

6. The method for intelligent defect detection of capacitor appearance based on YOLO-NAS and Transformer fusion according to claim 5 is characterized in that: The multi-head attention mechanism includes applying linear projection and position encoding to each scale feature map to generate query Q, key K and value V, and enhancing the feature fusion effect by calculating the cross attention of different feature maps.

7. The method for intelligent defect detection of capacitor appearance based on YOLO-NAS and Transformer fusion according to claim 1 is characterized in that: The YOLO detection head is responsible for detecting each target, including anchor box positioning, target category prediction and confidence scoring.

8. The method for intelligent defect detection of capacitor appearance based on YOLO-NAS and Transformer fusion according to claim 1 is characterized in that: The training process of S3 includes the following steps: S3.

1. Optimize object detection using improved loss functions, including EIoU and SIoU losses. S3.2, use adaptive hyperparameter optimization tools to automatically adjust learning rate and momentum; S3.

3. Apply feature distillation technology and use the high-order features of the teacher model to guide the student model, thereby improving the classification and positioning capabilities of the lightweight model.

Citation Information

Patent Citations

  • Method and system for detecting smoking in video based on improved network model

    CN117115700A

  • Cigarette appearance defect detection method based on variational bayesian inference

    WO2024159563A1