Light-weight-based tire flaw detection method

By introducing StarNet, C3k2-Star module and AFGC attention mechanism into the tire defect detection model, the YOLO11 framework is improved, and the accuracy and efficiency problems of the existing model in complex scenarios are solved, achieving high-precision and high-efficiency tire defect detection.

CN120182785APending Publication Date: 2025-06-20CHINA UNIV OF MINING & TECH +1
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510259816.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-06
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

When dealing with complex industrial inspection scenarios, existing tire defect detection models are prone to problems such as low accuracy and high missed detection rates, especially in the case of complex textures, irregular geometric shapes and diverse defect types on the tire surface.

Method used

Using a lightweight tire defect detection method, the YOLO11 framework is improved and the feature extraction and detection accuracy of the model is improved by introducing the StarNet network as the backbone network, combining the C3k2-Star module and the adaptive fine-grained channel attention mechanism AFGC.

Benefits of technology

It significantly improves the accuracy and efficiency of tire defect detection, reduces the missed detection rate, meets the high accuracy and efficiency requirements of industrial inspection, and reduces the computational complexity, making it suitable for deployment on resource-constrained equipment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182785A_ABST
    Figure CN120182785A_ABST
Patent Text Reader

Abstract

The invention discloses a tire flaw detection method based on light weight, and belongs to the technical field of industrial flaw detection, and the method comprises the following steps: S1, data acquisition and preprocessing: employing a camera for shooting, and taking a plurality of photos as a model for training; accurately marking the tire x-ray image by adopting a detection marking tool labmeli, wherein the marked content comprises the bounding box information of cord thread openings, foreign matters in the tire, cracks and flaw targets; s2, a target detection network is constructed and trained, StarNet is adopted as a backbone network, meanwhile, a C3k2-Star module is introduced, and then an adaptive fine-grained channel attention mechanism AFGC is introduced; and S3, generating an image detection frame. The tire flaw detection method based on light weight can adapt to the characteristics of tire detection scenes, the detection accuracy and efficiency are improved, meanwhile, the omission ratio is reduced, and the actual production requirement is met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of industrial defect detection, and in particular to a lightweight-based tire defect detection method. Background Art

[0002] Computer vision object detection aims to identify and locate target objects existing in images, which belongs to a classic task in the field of computer vision and has wide applications in fields such as industrial intelligence and automated detection. Especially in industrial product quality control, object detection technology can significantly improve the detection efficiency and accuracy, and gradually replace the traditional manual detection method. With the rapid development of deep learning technology, object detection tasks have been extended to multiple fields, gradually solving the problems of low efficiency, poor accuracy, time-consuming and laborious of the manual detection method.

[0003] In recent years, there have been multiple application scenarios of dense and complex object detection in industrial scenarios, such as defect detection of electronic components, detection of printing defects in product packaging, and detection of defects such as cracks and bubbles on the tire surface. These scenarios are usually accompanied by the following challenges: the targets are mutually occluded, the target textures are complex and the background interference is significant, resulting in a reduced resolvability of the target objects. When existing single-stage object detection models (such as YOLO, SDD) and two-stage detection models (such as FastR-CNN, FasterR-CNN) are used to process such complex industrial detection scenarios, problems such as low accuracy and high miss detection rate are likely to occur.

[0004] In view of the above problems, the existing Chinese patent publication number CN114511521B discloses a tire defect detection method based on multiple representations and multiple sub-domain adaptions. In this method, during the model training stage, the tire X-ray image is cropped and then input into a feature extraction module composed of a convolutional neural network. Further, a hedging feature representation module is used to extract multiple feature representations in the image, and based on this, multiple sub-domain adaption optimization is performed. Specifically, this method effectively overcomes the domain shift problem between images collected by different X-ray machines by minimizing the distance between the same-class sub-domains of the source domain and the target domain and maximizing the distance between different-class sub-domains, and finally obtains the detection result through a defect detection classifier. However, this method relies on a complex sub-domain adaption strategy and has a high computational cost. When facing scenarios with various complex backgrounds or diverse target shapes, the generalization ability of the model still has certain limitations, which may lead to a decrease in detection efficiency and an increase in the miss detection rate.

[0005] In tire defect detection, due to the complex texture structure, irregular geometric shape and diverse defect types (such as cracks, bulges, bubbles, etc.) on the tire surface, traditional object detection models are difficult to meet the high-precision and high-efficiency requirements of industrial detection. Summary of the Invention

[0006] The object of the present invention is to provide a lightweight-based tire defect detection method, which can adapt to the characteristics of the tire detection scenario, improve the accuracy and efficiency of detection, reduce the missed detection rate at the same time, and meet the actual production requirements.

[0007] To achieve the above object, the present invention provides a lightweight-based tire defect detection method, including the following steps:

[0008] Step S1, data collection and preprocessing, using a camera to take pictures, and taking several images as a data set for model training;

[0009] Using the detection annotation tool labelimg to accurately annotate the tire X-ray images, and the annotation content includes the boundary box information of the cord opening, foreign objects inside the tire, cracks and defect targets;

[0010] Step S2, target detection network construction and training, using StarNet as the backbone network, introducing the C3k2-Star module at the same time, and then introducing the adaptive fine-grained channel attention mechanism AFGC to improve the YOLO11 framework;

[0011] Step S3, generating an image detection frame.

[0012] Preferably, in step S1, the data set is enhanced, including rotating the image angle; at the same time, the original image with a resolution of 2688×2688 is cropped into sub-images with a size of 1024×1024.

[0013] Preferably, before training the data set in step S1, the self-built data set is randomly divided into a training set, a validation set and a test set according to a ratio of 7:2:1.

[0014] Preferably, the star operation in StarNet in step S2 is specifically:

[0015]

[0016] Among them, W1 and W2 are two different weight matrices (or filters), X is the input feature map, and T represents the transpose of the matrix. That is, the input feature map is linearly transformed by two different weight matrices, and then the transformed feature maps are multiplied element by element to capture the interaction information between the features.

[0017] Preferably, the C3k2-Star module adopts a feature segmentation and splicing strategy, divides the input features into two parts, one part is passed through a conventional convolution operation, and the other part is subjected to deep feature extraction through multiple StarBlock layers or bottleneck structures, and finally the two parts of the features are spliced and fused through a 1x1 convolution.

[0018] Preferably, the adaptive fine-grained channel attention mechanism AFGC first extracts global information from the feature map of each channel through global average pooling GAP to form a channel descriptor D:

[0019]

[0020] where F represents the input feature map; C represents the number of channels; H and W represent the height and width of the feature map respectively; F c (i, j) represents the value of the c-th channel at position (i, j);

[0021] Then, the adaptive fine-grained channel attention mechanism AFGC uses a banded matrix B and a diagonal matrix A to capture local and global channel interaction information respectively. AFGC measures the correlation between global and local information through a cross-correlation operation to generate a correlation matrix C, and its formula is as follows:

[0022]

[0023] where ⊙ represents element-wise multiplication; B represents the banded matrix for capturing local information; A represents the diagonal matrix for capturing global information; * is the cross-correlation operation; D local represents local channel information, captured by the banded matrix, for capturing the interaction relationship between local channels; D glocal represents global channel information, captured by the diagonal matrix, for capturing the dependency relationship between global channels;

[0024] Next, the adaptive fine-grained channel attention mechanism AFGC combines global and local weight vectors and dynamically adjusts the channel weights of the feature map through an adaptive fusion strategy:

[0025] W = σ(C) ⊙ (softmax(D global ) + softmax(D local ));

[0026] where σ represents the sigmod activation function, used to generate the weight of each channel; the softmax function is used to normalize the weight vector;

[0027] Finally, the adaptive fine-grained channel attention mechanism AFGC applies the adjusted weight W to the original feature map F to enhance the extraction of relevant features F enhanced :

[0028] F enhanced = F ⊙ W.

[0029] Preferably, in step S3, based on step S2, detection boxes are generated through test images, and finally the detection results are displayed.

[0030] Therefore, the present invention adopts the above-mentioned lightweight-based tire defect detection method, improves the YOLO11 model, introduces the StarNet network model as the backbone network, efficiently processes high-dimensional features through star operations, and improves the C3k2 module to optimize the feature extraction process, significantly reducing the number of model parameters while reducing redundant calculations. Further, the improved model introduces an adaptive fine-grained channel attention (AFGCattention) mechanism in the head part of YOLO11, improving the performance and accuracy of the model through more refined feature selection and weight allocation. At the same time, the computational complexity is reduced while maintaining the improvement of detection accuracy, thus effectively meeting the high-precision and high-efficiency detection requirements in complex scenarios.

[0031] The following will further describe the technical solutions of the present invention in detail through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] Figure 1 is a technical block diagram of an embodiment of a lightweight-based tire defect detection method of the present invention;

[0033] Figure 2 is a structural diagram of the StarNet network of an embodiment of a lightweight-based tire defect detection method of the present invention;

[0034] Figure 3 is a structural diagram of the C3k2-Star module network of an embodiment of a lightweight-based tire defect detection method of the present invention;

[0035] Figure 4 is a structural diagram of the adaptive fine-grained channel attention mechanism AFGC network of an embodiment of a lightweight-based tire defect detection method of the present invention;

[0036] Figure 5 is an improved model diagram of YOLO11 of an embodiment of a lightweight-based tire defect detection method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0037] The following further illustrates the technical solutions of the present invention through the accompanying drawings and embodiments.

[0038] Unless otherwise defined, the technical terms or scientific terms used in the present invention shall have the ordinary meanings understood by those of ordinary skill in the art to which the present invention belongs.

[0039] Embodiment 1

[0040] As Figure 1 shown, the present invention provides a lightweight-based tire defect detection method, including the following steps:

[0041] Step S1: Data acquisition and preprocessing. Use a camera to take pictures, and select several images as a dataset for model training;

[0042] Use the detection annotation tool labelimg to accurately annotate the tire X-ray images. The annotation content includes the bounding box information of cord openings, foreign objects inside the tire, cracks, and defect targets;

[0043] To further improve the detection accuracy, various enhancement processes are performed on the dataset, including methods such as rotation angle to expand the diversity of samples, thereby enhancing the model's adaptability to targets with different orientations. At the same time, the original image with a resolution of 2688×2688 is cropped into sub-images with a size of 1024×1024.

[0044] Before training the dataset in Step S1, the self-built dataset is randomly divided into a training set, a validation set, and a test set in a ratio of 7:2:1.

[0045] Step S2: Target detection network construction and training. Use StarNet as the backbone network, introduce the C3k2-Star module, and then introduce the adaptive fine-grained channel attention mechanism AFGC to improve the yolo11 framework. The improved model diagram of YOLO11 is as Figure 5 shown;

[0046] StarNet is an efficient convolutional neural network that inherits the advantages of traditional convolutional neural networks and significantly enhances the high-dimensionality and non-linearity of feature representation through the innovative "star operation". Its architecture consists of a basic convolutional layer and a StarBlock that integrates the "star operation".

[0047] The "star operation" is a feature mapping method based on element-wise multiplication that can map image features to an implicit high-dimensional non-linear space, thereby significantly enhancing the feature expressiveness. This operation can achieve efficient feature extraction and fusion without increasing the network width, enabling the model to have stronger pattern expression ability while maintaining a compact structure.

[0048] The core advantage of StarNet lies in the implicit construction of a high-dimensional feature space through simple element-wise multiplication, which not only increases the feature dimension but also significantly improves the representation ability of complex features without significantly increasing the computational complexity. This design makes StarNet perform excellently in processing complex scenarios, can run efficiently, and is suitable for low-latency inference under limited computing resources to meet the needs of real-time applications.

[0049] After replacing the backbone network of YOLO11 with StarNet, the model is made lightweight. This is particularly important for deploying the model on resource-constrained embedded devices, as these devices usually cannot handle overly complex computational tasks. By optimizing the network structure, the total number of model parameters is significantly reduced, and the computational complexity is lowered. Secondly, its feature extraction ability is enhanced. The efficient feature expression ability of StarNet ensures the detection accuracy of the model in complex scenarios. Even when faced with small targets with variable directions and complex backgrounds, it can still achieve excellent performance. The StarNet network model structure is as Figure 2 shown.

[0050] The C3K2 module is a key feature extraction component in the YOLO11 model and is an improvement based on the traditional C3 module. It enhances the feature extraction ability by introducing variable convolution kernels (such as 3x3, 5x5, etc.) and channel separation strategies, and is particularly suitable for complex scenarios and deep feature extraction tasks. The structural characteristics of the C3K2 module include the variable convolution kernel design, which dynamically adjusts the receptive field through convolution kernels of different sizes and can adapt to object detection at different scales, especially performing outstandingly in scenarios with complex backgrounds or large variations in object sizes.

[0051] The C3K2 module adopts a feature segmentation and splicing strategy, dividing the input features into two parts. One part is passed through conventional convolution operations, and the other part undergoes deep feature extraction through multiple StarBlock layers or bottleneck structures. Finally, the two parts of the features are spliced and fused through 1x1 convolution. This design maintains the lightweight of the network while enhancing the deep feature extraction ability. In terms of enhancing feature extraction, the C3K2 module significantly improves the detection accuracy in object boundaries and complex backgrounds through multi-scale convolution kernels.

[0052] However, the C3K2 module also has some disadvantages. First, although multi-scale convolution kernels are used to increase the receptive field, this design may lead to an increase in computational volume, especially when processing high-resolution images. Secondly, although the operations of feature segmentation and splicing enhance the deep feature extraction ability, they may also bring certain computational redundancy. Especially when splicing low-level features and high-level features, it is easy to introduce redundant information, thus affecting the inference speed.

[0053] In addition, the C3K2 module may show relatively low accuracy in some small target scenarios (such as detecting bubble defects inside tires), because it mainly captures features through a large receptive field, and the detailed information of small objects may be ignored.

[0054] Therefore, the C3k2-Star module is introduced in YOLO11. By replacing the traditional C3K layer with the StarBlock layer, redundant calculations are effectively reduced while maintaining a strong feature extraction ability. This improvement not only enhances the object detection performance and accuracy of the model but also strengthens the feature fusion ability, enabling the model to better capture complex features in images. The introduction of the C3k2-Star module also helps to accelerate the training and inference speeds of the model, which is particularly important for application scenarios that require quick responses.

[0055] In addition, due to the advantages of the C3k2-Star module in reducing the number of parameters and computational load, it makes the model more lightweight and suitable for deployment on resource-constrained devices. These characteristics of the C3k2-Star module work together to improve the model's performance in multi-scale object detection tasks, especially its ability in real-time precise detection, providing strong support for various practical applications. The network structure diagram of the C3k2-Star module is as Figure 3 shown.

[0056] The specific star operation in the StarNet of step S2 is as follows:

[0057]

[0058] Among them, W1 and W2 are two different weight matrices (or filters), X is the input feature map, and T represents the transpose of the matrix. That is, the input feature map is linearly transformed by two different weight matrices, and then the transformed feature maps are multiplied element by element to capture the interaction information between features.

[0059] The advantages of the Adaptive Fine-Grained Channel Attention mechanism (AFGC) in the image dehazing task are reflected in multiple aspects, significantly improving the dehazing performance and optimizing the model efficiency. Through fine-grained feature selection and effective interaction between global and local information, the feature representation ability of the network is enhanced, helping the model to better distinguish and restore details and textures in the image.

[0060] Secondly, AFGC realizes adaptive feature weight allocation, dynamically adjusts the weights of each feature map according to the image content, effectively highlights important features and suppresses interference information, thereby improving the dehazing effect. Through fine-grained feature processing and adaptive weight allocation, AFGC greatly improves the quality of dehazed images and enhances the robustness of the model in complex environments.

[0061] Despite providing powerful feature processing capabilities, the AFGC design still focuses on keeping the model lightweight to ensure its efficient operation on resource-constrained devices. In addition, AFGC optimizes the computational efficiency and further improves the response speed in real-time tasks by reducing redundant calculations. AFGC also has strong adaptability and can handle images with different fog densities, non-uniform haze, and difficult-to-distinguish targets, so it can maintain excellent performance in various environments.

[0062] The Adaptive Fine-Grained Channel Attention Mechanism (AFGC) first extracts global information from the feature map of each channel through Global Average Pooling (GAP) to form a channel descriptor D:

[0063]

[0064] where F represents the input feature map; C represents the number of channels; H and W represent the height and width of the feature map respectively; F c (i, j) represents the value of the c-th channel at position (i, j);

[0065] Then, the AFGC uses a banded matrix B and a diagonal matrix A to capture local and global channel interaction information respectively. AFGC measures the correlation between global and local information through a cross-correlation operation to generate a correlation matrix C, and its formula is as follows:

[0066]

[0067] where ⊙ represents element-wise multiplication; B represents the banded matrix used to capture local information; A represents the diagonal matrix used to capture global information; * is the cross-correlation operation; D local represents local channel information, captured by the banded matrix, for capturing the interaction relationship between local channels; D glocal represents global channel information, captured by the diagonal matrix, for capturing the dependency relationship between global channels.

[0068] Next, the AFGC combines the global and local weight vectors and dynamically adjusts the channel weights of the feature map through an adaptive fusion strategy:

[0069] W = σ(C) ⊙ (softmax(D global ) + softmax(D local ));

[0070] where σ represents the sigmod activation function, used to generate the weight of each channel; the softmax function is used to normalize the weight vector;

[0071] Finally, the Adaptive Fine-Grained Channel Attention mechanism AFGC applies the adjusted weight W to the original feature map F to enhance the extraction of relevant features F enhanced :

[0072] F enhanced = F ⊙ W.

[0073] The AFGC mechanism realizes the effective weighting and optimization of the feature map through these key steps, thereby improving the performance and effect of the dehazing network. Its model structure diagram is as Figure 4 shown

[0074] Step S3: Generate image detection boxes.

[0075] Based on Step S2, detection boxes are generated through the test image, and finally the detection results are displayed.

[0076] The present invention first preprocesses the data: performs data augmentation rotation and other methods, and then obtains appropriate training data, improving the generalization ability of the network model. In terms of model improvement, StarNet is first used as the backbone network. This network maps image features to a high-dimensional non-linear space through simple element multiplication, improving the efficiency of feature extraction and fusion.

[0077] Secondly, in terms of feature fusion, the C3k2-Star module is introduced to replace the traditional C3K2 module. The C3K2-Star module enhances the ability to extract deep features in complex scenes through the combination of variable convolution kernels and the design of the StarBlock layer. Compared with the traditional C3K module, C3K2-Star improves the feature expression ability while reducing redundant calculations and the number of parameters, thus maintaining the efficiency of the network. This module can effectively capture complex features in images through multi-scale receptive fields and feature segmentation and stitching strategies.

[0078] Finally, the AFGC attention mechanism is introduced behind the feature extraction module C3k2_star. By connecting the attention mechanism at the back, the model can focus more on the features of small targets and avoid the interference of large targets. This makes the details of small targets more prominent, thereby improving the accuracy of small target detection.

[0079] Combining the advantages of the above methods and applying them to object detection can reduce the difficulty and cost of model training, significantly reduce the number of parameters and computational complexity of the model, and maintain a lightweight design. This enables the model to run efficiently on resource-constrained devices and accelerates the inference speed, making it suitable for real-time application scenarios.

[0080] Therefore, the present invention adopts the above-mentioned lightweight-based tire defect detection method, which can adapt to the characteristics of the tire detection scenario, improve the accuracy and efficiency of detection, reduce the missed detection rate at the same time, and meet the actual production requirements.

[0081] It should be noted that the content not elaborated in detail in the present invention is all prior art and well-known to those skilled in the art.

[0082] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that they can still modify or equivalently replace the technical solutions of the present invention, and these modifications or equivalent replacements cannot make the modified technical solutions deviate from the spirit and scope of the technical solutions of the present invention.

Claims

1. A lightweight tire defect detection method, characterized in that: The following steps are involved: Step S1, data collection and preprocessing, using a camera to take pictures and take several images as a data set for model training; The detection and annotation tool labelimg is used to accurately annotate tire X-ray images, including the bounding box information of cord openings, foreign objects in the tire, cracks, and defect targets; Step S2: Build and train the target detection network. StarNet is used as the backbone network. The C3k2-Star module is introduced. Then the adaptive fine-grained channel attention mechanism AFGC is introduced to improve the yolo11 framework. Step S3: Generate an image detection frame.

2. A lightweight tire defect detection method according to claim 1, characterized in that: In step S1, the data set is enhanced, including rotating the image angle; at the same time, the original image with a resolution of 2688×2688 is cropped into a sub-image with a size of 1024×1024.

3. The lightweight tire defect detection method according to claim 2, characterized in that: In step S1, before training the data set, the self-built data set is randomly divided into a training set, a validation set, and a test set in a ratio of 7:2:

1.

4. The lightweight tire defect detection method according to claim 3, characterized in that: The star operation in step S2StarNet is as follows: Among them, W1 and W2 are two different weight matrices, X is the input feature map, and T represents the transpose of the matrix.

5. The lightweight tire defect detection method according to claim 4, characterized in that: The C3k2-Star module adopts a feature segmentation and splicing strategy to divide the input features into two parts. One part is passed through conventional convolution operations, and the other part is deep feature extracted through multiple StarBlock layers or bottleneck structures. Finally, the two parts of the features are spliced ​​and fused through 1x1 convolution.

6. The lightweight tire defect detection method according to claim 5, characterized in that: The adaptive fine-grained channel attention mechanism AFGC first extracts global information from the feature map of each channel through global average pooling GAP to form a channel descriptor D: Among them, F represents the input feature map; C represents the number of channels; H and W represent the height and width of the feature map respectively; F c (i, j) represents the value of the cth channel at position (i, j); Then the adaptive fine-grained channel attention mechanism AFGC uses the band matrix B and the diagonal matrix A to capture the local and global channel interaction information respectively. AFGC measures the correlation between global and local information through cross-correlation operations and generates a correlation matrix C, whose formula is as follows: Among them, ⊙ represents element-wise multiplication; B represents a band matrix used to capture local information; A represents a diagonal matrix used to capture global information; ★ is a cross-correlation operation; D local Represents local channel information; D glocal Represents global channel information; Then the adaptive fine-grained channel attention mechanism AFGC combines the global and local weight vectors and dynamically adjusts the channel weights of the feature map through an adaptive fusion strategy: W=σ(C)⊙(softmax(D global )+softmax(D local )); Among them, σ represents the sigmoid activation function, which is used to generate the weight of each channel; the softmax function is used to normalize the weight vector; Finally, the adaptive fine-grained channel attention mechanism AFGC applies the adjusted weights W to the original feature map F to enhance the extraction of relevant features F enhanced : F enhanced =F⊙W。 7. The lightweight tire defect detection method according to claim 6, characterized in that: In step S3, a detection frame is generated based on step S2 and the test image, and the detection result is finally displayed.

Citation Information

Patent Citations

  • Tire defect detection method based on multiple representations and multiple subdomains adaptation

    CN114511521B

  • Tire X-ray image impurity defect detection method

    CN113066047A

  • Insulator detection method based on target detection algorithm and attention mechanism

    CN116895030A

  • Lightweight parking detection method based on multi-scale attention mechanism

    CN119314141A

  • Lightweight fabric surface defect detection method and device

    CN119399093A