Lightweight bearing surface defect small target detection method based on improved YOLOv8

By improving the YOLOv8 network, combining dynamic detection heads and ADown modules, the feature expression and information fusion are enhanced, and the accuracy and model scale in bearing surface defect detection are solved, achieving efficient and lightweight detection effects, suitable for industrial real-time detection.

CN120235853APending Publication Date: 2025-07-01EAST CHINA UNIV OF TECH
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510658950.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-21
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

The existing bearing surface defect detection methods have shortcomings in terms of accuracy, model scale and small object detection capabilities, especially in complex contexts, the detection performance of small object defects is not ideal, and the equipment is expensive or depends on professional qualities.

Method used

Using an improved YOLOv8 network, a dynamic detection head with scale perception, spatial perception and task perception is introduced, an ADown module is introduced for downsampling, and a small object detection layer is added to the neck of the network to enhance feature expression and information fusion, and reduce the amount of model parameters and calculation complexity.

Benefits of technology

It improves the detection accuracy and recall of small defects on the bearing surface, reduces the model scale and calculation complexity, and is suitable for lightweight equipment with limited resources to meet the needs of real-time industrial inspection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120235853A_ABST
    Figure CN120235853A_ABST
Patent Text Reader

Abstract

The invention provides a lightweight bearing surface defect small target detection method based on improved YOLOv8. According to the method, based on a YOLOv8n network architecture, a dynamic detection head is constructed, an ADown module is introduced, a small target detection layer is additionally arranged, and finally an improved YOLOv8n-DAS model is established; the method comprises the following steps: firstly, constructing a dynamic detection head, fusing scale, space and task perception attention mechanisms, intensifying feature expression ability in all directions, and accurately capturing complex tiny defect features on the surface of a bearing; secondly, the ADown module highlights edge defect information, simplifies the model parameter scale and reduces consumption of computing resources by combining traditional convolution, average pooling and maximum pooling operations; finally, small target detection layers are arranged at the neck and the head of the network, the small target detection layers are added, an extra small target detection head is arranged, shallow details and deep semantic information are deeply integrated, key details are reserved to the maximum extent, and the recognition capacity of small targets is improved. According to the method, the detection performance of the model on irregular and tiny defects and the detection capability of the model on small target defects are remarkably improved, the parameter quantity and the calculation complexity of the model are reduced, and deployment and application on lightweight equipment with limited resources are facilitated. And the method can be expanded to defect detection of industrial parts such as gears and blades by replacing training data, has the characteristics of light weight and high universality, and meets multi-scene deployment requirements.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology in industrial defect detection, target recognition, image feature extraction and characterization, and particularly relates to a lightweight small target detection method for bearing surface defects based on improved YOLOv8. Background Art

[0002] The background art of the present invention is as follows:

[0003] Traditional bearing defect detection methods are mainly divided into two categories: manual visual inspection and equipment inspection. Manual visual inspection has low efficiency, and due to the physiological limitations of the human eye and visual fatigue problems, it is extremely easy to miss small target defects. For equipment inspection methods, such as optical instruments, ultrasonic detection, and infrared thermal imaging, although the detection accuracy is improved, these devices are usually expensive, require high professional qualities of operators, and have limited detection capabilities for small target defects in complex backgrounds.

[0004] In recent years, the rapid development of deep learning technology has brought new breakthroughs to the field of industrial defect detection. Detection algorithms based on deep learning are mainly divided into two-stage detection algorithms and one-stage detection algorithms. Two-stage detection algorithms, represented by R-CNN, Fast RCNN, and Faster R-CNN, generate candidate bounding boxes through sub-networks, with high accuracy but slow speed. While one-stage detection algorithms such as the YOLO series of algorithms directly generate candidate bounding boxes from feature maps, having great advantages in execution speed, being more suitable for industrial real-time detection scenarios, and their detection accuracy has reached or even exceeded the level of two-stage algorithms in continuous optimization.

[0005] Although existing detection models and improvement methods have achieved good detection effects on their respective detection objects, there are still some problems such as insufficient accuracy of the model, large model volume caused by excessive number of model parameters and computational complexity, and unsatisfactory detection performance for small target defects in complex backgrounds. In view of these problems, the present invention proposes a lightweight small target detection method for bearing surface defects based on improved YOLOv8, which improves the detection performance of the model for irregular and tiny defects and the detection ability for small target defects, reduces the scale of the model, is conducive to deployment and application on lightweight devices with limited resources, and meets the real-time detection requirements of industrial scenarios. Summary of the Invention

[0006] Objective of the Invention: Aiming at the deficiencies of the existing bearing surface defect detection methods in terms of accuracy, model scale, and small target detection ability, the present invention proposes a lightweight small target detection method for bearing surface defects based on improved YOLOv8. By combining a dynamic detection head that perceives scale, space, and tasks, the feature expression ability is enhanced; the ADown module is used to replace the traditional convolution for downsampling, highlighting the edge defect features while reducing the model scale; a small target detection layer is added to the network neck to enhance the fusion of shallow and deep semantic information, retain more detailed information, and improve the detection accuracy of small target defects. Ensure that while maintaining light weight and real-time performance, the detection accuracy, recall rate, and average precision are improved to meet the requirements of industrial real-time detection.

[0007] Technical Solution: To achieve the above objective, the technical solution adopted by the present invention is a lightweight small target detection method for bearing surface defects based on improved YOLOv8, including the following steps:

[0008] First, use the LableImg software to annotate the data set, including three common defects: grooves, scratches, and abrasions. To enhance the generalization ability of the model and reduce the risk of overfitting, a series of augmentation operations such as increasing and decreasing brightness, Gaussian blur, and Gaussian noise are performed on the data set. The augmented data set is randomly divided into a training set, a validation set, and a test set according to the ratio of 8:1:1 to meet the different needs of model training, validation, and testing.

[0009] Second, use YOLOv8n as the baseline model and introduce the DyHead Block to design a dynamic detection head, which integrates scale perception, space perception, and task perception attention mechanisms, enabling the detection head to dynamically adjust the focus of attention according to different scale, position, and task features, thereby enhancing the feature expression ability for tiny defects on the bearing surface.

[0010] Third, use the ADown module to replace the traditional convolution for downsampling operations. The ADown module cleverly combines average pooling, traditional convolution, and max pooling to highlight the edge defect features, reduce the number of model parameters and computational complexity, and achieve model lightweight.

[0011] Fourth, add a small target detection layer to the network neck, focusing on capturing and retaining the feature information of small target defects. By fusing shallow detail features and deep semantic features, the problem of easy loss of small target defect feature information during network propagation is solved, and the detection accuracy of small target defects is improved.

[0012] Fifth, start the model training process, set the training cycle, select the optimization algorithm, determine the initial learning rate, and configure the momentum.

[0013] VI. Precision, Recall, and mean Average Precision (mAP) are used as the main evaluation metrics. At the same time, the number of parameters, weight size, and Frames Per Second (FPS) are considered to comprehensively evaluate the detection accuracy, efficiency, and resource consumption of the model. Based on these metrics, the performance of the model in the bearing surface defect detection task is objectively measured to verify its effectiveness and practicality.

[0014] Beneficial effects:

[0015] In response to the challenge of detecting tiny defects on the bearing surface, the present invention proposes a lightweight small target detection method for bearing surface defects based on improved YOLOv8. The main beneficial effects are as follows: First, by introducing a dynamic detection head and an ADown module, the feature expression ability and the edge feature extraction efficiency are enhanced respectively, effectively solving the problem of easy omission of small target defects; Second, a small target detection layer is used to strengthen the fusion of shallow and deep semantic information, further improving the recognition ability of small-sized defects; Third, by combining the attention mechanism and a lightweight feature extraction network, the detection accuracy and the positioning accuracy are improved, and the number of model parameters and the computational complexity are reduced, which is beneficial to deployment and application on lightweight devices with limited resources and meets the real-time detection requirements of industrial scenarios. Description of the Drawings

[0016] Figure 1 is the implementation flowchart of the present invention;

[0017] Figure 2 is the model structure diagram of a lightweight small target detection method for bearing surface defects based on improved YOLOv8 of the present invention;

[0018] Figure 3 is the structure diagram of the DyHead Block module of the present invention;

[0019] Figure 4 is the structure diagram of the ADown module of the present invention;

[0020] Figure 5 is the bearing defect category diagram of the present invention, where (a) is a groove, (b) is a scratch, and (c) is a bruise;

[0021] Figure 6 is a partial display of the dataset expansion of the present invention, where (d) is the original image, (e) is the image processed by increasing the brightness, (f) is the image processed by decreasing the brightness, (g) is the image processed by Gaussian blur, and (h) is the image processed by Gaussian noise;

[0022] Figure 7Comparison of the heat map and detection results of the present invention, where (i) is the ground truth box, (j) is the YOLOv8n heat map, (k) is the YOLOv8n-DAS heat map, (l) is the YOLOv8n detection result map, and (m) is the YOLOv8n-DAS detection result map. Detailed implementation manners

[0023] The present invention will be further described below with reference to the accompanying drawings.

[0024] The present invention discloses a lightweight small target detection algorithm for bearing surface defects based on improved YOLOv8, which includes the following steps:

[0025] I. The experimental dataset of the present invention consists of 1,200 bearing pictures collected from the workbench scene, including three common defects: grooves, scratches, and abrasions. The pictures are labeled using the LableImg software. Since the scale of the original dataset is small, in order to enhance the generalization ability of the model and reduce the risk of overfitting, a series of augmentation operations such as brightness adjustment, Gaussian blur, and Gaussian noise are performed on the dataset. Finally, the augmented dataset contains 6,000 pictures and is randomly divided into a training set, a validation set, and a test set according to the ratio of 8:1:1 to meet the different needs of model training, validation, and testing.

[0026] II. Taking YOLOv8n as the baseline model, a dynamic detection head is designed using the DyHead Block. The structural diagram of the DyHead Block module is as Figure 3 shown. Given a four-dimensional feature vector from the feature pyramid, after reshaping it into a three-dimensional feature vector, it is sequentially passed through scale-aware attention, spatial-aware attention, and channel-aware attention to enhance the features; the scale-aware attention uses a linear function implemented by a 1×1 convolutional layer and a hard-sigmoid non-linear function with small computational complexity and high efficiency to map the input value to the range of 0 to 1, and dynamically fuse the features according to the semantic importance of different scales, enhancing the adaptability and detection ability of the model to defects of different scales. The formula for scale-aware attention is as follows:

[0027]

[0028] where π L (F) is the scale-aware attention weight, F is the input feature map, σ is the Hard-Sigmoid activation function, f is a linear transformation function implemented by a 1×1 convolutional layer, is the normalization factor, ∑ S,CF performs global average pooling on the feature map in the spatial dimension (S = H×W) and the channel dimension (C), compressing it into global statistics of L×1×1×1. L represents the number of layers of the pyramid network, and H, W, and C represent the height, width, and number of channels of the feature layer respectively. The spatial perception attention uses deformable convolution. Through self-learned spatial offsets and importance scalars, it selectively focuses on discriminative regions of the input features, assigns importance weights to different positions, enables the model to concentrate on key parts, improves the utilization efficiency of spatial position information, and better captures defect features. The formula for spatial perception attention is as follows:

[0029]

[0030] where π s (F)·F is the feature map weighted by spatial perception attention, L is the total number of layers of the feature pyramid (Layer), K is the number of sparse sampling points, w l,k is the weight of the k-th sampling point in the l-th layer of the feature map, F(l;p k +Δp k ;c) is the feature value at position p k in the l-th layer of the feature map after being adjusted by the self-learned offset Δp k , c represents the channel dimension, and Δm k represents the importance scalar of the k-th sampling point. Channel perception attention compresses channel features through global average pooling, establishes relationships between channels through a fully connected layer, and then adjusts the weights using a non-linear mapping to enhance or suppress channel features, effectively integrating information and highlighting key features. The calculation formula for channel perception attention is as follows:

[0031]

[0032] where π C (F): channel perception attention weight, F is the input feature map, F c is the feature slice on the C-th channel, α1, α2: channel scaling coefficients, β1, β2: channel offset coefficients, max(·): takes the maximum value of two different channel adjustment strategies (α 1 ·F C +β 1 and α 2 ·F C +β 2 ).

[0033] Third, the ADown module cleverly combines average pooling, traditional convolution and maximum pooling. First, average pooling is performed on the features to reduce background information interference, and then the features are split in half by channel, and input into the traditional convolution and maximum pooling layers for downsampling respectively, and finally the results are spliced. This structure allows the network to retain more edge feature information while learning complex features, reducing the risk of overfitting, and making the model lighter without losing accuracy. The ADown module structure is as follows: Figure 4 shown.

[0034] Fourth, a small target detection layer is added at the neck of the network, focusing on capturing and retaining the feature information of small target defects. The feature map after the second upsampling is further upsampled for the third time and fused with the shallow features of the third layer of the backbone network. By fusing shallow detail features with deep semantic features, the problem of small target defect feature information being easily lost during network transmission is solved, thereby improving the detection accuracy of small target defects.

[0035] 5. Start the model training process, set the training cycle to 300 epochs, select stochastic gradient descent (SGD) as the optimization method, set the initial learning rate to 0.01, and configure the momentum to 0.937. During the training process, the model is continuously iteratively updated to gradually learn the characteristic information of bearing surface defects to achieve effective detection of defect targets.

[0036] 6. Using precision, recall, and mean average precision (mAP) as the main evaluation indicators, while considering the number of parameters, weights, and frames per second (FPS), the model's detection accuracy, efficiency, and resource consumption are comprehensively evaluated. Based on these indicators, the performance of the model in the bearing surface defect detection task is objectively measured to verify its effectiveness and practicality. The experimental results are shown in Table 1.

[0037] Table 1 Performance comparison of different methods

[0038]

[0039] The experimental results show that the YOLOv8n-DAS method proposed in this invention demonstrates significant advantages compared to other lightweight object detection methods. By comparing the performance with popular methods such as YOLOv5s, YOLOv7-tiny, YOLOv8s, YOLOv9s, YOLOv10s, etc., YOLOv8n-DAS outperforms these methods in key metrics such as Precision, Recall, and mean Average Precision (mAP). This indicates that YOLOv8n-DAS has higher accuracy and reliability in detecting bearing surface defects. YOLOv8n-DAS is also more lightweight in terms of the number of parameters and model size, which means it requires less storage space and computing resources in practical applications, reducing the hardware requirements and operating costs. Although the detection speed (FPS) of YOLOv8n-DAS is slightly lower than some other methods, it can still meet the requirements of industrial real-time detection, ensuring rapid and efficient detection of bearings on the production line. These results fully demonstrate the superior performance and practical application potential of the YOLOv8n-DAS method in the field of bearing surface defect detection.

[0040] The above is only a preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A lightweight bearing surface defect small target detection method based on improved YOLOv8, characterized by: First, the dataset is preprocessed, including a series of expansion operations such as increasing or decreasing brightness, Gaussian blur, and Gaussian noise to reduce the risk of overfitting, and then randomly divided into training set, validation set, and test set in a ratio of 8:1:

1. A dynamic detection head combined with an attention mechanism is used to enhance the feature expression capability by unifying scale perception, spatial perception, and task perception. The ADown module is used for downsampling to highlight edge defect features and reduce the model size. A small target detection layer is added to the neck to strengthen the fusion of shallow semantic information and deep semantic information and retain more detailed information, thereby improving the detection accuracy, recall rate and average precision of small targets with bearing surface defects. YOLOv8n is used as the baseline model, and the training cycle is set, the optimization algorithm is selected, the initial learning rate is determined, and the momentum is configured to start the model training process. Precision, recall rate, and mean average precision (mAP) are used as the main evaluation indicators to evaluate the detection accuracy, efficiency, and resource consumption of the model.

2. According to claim 1, a lightweight bearing surface defect small target detection method based on improved YOLOv8 is characterized in that: The dataset was annotated using LableImg software, including three common defects: grooves, scratches, and abrasions. In order to enhance the generalization ability of the model and reduce the risk of overfitting, the dataset was expanded by a series of operations such as increasing and decreasing brightness, Gaussian blur, and Gaussian noise. Finally, the expanded dataset was randomly divided into training set, validation set, and test set in a ratio of 8:1:1 to meet the different needs of model training, validation, and testing.

3. According to claim 1, a lightweight bearing surface defect small target detection method based on improved YOLOv8 is characterized in that: The dynamic detection head adopts DyHead Block. Given a four-dimensional feature vector from a feature pyramid, it is reshaped into a three-dimensional feature vector and then reinforced with scale-aware attention, space-aware attention, and channel-aware attention in sequence. The scale-aware attention is a linear function implemented by a 1×1 convolutional layer and a hard-sigmoid nonlinear function with low computational complexity and high efficiency. The input value is mapped to a range of 0 to 1, and features are dynamically fused according to the semantic importance of different sizes to enhance the model's adaptability and detection capabilities for defects of different scales. The scale-aware attention formula is as follows: where π L (F) is the scale-aware attention weight, F is the input feature map, σ is the Hard-Sigmoid activation function, f is the linear transformation function implemented by the 1×1 convolutional layer, is the normalization factor, ∑ S,CF To perform global average pooling on the feature map in the spatial dimension (S = H × W) and channel dimension (C), it is compressed into a global statistic of L × 1 × 1 × 1, where L represents the number of layers of the pyramid network, and H, W, and C represent the height, width, and number of channels of the feature layer, respectively; spatial perception attention uses deformable convolution to selectively focus on the discriminable area of ​​the input feature through self-learning spatial offset and importance scalar, and assign importance weights to different positions, so that the model can focus on key parts, improve the utilization efficiency of spatial position information, and better capture defect features. The formula for spatial perception attention is as follows: where π s (F) F is the feature map after spatial perception attention weighting, L is the total number of layers of the feature pyramid, K is the number of sparse sampling points, and w l,k is the weight of the kth sampling point in the lth layer feature map, F(l; p k+Δ p k ; c) is the position p in the feature map of the first layer k After self-learning offset Δp k The adjusted eigenvalue, c represents the channel dimension, Δm k Represents the importance scalar of the kth sampling point; channel-aware attention compresses channel features through global average pooling, establishes the relationship between channels through the fully connected layer, and then uses nonlinear mapping to adjust the weights to enhance or suppress channel features, effectively integrate information, and highlight key features. The calculation formula of channel-aware attention is as follows: where π C (F): channel-aware attention weight, F input feature map, F c is the feature slice on the Cth channel, α1, α2: channel scaling coefficients, β1, β2: channel offset coefficients, max(·): for two different channel adjustment strategies (α 1 ·F C +β 1 and α 2 ·F C +β 2 ) takes the maximum value.

4. According to claim 1, a lightweight bearing surface defect small target detection method based on improved YOLOv8 is characterized in that: The ADown module first averages the features to reduce background information interference, then splits the features in half by channel, inputs them into the traditional convolution and maximum pooling layer downsampling, and finally concatenates the results. This structure enables the network to retain more edge feature information while learning complex features, reduces the risk of overfitting, and makes the model lighter without losing accuracy.

5. According to claim 1, a lightweight bearing surface defect small target detection method based on improved YOLOv8 is characterized in that: The small target detection layer is set at the neck and head of the network, and a 104×104 scale small target detection layer is added and equipped with an additional small target detection head. The feature map after the second upsampling is further upsampled for the third time and fused with the shallow features of the third layer of the backbone network. One side is sent to the small target detection head, and the other side is further downsampled and passed to other detection heads, which improves the efficiency of feature fusion and realizes accurate detection of small target defects.

Citation Information

Cited By

  • Semi-supervised surface defect detection method based on self-attention and task alignment

    CN121053068A

  • Appearance defect detection method for elevator polyurethane buffer

    CN121053520A

  • Visible light ship image target detection method based on improved YOLOv8

    CN121505236A

  • Small mechanical part defect visual detection system based on YOLO lightweight

    CN121921281A