Bearing surface defect detection method based on Yolov8s-DDC model

By improving the Yolov8s-DDC model and combining it with depthwise separable convolution, diversified branch modules, and Monte Carlo attention mechanism, the challenges of complex background and small target detection in bearing surface defect detection are solved, achieving high-precision and high-efficiency industrial inspection results.

CN120634949APending Publication Date: 2025-09-12ZHEJIANG SCI-TECH UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510515654.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Existing machine vision technology has difficulty in effectively identifying diverse defects in complex backgrounds during bearing surface defect detection, especially in detecting small targets. The model complexity and efficiency need to be optimized, making it difficult to meet the high-precision and high-efficiency requirements of industrial online detection.

Method used

A Yolov8s-DDC model was constructed to improve feature extraction and small object detection capabilities by introducing a depthwise separable convolution module into the backbone network, a diversified branch module into the neck network, and a Monte Carlo attention mechanism (CMA) module at the end of the neck network.

Benefits of technology

The accuracy and speed of bearing surface defect detection have been significantly improved to meet the requirements of industrial online detection. The model is lightweight and has enhanced recognition capabilities for complex backgrounds and subtle defects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120634949A_ABST
    Figure CN120634949A_ABST
Patent Text Reader

Abstract

The invention provides a bearing surface defect detection method based on a Yov8s-DDC model, and the method comprises the following steps: constructing a Yov8s-DDC model for defect detection, the Yov8s-DDC model comprising a backbone network, a neck network and a head network; wherein the backbone network is obtained by introducing a deep separable convolution module to replace a part of standard convolution modules; the neck network is based on a neck network structure of a Yolov8s model, a diversified branch module is introduced to replace a part of standard convolution modules, and a CMA module combined with a Monte Carlo attention mechanism is introduced at the tail end of the neck network to obtain the neck network; a bearing surface defect data set is constructed, the Yov8s-DDC model is trained, and the trained Yov8s-DDC model is deployed; and acquiring a surface image of a to-be-detected bearing, inputting the surface image of the to-be-detected bearing into the trained Yolov8s-DDC model, and outputting defect information of the surface of the to-be-detected bearing. Aiming at the detection difficulties of complex surface texture, oil stain interference, various defects and different sizes of the bearing, the detection precision and speed are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a bearing surface defect detection method, in particular to a bearing surface defect detection method based on a Yolov8s-DDC model, and belongs to the technical field of machine vision defect detection. Background Art

[0002] Bearings are essential core components in modern machinery. Their quality directly impacts the performance, efficiency, and service life of the entire equipment, and can even lead to serious safety incidents. Therefore, during the bearing manufacturing process, rigorous surface quality testing is crucial to promptly identify and eliminate defective products.

[0003] A variety of defects may occur on the bearing surface during production, assembly, transportation, and other links. Common ones include black spots (usually caused by oxidation or improper anti-rust treatment), scratches (caused by mechanical contact or improper operation), bumps (caused by impact), scrap (caused by the material itself or forging problems), and wear (caused by poor lubrication, overload, etc.). These defects vary in shape and size. In addition, the surface of the bearing ring itself usually has a complex processing texture and may retain clean oil or impurities, which interferes with the accurate identification of defects. In particular, oil droplets, impurities, and black spot defects may be visually similar, shallow scratches may not be obvious in two-dimensional images, and changes in lighting conditions may also affect the visibility of defects, all of which increase the difficulty of detection.

[0004] Traditional bearing defect detection relies primarily on manual visual inspection, a method that is inefficient, labor-intensive, subject to subjective influences, and prone to missed and false detections. This method struggles to meet the high-quality, efficient inspection requirements of modern, large-scale production. While machine vision technology has been introduced, simple automated systems often struggle to cope with complex backgrounds and varying defect morphologies.

[0005] In recent years, artificial intelligence technologies, particularly deep learning, have made significant progress in the field of object detection. Both single-stage object detection algorithms (such as the YOLO series and SSD) and two-stage object detection algorithms (such as R-CNN, Faster R-CNN, and Mask R-CNN) have been widely explored for industrial defect detection. Two-stage algorithms generally offer higher accuracy but are slower and require larger models. Single-stage algorithms offer speed advantages and are more suitable for online detection scenarios. The YOLO series of algorithms is a representative example, with continuous iterative optimizations achieving a good balance between speed and accuracy. Yolov8, a newer version of the YOLO series, excels in multiple vision tasks and offers models of varying sizes to suit diverse needs. However, even the advanced Yolov8 model can still suffer from issues such as insufficient sensitivity for small object detection, inadequate feature extraction, and the need for further optimization of model complexity and efficiency when directly applied to specific industrial scenarios, such as bearing surfaces, which feature low contrast, complex textures, oil contamination, and small defects. Therefore, it is of great significance to adaptively improve the existing model and develop a detection method with higher accuracy and efficiency in view of the actual difficulties and needs of bearing surface defect detection. Summary of the Invention

[0006] Based on the above background, the purpose of the present invention is to provide a bearing surface defect detection method based on the Yolov8s-DDC model, which aims to improve the detection accuracy and speed of bearing surface defects with complex texture, oil pollution interference, and various defects of different sizes, and meet the requirements of industrial online detection.

[0007] In order to achieve the above-mentioned object of the invention, the present invention provides the following technical solutions:

[0008] A bearing surface defect detection method based on the Yolov8s-DDC model, the method comprising the following steps:

[0009] A Yolov8s-DDC model for defect detection is constructed. The Yolov8s-DDC model includes a backbone network, a neck network, and a head network. The backbone network is based on the backbone network structure of the Yolov8s model and is obtained by introducing depthwise separable convolution modules to replace some standard convolution modules. The neck network is based on the neck network structure of the Yolov8s model and is obtained by introducing diversified branch modules to replace some standard convolution modules, and a CMA module combined with a Monte Carlo attention mechanism is introduced at the end of the neck network. The head network follows the head network structure of the Yolov8s model and is used to perform defect classification prediction and position regression prediction based on the multi-scale feature maps output by the neck network.

[0010] Acquire bearing surface images, construct a bearing surface defect dataset containing multiple preset defect types, divide the bearing surface defect dataset into a training set, a validation set, and a test set, train the Yolov8s-DDC model using the training set and the validation set, and deploy the trained Yolov8s-DDC model;

[0011] Obtain a surface image of the bearing to be inspected, input the surface image of the bearing to be inspected into the trained Yolov8s-DDC model, and output defect information of the bearing surface to be inspected based on the processing results of the Yolov8s-DDC model.

[0012] Preferably, the backbone network includes an initial convolution module, multiple repeated feature extraction units and a spatial pyramid pooling fast module connected in sequence, each feature extraction unit includes a depth-separable convolution module and a C2f module; in the neck network, the diversified branch module is replaced by the standard convolution module on the feature fusion path in the neck network structure of the Yolov8s model, and the neck network is provided with three CMA modules, each CMA module is respectively connected to a detection head in the head network.

[0013] Preferably, the CMA module is constructed based on the C2f module of the Yolov8s model and is obtained by adding a Monte Carlo attention layer after the last convolutional layer within the C2f module.

[0014] Preferably, the Monte Carlo attention layer generates an attention map by a random sampling pooling operation, wherein the random sampling pooling operation randomly selects attention map elements from at least two pooling tensors of different scales for combination, and the pooling tensors of different scales are obtained by applying an average pooling function of different pooling kernel sizes to the input features.

[0015] Preferably, the attention map generated by the Monte Carlo attention layer is calculated by the following formula:

[0016]

[0017] Where A m (x) represents the attention map, x represents the input tensor, i represents the output size of the attention map, n represents the number of output pooling tensors, f(x,i) represents the average pooling function, P1(x,i) represents the association probability, and satisfies the condition and

[0018] Preferably, the depthwise separable convolution module includes a 3x3 depthwise convolution layer with a stride of 2, a 1x1 point-by-point convolution layer with a stride of 2, a batch normalization layer and a GELU activation function layer, which are connected in sequence.

[0019] Preferably, during the training phase, the diversified branch module includes multiple parallel branches with different structures or parameters, and the branches are selected from at least two of convolutional layers and average pooling layers with different convolution kernel sizes; during the inference phase, the multiple parallel branches of the diversified branch module in the training phase are equivalently converted into a single convolutional layer.

[0020] Preferably, when training the Yolov8s-DDC model, the hyperparameters set include: 200 training rounds, 8 batch size, 0.01 initial learning rate, and stochastic gradient descent optimizer.

[0021] Preferably, the preset defect types in the bearing surface defect data set include black spots, scratches, bumps, material waste, and abrasions.

[0022] Compared with the prior art, the present invention has the following advantages:

[0023] The present invention discloses a bearing surface defect detection method based on the Yolov8s-DDC model. The method makes targeted improvements based on the Yolov8s model, significantly improves the detection accuracy while maintaining a high detection speed for bearing surface defects, and meets the accuracy and efficiency requirements of industrial online detection. The present invention introduces depthwise separable convolution into the backbone network, effectively reducing the number of parameters and floating-point operations of the model, making the model more lightweight. The present invention introduces a diversified branch module into the neck network. Through its multi-branch structure in the training phase, it can capture richer and more diverse feature representations, improve the model's ability to recognize complex backgrounds and subtle defect features, and maintain high efficiency in the reasoning phase. The present invention introduces a CMA module combined with a Monte Carlo attention mechanism (MCA) at the end of the neck network, and utilizes the random sampling pooling and cross-scale information aggregation capabilities of the MCA to enhance the model's attention to important areas and targets of different scales in the image, thereby significantly improving the detection accuracy of tiny defects on the bearing surface. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.

[0025] Figure 1 It is a schematic structural diagram of the Yolov8s-DDC model in the present invention;

[0026] Figure 2Schematic diagram of the structure of the depthwise separable convolution module (DSConv) in the present invention;

[0027] Figure 3 Schematic diagram of the structure of the diversified branch module (DBB) in the present invention;

[0028] Figure 4 It is a conversion method of the diversified branch module in the present invention from the parallel branches in the training phase to the conventional convolutional layer in the inference phase;

[0029] Figure 5 It is a structural diagram of the CMA module in the present invention;

[0030] Figure 6 Schematic diagram of the Monte Carlo Attention Mechanism (MCA) in the present invention;

[0031] Figure 7 Schematic diagram of the types of bearing surface defects in the present invention;

[0032] Figure 8 are the training loss and validation loss of the Yolov8s-DDC model in this invention;

[0033] Figure 9 This is a comparison of the mAP curves of the Yolov8s-DDC model and the Yolov8s model in the present invention;

[0034] Figure 10 This is a comparative experimental detection effect diagram on the bearing surface defect data set in the present invention;

[0035] Figure 11 This is a comparison test result diagram of the present invention on the hot-rolled strip surface defect data set. DETAILED DESCRIPTION

[0036] The technical solution of the present invention will be further described in detail below through specific embodiments and in conjunction with the accompanying drawings. It should be understood that the implementation of the present invention is not limited to the following embodiments, and any form of modification and / or change made to the present invention will fall within the scope of protection of the present invention.

[0037] In the present invention, unless otherwise specified, all parts and percentages are by weight. The equipment and raw materials used are commercially available or commonly used in the art. The methods in the following embodiments, unless otherwise specified, are conventional methods in the art. The components or equipment in the following embodiments, unless otherwise specified, are all universal standard parts or components known to those skilled in the art. Their structures and principles are known to those skilled in the art through technical manuals or routine experimental methods.

[0038] The following detailed description of the embodiments of the present invention is made in conjunction with the accompanying drawings. In the following detailed description, for ease of explanation, many specific details are set forth to provide a comprehensive understanding of the embodiments of the present invention. However, one or more embodiments may be implemented by those skilled in the art without these specific details.

[0039] An embodiment of the present invention discloses a bearing surface defect detection method based on the Yolov8s-DDC model, the method comprising the following steps:

[0040] A Yolov8s-DDC model for defect detection was constructed. The Yolov8s-DDC model consists of a backbone network, a neck network, and a head network. The backbone network is based on the backbone network structure of the Yolov8s model and introduces depthwise separable convolution modules to replace some standard convolution modules. The neck network is based on the neck network structure of the Yolov8s model and introduces diversified branch modules to replace some standard convolution modules. A CMA module combined with the Monte Carlo attention mechanism is introduced at the end of the neck network. The head network uses the head network structure of the Yolov8s model and is used to perform defect classification prediction and position regression prediction based on the multi-scale feature maps output by the neck network.

[0041] Acquire bearing surface images and construct a bearing surface defect dataset containing multiple preset defect types. Divide the bearing surface defect dataset into training, validation, and test sets. Use the training and validation sets to train the Yolov8s-DDC model, and deploy the trained Yolov8s-DDC model.

[0042] The surface image of the bearing to be inspected is obtained and input into the trained Yolov8s-DDC model. Based on the processing results of the Yolov8s-DDC model, the defect information of the bearing surface to be inspected is output.

[0043] Reference Figure 1 The Yolov8s-DDC model constructed in this embodiment has the following key improvements based on the Yolov8s model:

[0044] The standard convolutional module for feature extraction in the original backbone network structure of the Yolov8s model is replaced with a depthwise separable convolutional module (DSConv). After the improvement, the backbone network consists of an initial convolutional module, four repeated feature extraction units, and a spatial pyramid pooling fast module (SPPF) connected in sequence. Each feature extraction unit includes a depthwise separable convolutional module and a C2f module.

[0045] The two standard convolutional modules in the original neck network structure (PAN-FPN structure) of the Yolov8s model are replaced with a Diverse Branch Block (DBB), and the three C2f modules at the end of the original neck network of the Yolov8s model, i.e., before the final output to the head network, are replaced with a CMA module. The CMA module is based on the structure of the C2f module and embeds the Monte Carlo Attention (MCA) mechanism in it.

[0046] By introducing depthwise separable convolutions in the backbone network, the model reduces computational complexity and parameter count, while leveraging its properties to more effectively capture spatial and channel-wise information during feature extraction. The introduction of a diversified branch module in the neck network significantly improves the model's feature representation capabilities. As a structural reparameterization technique designed to enhance the performance of convolutional neural networks, the diversified branch module enriches the feature space by introducing diverse branches that combine different scales and complexities. This approach not only improves the model's feature extraction capabilities but also effectively controls parameter increase, ensuring detection efficiency, particularly when processing high-resolution images and complex scenes. Compared to traditional convolutional layers, the diversified branch module reduces information loss and feature redundancy. Secondly, at the end of the neck network, a new module, the Monte Carlo Attention Mechanism (CMA), is introduced, further enhancing the model's flexibility. This mechanism, based on a randomly sampled pooling operation, generates a scale-independent attention map, enabling the model to capture relevant information at different scales, making it particularly effective in detecting small objects. Small objects are often obscured or blurred in the feature layer, making them difficult to effectively identify using traditional attention mechanisms. Monte Carlo attention enhances the recognition of important features through adaptive feature weighting, thereby improving the recognizability of small objects. Combining the Diversified Branch Module with CMA in the neck network not only enhances feature diversity and the model's expressiveness, but also optimizes small object recognition performance. In summary, the introduction of the depthwise separable convolutional module, the Diversified Branch Module, and the CMA module significantly improves the overall performance of the model.

[0047] The depthwise separable convolution module, diversified branch module and CMA module are described in detail below.

[0048] Reference Figure 2The depthwise separable convolution module consists of a 3x3 depthwise convolution layer with a stride of 2, a 1x1 pointwise convolution layer with a stride of 2, a batch normalization layer, and a GELU activation function layer connected in sequence. Its calculation process mainly includes two steps: depthwise convolution and pointwise convolution. Depthwise convolution performs independent convolution on each channel of the input feature map. For each channel, a small convolution kernel (3x3) is used to operate, which can effectively extract spatial features. Since each channel is processed independently, the amount of calculation is relatively small. Pointwise convolution uses a 1×1 convolution kernel to process the output of depthwise convolution. The purpose of pointwise convolution is to combine the features of each channel to generate a new output feature map. In this way, the model can learn the relationship between channels and feature interactions.

[0049] Reference Figure 3 During the training phase, the Diversified Branch Module consists of convolutional layers of varying sizes and average pooling layers. These layers are arranged in parallel in a complex manner, and their outputs are ultimately merged. Upon completion of training, this complex structure is converted into a single convolutional layer for the model's inference phase. Specifically, during inference, the multiple parallel branches of the Diversified Branch Module during training are equivalently converted into a single convolutional layer. This conversion allows the Diversified Branch Module to increase its microstructural complexity during training while maintaining its macroarchitecture, effectively improving model performance.

[0050] Reference Figure 4 There are typically six transformation methods used to convert the parallel branches of the Diversified Branch Module during the training phase into regular convolutional layers during the inference phase. As shown in the figure, Transform I fuses convolutional layers with batch normalization to reduce model complexity; Transform II merges the outputs of convolutional layers with the same configuration to further simplify computation; Transform III merges sequential convolutional layers to improve feature extraction efficiency; Transform IV merges multiple convolutional layers through deep concatenation to enhance feature diversity; Transform V incorporates average pooling operations into convolution operations to reduce computational burden and improve feature extraction capabilities; and Transform VI combines convolutional layers of different scales to enhance the model's ability to capture multi-scale features. Through these transformation methods, the Diversified Branch Module can improve model performance and efficiency without increasing computational cost during inference.

[0051] Reference Figure 5The CMA module is built based on the C2f module of the Yolov8s model and is obtained by adding a Monte Carlo attention layer after the last convolutional layer inside the C2f module. The C2f module effectively improves the diversity and richness of feature extraction by fusing feature information of different scales. The Monte Carlo Attention (MCA) mechanism randomly samples the feature maps and uses the Monte Carlo method to optimize the attention allocation process, automatically capturing important spatial and channel information in the image, thereby effectively enhancing the model's ability to focus on key areas. By embedding the Monte Carlo attention layer into the C2f module, the network's ability to model complex image features can be improved. In particular, when dealing with visual tasks with high noise or background interference, the CMA module can focus on valuable feature areas more accurately.

[0052] The Monte Carlo attention layer generates an attention map through a random sampling pooling operation, which randomly selects attention map elements from at least two pooling tensors of different scales for combination. Pooling tensors of different scales are obtained by applying an average pooling function with different pooling kernel sizes to the input features. Figure 6 , MCA (indicated by the blue block in the figure) generates an attention map by randomly selecting 1×1 attention maps from three pooled tensors of different scales (3×3, 2×2, 1×1). The attention map generated by the Monte Carlo attention layer is calculated by the following formula:

[0053]

[0054] Where A m (x) represents the attention map, x represents the input tensor, i represents the output size of the attention map, n represents the number of output pooling tensors, f(x,i) represents the average pooling function, P1(x,i) represents the association probability, and satisfies the condition and

[0055] The bearing surface defect dataset constructed in this embodiment, which contains multiple preset defect types, is derived from various bearing surface defect images collected in actual industrial sites. Figure 7 Bearing surface defects are primarily categorized into five types: black spots, scratches, bumps, scrap, and abrasions. We collected bearing surface defect images at resolutions of 5472×3468 and 2024×2020 and manually cropped them to a uniform size of 640×640. Due to the uneven number of bearing defect images collected in actual industrial settings, image processing and data augmentation techniques were used to expand the dataset to 5148 images to ensure training rationality and balance between different defect types. The number of bearing surface defect images after augmentation is shown in Table 1.

[0056] Table 1: Quantity of images of various types of bearing surface defects after expansion

[0057] dark spots scratches bump Waste abrasion quantity 1049 1029 1023 1020 1027

[0058] The expanded dataset is divided into training, validation, and test sets in a 6:2:2 ratio. All images are annotated, with each image corresponding to a label that records the defect location and type. The dataset division is shown in Table 2.

[0059] Table 2 Dataset partitioning table

[0060]

[0061]

[0062] For Yolov8s-DDC model training, we set the training epochs to 200, the batch size to 8, the learning rate to 0.01, and the SGD optimizer. The input images were uniformly sized at 640×640. Model performance was evaluated using five metrics: precision, recall, mAP, frames per second (FPS), and GFLOPs.

[0063] To verify the effectiveness of the improved Yolov8s-DDC model, we conducted an ablation experiment. The results are shown in Table 3. Here, √ indicates the use of this module, DSC stands for the depthwise separable convolution module, and DBB stands for the diversified branching module.

[0064] Table 3 Ablation experiments

[0065]

[0066] As shown in Table 3, the initial Yolov8s model achieved a mAP of 95.4% and an FPS of 128 frames per second. Adding DSC to the backbone network improved mAP to 95.9%, FPS to 131 frames per second, and parameter count to 9.74M, improving both detection accuracy and efficiency. Introducing DBB to the neck network further improved mAP to 96%. Adding CMA to the end of the neck further increased mAP to 96%, while reducing GFLOPs to 28.4. When these improvements were applied in pairs, mAP increased by 0.9%, 1.1%, and 1.1%, respectively. When all improvements were applied in combination, mAP increased by 1.5%, while the parameter count remained essentially the same as the original model. Despite a drop in FPS to 106 frames per second and GFLOPs to 26.6, the overall accuracy and detection efficiency still met practical industrial requirements.

[0067] During the training process of the Yolov8s-DDC model in this example, the training loss and validation loss refer to Figure 8 It can be seen that in the first 50 rounds, the training loss and validation loss converge rapidly until they are fully converged at 200 rounds. Figure 9 .

[0068] To further verify the effect of the Yolov8s-DDC model, we compared it with the current mainstream target algorithms Yolov5, Yolov6, Yolov8, and Yolov10, and compared the detection results with reference to Figure 10 It can be seen that the Yolov8s-DDC model has a high detection accuracy for all types of defects, basically surpassing other target detection algorithms. The comparative experiments are shown in Table 4.

[0069] Table 4 Comparative experiment

[0070] Yolov5s Yolov6s Yolov8s Yolov10s Yolov8s-DDC dark spots 95.4% 93.5% 92.9% 91.2% 96.7% scratches 96.2% 99% 97.3% 94.5% 98.5% bump 88.4% 84.9% 88.5% 86.9% 90.4% Waste 98.5% 98.4% 99.5% 99% 99.5% abrasion 93.5% 95.8% 98.6% 97.4% 99.5% P 92.5% 90.4% 94.9% 91.2% 96.8% R 89.9% 86% 92.4% 87.2% 93.6% mAP 94.4% 94.3% 95.4% 93.8% 96.9% FPS 122 115 128 80 106

[0071] As can be seen, Yolov8s-DDC achieves improvements in mAP by 2.5%, 2.6%, 1.5%, and 3.1%, respectively, compared to other Yolo series object detection algorithms. Compared to the original Yolov8 model, mAP is improved by 1.5%, with the detection accuracy of black spots and bump defects increasing by 3.8% and 1.9%, respectively. Furthermore, precision and recall are also improved by 1.9% and 1.2%, respectively. Despite a slight decrease in FPS, the improved algorithm still meets the requirements of industrial applications in terms of detection accuracy and efficiency.

[0072] In order to further evaluate the effect of the Yolov8s-DDC model, a comparative experiment was conducted on the Northeastern University hot-rolled strip surface defect dataset. The dataset contains 1,800 images, divided into 6 defect categories: crazing, inclusion, patches, pitted surface, rolled-in scale, and scratches. The dataset is divided into training set, validation set, and test set in a ratio of 6:2:2. Performance comparisons were conducted based on the YOLOv5, YOLOv6, YOLOv8, and YOLOv10 models. The detailed data summary of the comparative experiment is shown in Table 5, and the detection effect of the experimental results is shown in the following figure. Figure 11 shown.

[0073] Table 5 Comparative test on hot-rolled strip surface defect dataset

[0074] Yolov5s Yolov6s Yolov8s Yolov10s Yolov8s-DDC crazing 38.7% 35.8% 42.9% 35% 42.3% inclusion 77.9% 79.4% 84.5% 79.4% 84.5% patches 89.8% 90.3% 94.2% 88.7% 93.8% Pitted surface 78% 82.9% 89.7% 73.4% 84.2% rolled-inscale 49.5% 56.2% 60.4% 55.3% 65.6% scratches 92.9% 95.3% 84.9% 80.8% 90.3% mAP 71.1% 73.3% 76.1% 68.8% 76.8% FPS 115 113 135 76 102

[0075] It can be seen that Yolov8s-DDC shows obvious advantages over the original Yolov8 model and has also achieved significant improvements in comparison with other detection algorithms. This shows that the Yolov8s-DDC model is not only effective, but also has strong applicability and robustness.

[0076] This document uses specific examples to illustrate the principles and implementation methods of the present invention. The above examples are only intended to help understand the method and core concept of the present invention. It should be noted that those skilled in the art can make various improvements and modifications to the present invention without departing from the principles of the present invention, and such improvements and modifications also fall within the scope of protection of the claims of the present invention.

Claims

1. A bearing surface defect detection method based on the Yolov8s-DDC model, characterized by: The method comprises the following steps: A Yolov8s-DDC model for defect detection is constructed. The Yolov8s-DDC model includes a backbone network, a neck network, and a head network. The backbone network is based on the backbone network structure of the Yolov8s model and is obtained by introducing depthwise separable convolution modules to replace some standard convolution modules. The neck network is based on the neck network structure of the Yolov8s model and is obtained by introducing diversified branch modules to replace some standard convolution modules, and a CMA module combined with a Monte Carlo attention mechanism is introduced at the end of the neck network. The head network follows the head network structure of the Yolov8s model and is used to perform defect classification prediction and position regression prediction based on the multi-scale feature maps output by the neck network. Acquire bearing surface images, construct a bearing surface defect dataset containing multiple preset defect types, divide the bearing surface defect dataset into a training set, a validation set, and a test set, train the Yolov8s-DDC model using the training set and the validation set, and deploy the trained Yolov8s-DDC model; Obtain a surface image of the bearing to be inspected, input the surface image of the bearing to be inspected into the trained Yolov8s-DDC model, and output defect information of the bearing surface to be inspected based on the processing results of the Yolov8s-DDC model.

2. The bearing surface defect detection method based on the Yolov8s-DDC model according to claim 1, characterized in that: The backbone network includes an initial convolution module, multiple repeated feature extraction units and a spatial pyramid pooling fast module connected in sequence, each feature extraction unit includes a depth-separable convolution module and a C2f module; in the neck network, the diversified branch module replaces the standard convolution module on the feature fusion path in the neck network structure of the Yolov8s model, and the neck network is equipped with three CMA modules, each of which is connected to a detection head in the head network.

3. The bearing surface defect detection method based on the Yolov8s-DDC model according to claim 1, characterized in that: The CMA module is built based on the C2f module of the Yolov8s model and is obtained by adding a Monte Carlo attention layer after the last convolutional layer inside the C2f module.

4. The bearing surface defect detection method based on the Yolov8s-DDC model according to claim 3 is characterized in that: The Monte Carlo attention layer generates an attention map through a random sampling pooling operation, which randomly selects attention map elements from at least two pooling tensors of different scales for combination, and the pooling tensors of different scales are obtained by applying an average pooling function with different pooling kernel sizes to the input features.

5. The bearing surface defect detection method based on the Yolov8s-DDC model according to claim 4 is characterized in that: The attention map generated by the Monte Carlo attention layer is calculated by the following formula: Where A m (x) represents the attention map, x represents the input tensor, i represents the output size of the attention map, n represents the number of output pooling tensors, f(x,i) represents the average pooling function, P1(x,i) represents the association probability, and satisfies the condition and 0.

6. The bearing surface defect detection method based on the Yolov8s-DDC model according to claim 1, characterized in that: The depthwise separable convolution module includes a 3x3 depthwise convolution layer with a stride of 2, a 1x1 point-by-point convolution layer with a stride of 2, a batch normalization layer, and a GELU activation function layer, which are connected in sequence.

7. The bearing surface defect detection method based on the Yolov8s-DDC model according to claim 1, characterized in that: During the training phase, the diversified branch module includes multiple parallel branches with different structures or parameters, and the branches are selected from at least two of convolutional layers and average pooling layers with different convolution kernel sizes; during the inference phase, the multiple parallel branches of the diversified branch module in the training phase are equivalently converted into a single convolutional layer.

8. The bearing surface defect detection method based on the Yolov8s-DDC model according to claim 1, characterized in that: When training the Yolov8s-DDC model, the hyperparameters set include: 200 training rounds, 8 batch size, 0.01 initial learning rate, and stochastic gradient descent optimizer.

9. The bearing surface defect detection method based on the Yolov8s-DDC model according to claim 1, characterized in that: The preset defect types in the bearing surface defect dataset include black spots, scratches, bumps, material scrap, and abrasions.

Citation Information

Cited By

  • Crushed material and waste material sorting method and system based on improved YOLOv8

    CN122265663A