A high-precision intelligent detection method for surface defects of bearing rings

By improving the YOLOv5 target detection network, combined with SPD, C2f and CARAFE modules, a high-precision bearing ring surface defect detection model was built, solving the problem of low detection accuracy in the prior art, and achieving efficient detection of complex backgrounds and multiple defect types.

CN116645328BActive Publication Date: 2025-06-06ZHEJIANG SCI-TECH UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310454335.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-25
Publication Date
2025-06-06
Estimated Expiration
2043-04-25

AI Technical Summary

Technical Problem

The prior art is difficult to detect bearing ring surface defects with high accuracy, especially in cases of complex backgrounds, multiple defect types and low resolution images.

Method used

Based on the improved YOLOv5 target detection network, the bearing ring surface defect detection model is constructed, and the SPD module is used to replace the convolution step size to increase the number of feature map channels; the C2f module is used for feature extraction and fusion, and the feature map quality and multi-scale capabilities are improved; the CARAFE lightweight universal upsampling module is introduced into the neck network to enrich context information and reduce information loss.

Benefits of technology

The expression and generalization capabilities of the bearing ring surface defect detection model are improved, the detection capabilities of low-resolution images and small objects are enhanced, and high-precision detection is achieved, with detection accuracy of more than 97%.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116645328B_ABST
    Figure CN116645328B_ABST
Patent Text Reader

Abstract

The present invention provides a high-precision intelligent detection method for surface defects of bearing rings, which comprises the following steps: constructing a bearing ring surface defect detection model based on an improved YOLOv5 target detection network, comprising a backbone network, a neck network, and a target detection head module, wherein the backbone network is provided with five downsampling units in sequence, wherein the first downsampling unit comprises a CBS module, and the other four downsampling units comprise a CBS module and an SPD module which are arranged in sequence, and a C2f module is arranged after the other four downsampling units, wherein the SPD module adopts an SPD layer to replace the step convolution in the YOLOv5 target detection network, and the C2f module is used to realize separation convolution and splicing operations, and the neck network is provided with a CARAFE lightweight general upsampling module; setting a loss function of the bearing ring surface defect detection model, collecting a data set and dividing it into a training set, a verification set, and a test set, and performing multiple rounds of training. The present invention realizes high-precision detection of surface defects of bearing rings, and the detection accuracy reaches more than 97%.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to a bearing defect detection method, in particular to a high-precision bearing ring surface defect intelligent detection method, and belongs to the technical field of industrial visual detection. Background Art

[0002] As a component that plays a role in fixing and reducing load friction in the process of mechanical transmission, bearings are widely used to guide the rotation of shaft parts and bear the load transmitted to the frame by the shaft. The quality of bearings will seriously affect the stability of the entire mechanical equipment. However, during the production and assembly of bearings, due to the influence of factors such as materials, processing, assembly, and transportation, some defects will inevitably occur on the surface of the bearings. The formation of surface defects of bearing rings is mainly in the raw material processing link. The defects mainly include car scrap, forging scrap, black spots, bumps, scratches, etc. In the forging link, temperature differences can cause forging scrap. In the turning process, due to the high-speed operation of the machine, a slight misalignment will cause excessive cutting and lead to car scrap or scratches. In the anti-rust treatment process, due to the uneven application of anti-rust oil, the humid environment at the production site will cause the bearings to rust. During transportation, surface bumps will cause bumps and scratches.

[0003] In recent years, with the development of machine vision and deep learning technology, many defect detection methods based on machine vision and deep learning have been widely used in various industrial scenarios. However, visual detection methods for surface defects of bearing rings are rare. The main reason is that the background texture of bearing rings is complex, the defects vary in size and type, and the brightness is uneven. In addition, there is a lot of oil and dust in the production environment of bearings, which will interfere with the collected bearing images. Summary of the invention

[0004] Based on the above background, the purpose of the present invention is to provide a high-precision bearing ring surface defect intelligent detection method to solve the problems described in the background technology.

[0005] In order to achieve the above-mentioned object of the invention, the present invention provides the following technical solutions:

[0006] A high-precision bearing ring surface defect intelligent detection method, the method comprising the following steps:

[0007] A bearing ring surface defect detection model is constructed based on an improved YOLOv5 target detection network. The bearing ring surface defect detection model includes a backbone network for realizing feature extraction, a neck network for multi-scale fusion of different-level features extracted by the backbone network, and a target detection head module for performing target detection and classification. The backbone network is provided with five downsampling units in sequence, the first downsampling unit includes a CBS module, the other four downsampling units include a CBS module and an SPD module arranged in sequence, and a C2f module is arranged after the other four downsampling units. The CBS module includes a convolution layer, a normalization layer and an activation function layer. The SPD module uses an SPD layer to replace the stride convolution in the YOLOv5 target detection network. The C2f module is used to realize separation convolution and splicing operations. The neck network is provided with a CARAFE lightweight general upsampling module. The neck network fuses the feature maps output by the third downsampling unit, the fourth downsampling unit and the fifth downsampling unit in the backbone network. When performing feature fusion, shallow semantic information is transmitted from top to bottom, and deep semantic information is transmitted from bottom to top.

[0008] Setting a loss function of the bearing ring surface defect detection model, collecting a data set and dividing it into a training set, a validation set, and a test set, and performing multiple rounds of training on the bearing ring surface defect detection model according to the set training parameters;

[0009] The image of the bearing ring under test is input into the trained bearing ring surface defect detection model, and the surface defect detection result of the bearing ring under test is output.

[0010] The bearing ring surface defect detection model constructed by this method uses the SPD module to replace the Conv module of the previous layer for downsampling on the basis of the YOLOv5 target detection network, which increases the number of channels of the feature map. While keeping the resolution of the feature map unchanged, the detection capability of low-resolution images and small objects is improved, thereby improving the expression and generalization capabilities of the bearing ring surface defect detection model; the C2f module is used to extract features, realize the fusion of feature maps with different numbers of channels, improve the quality and efficiency of feature maps, improve the receptive field and multi-scale capability of feature maps, and obtain more global and higher semantic features; the neck network uses the CARAFE lightweight general upsampling module for upsampling, which has a larger receptive field and better semantic adaptability, while only introducing a small amount of parameters and calculations, retaining more feature details and structural information, and improving the quality and accuracy of upsampling.

[0011] Preferably, the SPD layer is used to halve the height and width of the input feature map and increase the number of channels of the input feature map by four times.

[0012] Preferably, the C2f module includes two CBS modules, a Split module and several Bottleneck modules. The feature map input to the first CBS module is split into two sub-feature maps through the Split module, wherein a sub-feature map split by the Split module is output through several Bottleneck modules, and after the sub-feature map output by each Bottleneck module is spliced ​​with another sub-feature map split by the Split module, the feature map after the separation convolution and splicing operations is output through the second CBS module, and the feature map has the same size as the feature map input to the first CBS module.

[0013] Preferably, the CARAFE lightweight universal upsampling module includes a kernel prediction module and a content-aware reorganization module, the kernel prediction module is used to generate weights on the kernel for reorganization calculation, and the content-aware reorganization module is used to reorganize features according to the calculated weights.

[0014] Preferably, the mathematical expression of the loss function is:

[0015] LOSS=w box L box +w obj L obj +w cls L cls

[0016] Where, L box is the positioning error function, L obj is the confidence loss function, L cls is the classification loss function, w box 、w obj 、w cls are the weight coefficients corresponding to the above functions respectively.

[0017] Preferably, the mathematical expression of the positioning error function is:

[0018]

[0019] Where IOU is the intersection-over-union ratio of the predicted box B and the true box A, ρ is the Euclidean distance between the center coordinates of the true box A and the predicted box B, c is the diagonal distance of the minimum box surrounding the center coordinates of the true box A and the predicted box B, α is the weight coefficient, and v is a parameter to measure the consistency of the aspect ratio of A and B.

[0020] Preferably, the classification loss function and the confidence loss function both adopt a binary cross entropy loss function, and the mathematical expression of the binary cross entropy loss function is:

[0021]

[0022] In the formula, n represents the number of input samples, y i represents the target value, x i Represents the predicted output value.

[0023] Preferably, the collecting of data sets includes collecting data sets and dividing the data sets into vehicle scrap, forging scrap, black spots, bumps and scratches according to the types of defects.

[0024] Preferably, the set training parameters include a batch size value of 32, a dynamic parameter of 0.937, a learning rate of 0.01, a cosine annealing learning rate of 0.1, a data enhancement value of 1.0, an image size of 640×640, and a training number of 100 times.

[0025] Preferably, a mosaic data enhancement method is used in multiple rounds of training of the bearing ring surface defect detection model.

[0026] Compared with the prior art, the present invention has the following advantages:

[0027] A high-precision bearing ring surface defect intelligent detection method of the present invention builds a bearing ring surface defect detection model based on an improved YOLOv5 target detection network, replaces the C3 module in the backbone network with a C2f module, effectively reduces the number of network parameters and calculations, and can obtain more global and higher semantic level features. The SPD module is used to effectively improve the model's ability to detect low-resolution images and small object images. The CARAFE lightweight general upsampling module is used to improve the neck network, enrich contextual information, and reduce information loss during transmission. This not only improves the model's defect detection capability, but also improves the model network diversity and robustness, so that the model can be adapted to different instances and scenarios. The high-precision bearing ring surface defect intelligent detection method realizes high-precision detection of bearing ring surface defects based on the above-mentioned bearing ring surface defect detection model, and the detection accuracy reaches more than 97%. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.

[0029] Figure 1 It is a flow chart of a high-precision bearing ring surface defect intelligent detection method of the present invention;

[0030] Figure 2 It is a network structure diagram of the bearing ring surface defect detection model in the present invention;

[0031] Figure 3 It is a structural schematic diagram of the SPD module in the present invention;

[0032] Figure 4 It is a schematic diagram of the structure of the traditional C3 module;

[0033] Figure 5 It is a schematic diagram of the structure of the C2f module in the present invention;

[0034] Figure 6 This is a schematic diagram of the principle of the upsampling module in the traditional YOLOv5 target detection network;

[0035] Figure 7 It is a structural schematic diagram of the CARAFE lightweight universal upsampling module in the present invention;

[0036] Figure 8 It is a graph of training loss, verification loss, and mAP curves during the training process of the bearing ring surface defect detection model in the present invention;

[0037] Fig. 9 is the test result of the bearing ring surface defect data set on each model;

[0038] Fig.10 This is a comparison chart of the test results of the bearing ring surface defect detection model of the present invention and the YOLOv5 model on the fabric data set. DETAILED DESCRIPTION

[0039] The technical solution of the present invention is further described in detail below through specific embodiments and in conjunction with the accompanying drawings. It should be understood that the implementation of the present invention is not limited to the following embodiments, and any form of modification and / or change made to the present invention will fall within the protection scope of the present invention.

[0040] In the present invention, unless otherwise specified, all parts and percentages are weight units, and the equipment and raw materials used can be purchased from the market or are commonly used in the art. The methods in the following embodiments, unless otherwise specified, are conventional methods in the art. The components or equipment in the following embodiments, unless otherwise specified, are universal standard parts or components known to those skilled in the art, and their structures and principles are known to those skilled in the art through technical manuals or conventional experimental methods.

[0041] refer to Figure 1 The embodiment of the present invention discloses a high-precision bearing ring surface defect intelligent detection method, the method comprising the following steps:

[0042] S1. A bearing ring surface defect detection model is constructed based on the improved YOLOv5 target detection network. The bearing ring surface defect detection model includes a backbone network for realizing feature extraction, a neck network for multi-scale fusion of different-level features extracted by the backbone network, and a target detection head module for performing target detection and classification. The backbone network is provided with five downsampling units in sequence. The first downsampling unit includes a CBS module, and the other four downsampling units include a CBS module and an SPD module set in sequence, and a C2f module is set after the other four downsampling units. The CBS module includes a convolution layer, a normalization layer, and an activation function layer. The SPD module uses an SPD layer to replace the stride convolution in the YOLOv5 target detection network. The C2f module is used to realize separation convolution and splicing operations. The neck network is provided with a CARAFE lightweight general upsampling module. The neck network fuses the feature maps output by the third downsampling unit, the fourth downsampling unit, and the fifth downsampling unit in the backbone network. When performing feature fusion, shallow semantic information is transmitted from top to bottom, and deep semantic information is transmitted from bottom to top;

[0043] S2. Setting the loss function of the bearing ring surface defect detection model, collecting the data set and dividing it into a training set, a validation set, and a test set, and performing multiple rounds of training on the bearing ring surface defect detection model according to the set training parameters;

[0044] S3. Input the image of the bearing ring to be tested into the trained bearing ring surface defect detection model, and output the surface defect detection result of the bearing ring to be tested.

[0045] The embodiments of the present invention are described in detail below in conjunction with the accompanying drawings. In the following detailed description, for the convenience of explanation, many specific details are set forth to provide a comprehensive understanding of the embodiments of the present invention.

[0046] The traditional YOLOv5 target detection network includes a backbone network, a neck network and a target detection head module. CSPDarknet53 is used as the backbone network, and the feature pyramid network (FPN) and the path aggregation network (PAN) are combined as the neck network to fuse the features extracted from the backbone. The main body of the target detection head module is three Detect detectors, which use grid-based anchors to perform target detection on feature maps of different scales.

[0047] In the visual inspection of bearing ring surface defects, since there are many types of bearing ring surface defects and the situation is complicated, there are defects of various sizes and shapes such as machine scrap, forging scrap, black spots, bumps, scratches, etc., it is necessary to combine shallow and high-level semantic information and effectively integrate different scale features. However, the traditional YOLOv5 target detection network cannot achieve high-precision and high-efficiency detection of bearing ring surface defects.

[0048] This high-precision bearing ring surface defect intelligent detection method is based on the traditional YOLOv5 target detection network to build a bearing ring surface defect detection model. The model network structure is referenced Figure 2 The bearing ring surface defect detection model includes a backbone network for feature extraction, a neck network for multi-scale fusion of different levels of features extracted by the backbone network, and a target detection head module for performing target detection and classification.

[0049] The backbone network is provided with five downsampling units in sequence, the first downsampling unit includes a CBS module, the other four downsampling units include CBS modules and SPD modules arranged in sequence, and the other four downsampling units are followed by C2f modules, and finally connected to the neck network through the SPPF module. Among them, the CBS module includes a convolution layer, a normalization layer and an activation function layer, which is a prior art and will not be described here.

[0050] The backbone network extracts feature maps of different sizes from the input image. The input image size is 640×640 pixels. After 2, 4, 8, 16, and 32 downsampling, the backbone network generates five layers of feature maps, whose sizes are 320×320 pixels, 160×160 pixels, 80×80 pixels, 40×40 pixels, and 20×20 pixels, respectively, and the number of channels are 32, 64, 128, 256, and 512, respectively.

[0051] The neck network fuses the feature maps output by the third downsampling unit, the fourth downsampling unit, and the fifth downsampling unit in the backbone network. Specifically, a CBS module in the neck network is connected to the SPPF module of the backbone network, a concat module in the neck network is connected to the third C2f module of the backbone network, and another concat module in the neck network is connected to the second C2f module of the backbone network. During the fusion process, the FPN structure transmits shallow semantic information from top to bottom, while the PAN structure transmits deep semantic information from bottom to top. The FPN structure and the PAN structure jointly enhance the feature fusion capability of the neck network. After feature fusion, three new feature maps are generated through three output layers. The three output layers are shallow, medium, and deep layers. The sizes of the output feature maps are 80×80 pixels, 40×40 pixels, and 20×20 pixels, respectively, and the number of channels is 128, 256, and 512, respectively. The smaller the feature map, the larger the image area corresponding to each grid unit in the feature map. Among the output feature maps of the above three output layers, the shallow feature map is suitable for detecting small targets, the medium feature map is suitable for detecting medium targets, and the deep feature map is suitable for detecting large targets.

[0052] Based on the above new feature maps, the object detection head module performs object detection and classification.

[0053] Compared with the traditional YOLOv5 target detection network, the bearing ring surface defect detection model uses the SPD module instead of the previous layer of Conv for downsampling, which can increase the number of channels of the feature map. While keeping the resolution of the feature map unchanged, it improves the detection ability of low-resolution images and small objects, thereby improving the expression and generalization capabilities of the model; the C2f module is used to extract features. Compared with the C3 module, C2f uses separate convolution and splicing operations to achieve the fusion of feature maps with different numbers of channels; the CARAFE lightweight general upsampling module is used in the neck network for upsampling, which can retain more feature details and structural information, and improve the quality and accuracy of upsampling.

[0054] The SPD module, C2f module and CARAFE lightweight universal upsampling module are described in detail below.

[0055] In the task of bearing surface defect detection with low image resolution or small objects, the performance of convolutional neural network models in computer vision tasks such as image classification and target detection will drop rapidly. The reason is that the convolution step and pooling layer are used in the existing model structure, which will lead to the loss of fine-grained information and the learning of less efficient feature representation.

[0056] By introducing the SPD module, the problem of information loss and performance degradation caused by the traditional convolution step or pooling layer when processing low-resolution images and small objects is solved. The SPD module can improve the receptive field and positioning accuracy because it retains all the information of the input feature map without losing some details like the convolution step or pooling layer. The SPD module is used for downsampling by changing the step size of the Conv layer above the SPD module from 2 to 1. Through the SPD module, the original image or intermediate feature map is divided into a series of sub-feature maps and stacked together, thereby increasing the number of channels and receptive field while reducing the spatial size. For example, for any intermediate feature map of size S×S×C Figure X , the series of sub-feature graphs cut out are:

[0057] f (0,0) =X[0:S:scale, 0:S:scale]

[0058] f (scale-1,0) =X[scale-1:S:scale,:S:scale]

[0059] f (0,scale-1) =X[0:S:scale,scale-1:S:scale]

[0060] f (scale-1,scale-1) =X[scale-1: S: scale, scale-1: S: scale]

[0061] In general, given any original feature Figure X , subgraph f (x,y) It is formed by dividing all X(i+y) by i+x and i+y in proportion. Therefore, each sub-image downsamples X by the scale factor scale. Figure 3 , when scale = 2, 4 sub-feature maps f are obtained (0,0) , f (0,1) , f (1,0) , f (1,1) , the size of each feature map is (S / 2,S / 2,C). At the same time, X is downsampled by 2 times and these sub-feature maps are connected along the channel dimension to obtain a new feature Figure X ′, the spatial dimension of the feature map is half of the original X, and the channel dimension is 4 times of the original X.

[0062] Therefore, the SPD module can halve the height and width of the input feature map, while increasing the number of channels by four times, increasing the depth of the feature map and keeping the total number of elements in the feature map unchanged, thereby improving the expressiveness of the features and the multi-scale fusion effect. In this way, the subsequent C2f module can be calculated on a smaller spatial size without losing information.

[0063] The C2f module is obtained by reducing one CBS module from the traditional C3 module and adjusting the splicing method of multiple Bottleneck modules. Figure 4 The traditional C3 module contains three standard convolutional layers (CBS modules) and n Bottlenneck modules. When set in the backbone network, the Bottlenneck module uses shortcuts, and when set in the neck network, the Bottlenneck module does not use shortcuts. The input feature map enters two branches, one of which is spliced ​​through multiple Bottleneck modules and one standard convolutional layer to obtain a sub-feature map, and the other is only through one standard convolutional layer to obtain another sub-feature map. Finally, the two sub-feature maps are spliced ​​and output.

[0064] Different from the traditional C3 module, the C2f module reduces a standard convolutional layer and concatenates the sub-feature maps output by each Bottleneck module, thereby obtaining richer gradient flow information while ensuring lightweight. Figure 5 The C2f module includes two CBS modules, a Split module and multiple Bottleneck modules. The input feature map size is h×w×c in After passing through the CBS module, the output size is h×w×c outThe feature map input to the first CBS module is split into two sub-feature maps through the Split module, where a sub-feature map split by the Split module is output through several Bottleneck modules, and each sub-feature map output by the Bottleneck module is concatenated with another sub-feature map split by the Split module, and then the feature map after the separation convolution and concatenation operation is output through the second CBS module. The setting of the Bottleneck module in the C2f module is the same as that in the traditional C3 module. When set in the backbone network, the Bottlenneck module uses a shortcut, and when set in the neck network, the Bottlenneck module does not use a shortcut.

[0065] Compared with the traditional C3 module, the C2f module is lighter, uses fewer parameters and computations, while maintaining high accuracy and speed. The C2f module extracts feature maps in the backbone network and implements feature fusion and channel separation through the CSP structure, improving the quality and efficiency of feature maps, improving the receptive field and multi-scale capabilities of feature maps, and obtaining more global and higher semantic features.

[0066] CARAFE lightweight universal upsampling module is an upsampling module. The function of the upsampling module is to expand a small-resolution image or feature map into a high-resolution image or feature map so that it can be displayed on a higher-resolution display device or improve the performance of subsequent tasks. The upsampling module can be used as an intermediate layer in a convolutional network to expand the feature map size and facilitate tensor splicing. There are many implementation methods for the upsampling module, such as nearest neighbor interpolation, bilinear interpolation, bicubic interpolation, trilinear interpolation, anti-pooling, transposed convolution, etc. Almost all of them use the interpolation method, that is, based on the original image pixels, a suitable interpolation algorithm is used to insert new elements between pixels. In the traditional YOLOv5 target detection network, the nearest neighbor interpolation is used as the algorithm of the upsampling module. Its implementation is to map each pixel in the target image to the original image through coordinate transformation, and then take the grayscale value of the closest original image pixel as the grayscale value of the target image pixel. However, referring to Figure 6 As shown, the missing pixels are generated by directly using the original color closest to them, that is, copying the pixels next to them, which produces obvious aliasing.

[0067] The CARAFE lightweight universal upsampling module uses a small convolutional network to generate an adaptive upsampling kernel, and then performs a dot product with the corresponding neighborhood pixels in the input feature map to obtain the upsampled feature map. It has a larger receptive field and better semantic adaptability, while introducing only a small amount of parameters and computation. Figure 7The CARAFE lightweight universal upsampling module includes a kernel prediction module and a content-aware reassembly module. The kernel prediction module is used to generate weights on the kernel for reassembly calculation, and the content-aware reassembly module is used to reassemble features according to the calculated weights. Figure 7 As shown in , the feature size is C×H×W Figure X The CARAFE lightweight universal upsampling module is upsampled by a factor of σ. For each position l = (i, j), there is a kernel for predicting the upsampling kernel used for reorganization. First, the kernel prediction module compresses the channel into C by the channel compression module. m , reducing the amount of subsequent calculations, which allows the use of a larger upsampling kernel for subsequent upsampling; then based on the size of the compressed feature map, a kernel of size k is used encoder The convolutional layer generates an upsampling kernel for reorganizing features, using a larger k encoder The receptive field will be expanded, and the channel will become Then reorganize the newly obtained feature map into The feature map is normalized using the softmax function for all channels at each position. The mathematical expression is:

[0068] W l′ =ψ(N(X l , k encoder )

[0069] X l ′=φ(N(X l , k up ), W l′ )

[0070] For any position of the output X′, there is a corresponding source position l = (i, j) at the input X, where i = (i′ / σ) and j = (j′ / σ). l , k up ) is represented by the k of X centered at position l up ×k up Sub-region, the kernel prediction module ψ according to X l The sub-region predicts the position kernel W of each position l′ l′ The content-aware reorganization module φ is based on X l The sub-region and position kernel W l′ Reorganize to get X l ′.

[0071] After the CARAFE lightweight universal upsampling module is introduced into the traditional YOLOv5 target detection network, different upsampling kernels can be dynamically generated through different positions of the input feature map to adapt to targets of different scales and shapes, thereby adapting to different instances and scenarios. Afterwards, the inner product operation is performed with the local neighborhood of the input feature map to obtain a new upsampled feature map, so that the upsampled feature map has higher resolution and richer detail information, thereby improving the recognition and positioning capabilities of different targets in the target detection task.

[0072] The mathematical expression of the loss function used to train the bearing ring surface defect detection model is:

[0073] LOSS=w box L box +w obj L obj +w cls L cls

[0074] Where, L box is the positioning error function, L obj is the confidence loss function, L cls is the classification loss function, w box 、w obj 、w cls are the weight coefficients corresponding to the above functions respectively.

[0075] The mathematical expression of the positioning error function is:

[0076]

[0077] Where IOU is the intersection-over-union ratio of the predicted box B and the true box A, ρ is the Euclidean distance between the center coordinates of the true box A and the predicted box B, c is the diagonal distance of the minimum box surrounding the center coordinates of the true box A and the predicted box B, α is the weight coefficient, and ν is a parameter to measure the consistency of the aspect ratio of A and B.

[0078] The mathematical expression of IOU is:

[0079]

[0080] Among them, A is the real box, B is the predicted box, A∩B represents the intersection of A and B, and A∪B represents the union of A and B.

[0081] The mathematical expressions of α and ν are,

[0082]

[0083]

[0084] Both the classification loss function and the confidence loss function use the binary cross entropy loss function. The mathematical expression of the binary cross entropy loss function is:

[0085]

[0086] In the formula, n represents the number of input samples, y i represents the target value, x i Represents the predicted output value.

[0087] The data set in this high-precision bearing ring surface defect intelligent detection method is collected on the bearing ring production line by an industrial camera. The image resolution is 5472×3468, and the size of each image is about 19M. The defective images are manually cut into windows, each window is 640×640 pixels in size, and the images containing defects are selected. At the same time, the data set is divided into car scrap, forging scrap, black spots, bumps, and scratches according to the type of defects. Since the number of each defect type in actual production varies, in order to ensure the rationality of training and the balance between each defect type, the number of each defect has been expanded to 5660. The statistical data of various defects after expansion are shown in Table 1.

[0088] Table 1 Expanded defect dataset

[0089] Car waste Forging waste Dark spots Bumps and bruises Scratching quantity 1140 1085 1148 1120 1167

[0090] Before sending the data set into the network for training, the data set needs to be divided. According to the number of data set samples and the rationality of training, the present invention divides each defect sample into training set, verification set and test set, with a division ratio of 6:2:2. The results are shown in Table 2.

[0091] Table 2 Statistics of training set, validation set and test set of bearing defect images

[0092]

[0093]

[0094] The hardware environment and software version of the bearing ring surface defect detection model are shown in Table 3.

[0095] Table 3 Hardware environment and software version

[0096]

[0097] The training parameters set for the bearing ring surface defect detection model are shown in Table 4.

[0098] Table 4 Training parameters

[0099] Training parameters value Batch size 32 Dynamic parameters 0.937 Learning Rate 0.01 Cosine annealing learning rate 0.1 Data Augmentation 1.0 Image size 640×640 Number of training sessions 100

[0100] In order to enrich the information of the detected target and improve the robustness of the model, a data enhancement method is used during the training process. GridMask randomly generates a grid-shaped occluder on the image, with a pixel value of 0 inside the occluder, and the classification result remains unchanged. This method may reduce the clarity and quality of the image. RandAugment is an automatic data enhancement method that randomly selects two transformations from a predefined transformation set and applies them to the image with a random amplitude, but it may introduce some overly strong or inappropriate transformations, such as color distortion and object deformation, thereby reducing the recognizability of the image. Therefore, the bearing ring surface defect detection model in the present invention adopts a mosaic data enhancement method during the training process. The mosaic data enhancement method is a method of splicing four images into one image, by selecting 4 images and scaling them to the same size, then randomly selecting a cutting point, cutting each image into four parts, and then splicing the parts of different images into a new image, retaining the label of the original image, and finally performing other data enhancement operations on the new image, such as random rotation, cropping, scaling, adjusting brightness, etc. Mosaic data augmentation methods can improve the performance of object detection tasks, especially for small objects and dense scenes, enriching the dataset of small objects. It can also increase the diversity and complexity of training images, thereby improving the generalization ability of the model.

[0101] In order to verify the effectiveness of the bearing ring surface defect detection model, the present invention uses mean average precision (mAP), average precision (AP) and FPS (frames per second) as measurement indicators, and the confusion matrix is ​​shown in Table 5.

[0102] Table 5 Confusion matrix

[0103]

[0104] In Table 5, TP (True Positive) indicates the number of positive samples that are predicted correctly, FP (False Positive) indicates the number of negative samples that are predicted as positive samples; FN (False Negative) indicates the number of positive samples that are predicted as negative samples; TN (True Negative) indicates the number of negative samples that are predicted as negative samples.

[0105] FPS indicates the number of images that the target detection network can detect per second. The larger the FPS, the more images the target detection network can process per second, and the faster the processing speed.

[0106] The calculation formulas for precision and recall are as follows:

[0107]

[0108]

[0109] The mathematical expressions of AP and mAP are as follows:

[0110]

[0111]

[0112] AP is composed of the area of ​​the PR curve surrounded by precision and recall. mAP represents the average AP value of each category, which is used to measure the detection performance of the model for all categories.

[0113] Different training parameters will affect the performance of the model, including input image size, number of training times, batch size, learning rate, and optimizer used. The present invention uses the parameter training of exp1 in Table 6 in the experiment. In order to verify whether the above parameters are optimal, multiple experiments were conducted by adjusting the following parameters, and the performance changes on the model were observed based on the bearing surface defect dataset. The experimental results are shown in Table 6.

[0114] Table 6 Hyperparameter adjustment

[0115] experiment Input size Number of training sessions Batch size Learning Rate Optimizer mAP exp1 640 100 32 0.01 SGD 97.3% exp2 320 100 32 0.01 SGD 95.2% exp3 640 100 16 0.01 SGD 96.5% exp4 640 100 32 0.1 SGD 96% exp5 640 100 32 0.01 Adam 90.4% exp6 640 100 8 0.01 Adam 89.9% exp7 640 100 16 0.1 SGD 95.6% exp8 640 100 8 0.01 SGD 96.1%

[0116] In the experiment, it was found that when the number of training times was close to 100, the change of the loss function tended to be stable, so the present invention set the number of training times to 100. At the same time, it can be seen from Table 6 that the mAP of the parameter setting of exp1 is the highest, which also verifies that the setting of the experimental parameters of the present invention is reasonable. The mAP of exp6 is the lowest, indicating that the batch size and the choice of optimizer still have a greater impact on the experimental results.

[0117] In summary, the bearing ring surface defect detection model in the present invention makes three improvements to the traditional YOLOv5 target detection network. In order to verify the effectiveness of each improvement and the effectiveness of the combination of the two improvements, an ablation experiment was carried out. The experimental results are shown in Table 7.

[0118] Table 7 Ablation experiment results

[0119]

[0120]

[0121] It can be seen from Table 7 that the mAP of the YOLOv5 model is 95.8%, and the mAP of the improved YOLOv5 model with the C2f module is 96.5%, indicating that the C2f module is helpful for bearing surface defect detection; the mAP of the improved YOLOv5 model using CARAFE upsampling is 96.2%; the mAP of the improved YOLOv5 model with the SPD module is 96.5%; the mAP of the improved YOLOv5 model after using C2f and CARAFE at the same time is 96.9%, indicating that the combination of the two is also helpful to improve the detection of bearing surface defects; the mAP of the improved YOLOv5 model after combining the three modules is as high as 97.3%. The combination of the three not only improves the feature extraction of the backbone network, but also improves the quality and accuracy of upsampling. In the feature fusion stage, more semantic information is integrated into the pyramid layer, more feature details and structural information are retained, and the detection ability of low-resolution images and small objects is improved.

[0122] In order to further verify the effectiveness of the improved YOLOV5 defect detection model, this paper compares it with single-stage target detection methods such as YOLOV3, YOLOV5, YOLOV6 and YOLO7. The training loss, validation loss and mAP curves during the training process are shown in Figure 2. Figure 8 As shown in Table 8, the training and validation loss function curves converge quickly within the first 30 training times and complete convergence when the training times reach 100, while the mAP curve also shows an increasing trend with the increase of training times. The comparison results of each model are shown in Table 8.

[0123] Table 8 Model comparison

[0124] Model Car waste Forging waste Dark spots bump Scratching mAP FPS YOLOv3-tiny 91.20% 87.20% 92.20% 94.50% 96.50% 90.30% 384 YOLOv3 97.5% 88.3% 89.8% 94.8% 87.7% 91.6% 110 YOLOv5 99.5% 96.2% 92.8% 96.7% 93.9% 95.8% 106 YOLOv6n 97.4% 93.8% 90.9% 94.3% 94.6% 94.2% 120 YOLOv7-tiny 98.8% 83.6% 87.5% 88.4% 86.4% 88.9% 157 Model of the present invention 99.4% 96.5% 96.7% 98.9% 95% 97.3% 100

[0125] As can be seen from Table 8, the model of the present invention is significantly better than other target detection networks. The mean average precision on the YOLOv7-tiny model is the lowest, which is 88.9%, which does not meet the detection requirements; the mean average precision of the YOLOv5 model is 95.8%, while the mean average precision of the model of the present invention is 97.3%, and the overall precision is improved by 1.5%. Among them, the precision of black spots is improved by 3.9%, and the precision of bumps is improved by 4%, which greatly improves the defect detection accuracy of black spots and bumps, and only slightly decreases in FPS.

[0126] The present invention randomly selected 5 pictures to test on the above models, and the results are as follows: Fig. 9As shown in the figure, different models have different detection effects on the bearing ring surface defect dataset. YOLOv3-tiny did not detect the scrap defect, and the confidence level when detecting the scratch defect was only 0.45, indicating that the accuracy of the YOLOv3-tiny detection model is low. The YOLOv7-tiny model misdetected the black spot defect when detecting the bump, which also indirectly verifies that the detection accuracy of the YOLOv7-tiny model is low.

[0127] In order to further verify the effectiveness of the model of the present invention, a comparative experiment was conducted with a cloth data set. Similar to the experimental method for bearing ring surface defect detection, the collected cloth data set was first expanded. The size of each cloth image is 400*400 pixels, with a total of 878 images. The data set was expanded to 3317 images by horizontally flipping the images, changing the brightness, etc., and then the data set was divided into a training set, a validation set, and a test set in a ratio of 6:2:2. The comparison between the model of the present invention and the YOLOv5 model is shown in the figure. Fig.10 The comparison results between the model of the present invention and other models are shown in Table 9.

[0128] Table 9 Model comparison

[0129] Model Hole LLine SLine mAP FPS YOLOv3 99.5% 96% 97.5% 97.7% 116 YOLOv3-tiny 99.3% 81.5% 97.5% 92.8% 400 YOLOv5s 99.3% 97.9% 98.9% 98.7% 149 YOLOv6n 98% 95.1% 95.1% 96.3% 124 YOLOv7-tiny 98.7% 94.6% 98.2% 97.2% 164 Model of the present invention 99.5% 98.2% 99.4% 99% 124

[0130] As shown in Table 9, the models of the present invention all show the best results, with mAP reaching 99%, which shows that the models of the present invention are adaptable to objects of different scales and shapes, and thus to different instances and scenarios.

[0131] The principles and implementation methods of the present invention are described in this article using specific examples. The description of the above embodiments is only used to help understand the method and core ideas of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made to the present invention without departing from the principles of the present invention, and these improvements and modifications also fall within the scope of protection of the claims of the present invention.

Claims

1. A high-precision intelligent detection method for surface defects of bearing rings. Features: The method comprises the following steps: A bearing ring surface defect detection model is constructed based on an improved YOLOv5 target detection network. The bearing ring surface defect detection model includes a backbone network for realizing feature extraction, a neck network for multi-scale fusion of different-level features extracted by the backbone network, and a target detection head module for performing target detection and classification. The backbone network is provided with five downsampling units in sequence, the first downsampling unit includes a CBS module, the other four downsampling units include a CBS module and an SPD module arranged in sequence, and a C2f module is arranged after the other four downsampling units. The CBS module includes a convolution layer, a normalization layer and an activation function layer. The SPD module uses an SPD layer to replace the stride convolution in the YOLOv5 target detection network. The C2f module is used to realize separation convolution and splicing operations. The neck network is provided with a CARAFE lightweight general upsampling module. The neck network fuses the feature maps output by the third downsampling unit, the fourth downsampling unit and the fifth downsampling unit in the backbone network. When performing feature fusion, shallow semantic information is transmitted from top to bottom, and deep semantic information is transmitted from bottom to top. Setting a loss function of the bearing ring surface defect detection model, collecting a data set and dividing it into a training set, a validation set, and a test set, and performing multiple rounds of training on the bearing ring surface defect detection model according to the set training parameters; Input the image of the bearing ring to be tested into the trained bearing ring surface defect detection model, and output the surface defect detection result of the bearing ring to be tested; The SPD layer is used to halve the height and width of the input feature map and increase the number of channels of the input feature map by four times.

2. According to claim 1, a high-precision bearing ring surface defect intelligent detection method, Features: The C2f module includes two CBS modules, a Split module and several Bottleneck modules. The feature map input to the first CBS module is split into two sub-feature maps through the Split module, wherein a sub-feature map split by the Split module is output through several Bottleneck modules, and after the sub-feature map output by each Bottleneck module is spliced ​​with another sub-feature map split by the Split module, the feature map after the separation convolution and splicing operation is output through the second CBS module, and the feature map has the same size as the feature map input to the first CBS module.

3. According to claim 1, a high-precision bearing ring surface defect intelligent detection method, Features: The CARAFE lightweight universal upsampling module includes a kernel prediction module and a content-aware reorganization module. The kernel prediction module is used to generate weights on the kernel for reorganization calculation, and the content-aware reorganization module is used to reorganize features according to the calculated weights.

4. According to claim 1, a high-precision bearing ring surface defect intelligent detection method, Features: The mathematical expression of the loss function is: LOSS=w box L box +w obj L obj +w cls L cls Where, L box is the positioning error function, L obj is the confidence loss function, L cls is the classification loss function, w box 、w obj 、w cls are the weight coefficients corresponding to the above functions respectively.

5. A high-precision bearing ring surface defect intelligent detection method according to claim 4, Features: The mathematical expression of the positioning error function is: Where IOU is the intersection-over-union ratio of the predicted box B and the true box A, ρ is the Euclidean distance between the center coordinates of the true box A and the predicted box B, c is the diagonal distance of the minimum box surrounding the center coordinates of the true box A and the predicted box B, α is the weight coefficient, and v is a parameter to measure the consistency of the aspect ratio of A and B.

6. A high-precision bearing ring surface defect intelligent detection method according to claim 4, Features: The classification loss function and the confidence loss function both use a binary cross entropy loss function, and the mathematical expression of the binary cross entropy loss function is: In the formula, n represents the number of input samples, y i represents the target value, x i Represents the predicted output value.

7. A high-precision bearing ring surface defect intelligent detection method according to claim 1, Features: The collecting of data sets includes collecting data sets and classifying the data sets into vehicle scrap, forging scrap, black spots, bumps and scratches according to the types of defects.

8. A high-precision bearing ring surface defect intelligent detection method according to claim 1, Features: The set training parameters include a batch size value of 32, a dynamic parameter of 0.937, a learning rate of 0.01, a cosine annealing learning rate of 0.1, a data enhancement value of 1.0, an image size of 640×640, and a training number of 100 times.

9. A high-precision bearing ring surface defect intelligent detection method according to claim 1, Features: In multiple rounds of training of the bearing ring surface defect detection model, a mosaic data enhancement method is used.