A Visual Inspection Method for Surface Defects of Bearing Rings

By constructing the AS-YOLOv7 algorithm model, the RFL module and SDL module were introduced, the problem of insufficient multi-scale object recognition capability in surface defect detection of bearing rings was solved, and high-precision defect detection and positioning were achieved.

CN116664941BActive Publication Date: 2025-05-30ZHEJIANG SCI-TECH UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310670459.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-07
Publication Date
2025-05-30
Estimated Expiration
2043-06-07

AI Technical Summary

Technical Problem

In the detection of surface defects of bearing rings, the multi-scale target recognition capability is poor and the feature extraction capability is insufficient, resulting in more missed detection and missed detection of small targets.

Method used

AS-YOLOv7 algorithm model is constructed, and by introducing RFL modules at the end of the backbone network unit and SDL modules in the detection head unit, the effective receptive field is expanded, feature extraction capabilities are enhanced, and multi-scale object detection capabilities are improved.

Benefits of technology

The balance of the surface defect detection accuracy of bearing ring surface defects and model inference speed is achieved, the defect categories are accurately detected and the defect area is accurately positioned, which significantly improves the detection accuracy of defects such as forged waste, black spots, scratches and other types.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure SMS_1
    Figure SMS_1
  • Figure SMS_2
    Figure SMS_2
  • Figure SMS_3
    Figure SMS_3
Patent Text Reader

Abstract

The present invention provides a visual inspection method for surface defects of bearing rings, comprising the following steps: constructing an AS-YOLOv7 algorithm model for target detection and classification of bearing ring images, where the AS-YOLOv7 algorithm model includes a backbone network unit, a neck network unit, and a detection head unit; the backbone network unit is provided with an RFL module, the RFL module is located at the end of the backbone network unit, the RFL module includes an ECA-Net module, a RepLKNet module, and a CBS module, the ECA-Net module and the RepLKNet module are arranged in parallel and then serially arranged with the CBS module; the detection head unit is provided with an SDL module; training the AS-YOLOv7 algorithm model multiple times through a data set according to set parameters; inputting the image of the bearing ring to be measured into the trained AS-YOLOv7 algorithm model, and outputting the surface defect detection result of the bearing ring to be measured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to a visual detection method, in particular to a bearing ring surface defect visual detection method, belonging to the technical field of image recognition. Background Art

[0002] In the process of product manufacturing, surface defect detection is an important part of industrial site quality control. As a mechanical component that fixes the rotating body in the mechanical structure and reduces the rotational friction coefficient, bearings are widely used in mechanical equipment to guide the rotation of shaft parts and bear the transmission of shafts to the frame. Their accuracy, performance, life and reliability will seriously affect the stability of the entire mechanical structure. However, in the actual bearing production process, due to the influence of factors such as materials, processing, assembly, and transportation, some defects will inevitably occur on the surface of the bearing ring, such as spiral patterns, forging waste, black spots, dents, scratches, etc. These defects not only affect the service life and performance of the bearing, but once the defective bearings are assembled into mechanical equipment, they may even cause damage to the mechanical equipment. Therefore, surface defect detection after bearing production is necessary.

[0003] At present, most domestic bearing manufacturers rely mainly on manual inspection to detect defects on the surface of bearing rings, but the accuracy and speed of manual inspection will decrease as the inspector's working hours increase. With the development of machine vision and deep learning technology, due to its powerful feature expression ability, generalization ability and cross-scenario ability, it is widely used in defect detection of solar panels, optical films, LCD screens, magnetic tiles, textiles and other products. Using machine vision and deep learning technology for automatic optical inspection of bearing ring surface defects can greatly improve the accuracy and speed of defect detection and optimize the production process of bearing ring surface.

[0004] The texture background of the bearing ring surface is relatively complex, and the black spots and depressions on the bearing ring surface are small target defects. The resolution of spiral patterns, forging scraps, and scratch defects is much greater than that of black spots. The visual detection method for bearing ring surface defects in the existing technology has poor recognition ability for multi-scale targets, and the feature extraction ability needs to be improved, resulting in many small target missed detections and false detections. Summary of the invention

[0005] Based on the above background, the purpose of the present invention is to provide a method for visual detection of bearing ring surface defects, to achieve a balance between bearing ring surface defect detection accuracy and model reasoning speed, to accurately detect defect categories and precisely locate defect areas.

[0006] In order to achieve the above-mentioned object of the invention, the present invention provides the following technical solutions:

[0007] A method for visually detecting surface defects of a bearing ring, the method comprising the following steps:

[0008] Construct an AS-YOLOv7 algorithm model for object detection and classification of bearing ring images. The AS-YOLOv7 algorithm model is improved based on the YOLOv7 model. The AS-YOLOv7 algorithm model includes a backbone network unit for feature extraction, a neck network unit for multi-scale fusion of different-level features extracted by the backbone network unit, and a detection head unit for performing object detection and classification. The backbone network unit is provided with an RFL module, which is located at the end of the backbone network unit. The RFL module includes an ECA-Net module, a RepLKNet module, and a CBS module. The ECA-Net module and the RepLKNet module are arranged in parallel and then serially arranged with the CBS module. The detection head unit is provided with an SDL module, which includes an SPDConv module, a CBS module, and an ODConv module arranged in series.

[0009] Set the parameters of the AS-YOLOv7 algorithm model, collect the data set, and perform multiple rounds of training on the AS-YOLOv7 algorithm model through the data set according to the set parameters until the AS-YOLOv7 algorithm model reaches the set measurement index, and the training is completed.

[0010] Input the image of the bearing ring to be measured into the trained AS-YOLOv7 algorithm model, and output the surface defect detection result of the bearing ring to be measured.

[0011] In this visual detection method for surface defects of bearing rings, an AS-YOLOv7 algorithm model suitable for object detection and classification of bearing ring images is obtained by improving the YOLOv7 model. By introducing an RFL module at the end of the model backbone network unit, the effective receptive field of the model is expanded, and the feature extraction ability of the model is enhanced. Specifically, the RepLKNet module in the RFL module introduces a 31×31 super-large convolution kernel to expand the effective receptive field of the model, while the ECA-Net module reduces the model complexity and avoids the interference of invalid information such as the surface background texture of the bearing ring in the bearing ring image. By using the SDL module to replace the original detection head unit of the YOLOv7 model, the downstream task performance of the model is optimized, the expression ability of the model is enhanced, and the detection ability of the model for multi-scale targets is improved, thus effectively solving the problems of large resolution span of surface defects and large proportion of small target defects in bearing ring images. Specifically, the combination of the SPDConv module and the CBS module in the SDL module reduces the loss of fine-grained feature information, improves the detection ability of the model for small targets, and amplifies the channel number of the output features to 4C, and then inputs them into the ODConv module. The ODConv module can reduce the additional parameters of the model, reduce the model complexity, and improve the expression ability of the model.

[0012] Preferably, the ECA-Net module realizes spatial feature compression by performing global average pooling on the input feature image in the spatial dimension, then captures cross-channel interaction information through one-dimensional convolution on the compressed feature image and assigns different channel weights, generates a new feature image through an activation function, and finally multiplies the generated new feature image with the original input feature image channel by channel to obtain the feature image of the final dimension; the RepLKNet module includes a Stem sub-module, four Stage sub-modules, and three Transition sub-modules arranged in series. Among them, one Stage sub-module is connected to the Stem sub-module, and two adjacent Stage sub-modules are connected by one Transition sub-module. The Stem sub-module is used for upsampling and size reduction of the input image, the Transition sub-module is used for image downsampling, and the Stage sub-module is stacked by RepLK Block layers and ConvFFN layers.

[0013] Preferably, the SPDConv module includes a depth convolution layer and a non-strided convolution layer arranged in series; the ODConv module is a full-dimensional dynamic convolution module, and the ODConv module learns along all four dimensions of the kernel space at any convolution layer through a multi-dimensional attention mechanism and a parallel strategy.

[0014] Preferably, the backbone network unit is further provided with a plurality of CBS modules, a plurality of ELAN modules, a plurality of MPconv modules, and one SPPCSPC module. The RFL module is located after the last ELAN module arranged in series and before the SPPCSPC module.

[0015] Preferably, setting the parameters of the AS-YOLOv7 algorithm model includes setting the training parameters of the AS-YOLOv7 algorithm model. The training parameters include: initial learning rate 0.1, minimum learning rate 0.01, batch size value 32, dynamic parameter 0.937, weight decay parameter 0.0005, optimizer SGD, and number of training epochs 300.

[0016] Preferably, setting the parameters of the AS-YOLOv7 algorithm model further includes setting the loss function of the AS-YOLOv7 algorithm model. The mathematical expression of the loss function is,

[0017] LOSS = w box L box + w obj L obj + w cls L cls

[0018] In the formula, L box is the positioning error function, Lobj is the confidence loss function, L cls is the classification loss function, w box , w obj , w cls are the weight coefficients corresponding to the above functions respectively;

[0019] The mathematical expression of the positioning error function is

[0020]

[0021] In the formula, IOU is the intersection over union of the predicted bounding box B and the ground truth bounding box A, ρ is the Euclidean distance between the center point coordinates of the ground truth bounding box A and the predicted bounding box B, c is the diagonal distance of the smallest rectangle enclosing the center point coordinates of the ground truth bounding box A and the predicted bounding box B, α is the weight coefficient, and v is a parameter for measuring the aspect ratio consistency between A and B;

[0022] Both the classification loss function and the confidence loss function adopt the binary cross-entropy loss function, and the mathematical expression of the binary cross-entropy loss function is

[0023]

[0024] In the formula, n represents the number of input samples, y i represents the target value, and x i represents the predicted output value.

[0025] Preferably, when the AS-YOLOv7 algorithm model is trained in multiple rounds through the dataset according to the set parameters, the dataset is divided into a training set, a validation set, and a test set in a ratio of 7:2:1, and mosaic data augmentation processing is performed to enrich the training set.

[0026] Preferably, the dataset is divided according to the defect type, and the defect types include spiral marks, forging scraps, black spots, dents, and scratches.

[0027] Preferably, the mosaic data augmentation processing includes: randomly extracting 4 pictures from the training set, performing random scaling, random cropping, and random arrangement transformations on the pictures, and randomly selecting a picture splicing point, and splicing the transformed pictures into the same window according to the picture splicing point to form a new spliced picture.

[0028] Preferably, the measurement indicators include mean average precision mAP, average precision AP, and frames per second FPS.

[0029] Compared with the prior art, the present invention has the following advantages:

[0030] A visual inspection method for surface defects of bearing rings, based on the improvement of the YOLOv7 model, obtains the AS-YOLOv7 algorithm model suitable for target detection and classification of bearing ring images, realizes the balance between the detection accuracy of surface defects of bearing rings and the model inference speed, can not only accurately detect the defect categories, but also achieve precise positioning of the defect areas, especially significantly improving the detection accuracy for surface defects of bearing rings such as forging waste, black spots, scratches, etc.;

[0031] In the present invention, by introducing the RFL module at the end of the model backbone network unit, the effective receptive field of the model is expanded, the feature extraction ability of the model is enhanced, and the problem of complex background texture and difficult feature extraction on the surface of bearing rings is solved;

[0032] In the present invention, by using the SDL module to replace the original detection head unit of the YOLOv7 model, the downstream task performance of the model is optimized, the expression ability of the model is enhanced, and the detection ability of the model for multi-scale targets is improved, thereby effectively solving the problems of large resolution span of surface defects and large proportion of small target defects in bearing ring images. Brief Description of the Drawings

[0033] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained according to the provided drawings.

[0034] Figure 1 is a schematic flow diagram of the visual inspection method for surface defects of bearing rings of the present invention;

[0035] Figure 2 is a schematic structural diagram of the AS-YOLOv7 algorithm model in the present invention;

[0036] Figure 3 is a schematic structural diagram of the ECA-Net module in the present invention;

[0037] Figure 4 is a schematic structural diagram of the RepLKNet module in the present invention;

[0038] Figure 5 is a schematic structural diagram of the RFL module in the present invention;

[0039] Figure 6 is a schematic working principle diagram of the SPDConv module in the present invention;

[0040] Figure 7 is a schematic working principle diagram of the ODConv module in the present invention;

[0041] Figure 8 It is a schematic structural diagram of the SDL module in the present invention;

[0042] Figure 9 It is a schematic diagram of a sample of the defect type in the present invention, Figure 9 (a) is a schematic diagram of a spiral thread defect, Figure 9 (b) is a schematic diagram of a forging waste defect, Figure 9 (c) is a schematic diagram of a black spot defect, Figure 9 (d) is a schematic diagram of a dent defect, Figure 9 (e) is a schematic diagram of a scratch defect;

[0043] Figure 10 It is a schematic flow diagram of mosaic data enhancement processing in the present invention;

[0044] Figure 11 It is a curve graph of the training results of the AS-YOLOv7 algorithm model in the present invention;

[0045] Figure 12 It is a comparison diagram of the visual detection effects of the AS-YOLOv7 algorithm model and other existing models on the surface defects of bearing rings in the present invention. Specific Embodiments

[0046] The technical solution of the present invention will be further specifically described below through specific embodiments and in conjunction with the accompanying drawings. It should be understood that the implementation of the present invention is not limited to the following embodiments, and any form of modification and / or change made to the present invention will fall within the protection scope of the present invention.

[0047] In the present invention, unless otherwise specified, all parts and percentages are in weight units, and the equipment and raw materials used can be purchased from the market or are commonly used in the art. The methods in the following embodiments are conventional methods in the art unless otherwise specified. The components or equipment in the following embodiments are universal standard parts or components known to those skilled in the art unless otherwise specified, and their structures and principles can all be known by those skilled in the art through technical manuals or through conventional experimental methods.

[0048] The following will make a detailed description of the embodiments of the present invention in conjunction with the accompanying drawings. In the following detailed description, for the sake of explanation, many specific details are set forth to provide a comprehensive understanding of the embodiments of the present invention. However, one or more embodiments can also be implemented by those skilled in the art without these specific details.

[0049] As Figure 1 shown, the present invention discloses a method for visual detection of surface defects of bearing rings, and the method includes the following steps:

[0050] S1. Construct an AS-YOLOv7 algorithm model for object detection and classification of bearing ring images. The AS-YOLOv7 algorithm model is improved based on the YOLOv7 model. The AS-YOLOv7 algorithm model includes a backbone network unit for feature extraction, a neck network unit for multi-scale fusion of different-level features extracted by the backbone network unit, and a detection head unit for object detection and classification. The backbone network unit is equipped with an RFL module, which is located at the end of the backbone network unit. The RFL module includes an ECA-Net module, a RepLKNet module, and a CBS module. The ECA-Net module and the RepLKNet module are set in parallel and then serially set with the CBS module. The detection head unit is equipped with an SDL module, which includes an SPDConv module, a CBS module, and an ODConv module set in series.

[0051] S2. Set the parameters of the AS-YOLOv7 algorithm model, collect the data set, and train the AS-YOLOv7 algorithm model through the data set according to the set parameters for multiple rounds until the AS-YOLOv7 algorithm model reaches the set measurement index, and the training is completed.

[0052] S3. Input the image of the bearing ring to be measured into the trained AS-YOLOv7 algorithm model, and output the surface defect detection result of the bearing ring to be measured.

[0053] In this visual detection method for surface defects of bearing rings, an AS-YOLOv7 algorithm model suitable for object detection and classification of bearing ring images is obtained by improving the YOLOv7 model. This method expands the effective receptive field of the model and enhances the feature extraction ability of the model by introducing an RFL module at the end of the backbone network unit of the model. This method optimizes the downstream task performance of the model, enhances the expression ability of the model, and improves the detection ability of the model for multi-scale targets by using an SDL module to replace the original detection head unit of the YOLOv7 model, thus effectively solving the problems of large resolution span of surface defects and large proportion of small target defects in bearing ring images.

[0054] The following will gradually elaborate on this method in detail.

[0055] 1. Based on the improvement of the YOLOv7 model, construct an AS-YOLOv7 algorithm model for object detection and classification of bearing ring images

[0056] This method uses the YOLOv7 model as the baseline neural network model. As one of the latest basic models in the YOLO model series, the YOLOv7 model exceeds most known real-time object detectors in terms of detection speed and accuracy in the range of 5 FPS to 160 FPS. Among the known real-time object detectors with more than 30 frames per second, the YOLOv7 model also has the highest accuracy rate.

[0057] This method constructs the AS-YOLOv7 algorithm model as shown in Figure 2 by introducing the RFL module into the backbone network unit and introducing the SDL module to replace the detection head unit of the traditional YOLOv7 model. The backbone network unit is provided with multiple CBS modules, multiple ELAN modules, multiple MPconv modules, and one SPPCSPC module. The RFL module is located after the last ELAN module arranged in series and before the SPPCSPC module. The basic task of the backbone network unit is to extract image features and transmit the extracted image features to the neck network unit. The neck network unit is the same as the traditional YOLOv7 model and adopts the PAFPN structure for stack scaling in the neck network unit. The neck network unit obtains three types of features with large, medium, and small sizes through the fusion processing of high-level features and low-level features, and transmits the fused features to the detection head unit respectively to achieve the integration of high-resolution information and high-semantic information. The basic task of the detection head unit with the SDL module is to decouple the high-resolution information transmitted by the neck network unit and detect the category and location of the target. The RFL module and the SDL module are described in detail below.

[0058] 1.1. ECA-Net Module

[0059] The ECA-Net module realizes spatial feature compression by performing global average pooling on the input feature image in the spatial dimension, then captures cross-channel interaction information through one-dimensional convolution on the compressed feature image and assigns different channel weights, generates a new feature image through the activation function, and finally multiplies the generated new feature image with the original input feature image channel by channel to obtain the feature image of the final dimension.

[0060] Specifically, the structure of the ECA-Net module is as shown in Figure 3As shown below. First, the feature image with an input dimension of C×H×W is subjected to global average pooling (GAP) in the spatial dimension to obtain a feature image of 1×1×C, realizing spatial feature compression. Secondly, the compressed feature image is passed through a one-dimensional convolution k. In this embodiment, k = 5, which captures interaction information across channels, assigns weights to different channels, and generates a feature image of size 1×1×C through the activation function σ. Finally, the generated feature image of 1×1×C is multiplied channel by channel with the original input feature image C×H×W to obtain the final feature image with a dimension of C×H×W. The ECA-Net module captures interaction information across channels, obtains higher accuracy with lower complexity, and can improve the performance of various deep CNN architectures. The size k of its convolution kernel is adaptively determined by the channel dimension, and the calculation formula of the convolution kernel k is as shown in the formula:

[0061]

[0062] where c represents the channel dimension, a represents the constant 2, b represents the constant 1, and |n| odd The odd number closest to n.

[0063] 1.2. RepLKNet Module

[0064] The RepLKNet module, as a pure CNN architecture module, has a convolution kernel size of 31×31, and its specific structure is as Figure 4 shown. The RepLKNet module includes a Stem sub-module, four Stage sub-modules, and three Transition sub-modules arranged in series. Among them, one Stage sub-module is connected to the Stem sub-module, and two adjacent Stage sub-modules are connected by a Transition sub-module. The Stem sub-module is used to increase the dimension and reduce the size of the input image, and the Transition sub-module is used for image downsampling. The Stage sub-module is stacked by RepLK Block layers and ConvFFN layers. In the RepLKNet module, except for the depth-wise (DW) super-large convolution, other modules including DW3×3 convolution, 1×1 convolution, and batch normalization modules mostly have small convolution kernels, with simple structures and few parameters. In addition, the RepLKNet module adopts a reparameterization structure to increase the convolution kernel size, expand the effective receptive field and shape deviation, and at the same time introduces a short_cut layer, which ensures the detection efficiency, effectively improves the detection accuracy, and enhances the performance of the network's downstream tasks.

[0065] 1.3. RFL Module

[0066] The structure of the RFL module is as Figure 5As shown, it is composed of an ECA-Net module, a RepLKNet module, and a 3×3 convolution Conv. The number of channels of Conv is 256. The RepLKNet module in the RFL module introduces a 31×31 super-large convolution kernel to expand the effective receptive field of the model, while the ECA-Net module reduces the model complexity and avoids the interference of invalid information such as the surface background texture of the bearing ring in the bearing ring image.

[0067] 1.4. SPDConv Module

[0068] The SPDConv module includes a depth convolution layer and a non-strided convolution layer arranged in series. It replaces the strided convolution and pooling layer operations in the traditional CNN architecture, retains all information in the channel dimension, and effectively avoids the phenomenon of loss of fine-grained feature information and decline in network feature expression ability caused by the use of strided convolution and pooling layers in the traditional CNN architecture. The working principle of the SPDConv module is as Figure 6 shown. Taking the case of scale = 2 as an example, consider the intermediate feature map X of size S×S×C 1 . A series of sub-feature maps are sliced out through depth convolution. The definition of the sub-feature map is shown in the following formula:

[0069] f scale-1,scale-1 = X[scale - 1:S:scale - 1:S:scale]

[0070] Four sub-feature maps f 0,0 , f 1,0 , f 0,1 , f 1,1 are obtained through the above formula. The sub-feature maps are interconnected through the channel dimension to obtain a feature map X 0 , and are input into the non-strided convolution layer. The feature map X 0 is further transformed into a feature map through the C2 filter . The finally transformed map X″ outputs SPDConv.

[0071] 1.5. ODConv Module

[0072] The ODConv module is a full-dimensional dynamic convolution module. The ODConv module learns along all four dimensions of the kernel space in any convolution layer through a multi-dimensional attention mechanism and a parallel strategy. In the ODConv module, four different types of attention mechanisms are respectively added to four different dimensions. Its attention mechanism is as Figure 7 shown. Figure 7 (a) represents the spatial coordinate multiplication operation along the spatial dimension, Figure 7 (b) represents the channel multiplication operation along the input channel dimension, Figure 7(c) represents the filter multiplication operation along the output channel dimension, Figure 7 (d) represents the convolution kernel dimension multiplication operation along the spatial dimension of the convolution kernel. By introducing the above four attention mechanisms, the additional parameters of the network can be reduced, the representation ability of the network can be improved, and the feature extraction ability of the basic convolution operation can be enhanced. The definition of the ODConv module is shown as follows:

[0073] y = (α w1 ⊙α f1 ⊙α c1 ⊙α s1 ⊙W 1 +…+α wn ⊙α fn ⊙α cn ⊙α sn ⊙W n ) * x

[0074] where α w1 ∈R represents the attention scalar of the convolution kernel W 1 , α s1 ∈R k ×k, α ci ∈R cin and α f1 ∈R cout represent the attention mechanisms calculated along the spatial dimension, input channel dimension, and output channel dimension respectively, and ⊙ represents the multiplication operation of different dimensions of the convolution kernel space.

[0075] 1.6, SDL Module

[0076] The structure of the SDL module is as Figure 8 shown. It combines the SPDConv module, the 3×3 convolution Conv, and the ODConv module. The combination of the SPDConv module in the SDL module and the 3×3 convolution Conv as the CBS module reduces the loss of fine-grained feature information, improves the detection ability of the model for small targets, and amplifies the number of output feature channels to 4C, and then inputs them into the ODConv module. The ODConv module can reduce the additional parameters of the model, reduce the model complexity, and improve the expression ability of the model.

[0077] 2. Preparation of the Dataset

[0078] 2.1. Dataset Source

[0079] At the end of the bearing ring production line in the industrial field, a area array camera is used to collect the surface pictures of the completed bearing rings and send them to the host computer as the dataset and prepare the dataset division.

[0080] 2.2. Dataset Division

[0081] In this embodiment, the defect types are divided into five types: spiral marks, forging waste, black spots, dents, and scratches. The data set is divided based on these, and the statistical data of various defect samples after division are shown in Table 1.

[0082] Table 1 Defect type samples of the data set

[0083] Spiral thread Forging waste Black spot Indentation Scratch Quantity 576 225 491 634 485

[0084] According to the number of samples in the data set and the rationality of training, each defect sample is divided into a training set, a validation set, and a test set, and the division ratio is 7:2:1. The results are shown in Table 2.

[0085] Table 2 Division of the data set

[0086] Training Verification Testing Total Spiral thread 405 114 57 576 Forging waste 159 44 22 225 Black spot 344 98 49 491 Indentation 445 126 63 634 Scratch 341 96 48 485

[0087] This data set contains a total of 2411 defect images, and is divided into five types of defects: spiral marks, forging waste, black spots, dents, and scratches according to the defect types. Among them, there are 576 spiral mark defect images, 225 forging waste defects, 491 black spot defects, 634 dent defects, and 485 scratch defects. Typical defect sample examples in the data set are as Figure 9 shown, where Figure 9 (a) is a spiral mark defect, Figure 9 (b) is a forging waste defect, Figure 9 (c) is a black spot defect, Figure 9 (d) is a dent defect, Figure 9 (e) is a scratch defect.

[0088] 2.3. Making data set labels

[0089] In this embodiment, it is necessary to make labels for the data set. Use LabelImg to make labels, mark the defect positions and defect types on LabelImg, and then export to generate label files.

[0090] 3. Set model parameters and train the model

[0091] 3.1. Mosaic data augmentation processing

[0092] During the training process, in order to improve the robustness of the model and the detection accuracy, data augmentation can be added to the input end of the model. This embodiment adopts mosaic data augmentation processing. Specifically, as Figure 10As shown in the figure, by randomly selecting 4 pictures from the training set, the selected pictures are randomly scaled, randomly cropped, randomly arranged, and a random picture splicing point is selected. Finally, the transformed pictures are spliced into the same window according to the splicing point. Through steps such as random scaling, this processing method adds more small samples to the network, makes the sample distribution more uniform, and improves the robustness of the network; and splicing 4 pictures in one window speeds up the convergence speed of the network.

[0093] 3.2. Loss Function

[0094] In this embodiment, the loss function of the AS-YOLOv7 algorithm model is set, and the mathematical expression of the loss function is

[0095] LOSS = w box L box + w obj L obj + w cls L cls

[0096] In the formula, L box is the positioning error function, L obj is the confidence loss function, L cls is the classification loss function, w box 、w obj 、w cls are the weight coefficients corresponding to the above functions respectively;

[0097] The mathematical expression of the positioning error function is

[0098]

[0099] In the formula, IOU is the intersection over union of the predicted box B and the ground truth box A, ρ is the Euclidean distance between the center point coordinates of the ground truth box A and the predicted box B, c is the diagonal distance of the smallest rectangle enclosing the center point coordinates of the ground truth box A and the predicted box B, α is the weight coefficient, and v is a parameter to measure the aspect ratio consistency of A and B;

[0100] Both the classification loss function and the confidence loss function adopt the binary cross-entropy loss function, and the mathematical expression of the binary cross-entropy loss function is

[0101]

[0102] In the formula, n represents the number of input samples, y i represents the target value, and x i represents the predicted output value.

[0103] 3.3. Training Parameters

[0104] During the training process, the hardware environment and software configuration are as follows: The processor is an Intel(R) Core(TM) i7-10750H CPU @ 2.60 GHz, the memory is 32 GB, the graphics card model is Nvidia GTX 3090Ti (single card), the video memory is 24 GB, and the disk size is 1T. The operating system is Windows 11 (64-bit), the Compute Unified Device Architecture (CUDA) version is 11.7, the cuDNN version is 8.6.0, the deep learning framework uses Pytorch 1.13.1, and the compiler is Python 3.7. The training parameters include: initial learning rate 0.1, minimum learning rate 0.01, batch size value 32, dynamic parameter 0.937, weight decay parameter 0.0005, the optimizer is SGD, and the number of training epochs is 300.

[0105] The training result curve for the training set is as Figure 11 shown. Figure 11 In it, the upper row of graphs is the precision curve and recall rate curve during training, and the lower row of graphs is the mean average precision curve (shown in two different calculation methods). Through Figure 11 it can be seen that the recall rate curve converges rapidly within the first 25 epochs, the precision function curve converges rapidly within the first 50 epochs, and both reach complete convergence at around 100 epochs. The mean average precision curve of the AS-YOLOv7 algorithm model reaches complete convergence at around 75 epochs, which proves the advantages of the model in this embodiment requiring less training amount and faster convergence.

[0106] To verify the effectiveness of the AS-YOLOv7 algorithm model, the mean average precision mAP, average precision AP, and frames per second FPS are used as measurement indicators.

[0107] The definitions of AP and mAP are as follows:

[0108] AP = ∫ 0 1 P(R)dR

[0109]

[0110] Among them, AP represents the area between the PR (Precision-Recall) curve and the coordinate axes, mAP represents the average value of AP for surface defects of different types of bearing rings, and N represents the number of classes of test samples. In this embodiment, N = 5 is set.

[0111] This example builds a deep learning environment based on Pytorch and runs on GPU to obtain data using the trained model. In order to further verify the effectiveness of the AS-YOLOv7 algorithm model, the model is compared with single-stage target detection method models such as YOLOv5 and YOLOv7. The comparative experimental results are shown in Table 5.

[0112] Table 5 Comparison of model effectiveness

[0113]

[0114]

[0115] As shown in Table 5, on the bearing ring surface data set, the detection accuracy of the YOLOv5s and YOLOv5l networks is not high, and the reasoning speed is slow, only 85FPS and 79FPS respectively, and the overall performance is not good; the reasoning speed of YOLOv7 is faster, 122FPS, but the detection accuracy of YOLOv7 is only 96.1%. Although YOLOv7-X has good detection accuracy, the reasoning speed is slow. The AS-YOLOv7 algorithm model of this embodiment has an overall detection accuracy of 2.1% higher than that of YOLOv7, reaching 98.2%, among which the detection accuracy of forging defects is increased by 3.2%, the detection accuracy of black spot defects is increased by 5.2%, and the detection accuracy of scratch defects is increased by 1.8%. The detection effect of small targets, multi-scale targets and low-contrast defects has been significantly improved. In addition, the speed of the AS-YOLOv7 algorithm model is only lower than that of YOLOv7, and the FPS reaches 114FPS, which is significantly higher than YOLOv5l, YOLOv5s and YOLOv7-X. It can be seen that compared with other models, the AS-YOLOv7 algorithm model of this embodiment has better detection performance.

[0116] To further verify the reliability of the AS-YOLOv7 algorithm model in this embodiment, 5 pictures were randomly selected for testing on different models. The results are as follows: Figure 12 shown. Figure 12 The first row shows 5 randomly selected sample images, the second row shows the real frame positions of the defect areas in the 5 images, and rows 3-7 show the prediction results of each model. The marked numbers are the confidence levels of the predictions. The higher the confidence level, the higher the possibility that the image is the target. Figure 11 From left to right in the figure are spiral marks, forging waste, black spots, dents, and scratches. The AS-YOLOv7 algorithm model of this embodiment is the last row, and its confidence levels are 0.96, 0.90, 0.95, 0.93, and 0.94, respectively.

[0117] Obviously, by Figure 12It can be seen that different models have different detection effects on the bearing ring defect dataset. YOLOv5 and YOLOv5l did not detect indentation defects, and the confidence of YOLOv7 for indentation defects and scratch defects is low, making it difficult to locate them. However, the AS-YOLOv7 algorithm model of this embodiment is significantly superior to other models in terms of overall detection effect.

[0118] In this paper, specific examples are used to elaborate on the principle and implementation mode of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention. It should be pointed out that for those of ordinary skill in the art of this technology, without departing from the principle of the present invention, several improvements and modifications can be made to the present invention, and these improvements and modifications also fall within the protection scope of the claims of the present invention.

Claims

1. A visual inspection method for surface defects of bearing rings, characterized in that: This method includes the following steps: Construct an AS-YOLOv7 algorithm model for object detection and classification of bearing ring images. The AS-YOLOv7 algorithm model is improved based on the YOLOv7 model. The AS-YOLOv7 algorithm model includes a backbone network unit for feature extraction, a neck network unit for multi-scale fusion of different-level features extracted by the backbone network unit, and a detection head unit for performing object detection and classification. The backbone network unit is provided with an RFL module, and the RFL module is located at the end of the backbone network unit. The RFL module includes an ECA-Net module, a RepLKNet module, and a CBS module. The ECA-Net module and the RepLKNet module are arranged in parallel and then serially arranged with the CBS module. The detection head unit is provided with an SDL module, and the SDL module includes an SPDConv module, a CBS module, and an ODConv module arranged in series. Set the parameters of the AS-YOLOv7 algorithm model, collect the data set, and perform multiple rounds of training on the AS-YOLOv7 algorithm model through the data set according to the set parameters until the AS-YOLOv7 algorithm model reaches the set measurement index, and the training is completed. Input the image of the bearing ring to be measured into the trained AS-YOLOv7 algorithm model, and output the surface defect detection result of the bearing ring to be measured.

2. A visual inspection method for surface defects of bearing rings according to claim 1, characterized in that: The ECA-Net module realizes spatial feature compression through global average pooling of the input feature image in the spatial dimension, then captures cross-channel interaction information through one-dimensional convolution for the compressed feature image and assigns different channel weights, generates a new feature image through an activation function, and finally multiplies the generated new feature image with the original input feature image channel by channel to obtain the feature image of the final dimension. The RepLKNet module includes a Stem sub-module, four Stage sub-modules, and three Transition sub-modules arranged in series. Among them, one Stage sub-module is connected to the Stem sub-module, and adjacent two Stage sub-modules are connected by a Transition sub-module. The Stem sub-module is used for dimension increase and size reduction of the input image, the Transition sub-module is used for image downsampling, and the Stage sub-module is stacked by RepLK Block layers and ConvFFN layers.

3. A visual inspection method for surface defects of bearing rings according to claim 1, characterized in that: The SPDConv module includes a depth convolution layer and a non-strided convolution layer arranged in series. The ODConv module is a full-dimensional dynamic convolution module, and the ODConv module learns along all four dimensions of the kernel space at any convolution layer through a multi-dimensional attention mechanism and a parallel strategy.

4. A method for visually detecting surface defects of a bearing ring according to claim 1, Features: The backbone network unit is also provided with a plurality of CBS modules, a plurality of ELAN modules, a plurality of MPconv modules and a SPPCSPC module, and the RFL module is located after the last ELAN module arranged in series and before the SPPCSPC module.

5. A method for visually detecting surface defects of a bearing ring according to claim 1, Features: Setting the parameters of the AS-YOLOv7 algorithm model includes setting the training parameters of the AS-YOLOv7 algorithm model, and the training parameters include: an initial learning rate of 0.1, a minimum learning rate of 0.01, a batch size value of 32, a dynamic parameter of 0.937, a weight decay parameter of 0.0005, an optimizer of SGD, and a number of training rounds of 300.

6. A method for visually detecting surface defects of a bearing ring according to claim 1 or 5, Features: Setting the parameters of the AS-YOLOv7 algorithm model also includes setting the loss function of the AS-YOLOv7 algorithm model. The mathematical expression of the loss function is: LOSS=w box L box +w obj L obj +w cls L cls where L box is the positioning error function, L obj is the confidence loss function, L cls is the classification loss function, and w box , w obj , and w cls are the weight coefficients corresponding to the above functions, respectively; The mathematical expression of the positioning error function is: Where IOU is the intersection-over-union ratio of the predicted box B and the real box A, ρ is the Euclidean distance between the center coordinates of the real box A and the predicted box B, c is the diagonal distance of the minimum box surrounding the center coordinates of the real box A and the predicted box B, α is the weight coefficient, and v is a parameter to measure the consistency of the aspect ratio of A and B; The classification loss function and the confidence loss function both use a binary cross entropy loss function, and the mathematical expression of the binary cross entropy loss function is: where n represents the number of input samples, and y i represents the target value, and x i represents the predicted output value.

7. A method for visually detecting surface defects of a bearing ring according to claim 1, Features: When the AS-YOLOv7 algorithm model is trained for multiple rounds through the data set according to the set parameters, the data set is divided into training set, validation set and test set in a ratio of 7:2:1, and mosaic data enhancement processing is performed to enrich the training set.

8. A method for visually detecting surface defects of a bearing ring according to claim 7, Features: The data set is divided by defect type, which includes spiral marks, forging scraps, black spots, dents and scratches.

9. A method for visually detecting surface defects of a bearing ring according to claim 7, Features: The mosaic data enhancement processing includes: randomly extracting 4 pictures from the training set, randomly scaling, randomly cropping, and randomly arranging the pictures, and randomly selecting a picture splicing point, and splicing the transformed pictures into the same window according to the picture splicing point to form a spliced ​​new picture.

10. A method for visually detecting surface defects of a bearing ring according to claim 1, Features: The measurement indicators include mean average precision (mAP), average precision (AP) and frames per second (FPS).