YOLO-Bush rubber bushing surface defect detection method and system
Through the improved YOLO-Bush model and enhanced feature extraction fusion network, the problems of rubber bushing detection accuracy and efficiency are solved, and high-precision and low missed detection rate rubber bushing surface defect detection is achieved, which is suitable for real-time online detection of rubber bushings.
Patent Information
- Application Number
- CN202510746769.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-05
- Publication Date
- 2025-09-16
AI Technical Summary
Existing rubber bushing detection methods perform poorly in terms of detection accuracy and efficiency, and are difficult to meet the needs of high reliability and real-time online detection. This is mainly due to the special structure of rubber bushings, the need for defects to appear under pressure, the complex end face structure, and the small defect targets.
A special test bench for rubber bushings was built, and an enhanced feature extraction and fusion network was constructed. The improved YOLO-Bush model was adopted to improve the feature extraction and fusion capabilities through multi-dimensional collaborative attention modules, multi-order gated aggregation networks, improved position-enhanced attention mechanisms, and sub-pixel downsampling convolutions.
High-precision and efficient detection of rubber bushing surface defects is achieved, with a detection accuracy of 99.4%, an F1 value of 0.99, and an FPS value of 27.264, meeting the real-time online detection needs and reducing the missed detection rate.
Smart Images

Figure CN120655599A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to but is not limited to the field of image processing detection technology, and in particular relates to a YOLO-Bush rubber bushing surface defect detection method and system. Background Art
[0002] As the role of automobiles in our daily lives becomes increasingly important, the reliability requirements for rubber bushings, as key cushioning and vibration isolation components of automobiles, are also increasing. In the actual production process, due to various factors such as equipment performance, environmental conditions, and process flow, defects such as surface bubbles and mid-piece crushing are prone to appear on the surface of rubber bushings. Such surface defects not only affect the appearance quality and commercial value of the product, but may also have an adverse effect on the comfort and driving safety of the entire vehicle. Therefore, efficient and accurate detection of surface defects of rubber bushings is of great practical significance. At present, the detection of surface defects of rubber bushings is still mainly based on manual visual inspection. This method has problems such as low detection efficiency, poor accuracy, lack of unified quality evaluation standards, and high management difficulty. It is difficult to meet the high standards of product quality control required by modern production enterprises.
[0003] With the rapid development of artificial intelligence (AI) technology and the continuous improvement of machine vision algorithms, machine vision has demonstrated remarkable performance advantages in fields such as face recognition, image classification, and object detection. In the field of industrial inspection, machine vision, as a key automation tool, employs a diverse range of detection methods and is increasingly widely used. For example, Liu Chun et al. proposed a detection process for rubber sealing ring defect detection based on image acquisition, improved median filtering for denoising, Canny edge detection, and morphological dilation operations. This approach achieves non-contact, online, and objective automatic defect recognition. Zhang Hong et al. analyzed the tilt angle changes of contour tangents to accurately locate and screen true burr points. This method, combined with the minimum bounding rectangle height to determine defects, effectively improved the detection accuracy of sealing ring burr defects. Although traditional image processing algorithms were proposed long ago and have achieved a high level of technical maturity, they have been widely used in the defect detection field and achieved certain results. However, these methods often rely on complex image processing processes and are susceptible to external environmental changes and parameter settings. Their robustness and generalization capabilities are weak, making them difficult to adapt to the high-precision inspection requirements of complex workpieces such as rubber bushings. In recent years, Cheng et al. designed a rubber bushing defect detection model based on lightweight YOLOv7. The model uses MobileNetv3 and RC-block structure for efficient feature extraction, and introduces BiFPN (bidirectional feature pyramid network) to achieve multi-scale feature fusion, thereby improving detection performance. The YOLO (You Only Look Once) series of algorithms are widely used in mobile terminals and embedded devices due to their advantages such as fast detection speed, lightweight structure, and easy deployment. However, this type of algorithm still has certain limitations in actual industrial inspection, such as high missed detection rate and low detection accuracy, which makes it difficult to meet the application requirements of high-reliability defect identification. Due to the special structure of rubber bushings, defects can only appear under pressure. At the same time, its end face structure is complex, the defect target is small, and the detection accuracy requirement is high. In response to the above problems, a rubber bushing surface defect detection method based on an improved YOLOxs model was studied. Summary of the Invention
[0004] The detection of existing rubber bushings faces the following problems: (1) Due to the special structure of rubber bushings, their main defects can only be seen under pressure. (2) The end face structure of rubber bushings is complex, with a variety of detailed textures and geometric features, which increases the difficulty of defect detection; (3) The defect targets are usually small in size. In response to the above difficult-to-solve problems, the existing detection methods have poor performance in detection accuracy and efficiency, and are difficult to meet the real-time online detection needs of enterprises. The present invention proposes a rubber bushing defect detection algorithm based on YOLO-Bush. First, a special detection test bench was built, and a rubber bushing defect data set was constructed. Then, by constructing an enhanced feature extraction and enhanced feature fusion network, the accuracy and real-time performance of defect detection are effectively improved, and the detection accuracy and efficiency of rubber bushing surface defect detection can be significantly improved. The present invention provides a rubber bushing surface defect detection method based on YOLO-Bush.
[0005] The present invention is achieved by a YOLO-Bush rubber bushing surface defect detection method, the method comprising:
[0006] Step S1: A special test bench for rubber bushings was built to simulate the extrusion conditions in an industrial environment;
[0007] Step S2: collecting defect images during the manufacturing process of the rubber bushing to obtain 294 original images of the rubber bushing;
[0008] Step S3: Use the image annotation tool Labeling to annotate the images collected in S2, use the data enhancement method to create a rubber bushing image dataset, and divide the dataset into a training set, a validation set, and a test set.
[0009] Step S4: Build a YOLO-Bush network structure model, improve the YOLOxs structure model, and obtain an improved YOLO-Bush network structure model;
[0010] Step S5: Use the training set in the rubber bushing surface defect dataset in S3 to train the YOLO-Bush network model in S4 to obtain a rubber bushing surface defect detection model;
[0011] Step S6: Model testing and evaluation: Use the rubber bushing surface defect detection model obtained in S5 to train the model using the test set in the rubber bushing surface defect dataset in S3. Finally, evaluate the model using the average accuracy, precision, and recall curves.
[0012] A further improvement of the technical solution of the present invention is that in step S1, the constructed intelligent platform for rubber bushing defect detection can simulate the extrusion conditions in an industrial environment.
[0013] A further improvement of the technical solution of the present invention is that in step S3, the ratio of the training set, the test set and the validation set is set to 8:1:1.
[0014] A further improvement of the technical solution of the present invention is that in step S3, the surface defects of the rubber bushing are labeled by using the image labeling tool Labeling to label each image in the surface defect data collected by the rubber bushing defect detection intelligent platform.
[0015] A further improvement to the technical solution of the present invention is that in step S4, a multi-dimensional collaborative attention module (MCA) and a multi-order gated aggregation network module (MOGA) are created in the backbone to construct an enhanced feature extraction network. The enhanced feature extraction network simultaneously infers attention in the channel, height, and width dimensions, while reducing redundant information, enhancing feature information content and discriminability, and adapting to the feature representation of small objects.
[0016] A further improvement of the technical solution of the present invention is that in step S4, the feature fusion network creates an improved position enhanced attention mechanism (ELA) and sub-pixel downsampling convolution (SPDConv), constructs an enhanced feature fusion network, achieves lightweight and reduces the missed detection rate.
[0017] A further improvement to the technical solution of the present invention is that in step S5, the initial learning rate for training is 0.001; the rate decays with each iteration to 0.95; the optimizer is Adam with an exponential decay rate of 0.9. The network is trained for a total of 200 iterations, divided into two phases: 50 rounds of frozen training with a batch size of 8, followed by 100 rounds of unfrozen training with a batch size of 4.
[0018] A further improvement of the technical solution of the present invention is that: in step S6, the trained model is used to test the test set to obtain the detection accuracy, F1, FPS parameters and average detection accuracy value of the trained model.
[0019] A further improvement of the technical solution of the present invention is that in step S6, the trained model is compared horizontally with newer mainstream target detection models such as YOLOv8, YOLOv9s, RT-DRTR, and the original model YOLOxs, and evaluated using four indicators: mAP50, F1 value, recall rate, and FPS.
[0020] In combination with the above technical solutions and the technical problems solved, the advantages and positive effects of the technical solutions to be protected by the present invention are as follows:
[0021] The method of the present invention combines automation technology, artificial intelligence detection technology and defect detection technology, and solves the problems that the rubber bushing has a special structure, defects can only appear under pressure, and its end face structure is complex, the defect target is small, and the detection accuracy requirement is high.
[0022] The method of the present invention increases the number of sample data sets and enhances the diversity of sample images by using data set enhancement technology, thereby solving the problem of insufficient samples of surface defects of rubber bushings and improving the robustness and practicability of the model.
[0023] The proposed method constructs an enhanced feature extraction network by creating a multi-dimensional collaborative attention module (MCA) and a multi-order gated aggregation network module (MOGA) in the backbone. The enhanced feature extraction network simultaneously infers attention in the channel, height, and width dimensions, while reducing redundant information, enhancing feature information content and discriminability, and adapting to small object feature representation.
[0024] In the feature fusion network, we created an improved enhanced location attention mechanism (ELA) and sub-pixel downsampling convolution (SPDConv), creating an enhanced feature fusion network. The improved ELA attention mechanism has a simple structure, enabling lightweight and precise localization of small objects, improving detection capabilities and reducing missed detections. The SPDConv downsampling factor replaces cross-row convolution and pooling layers with new CNN building blocks, enhancing small object detection and reducing missed detection rates. The feature fusion network significantly improves detection stability and small defect recognition capabilities. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 is a flow chart of a method in an embodiment of the present invention;
[0026] Figure 2 This is a picture of a defect-free rubber bushing in an embodiment of the present invention;
[0027] Figure 3 is the main surface defect of the rubber bushing in the embodiment of the present invention;
[0028] Figure 4 The actual operation process of current enterprises in detecting rubber bushing defects;
[0029] Figure 5 An intelligent platform for rubber bushing defect detection built in an embodiment of the present invention;
[0030] Figure 6 This is a flow chart of automatic detection of rubber bushings in an embodiment of the present invention;
[0031] Figure 7 Schematic diagram of the structure of a rubber bushing detection device in an embodiment of the present invention;
[0032] Figure 8 This is a diagram showing surface defects of a rubber bushing in an embodiment of the present invention;
[0033] Figure 9 This is a dataset enhancement diagram in an embodiment of the present invention;
[0034] Figure 10 Schematic diagram of the YOLO-Bush model in an embodiment of the present invention;
[0035] Figure 11 Schematic diagram of the MCA module in the enhanced feature extraction network in an embodiment of the present invention;
[0036] Figure 12 Schematic diagram of the MOGA module in the enhanced feature extraction network in an embodiment of the present invention;
[0037] Figure 13 Schematic diagram of improved ELA in the enhanced feature fusion network in an embodiment of the present invention;
[0038] Figure 14 Schematic diagram of SPDConv in the enhanced feature fusion network in an embodiment of the present invention;
[0039] Figure 15 Detection scatter plots of five mainstream network models in the embodiment of the present invention;
[0040] Figure 16 PR curves of five comparison models in the embodiment of the present invention;
[0041] Figure 17 1 is a graph of F1 of five comparative models in an embodiment of the present invention;
[0042] Figure 18 Schematic diagram of the detection results in an embodiment of the present invention. DETAILED DESCRIPTION
[0043] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0044] The present invention provides a rubber bushing surface defect detection method based on YOLO-Bush. First, a special detection test bench for rubber bushings is built to obtain rubber bushing surface defect detection images; then, the defect types and defect locations in the collected images are annotated; the data is expanded using data enhancement technology and made into a data set; based on the YOLOxs network model, improvements are made, and in the feature extraction network, a multi-dimensional collaborative attention module (MCA) is used to enhance the feature representation capability. Subsequently, a multi-order gated aggregation network module (MOGA) is used to replace the SPP structure in the original model, which enhances the ability to obtain contextual information and constructs an enhanced feature extraction network. In the feature fusion stage, an improved enhanced position attention mechanism (ELA) is added to achieve accurate positioning of key areas. At the same time, sub-pixel downsampling convolution (SPDConv) is used to replace deep 3×3 convolution for downsampling, retaining fine-grained information and improving feature expression capabilities, thereby constructing an enhanced feature fusion network. Finally, the model was tested on a test set. The results showed that the YOLO-Bush model achieved an average accuracy of 99.4% on the rubber bushing surface defect test set, an F1 value of 0.99, and an FPS value of 27.264. This invention provides a detection method for solving the difficult problems encountered in rubber bushing surface defect detection.
[0045] like Figure 1 As shown, the embodiment of the present invention discloses a YOLO-Bush rubber bushing surface defect detection method, which mainly includes the following steps:
[0046] Step S1: A special testing bench for rubber bushings was built.
[0047] S1.1: If Figure 2 The rubber bushing shown is defect-free; Figure 3 The main surface defects of rubber bushings include surface bubbles and center plate crushing.
[0048] S1.2: If Figure 4 It demonstrates the actual operation process of current enterprises in detecting rubber bushing defects.
[0049] During the company's production process, rubber bushings undergo a reduction process after injection molding. This process aims to reduce their outer diameter so they can be smoothly installed in the mating holes, thereby ensuring assembly accuracy, structural tightness, and airtightness. After the reduction is completed, the rubber bushing is placed on the limit mold so that its axis is precisely aligned with the center line of the pressure head. The operator then presses the start switch under the workbench to drive the pressure head downward and apply squeezing pressure to the rubber bushing. Under the action of the squeezing force, the main defects of the rubber bushing become apparent. With the help of light illumination, the operator can identify and judge its appearance defects with the naked eye.
[0050] S1.3: Due to the inherent properties of rubber, surface defects are often difficult to directly identify in an unextruded state. To meet the stringent requirements for defect detection in industrial inspections, the intelligent rubber bushing inspection platform must be able to apply extrusion pressure to the bushing. Furthermore, since rubber is mostly black, ensuring image quality requires the proper placement of the light source and camera.
[0051] S1.4: If Figure 5 As shown, the present invention builds an intelligent platform for detecting small defects in rubber bushings.
[0052] During the inspection process, a pressure rod, working in conjunction with a circular dot matrix light source, applies pressure to a rubber bushing placed on a PLC-controlled conveyor belt. Simultaneously, the dot matrix light illuminates the compressed surface of the rubber bushing, revealing previously subtle defects. CCD industrial cameras located on either side of the light source simultaneously capture images of the compressed surface. The rubber bushing is then manually flipped over, and pressure and image capture are repeated to obtain surface information from the other side.
[0053] S1.5: If Figure 6 The figure shows the flow chart of automatic detection of rubber bushings.
[0054] The automated inspection method and system for rubber bushings proposed in this invention are primarily used for the non-contact, automatic identification and screening of defects in post-molding rubber bushings. The process begins at the end of the molding process. After undergoing diameter reduction, the bushing passes through a conveyor and enters the inspection system. Upon reaching a designated position, a photoelectric sensor is triggered, and the inspection system begins operation. The inspection system consists of two main components: an image acquisition module and a software algorithm module. The image acquisition module, consisting of a light source, an industrial camera, a lens, and an extrusion device, ensures high-precision image capture of the rubber bushing under pressure, standard lighting, and a stable environment. The captured image data is processed and analyzed by the software algorithm module, which comprises both offline modeling and online detection. In the offline modeling phase, defect recognition logic is constructed by updating and training the detection algorithm model. In the online detection phase, the real-time inspection results are first reviewed and determined by the system. If a defective bushing is detected, the image and inspection information are stored, and the actuator (push rod) is controlled to push the bushing to the rejection and recovery area. If the bushing is defect-free, the product proceeds to the next process. The system as a whole realizes the automated closed-loop control of rubber bushings from molding to inspection and then to sorting. It has the advantages of non-contact, high efficiency and high precision. It is suitable for large-scale continuous production lines, effectively reducing manual inspection errors and labor costs, and improving product consistency and reliability.
[0055] S1.6: If Figure 7 The figure shows a simplified structural diagram of the rubber bushing detection device.
[0056] S2: Image acquisition.
[0057] S2.1: Collect images of the rubber bushing during the manufacturing process to obtain original images of the rubber bushing surface;
[0058] Step S3: Image dataset creation.
[0059] S3.1: Use image annotation tools to annotate the rubber bushing defect image obtained in step S2 to obtain a training sample set. The defects are mainly divided into two types, namely surface bubbles and mid-piece crush damage. Figure 8 As shown, the image annotation tool Labeling is used to annotate the rubber bushing surface defect image obtained in step S2.
[0060] S3.2: If Figure 9 As shown in the figure, the initial dataset of S3.1 is expanded using dataset enhancement methods including average blur (kernel size varies from 5×5 to 15×15), random cropping (cropping random parts of the image in the range of 50-80% of the original image size) and random rotation (rotating at random angles from -60° to 60°). Noise interference is also introduced to enrich the posture and scale of the rubber bushing surface defects, which is conducive to the subsequent training of a more stable model.
[0061] Step S4: Figure 10 As shown, build the YOLO-Bush network structure model.
[0062] S4.1: If Figure 11 As shown, a multi-dimensional collaborative attention module (MCA) is created in the feature extraction network. The MCA module implements the input tensor F∈R C×H×W The top branch first rotates F 90 degrees along the H axis and then inputs the squeeze transformation adaptive mechanism. This mechanism contains parallel global averaging and standard pooling layers to generate and Two feature descriptors, and Represent the mean feature and standard deviation feature of the mth channel, m∈{1,2…C}. The two are processed by the adaptive combination mechanism to obtain the aggregated feature map Model the relationship between channel dimension C and spatial dimension H. This process can be expressed as formula Where α and β are trainable floating parameters, which are greater than 0 and less than 0 respectively, and can be optimized by SGD. The incoming excitation transformation mechanism captures the interaction of features in space W and obtains the width feature weight Expression of interaction K using a gating mechanism C Mapping with channel dimension see formula Where λ and γ are two hyperparameters, K C is the kernel size. Next, an enhanced feature map is passed to the sigmoid activation function. Generate input-specific attention weights A in spatial dimension W W ∈R W×1×1 Then A is multiplied by element W Apply to Get the enhanced feature map F' W ∈R W×H×C Finally, rotate 90 degrees clockwise along the H axis to obtain the feature map F with the same shape as the original input. W ∈R C×H×W The middle branch and the lower-level branch have similar operations. In the integration stage, the outputs of the three branches are averaged and aggregated, and the weights of each dimension are calibrated to obtain the final refined feature map.
[0063] S4.2: If Figure 12 As shown, a multi-order gated aggregation network module (MOGA) is created in the enhanced feature extraction network. The MOGA module is based on X∈R C×HW As input, first pass DW 5×5,d=1 Deep convolution performs low-level feature interaction and then decomposes the output features into low-level features along the channel dimension. Intermediate and high-level Three sub-features, among which C l +C m +C h =C. Among them, X m and X h DW 5×5,d=2 and DW 7×7,d=3 For parallel processing, X l The processed features are concatenated to obtain Y C Finally, adaptive feature aggregation is achieved through the SiLU gating branch. This module not only retains the gating characteristics of Sigmoid but also ensures training stability. Finally, the gating weights are multiplied by the context features to make the network focus on key information.
[0064] S4.3: This paper proposes a multi-dimensional collaborative attention module (MCA) and a multi-order gated aggregation network module (MOGA), and combines the feature extraction network to construct an enhanced feature extraction network, which effectively improves the ability to pay attention to small targets of rubber bushings.
[0065] S4.4: If Figure 13As shown in Figure 1, in order to enhance the model's ability to perceive small objects, an improved enhanced position attention mechanism (ELA) is proposed. This module extracts position information in the horizontal and vertical directions respectively to achieve accurate modeling of spatial position features. First, for the input feature map x c Average pooling is performed on each channel in two spatial ranges along the horizontal direction (H, 1) and along the vertical direction (1, W), and spatial description information in the horizontal and vertical directions is obtained. and Then, a one-dimensional convolution with a kernel size of 7 is used for convolution operation. A larger kernel can provide a wider coverage of position information, effectively enhancing the interactive ability of the positioning information embedding, so that the improved model can accurately locate small targets. Then, the group normalization (GroupNorm) operation is used to process the enhanced position information. Finally, the Sigmoid activation function is used to obtain the horizontal and vertical representation y of the position attention. h and y w Finally, the original feature map x is multiplied by element c With these two attention maps y h and y w Fusion, get the output feature map Y
[0066]
[0067] y h =σ(G n (F h (z h ))) (3)
[0068] y w =σ(G n (F w (z w ))) (4)
[0069] Y=x c ×y h ×y w (5)
[0070] Where C, H, and W represent the number of channels, height, and width, respectively; x c (h,i) represents the eigenvalue of the cth channel at position (h,i); x c (j,w) represents the eigenvalue of the cth channel at position (j,w); F h (z h ) and F w (z w ) indicates a one-dimensional convolution operation with a kernel size of 7; G n Represents the GroupNorm operation; σ represents the Sigmoid activation function.
[0071] S4.5: If Figure 14 As shown in Figure 1, a sub-pixel downsampling convolution (SPDConv) is proposed in the network. SPDConv performs structured spatial rearrangement and channel enhancement on the original feature map to achieve effective downsampling while maintaining the discriminative features, thereby enhancing the model's ability to recognize small targets. Assume that the input feature is a square, denoted as X(S, S, C1), where S represents the height and width of the feature map, and C1 represents the number of input channels. SPDConv first performs spatial rearrangement and channel enhancement on the input feature map to achieve effective downsampling while maintaining the discriminative features, thereby enhancing the model's ability to recognize small targets. Figure X Perform structured segmentation and divide the feature map into scales based on the downsampling scale factor scale 2 non-overlapping sub-feature maps. Taking scale=2 as an example, split X into f 0,0 =X[0:S:scale,0:S:scale],f 1,0 =X[1:S:scale,0:S:scale],f 0,1 =X[0:S:scale,1:S:scale] and f 1,1 =X[1:S:scale,1:S:scale] and other sub-feature maps. Then, these sub-feature maps are connected along the channel dimension to obtain a new feature map The spatial dimension of X' is reduced by a scale factor, and the channel dimension is increased by a scale factor scale 2 Finally, a convolution without stride (stride=1) is used to transform Convert to Where C2 is the number of output channel rows. No cross-row convolution can retain all the discriminative feature information of small targets as much as possible and reduce the missed detection rate of small targets.
[0072] S4.6: We further constructed an enhanced feature fusion network by introducing an improved enhanced location attention mechanism (ELA) and sub-pixel downsampled convolution (SPDConv) into the feature fusion network of YOLO-Bush. The improved ELA attention mechanism has a simple structure, enabling lightweight and accurate localization of small objects, improving detection capabilities and reducing missed detections. The SPDConv downsampling factor replaces cross-row convolution and pooling layers with new CNN building blocks, enhancing small object detection and reducing missed detections. The feature fusion network significantly improves detection stability and small defect recognition capabilities.
[0073] Step S5: Train the YOLO-Bush network structure model to obtain a rubber bushing surface defect detection model.
[0074] S5.1: The dataset was divided into training, test, and validation sets in an 8:1:1 ratio. The experimental hardware configuration used the Ubuntu 20.0 operating system, the PyTorch 1.7.1 framework, OpenCV 2 and PyCharm code integrated development environments, and CUDA 11.0 GPU acceleration. The workstation hardware was an Intel 3.10GHz 64-core CPU, 256GB of memory, and two NVIDIA Titan XP GPUs.
[0075] S5.2: The initial learning rate is 0.001; the rate decays with each iteration to 0.95. The optimizer is Adam with an exponential decay rate of 0.9. The network is trained for a total of 200 iterations, divided into two stages: 50 epochs of frozen training with a batch size of 8, followed by 100 epochs of unfrozen training with a batch size of 4.
[0076] Step S6: Test and evaluate the YOLO-Bush model.
[0077] S6.1: Select the most commonly used indicators in the field of object detection: accuracy (Precision), recall (Recall), average precision (AP), mean average precision (mAP), number of parameters (Params), and processing speed for evaluation. Use the following formulas to calculate these evaluation indicators:
[0078]
[0079] TP (True Positive) is the number of true positives; FP (False Positive) is the number of false positives; FN (False Negative) is the number of false negatives; TN (True Negative) is the number of true negatives, P is the precision, and R is the recall.
[0080] S6.2: The present invention verifies the effects of introducing multi-dimensional collaborative attention module (MCA), multi-order gated aggregation network module (MOGA), improved position enhanced attention mechanism (ELA) and sub-pixel downsampling convolution (SPDConv) on the accuracy and parameter quantity of rubber bushing surface defect detection through ablation experiments.
[0081] Table 1
[0082]
[0083] By analyzing the contribution of each improvement strategy in Table 1 to the network of the present invention, it is found that each module has different degrees of improvement on the overall improvement of the model.
[0084] When MCA, MOGA, improved ELA, and SPDConv are fused individually, mAP50 improves to 97.2%, 97.4%, 97%, and 97.3%, respectively. F1 scores improve when MCA, MOGA, and improved ELA are fused, but remain at 0.89 when fused with SPDConv. This may be due to SPDConv's use of a non-cross-row convolution architecture, which enables the model to learn features more effectively, while cross-row convolutions help skip redundant pixel information. This unique structure may contribute to the lack of F1 improvement. However, SPDConv's strong stability and adaptability ensure that its F1 score remains at 0.89, showing no decline. Combining multiple improved methods further improves mAP50, F1, and FPS, while also reducing the number of parameters. Combining two improved methods achieves a maximum mAP50 of 98.3%, a maximum F1 of 0.98, and a maximum FPS of 26.401, while minimizing the number of parameters to 8.316. After integrating the three improvements, mAP50, F1, and FPS all improved, reaching a maximum of 98.9%, 0.98, and 27.263, respectively, while the number of parameters was reduced to a minimum of 7.984. Finally, after incorporating all improvements into the original model, mAP50 rose to 99.4%, F1 to 0.99, FPS to 27.264, and the number of parameters was reduced to 7.827.
[0085] In this way, the combination of the four can not only infer attention in the three dimensions of channel, width and height, reduce redundant information, but also achieve lightweight and accurate positioning of the area of interest, while retaining all the discriminant feature information of small targets as much as possible, reducing the missed detection rate of small targets, and reducing the number of parameters while improving detection accuracy.
[0086] S6.3: In order to determine the advantages of the improved YOLOxs model over the current mainstream detection network, this paper selected the current mainstream algorithms for comparative experiments, namely RT-DETR, YOLOxs, YOLOv8, YOLOv9s and the original YOLOxs network model. The results are shown in Table 2, and the scatter plot is shown in Figure 15 shown.
[0087] Table 2
[0088]
[0089] Table 2 shows that YOLO-Bush performs worse than RT-DETR in FPS, but achieves the best results in all other metrics. YOLO-Bush improves mAP50 by 0.1%, 2.1%, 0.7%, and 2.7% compared to RT-DETR, YOLOv8, YOLOV9s, and YOLOxs, respectively. Therefore, compared to other mainstream detection network models, YOLO-Bush maintains high detection speed and accuracy, meeting the real-time online detection requirements of rubber bushings.
[0090] The PR curve is a key indicator for evaluating model performance. AP is the abbreviation of the area of the PR curve. The larger the AP value, the better the performance of the model. That is, when the precision and recall rate are larger, and the PR curve is closer to the upper right corner, the better the model performance. Figure 16 Among the five PR curves, the improved model YOLO-Bush proposed in this paper is the best among all models. The F1 curve takes into account the precision and recall rate. The larger the F1 value, the better the detection performance of the model, as shown below. Figure 17 shown.
[0091] In this way, first, a rubber bushing detection device is built, and rubber bushing defect images are obtained under extrusion conditions in a simulated industrial environment; then the images are annotated and expanded using data enhancement technology to obtain a rubber bushing image dataset; secondly, the obtained dataset is used to pre-train the YOLO-Bush model to obtain a YOLO-Bush model for detecting rubber bushing surface defects; finally, the model is tested using the test set data to obtain a rubber bushing detection image, as shown in the figure. Figure 18 The trained model is then evaluated using model evaluation metrics. The left side shows the true label, and the right side shows the model's predicted probability for each defect type. When the defect type predicted by the model is consistent with the true label and the corresponding prediction probability is high, it indicates that the model has good recognition capabilities and can accurately determine the defect type. Comprehensive analysis of the training results using model evaluation metrics can effectively verify the model's performance and reliability, thereby determining its ability to replace manual inspection.
[0092] It should be noted that the embodiments of the present invention can be implemented by hardware, software, or a combination of software and hardware. The hardware portion can be implemented using dedicated logic; the software portion can be stored in a memory and executed by an appropriate instruction execution system, such as a microprocessor or dedicated design hardware. Those skilled in the art will appreciate that the above-mentioned devices and methods can be implemented using computer-executable instructions and / or contained in processor control code, for example, such as a carrier medium such as a disk, CD or DVD-ROM, a programmable memory such as a read-only memory (firmware), or a data carrier such as an optical or electronic signal carrier. The devices and modules of the present invention can be implemented by hardware circuits such as very large-scale integrated circuits or gate arrays, semiconductors such as logic chips, transistors, or programmable hardware devices such as field programmable gate arrays, programmable logic devices, etc., can also be implemented by software executed by various types of processors, or can be implemented by a combination of the above-mentioned hardware circuits and software, such as firmware.
[0093] The above description is only a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications, equivalent substitutions and improvements made by any technician familiar with this technical field within the technical scope disclosed by the present invention and within the spirit and principles of the present invention should be covered by the scope of protection of the present invention.
Claims
1. A YOLO-Bush rubber bushing surface defect detection method, characterized in that: The method includes: S1: A special test bench for rubber bushings was built to simulate the extrusion conditions in an industrial environment; S2: collecting images of the rubber bushing during the manufacturing process to obtain the original images of the rubber bushing; S3: Use the image annotation tool Labeling to annotate the images collected in S2, use the data enhancement method to create a rubber bushing image dataset, and divide the dataset into training set, validation set, and test set; S4: Build the YOLO-Bush network structure model, improve the YOLOxs structure model, and obtain the improved YOLO-Bush network structure model; S5: Use the training set in S3 to train the YOLO-Bush network model in S4 to obtain a rubber bushing surface defect detection model; S6: Model testing and evaluation: The rubber bushing surface defect detection model obtained in S5 is trained with the test set in the rubber bushing surface defect dataset in S3. Finally, the model is evaluated through average accuracy, precision, and recall curves.
2. The YOLO-Bush rubber bushing defect detection method according to claim 1, characterized in that: The special inspection test bench for rubber bushings described in step S1 is an automated inspection system for rubber bushings. During the inspection process, a pressure rod, in cooperation with a ring-shaped dot matrix light source, pressurizes the rubber bushing placed on a conveyor belt controlled by a PLC. At the same time, the dot matrix light source illuminates the surface of the rubber bushing after compression, thereby revealing defects that were originally difficult to detect. CCD industrial cameras located on both sides of the light source synchronously capture images of the compressed surface. Subsequently, the rubber bushing is manually flipped over and pressurized and imaged again to obtain surface information on the other side.
3. The YOLO-Bush rubber bushing surface defect detection method according to claim 1, characterized in that: In step S3, the ratio of the training set, test set and validation set is set to 8:1:
1.
4. The YOLO-Bush rubber bushing surface defect detection method according to claim 1, characterized in that: In step S3, the labeling of the surface defects of the rubber bushing is performed by using the image labeling tool Labeling to label each image in the surface defect data collected by the rubber bushing defect detection intelligent platform.
5. The YOLO-Bush rubber bushing surface defect detection method according to claim 1, characterized in that: The YOLO-Bush model described in step S4 is based on the YOLOxs model and builds an enhanced feature extraction network and an enhanced feature fusion network.
6. The YOLO-Bush rubber bushing surface defect detection method according to claim 1, characterized in that: In the backbone, we created a multi-dimensional collaborative attention module (MCA) and a multi-order gated aggregation network module (MOGA) to build an enhanced feature extraction network. Strengthen the feature extraction network to simultaneously infer attention in the channel, height, and width dimensions, while reducing redundant information, enhancing feature information content and discriminability, and adapting to small target feature representation; Create a multi-dimensional collaborative attention module (MCA) in the feature extraction network. The MCA module implements the input tensor F∈R C×H×W To the refined output tensor; the top branch first rotates F 90 degrees along the H axis and then inputs the squeeze transformation adaptive mechanism; the mechanism contains parallel global averaging and standard pooling layers to generate and Two feature descriptors; the two are processed by an adaptive combination mechanism to obtain an aggregated feature map The relationship between modeling channels C and spatial dimension H; This process can be expressed as the formula Where α and β are trainable floating parameters, which are greater than 0 and less than 0 respectively, and can be optimized by SGD; then The incoming excitation transformation mechanism captures the interaction of features in space W and obtains the width feature weight Expression of interaction K using a gating mechanism C Mapping with channel dimension see formula Where λ and γ are two hyperparameters, K C is the kernel size; next, an enhanced feature map is passed to the sigmoid activation function Generate input-specific attention weights A in spatial dimension W W ∈R W×1×1 ; Then A is multiplied by element W Apply to Get the enhanced feature map F' W ∈R W×H×C ;Finally, rotate 90 degrees clockwise along the H axis to obtain the feature map F with the same shape as the original input W ∈R C×H×W ; The middle branch and the lower-level branch have similar operations; in the integration stage, the outputs of the three branches are averaged and aggregated, and the weights of each dimension are calibrated to obtain the final refined feature map; Create a multi-order gated aggregation network module (MOGA) in the feature extraction network, and set the input to X∈R C×HW ;DW 5×5,d=1 First, it is used to process low-order feature interactions; then, the obtained output features are decomposed along the channel dimension into and DW 5×5,d=2 With DW 7×7,d=3 Parallel Processing X m With X h , X l After identity mapping, the processed features are concatenated to obtain Y C ; The gating branch uses SiLU to adaptively aggregate features, which has both Sigmoid gating effect and stable training; finally, the gating is multiplied with the context branch features to make the network focus on useful information.
7. The YOLO-Bush rubber bushing surface defect detection method according to claim 1, characterized in that: The YOLO-Bush feature fusion network uses an improved enhanced location attention mechanism (ELA) and sub-pixel downsampling convolution (SPDConv) to build an enhanced feature fusion network. The improved ELA attention mechanism has a simple structure, enabling lightweight and accurate positioning of small targets, improving detection capabilities and reducing missed detections. The SPDConv downsampling factor replaces cross-row convolution and pooling layers with new CNN building blocks, enhancing small target detection and reducing missed detection rates. The feature fusion network significantly improves detection stability and small defect recognition capabilities. The improved ELA attention mechanism structure combines one-dimensional convolution with group normalization, first performs horizontal and vertical average pooling on the channel, and obtains and Two outputs; then use 7-core one-dimensional convolution to enhance the positioning information, after group normalization, after Sigmoid, use element-by-element multiplication to convert the original feature map x c The output Y is fused with the two enhanced attention maps. The process is shown in the formula: y h =σ(G n (F h (z h ))) y w =σ(G n (F w (z w ))) Y=x c ×y h ×y w Create SPDConv downsampling factors in the network. SPDConv consists of spatial-to-depth layers and non-cross-row convolution layers. It has good performance in small target detection, strong stability, adaptability to various inputs, stable training, and good application effects. SPDConv downsamples the internal and overall feature maps of CNN. Input X(S,S,C1) is split into sub-feature maps by SPD and connected along the channel dimension to obtain The spatial dimension is reduced and the channel dimension is increased; finally, after the non-cross-row convolution, we get Retain the discriminative features of small targets and reduce the missed detection rate.
8. The YOLO-Bush rubber bushing surface defect detection method according to claim 1, characterized in that: In step S5, the training results are as follows: the initial learning rate of training is 0.001; each iteration decays gradually, and the decay rate is 0.95; the optimizer is Adam, and the exponential decay rate is 0.9; the network is trained for a total of 200 iterations, which is divided into two stages, first 50 rounds of frozen training with a batch size of 8, and then 100 rounds of unfrozen training with a batch size of 4.
9. The YOLO-Bush rubber bushing surface defect detection method according to claim 1, characterized in that: In step S6, the trained model is used to test the test set to obtain the detection accuracy, F1, FPS parameters and average detection accuracy value of the trained model.
10. The YOLO-Bush rubber bushing surface defect detection method according to claim 1, characterized in that: In step S6, the trained model is compared with the newer mainstream object detection models YOLOv8, YOLOv9s, RT-DRTR and the original model YOLOxs, and evaluated using four indicators: mAP50, F1 value, recall rate and FPS.