Oil seal defect intelligent detection method based on improved YOLOv12
By improving the YOLOv12 model and adding a small target detection layer, a dynamic upsampling module, and an MSAF module, the problems of poor consistency and insufficient robustness in oil seal defect detection are solved, and high-precision and real-time oil seal defect detection is achieved.
Patent Information
- Application Number
- CN202511180676.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-22
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2045-08-22
AI Technical Summary
The existing oil seal defect detection method relies on manual inspection, which has poor consistency, low efficiency and high cost. In addition, the existing YOLO model lacks detection accuracy and robustness in complex environments, making it difficult to meet industrial production needs.
The YOLOv12 model is improved by adding a small object detection layer, a dynamic upsampling module, an MPDIoU loss function, and an MSAF module to build the MPD-YOLO model, which improves the detection accuracy and robustness of small object defects.
It significantly improves the detection accuracy and real-time performance of oil seal surface defects, enhances the model's ability to identify small-size defects, solves the detection challenges in complex environments, and meets industrial production needs.
Smart Images

Figure CN120707566A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of oil seal defect detection, and in particular to an intelligent oil seal defect detection method based on improved YOLOv12. Background Art
[0002] In industrial production, oil seals, as a key sealing component, are widely used in various types of mechanical equipment. Their function is to prevent the leakage of lubricating oil or grease, while also preventing external contaminants from entering the equipment. They play a vital role in ensuring the normal operation and service life of mechanical equipment. Surface defects such as scratches, burrs, and dents on oil seals can cause lubricating oil leakage, which in turn affects the normal operation of the equipment. In severe cases, this can lead to equipment failure, resulting in production halts and significant economic losses.
[0003] Traditional methods for detecting surface defects in oil seals rely primarily on manual visual inspection, a method with numerous drawbacks. Firstly, manual inspection is significantly influenced by subjective factors, and the experience and judgment standards of different inspectors vary, resulting in inconsistent and inaccurate test results. Secondly, prolonged and repetitive manual inspections can easily lead to visual fatigue, resulting in missed and incorrect detections. Furthermore, manual inspection is inefficient, unable to meet the inspection needs of large-scale industrial production, and is also costly.
[0004] With the development of science and technology, oil seal surface defect detection methods based on image processing have gradually been applied. For example, by dividing the oil seal grayscale image into blocks and analyzing the pixel features in the blocks, the contrast between dirty areas and other areas is enhanced to achieve accurate segmentation of dirty areas. Oil seal defect image edge detection methods based on threshold segmentation and chain code method use the least squares method to fit the circular contour of the oil seal lip for defect detection. There are also oil seal defect detection methods based on entropy clustering segmentation and annular difference. However, because these detection methods require manual acquisition of oil seal defect features, this method of manually extracting defect features has certain limitations, is cumbersome, and requires high technical personnel.
[0005] In recent years, deep learning technology has flourished, and object detection algorithms based on convolutional neural networks have demonstrated strong performance in many fields. The YOLO family of algorithms, in particular, has been widely adopted for their ability to quickly and efficiently identify objects in images and excel in real-time object detection tasks. However, in the context of oil seal surface defect detection, existing YOLO models still face numerous challenges: 1) There are many types of defects in the oil seal production process, and it is difficult to obtain a large-scale, high-quality training dataset covering all defect types, which restricts the deep model's defect detection capability and accuracy.
[0006] 2) The actual industrial inspection environment is complex and changeable. Factors such as unstable lighting conditions and equipment vibration will interfere with the inspection results. As a result, the existing YOLO model is less efficient when processing high-definition image streams and maintaining high accuracy under changing lighting conditions. This makes it difficult to meet the industrial production needs for high-precision and high-reliability oil seal surface defect detection.
[0007] 3) The feature extraction capability of small-sized targets is limited and is easily interfered by complex background noise, resulting in reduced detection accuracy.
[0008] Therefore, developing a method that can effectively overcome the above problems and improve the efficiency and accuracy of oil seal surface defect detection has important practical significance and application value. Summary of the Invention
[0009] In response to the above-mentioned technical problems to be solved, the present invention provides an intelligent oil seal defect detection method based on an improved YOLOv12, which improves the detection accuracy and real-time detection capability of small target defects and the robustness of the model.
[0010] In order to solve the above technical problems, the technical solution proposed by the present invention is: An intelligent oil seal defect detection method based on improved YOLOv12 includes the following steps: Step S1, collecting image data of the oil seal and annotating it to obtain an oil seal defect dataset; Step S2, improving the network framework of the YOLOv12 model and constructing an oil seal surface defect detection model based on YOLOv12, namely the MPD-YOLO model; Step S3, training the MPD-YOLO model to obtain a trained MPD-YOLO model; Step S4: evaluating the trained MPD-YOLO model based on the oil seal defect test set, and adjusting the MPD-YOLO model based on the evaluation results; In step S2, improving the YOLOv12 model network framework includes the following operations: S2-1, adjust the head hierarchical structure of the YOLOv12 model, integrate shallow semantic information, and add a small object detection layer; S2-2, use the dynamic upsampling module to replace the nearest neighbor interpolation upsampling module of the YOLOv12 model; S2-3, use the MPDIoU loss function to replace the CIoU loss function of the YOLOv12 model; S2-4, use the MSAF module to replace the C3k2 module in the YOLOv12 model backbone network.
[0011] As a further improvement of the above technical solution: Preferably, the adjustment of the YOLOv12 model detection head hierarchical structure is to add a group of small target detection layers P2 on the basis of the three detection layers, and then add a P2 upsampling fusion module after two rounds of upsampling fusion modules of the first detection layer P3 and the second detection layer P4, and connect it to the small target detection layer P2.
[0012] Preferably, in the dynamic upsampling module, the input feature map X is first interpolated into a continuous feature map by bilinear interpolation; then, an offset O is generated by a linear layer, and the offset determines the position of the sampling point in the continuous space, and the sampling point is used to reconstruct the upsampled feature map X'.
[0013] Preferably, in the dynamic upsampling module, a sampling point generator based on a static range factor or a dynamic range factor is used to generate sampling points.
[0014] Preferably, the specific structure of the MSAF module is as follows: starting from the input features, a dual-branch parallel processing architecture is constructed, the upper branch of the dual branches uses dilation coefficients of 1, 2, 3, and 4 to obtain multi-scale features and splice them, and then adjust the dimension through channel shuffling and 1×1 convolution; the lower branch of the dual branches first implements multi-scale pooling through pyramid pooling, and after the pooling results are flattened and spliced, they are operated with the query vector extracted from the input features. The generated weights act on the value vector branch and are output through 1×1 convolution; finally, the results of the upper and lower branches are added element by element to output enhanced features that integrate multi-scale and attention mechanisms.
[0015] Preferably, the method of evaluating the trained MPD-YOLO model in step S4 is: The oil seal defect test set is input into the trained MPD-YOLO model, and the evaluation results are obtained based on the set key indicators: mean average precision, precision, and recall rate; The accuracy formula is: ; The recall formula is: ; In the formula, TP represents the number of samples correctly classified as positive; FP represents the number of samples incorrectly classified as positive; FN represents the number of samples incorrectly classified as negative. The calculation formula of the average precision mean is: ; Where, n represents the number of categories, i Represents the category index, Represents the functional form of the precision-recall curve, P Represents precision,R Represents the recall rate.
[0016] The intelligent oil seal defect detection method based on improved YOLOv12 provided by the present invention has the following advantages over the existing technology: (1) The intelligent oil seal defect detection method based on improved YOLOv12 of the present invention solves the problems of low small target detection rate and poor adaptability of existing methods. By adding a small target detection layer, using the MSAF module, combining dynamic upsampling and a new loss function, the detection accuracy, real-time performance and robustness of surface defects such as burrs, notches and scratches on oil seals are improved, effectively solving the problem of identifying small-sized defects in industrial inspection.
[0017] (2) The intelligent oil seal defect detection method based on improved YOLOv12 of the present invention introduces a small target detection layer and combines it with the P2 upsampling fusion block to effectively fuse shallow features, significantly enhancing the model's ability to detect small target defects and improving the accuracy and reliability of oil seal surface defect recognition.
[0018] (3) The intelligent oil seal defect detection method based on improved YOLOv12 of the present invention introduces a dynamic upsampling module, replaces the nearest neighbor interpolation upsampling method in the feature fusion step, and effectively avoids time-consuming dynamic convolution operations and additional sub-network construction through a content-based dynamic point sampling strategy, thereby achieving fine-grained offset control of small target areas, ensuring the accurate recovery of small-size defects and edge features during the upsampling process, and significantly improving the model's detection capability for small-size defects and multi-shape defects.
[0019] (4) The oil seal defect intelligent detection method based on the improved YOLOv12 of the present invention introduces MPDIoU to replace CIoU. MPDIoU calculates different loss function values for prediction boxes of different widths and heights, which can more effectively guide the prediction box to approach the real box, solving the problem of CIoU failure. MPDIoU not only simplifies the calculation process, but also stabilizes the convergence of the model, thereby improving the detection effect of small-size defects.
[0020] (5) The intelligent oil seal defect detection method based on the improved YOLOv12 of the present invention introduces the MSAF module to replace the C3k2 module of the YOLOv12 backbone network. The MSAF module can more comprehensively understand the input image by processing local features and global features at the same time. In the target detection task, this helps the model identify targets in the image and correctly locate them. Local features help capture the detailed information of the target such as edges, textures, etc., while global features provide the context of the target such as background, relationship between targets, etc. Through the dilated convolution and pooling operations of different sizes in the MSAF module, the module can extract features at multiple scales, thereby better processing multi-scale targets and avoiding the receptive field problem in traditional convolutional networks. Through the channel shuffling operation, the module can better fuse the features generated by each convolution operation, which not only avoids the information bottleneck, but also improves the network's ability to express details and complex patterns in the image. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 This is a flow chart of the oil seal defect intelligent detection method based on the improved YOLOv12 of the present invention.
[0022] Figure 2 Schematic diagram of the structure of the MPD-YOLO model of the present invention.
[0023] Figure 3 Schematic diagram of the structure of the dynamic upsampling module of the present invention.
[0024] Figure 4 This is a structural diagram of a sampling point generator based on a static range factor in the dynamic upsampling module of the present invention.
[0025] Figure 5 This is a structural diagram of a sampling point generator based on a dynamic range factor in the dynamic upsampling module of the present invention.
[0026] Figure 6 Schematic diagram of the structure of the MSAF module of the present invention. DETAILED DESCRIPTION
[0027] The following is a detailed description of the specific embodiments of the present invention. It should be understood that the specific embodiments described herein are only used to illustrate and explain the present invention and are not intended to limit the present invention.
[0028] like Figure 1As shown, the present invention is based on an intelligent oil seal defect detection method based on an improved YOLOv12. Based on the YOLOv12 model, a small target detection layer is introduced, the nearest neighbor interpolation upsampling module is replaced by a dynamic upsampling module, and the CIoU (Complete Intersection over Union) loss function is replaced by the MPDIoU (Minimum Perimeter Distance Intersection over Union) loss function. The C3k2 module in the YOLOv12 model backbone network is replaced by the MSAF module to form the MPD-YOLO oil seal defect detection model, which improves the detection accuracy and real-time detection capability of small target defects and the robustness of the model.
[0029] The present invention is based on an intelligent oil seal defect detection method based on an improved YOLOv12, and specifically includes the following steps: Step S1: Collect image data of oil seals, and perform labeling on the acquired oil seal image set to obtain training samples.
[0030] The specific steps include: S1-1 uses a high-definition camera to collect image data of industrial oil seals during the production process, and collects oil seal images from an open source dataset website to obtain the original dataset through these two methods.
[0031] S1-2, the original data set was manually screened to remove overly blurred and overexposed images, obtaining an oil seal image set with multiple scales and defect types; then data enhancement methods such as rotation, scaling, blurring, and brightness change were used to obtain sample data.
[0032] Data augmentation is performed on the oil seal images, expanding each image into multiple images. Specifically, data augmentation can be performed using the torchvision.transforms module (an image transformation module in the PyTorch visual toolkit) in the PyTorch platform. Images can be randomly rotated using the transforms.RandomRotation function; images can be randomly cropped and resized using the transforms.RandomResizedCrop function; images can be randomly transformed in brightness and contrast using the transforms.ColorJitter function; and images can be randomly Gaussian blurred using the transforms.GaussianBlur function. This increases the dataset size, improves model generalization, enhances the robustness of small object detection, and reduces the risk of overfitting.
[0033] S1-3, annotate the images and record the defect type and location for the defect area in each image; and manually review the annotation results to ensure the accuracy of the annotation; during the data set processing, use the LabelImg tool to perform detailed annotations on the collected oil seal images, marking the defect type and location; defects are divided into three categories: burr, dent, and scratch, with burrs being called "burr", dents being called "dent", and scratches being called "scratch".
[0034] S1-4, after the labeling is completed, the entire dataset is divided into training set, validation set, and test set in an 8:1:1 ratio to ensure the effectiveness and generalization ability of the model.
[0035] Step S2: Improve the network framework of the YOLOv12 model and build an oil seal surface defect detection model based on YOLOv12, namely the MPD-YOLO model.
[0036] In this embodiment, improving the YOLOv12 model network framework specifically includes the following operations: S2-1, adjust the hierarchical structure of the YOLOv12 model head, integrate shallow semantic information, and add a small object detection layer; First, the YOLOv12 model's head structure was adjusted to incorporate shallow semantic information. A small object detection layer, P2, was added to the three detection layers. After two rounds of upsampling fusion modules for the first and second detection layers, P3 and P4, a P2 upsampling fusion module was added and connected to the small object detection layer, P2. Combined with the P2 upsampling fusion block, this effectively integrates shallow features, significantly enhancing the model's ability to detect small object defects and improving the accuracy and reliability of oil seal surface defect recognition.
[0037] S2-2 replaces the nearest neighbor interpolation upsampling module in the YOLOv12 model with a dynamic upsampling module. This module replaces the nearest neighbor interpolation upsampling method in the feature fusion step. By adopting a content-based dynamic point sampling strategy, it effectively avoids time-consuming dynamic convolution operations and the construction of additional subnetworks, enabling fine-grained offset control of small target areas. This ensures accurate recovery of small defects and edge features during the upsampling process, significantly improving the model's detection capabilities for small and multi-shaped defects.
[0038] In this embodiment, in the dynamic upsampling module, the input feature map X is first interpolated into a continuous feature map through bilinear interpolation. Then, a linear layer is used to generate offsets O, which determine the positions of the sampling points from the continuous space. These sampling points are finally used to reconstruct the upsampled feature map X'. Specifically, Figure 3 As shown, given a The feature map X and a The point sampling set S is C, where C is the number of channels, H and B are the height and width of the feature map respectively, 2 represents the coordinates in the x and y directions, and g represents the number of groups. The feature map is divided into g groups along the channel range. The grid_sample function uses the positions in the point sampling set S to resample X and generate a size of The feature map X' of this upsampling process is shown in formula (1): (1) in, is a grid sampling function, X is the input feature, X' is the upsampled feature, and S is the sampling set. The sampling point generator generates a point sampling set S, given an upsampling scale factor s and a feature map X with a shape of C×H×B, using input and output channels of C and A linear layer is used to generate ×H×B is offset O, and then reshaped into .
[0039] Furthermore, in the dynamic upsampling module, a sampling point generator based on a static range factor or a dynamic range factor is used to generate sampling points.
[0040] (1) Static range factor: such as Figure 4 As shown in the figure, the input feature map X first generates low-dimensional features through a linear transformation, then multiplies it by a fixed factor of 0.25, and then performs a pixel rearrangement operation, which is expressed as: (2) Among them, O represents the offset, Represents a linear layer with C input channels and 2s² output channels.
[0041] The static range factor makes upsampling stable and controllable, and is suitable for processing scenes with consistent scale changes.
[0042] (2) Dynamic range factor: The introduction of dynamic range factor makes the sampling process more flexible. First, a range factor is generated, and then it is used to adjust the offset O. Here σ represents the Sigmoid function, which is used to generate the range factor. By dynamically adjusting the offset, dynamic upsampling can automatically adjust the sampling points according to the different input feature maps, thereby avoiding the limitations of traditional interpolation methods, which can be expressed as: (3) in, , represents two different linear layers, Used to generate values related to the dynamic range factor, Generates offset-related values.
[0043] Dynamic factors adaptively adjust the scale of different inputs, making them particularly suitable for processing multi-scale features or images with large resolution variations. This mechanism can better capture detailed information at different scales, improving the flexibility and adaptability of the model.
[0044] The above formula (3) represents the scheme adopted by the method proposed in the open book, such as Figure 5 As shown in the figure, the Sigmoid function and a static factor of 0.5 are used to limit the dynamic range to [0, 0.5], and a point-by-point dynamic range factor is generated with 0.25 as the center.
[0045] In this embodiment, further, the sampling set S is the sum of the offset O and the original sampling grid G. The calculation process is shown in formula (4): (4) This embodiment implements dynamic upsampling from low-resolution to high-resolution feature maps. The generation and position of sampling points are dynamically determined based on the content of the input feature map. This method helps to better extract feature information and ensures the efficiency and effectiveness of the upsampling process.
[0046] S2-3 replaces the CIoU loss function in the YOLOv12 model with the MPDIoU loss function. MPDIoU calculates different loss function values for prediction boxes of different widths and heights, more effectively guiding the predicted box closer to the ground truth box and resolving the CIoU failure issue. MPDIoU not only simplifies the calculation process but also stabilizes the model convergence, improving the detection of small defects.
[0047] The MPDIoU calculation formula is as follows: (5); Where w and h are the width and height of the prediction box respectively. is the distance between the upper left corner of the predicted box and the upper left corner of the real box, are the distances between the lower right corner of the predicted box and the lower right corner of the real box. and The calculation formula is as follows: (6); (7); in, is the coordinate of the upper left corner of the real box, is the coordinate of the lower right corner of the real box; is the coordinate of the upper left corner of the prediction box, The coordinates of the lower right corner of the prediction box.
[0048] The loss function calculation formula based on MPDIoU is as follows: (8);
[0049] The MPDIoU loss function comprehensively considers the offset in position and size between bounding boxes, and can effectively converge regardless of whether the predicted box overlaps or does not overlap with the annotated box. It also reduces the loss by minimizing the distance between the two predicted corner points. This point distance-based measurement method solves the problem that the CIoU loss function cannot effectively optimize the predicted bounding box and the true bounding box with the same aspect ratio but completely different length and width values. It thus more accurately measures the difference between the predicted box of the oil seal surface defect and the true box of the oil seal surface defect, and can effectively improve the model's detection accuracy of oil seal surface defects.
[0050] S2-4 replaces the C3k2 module in the YOLOv12 model backbone network with the Multi-Scale Attention Fusion (MSAF) module. The MSAF module simultaneously processes local and global features, enabling a more comprehensive understanding of the input image. In object detection tasks, this helps the model identify and correctly locate objects in the image. Local features help capture object details such as edges and textures, while global features provide context, such as background and relationships between objects. Through dilated convolutions and pooling operations of varying sizes in the MSAF module, the module extracts features at multiple scales, enabling better processing of multi-scale objects and avoiding the receptive field issues of traditional convolutional networks. Through channel shuffling, the module enables better fusion of features generated by individual convolution operations, avoiding information bottlenecks and improving the network's ability to represent details and complex patterns in the image.
[0051] like Figure 6 As shown in Figure 2, the specific structure of the MSAF module in S2-4 is as follows: This module uses the input features as the starting point and constructs a dual-branch parallel processing architecture. The upper branch uses dilated convolutions with dilation coefficients of 1, 2, 3, and 4 to obtain and concatenate multi-scale features. This is then shuffled through the channels and resized using 1×1 convolutions. The lower branch first performs multi-scale pooling using pyramid pooling. The pooling results are flattened and concatenated before being calculated with the query vector extracted from the input features. The generated weights are applied to the value vector branch and output through a 1×1 convolution. Finally, the results of the upper and lower branches are added element-by-element to output enhanced features that integrate multi-scale features and the attention mechanism.
[0052] In this embodiment, the MSAF module adopts multiple strategies to enhance the feature extraction capability of the network. It combines local feature extraction, global feature extraction and attention mechanism to improve the performance of target detection.
[0053] S2-4-1, Local Feature Capture: This module extracts local features through four convolutional layers with different expansion rates. These convolutional layers effectively expand the receptive field, capturing local information of varying sizes. Features from different convolutional layers are then mixed through channel shuffling to improve information flow.
[0054] S2-4-2, Global Feature Capture: To capture global context, the module uses a custom QKV module. This module generates query, key, and value tensors. It then calculates the similarity between the query and key to obtain attention weights. After softmax normalization, the calculated similarity is then added to the weighted value tensor to ultimately generate global features.
[0055] S2-4-3, Feature Fusion: Local and global features are combined at the output stage. First, the global features undergo a 1×1 convolution and are then added to the original input to obtain an enhanced feature representation. Finally, the local features are added to the global features to form the final output.
[0056] The local feature extraction formula is as follows: (9); (10); The global feature extraction formula is as follows: (11); (12); (13); (14); (15); (16); (17); (18); The final output formula is as follows: (19); in, represents the input feature map, 、 、 、 Represents the input feature map The results after convolution operations with different expansion rates (expansion rates are 1, 2, 3, and 4 respectively) are shown. Indicates splicing, For channel rearrangement operation, represents the feature map after channel rearrangement, represents a 1×1 convolution operation, is the weight of the 1×1 convolution operation when fusing the features after channel rearrangement, Representing local features , QKV_block Represents a custom query-key-value module, Q represents the query tensor, K represents the key tensor, V represents a value tensor, for The 1×1 convolution weight of the branch, the output dimension is recorded as , represents pyramid pooling, Indicates flattening, for The 1×1 convolution weight of the branch has the same output dimension as Consistent, that is , for V The 1×1 convolution weight of the branch, the output dimension is recorded as , T represents the matrix transpose, is the scaling factor, sim_map represents the similarity matrix, is a normalization function that normalizes the numerical vector into a probability distribution vector. dim represents the dimension. represents the normalized similarity matrix, W is the weight of the 1×1 convolution operation when fusing global features, Represents global features.
[0057] In this embodiment, Figure 2 As shown in the figure, an oil seal surface defect detection model based on YOLOv12 is constructed, specifically: The initialized input image is input to Conv module No. 0 (standard convolution module), the output of Conv module No. 0 is connected to the input of Conv module No. 1, the output of Conv module No. 1 is connected to the input of MSAF module No. 2 (multi-scale attention fusion module), the output of MSAF module No. 2 is connected to the input of Conv module No. 3 and the input of Concat module No. 16, the output of Conv module No. 3 is connected to the input of MSAF module No. 4, the output of MSAF module No. 4 is connected to the input of Conv module No. 5 and the input of Concat module No. 13, the output of Conv module No. 5 is connected to the input of A2C2f module No. 6 (attention mechanism and feature fusion module), the output of A2C2f module No. 6 is connected to the input of Conv module No. 7 and the input of Concat module No. 10, the output of Conv module No. 7 is connected to the input of A2C2f module No. 8, and the output of A2C2f module No. 8 is connected to the input of dynamic upsampling module No. 9 and the input of Concat module No. 25.
[0058] The output of the No. 9 dynamic upsampling module is connected to the input of the No. 10 Concat module, the output of the No. 10 Concat module is connected to the input of the No. 11 A2C2f module, the output of the No. 11 A2C2f module is connected to the input of the No. 12 dynamic upsampling module and the input of the No. 22 Concat module, the output of the No. 12 dynamic upsampling module is connected to the input of the No. 13 Concat module, the output of the No. 13 Concat module is connected to the input of the No. 14 A2C2f module, the output of the No. 14 A2C2f module is connected to the input of the No. 15 dynamic upsampling module and the input of the No. 19 Concat module, the output of the No. 15 dynamic upsampling module is connected to the input of the No. 16 Concat module, the output of the No. 16 Concat module is connected to the input of the No. 17 A2C2f module, the output of the No. 17 A2C2f module is connected to the input of the No. 18 Conv module and the first No. 27 Detect ( The output of the Conv module No. 18 is connected to the input of the Concat module No. 19, the output of the Concat module No. 19 is connected to the input of the A2C2f module No. 20, the output of the A2C2f module No. 20 is connected to the input of the Conv module No. 21 and the input of the second Detect module No. 27, the output of the Conv module No. 21 is connected to the input of the Concat module No. 22, the output of the Concat module No. 22 is connected to the input of the A2C2f module No. 23, the output of the A2C2f module No. 23 is connected to the input of the Conv module No. 24 and the input of the third Detect module No. 27, the output of the Conv module No. 24 is connected to the input of the Concat module No. 25, the output of the Concat module No. 25 is connected to the input of the C3k2 module No. 26, and the output of the C3k2 module No. 26 is connected to the input of the fourth Detect module No. 27.
[0059] The outputs of the first, second, third, and fourth 27 Detect modules constitute the output of the total feature.
[0060] Step S3: Use the training samples to train the oil seal surface defect detection model to obtain a trained MPD-YOLO model.
[0061] The MPD-YOLO model is trained based on the oil seal defect training set.
[0062] Step S4: evaluate the trained MPD-YOLO model and adjust the MPD-YOLO model according to the evaluation results.
[0063] The specific steps include: In S4-1, the oil seal defect test set is input into the trained MPD-YOLO model, and the evaluation results are obtained based on the set key indicators: mAP (mean average precision), Precision, and Recall.
[0064] When evaluating model performance on the validation set, detection accuracy (Precision, P), recall (Recall, R), and mean average precision (mAP) are used as evaluation indicators.
[0065] Among them, detection accuracy measures the proportion of targets that are actually defects among all targets predicted by the model to be defects, reflecting the accuracy of the model in defect identification. The calculation formula is as follows: (20); The recall rate indicates the proportion of defective targets successfully detected by the model among all actual defective targets. It measures the coverage of the model for real targets. The calculation formula is as follows: (twenty one); Where TP represents the number of samples correctly classified as positive (actually defects, but the model detection result is defects); FP represents the number of samples incorrectly classified as positive (actually normal, but the model detection result is defects); FN represents the number of samples incorrectly classified as negative (actually defects, but the model detection result is normal).
[0066] The mean average precision comprehensively considers the precision and recall performance of the model under different IoU thresholds. It is usually calculated by the area under the precision-recall curve and is an important indicator for evaluating the global performance of the model. The calculation formula is as follows: (twenty two);
[0067] Where, n represents the number of categories,i Represents the category index, Represents the functional form of the precision-recall curve, P Represents precision, R Represents the recall rate.
[0068] Through these evaluation indicators, the actual performance of the improved YOLOv12 model in oil seal surface defect detection can be comprehensively evaluated to ensure its optimization in terms of accuracy and recall.
[0069] S4-2, adjust the hyperparameters of the MPD-YOLO model according to the evaluation results to obtain the adjusted MPD-YOLO model.
[0070] The above examples are merely preferred embodiments of the present invention and are not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, they are not intended to limit the present invention. Therefore, any simple modifications, equivalent variations, and modifications to the above examples that do not depart from the technical solution of the present invention and are based on the technical essence of the present invention shall fall within the scope of protection of the technical solution of the present invention.
Claims
1. An intelligent oil seal defect detection method based on improved YOLOv12, characterized in that: The following steps are involved: Step S1, collecting image data of the oil seal and annotating it to obtain an oil seal defect dataset; Step S2, improving the network framework of the YOLOv12 model and constructing an oil seal surface defect detection model based on YOLOv12, namely the MPD-YOLO model; Step S3, training the MPD-YOLO model to obtain a trained MPD-YOLO model; Step S4: evaluating the trained MPD-YOLO model based on the oil seal defect test set, and adjusting the MPD-YOLO model based on the evaluation results; In step S2, improving the YOLOv12 model network framework includes the following operations: S2-1, adjust the head hierarchical structure of the YOLOv12 model, integrate shallow semantic information, and add a small object detection layer; S2-2, use the dynamic upsampling module to replace the nearest neighbor interpolation upsampling module of the YOLOv12 model; S2-3, use the MPDIoU loss function to replace the CIoU loss function of the YOLOv12 model; S2-4, use the MSAF module to replace the C3k2 module in the YOLOv12 model backbone network.
2. The oil seal defect intelligent detection method based on improved YOLOv12 according to claim 1 is characterized in that: The adjustment of the YOLOv12 model detection head hierarchical structure is to add a small target detection layer P2 on the basis of the three detection layers, and then add a P2 upsampling fusion module after two rounds of upsampling fusion modules of the first detection layer P3 and the second detection layer P4, and connect it to the small target detection layer P2.
3. The oil seal defect intelligent detection method based on improved YOLOv12 according to claim 2 is characterized in that: In the dynamic upsampling module, the input feature map X is first interpolated into a continuous feature map through bilinear interpolation; then, an offset O is generated through a linear layer, which determines the position of the sampling point in the continuous space, and the sampling point is used to reconstruct the upsampled feature map X'.
4. The oil seal defect intelligent detection method based on improved YOLOv12 according to claim 3 is characterized in that: In the dynamic upsampling module, a sampling point generator based on a static range factor or a dynamic range factor is used to generate sampling points.
5. The oil seal defect intelligent detection method based on improved YOLOv12 according to claim 4 is characterized in that: The specific structure of the MSAF module is as follows: starting from the input features, a dual-branch parallel processing architecture is constructed. The upper branch of the dual branch uses dilation convolutions with expansion coefficients of 1, 2, 3, and 4 to obtain multi-scale features and splice them, and then adjusts the dimension through channel shuffling and 1×1 convolution; the lower branch of the dual branch first implements multi-scale pooling through pyramid pooling. After the pooling results are flattened and spliced, they are operated with the query vector extracted from the input features. The generated weights are applied to the value vector branch and output through 1×1 convolution; finally, the results of the upper and lower branches are added element by element to output enhanced features that integrate multi-scale and attention mechanisms.
6. The oil seal defect intelligent detection method based on improved YOLOv12 according to claim 5 is characterized in that: The method for evaluating the trained MPD-YOLO model in step S4 is: The oil seal defect test set is input into the trained MPD-YOLO model, and the evaluation results are obtained based on the set key indicators: mean average precision, precision, and recall rate; The accuracy formula is: ; The recall formula is: ; In the formula, TP represents the number of samples correctly classified as positive; FP represents the number of samples incorrectly classified as positive; FN represents the number of samples incorrectly classified as negative. The calculation formula of the average precision mean is: ; Where, n represents the number of categories, i Represents the category index, Represents the functional form of the precision-recall curve, P Represents precision, R Represents the recall rate.
Citation Information
Patent Citations
Defect target detection method, system and equipment based on improved YOLOv5 and medium
CN119919646A
Penicillin bottle body defect detection method based on improved YOLOv8
CN120318607A
Cited By
Cell real-time detection method and device based on improved YOLOv12 and storage medium
CN121685440A