Mechanical part defect detection method and system based on improved YOLOv12 model

By improving the structure of the YOLOv12 model, using C3k2_RCB, GDSAFusion and A2C2f_STR modules, the accuracy and stability problems of the YOLOv12 model in mechanical parts detection are solved, and efficient and accurate defect detection is achieved.

CN120279020AActive Publication Date: 2025-07-08JIANGXI SCI & TECH NORMAL UNIV

Patent Information

Application Number
CN202510759132.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-09
Publication Date
2025-07-08
Estimated Expiration
2045-06-09

AI Technical Summary

Technical Problem

The existing YOLOv12 model has insufficient detection accuracy and recall in the detection of visual defects of mechanical parts, especially in the poor performance when dealing with small goals, complex backgrounds and multi-scale goals. Data imbalance and noise interference affect the generalization ability and stability of the model in industrial production environments.

Method used

By improving the YOLOv12 model, the original C3k2 module was replaced by the C3k2_RCB feature extraction module was used to replace the original C3k2 module, the gated dynamic space aggregator GDSAFusion was introduced to replace the Concat module, and the A2C2f_STR module was used in the backbone network and the neck network, and the model structure was optimized to enhance feature extraction and anti-interference capabilities.

Benefits of technology

It significantly improves the accuracy and efficiency of mechanical parts defect detection, and can accurately identify various complex and subtle defects in complex industrial environments, meeting the high-quality inspection needs of modern manufacturing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279020A_ABST
    Figure CN120279020A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of image processing and the field of industrial detection, and discloses a mechanical part defect detection method and system based on an improved YOLOv12 model. Then the preprocessed mechanical part image data set is used for training an improved YOLOv12 mechanical part defect detection model to obtain an optimized model, and improvement comprises the steps that a C3k2RCB feature extraction module is used for replacing an original C3k2 module in a backbone network and a neck network; replacing an original Concat module by using a GDSAFusion (Gated Dynamic Spatial Aggregate) in a neck network; the original A2C2f module is replaced by the A2C2fSTR in the backbone network and the neck network; and finally, performing real-time defect detection on the mechanical part by using the optimization model. According to the method, the structure of the YOLOv12 model is optimized, so that the defect detection precision and efficiency of the YOLOv12 model on the mechanical part in a complex industrial environment are remarkably improved, and particularly, the detection performance under tiny defects and a complex background is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of image processing and industrial inspection, and particularly relates to a mechanical part defect detection method and system based on an improved YOLOv12 model. Background Technique

[0002] In modern manufacturing, the quality of mechanical parts directly affects the performance and reliability of the entire product. Therefore, it is crucial to perform high-precision and high-efficiency visual defect detection on mechanical parts. Traditional mechanical part visual defect detection methods, such as image processing algorithms based on threshold segmentation and edge detection, often rely on manually setting parameters, have poor adaptability to complex lighting conditions, differences in part surface textures, and the detection of tiny defects, and have low detection efficiency, making it difficult to meet the needs of large-scale industrial production.

[0003] In recent years, deep learning has demonstrated powerful performance in the field of object detection. The YOLO (You Only Look Once) series of algorithms have been widely applied to industrial inspection scenarios due to their fast and efficient detection capabilities. As an advanced version of this series, the YOLOv12 model further optimizes the network structure and detection algorithm on the basis of inheriting the end-to-end and fast detection advantages of the YOLO series, significantly improving the detection accuracy and speed. However, in the actual application of mechanical part visual defect detection, the YOLOv12 model still has some problems. Mechanical parts have diverse shapes and sizes, and the surface defect features of some parts are subtle. The existing YOLOv12 model needs to further improve the detection accuracy and recall rate when dealing with small target detection, feature extraction in complex backgrounds, and multi-scale target detection. At the same time, data collection in the industrial production environment has problems such as data imbalance and noise interference, which also affect the generalization ability and stability of the model.

[0004] Therefore, it is urgent to propose a mechanical part defect detection method based on an improved YOLOv12 model. By optimizing and improving aspects such as the model structure, feature extraction method, and training strategy, the detection accuracy and efficiency of mechanical parts are improved, and the adaptability of the model in a complex industrial environment is enhanced to meet the needs of high-quality detection of mechanical parts in modern manufacturing. Summary of the Invention

[0005] The purpose of the present invention is to provide a mechanical part defect detection method and system based on an improved YOLOv12 model to solve the above problems and improve the detection accuracy and efficiency of mechanical parts.

[0006] In the first aspect, the present invention provides a mechanical part defect detection method based on an improved YOLOv12 model, including the following steps: Collect a defective mechanical part image dataset, and preprocess the mechanical part image dataset. The preprocessing includes annotating the defective parts of each mechanical part image, as well as performing image data augmentation and enhancement; Use the preprocessed mechanical part image dataset to train an improved YOLOv12 mechanical part defect detection model to obtain an optimized model. The improved YOLOv12 mechanical part defect detection model includes: In the backbone network and the neck network, use the C3k2_RCB feature extraction module to replace the original C3k2 module; In the neck network, use the gated dynamic spatial aggregator GDSAFusion to replace the original Concat module; In the backbone network, introduce the visual foundation model OverLoCK based on the dynamic convolution kernel and the context mixing mechanism; Use the optimized model to perform real-time defect detection on mechanical parts.

[0007] As an optional implementation manner of the first aspect of the present application, in the step of using the C3k2_RCB feature extraction module to replace the original C3k2 module in the backbone network and the neck network, the implementation process of the C3k2_RCB feature extraction module includes: initializing the C3k2_RCB module, including receiving input channel number, output channel number, the repetition number n of the RepConvBlock module, a judgment flag, a channel scaling factor, the number of convolutional feature extraction layers, and a flag parameter indicating whether to use residual connection; inputting the feature map into the processing unit inside the C3k2_RCB module; the processing unit dynamically constructs a core processing structure according to the judgment flag, and the core processing structure is composed of n RepConvBlock modules. Specifically, it includes: passing the feature map through the convolutional feature extraction layer, and every time it passes through a convolutional feature extraction layer, a feature splitting operation is used to divide it into two feature maps. The feature map obtained by passing one of the feature maps through n RepConvBlock modules is combined with the other feature map separated by the feature splitting operation using the feature splicing operation, and finally the features are integrated through a convolutional fusion layer to obtain the final output feature map.

[0008] As an alternative implementation of the first aspect of the present application, the implementation process of the RepConvBlock module includes: extracting local features from the input feature map through a first 3×3 depthwise separable convolution to obtain a local feature map; processing the local feature map through a multi-branch projection module, the projection module sequentially includes a normalization layer, a reparameterizable dilated convolution, a batch normalization layer, a channel attention module, a first 1×1 convolution layer, a GELU activation function, a second 3×3 depthwise separable convolution, global response normalization, and a second 1×1 convolution layer; if residual scaling is enabled, the input feature map is scaled and then added to the output of the projection module to obtain the output feature; otherwise, using a residual connection, directly taking the output of the projection module as the output feature.

[0009] As an alternative implementation of the first aspect of the present application, in the step of using the gated dynamic spatial aggregator GDSAFusion to replace the original Concat module in the neck network, the implementation process of the gated dynamic spatial aggregator GDSAFusion includes: concatenating the input feature map and the context feature in channels to form a fused feature; extracting local features from the fused feature through a 3×3 depthwise separable convolution, and then stabilizing the feature distribution through a normalization layer; calculating the spatial attention weight through a query-key mechanism, and the query-key mechanism uses a relative position bias; capturing multi-scale context information from the fused feature through a reparameterizable dilated convolution, and then performing channel adaptive calibration through a channel attention module; controlling the information flow through a gating mechanism; performing weighted fusion on the feature processed by the gating mechanism and the feature of the residual path; processing the weighted fused feature with a two-level scaling in combination with stochastic depth dropout to obtain an enhanced feature map.

[0010] As an alternative implementation of the first aspect of the present application, in the step of replacing the original A2C2f module with the A2C2f_STR module in the backbone network and the neck network, the implementation process of the A2C2f_STR module includes: connecting the input feature tensor to the initial convolutional layer, and the initial convolutional layer compresses the channel dimension of the input feature tensor through low-rank mapping and extracts the feature information with basic representation ability to obtain the feature map after the initial convolution process; temporarily storing the feature map after the initial convolution process in the list data structure, and entering the loop processing link composed of the visual modeling module that fuses cross-index temporal interaction. In the loop processing link, using the tail element of the list data structure as the input, sequentially passing it into each visual modeling module that fuses cross-index temporal interaction for processing to obtain the feature vector generated after being processed by the visual modeling module that fuses cross-index temporal interaction; fusing the feature vector generated after being processed by the visual modeling module that fuses cross-index temporal interaction with the original input feature tensor through the residual mapping mechanism to obtain the finally output fused feature.

[0011] As an alternative implementation of the first aspect of the present application, the implementation process of the visual modeling module that fuses cross-index temporal interaction includes: performing the first layer of normalization processing on the feature map after the initial convolution process to obtain the first normalized feature; realizing the adaptive weight fusion of the first normalized feature and the feature map after the initial convolution process through the learnable parameter to obtain the adaptive weight fusion feature; inputting the adaptive weight fusion feature into the cross-index temporal interaction modeling component, and the cross-index temporal interaction modeling component uses the scan index and the inverse scan index to perform the temporal feature recombination and modeling in the cross-space dimension to obtain the feature processed by the cross-index temporal interaction modeling component; performing the second layer of normalization processing on the feature processed by the cross-index temporal interaction modeling component to obtain the second normalized feature; entering the second normalized feature into the feed-forward convolutional block to complete the non-linear transformation to obtain the feed-forward output; fusing the feature processed by the cross-index temporal interaction modeling component with the feed-forward output through the residual connection to obtain the feature vector generated after being processed by the visual modeling module that fuses cross-index temporal interaction.

[0012] As an alternative implementation of the first aspect of the present application, the implementation process of the cross-index temporal interaction modeling component includes: generating index pairs for the adaptive weight fusion features input to the cross-index temporal interaction modeling component through an index generation module, where the index pairs include a scan index and an inverse scan index; using the scan index to rearrange the features of the adaptive weight fusion features input to the cross-index temporal interaction modeling component to obtain rearranged one-dimensional temporal features; mapping the rearranged one-dimensional temporal features to dynamic parameters through a temporal parameter projection layer; performing temporal modeling on the dynamic parameters by a selective scan engine based on a state space model to output one-dimensional sequence features after temporal modeling; restoring the one-dimensional sequence features after temporal modeling to two-dimensional spatial features through the inverse scan index; generating gating weights through an adaptive gating unit, and multiplying the gating weights element-wise with the two-dimensional spatial features to output the finally processed features of the cross-index temporal interaction modeling component.

[0013] In a second aspect, an embodiment of the present application provides a mechanical part defect detection system based on an improved YOLOv12 model, including: An image acquisition module, configured to collect a defective mechanical part image dataset and preprocess the mechanical part image dataset, where the preprocessing includes annotating the defective parts of each mechanical part image, as well as performing image data augmentation and enhancement; A model training module, configured to use the preprocessed mechanical part image dataset to train an improved YOLOv12 mechanical part defect detection model to obtain an optimized model, where the improved YOLOv12 mechanical part defect detection model includes: Using a C3k2_RCB feature extraction module to replace the original C3k2 module in the backbone network and the neck network; Using a gated dynamic spatial aggregator GDSAFusion to replace the original Concat module in the neck network; Using A2C2f_STR to replace the original A2C2f module in the backbone network and the neck network; A defect detection module, configured to perform real-time defect detection on mechanical parts using the optimized model.

[0014] In a third aspect, an embodiment of the present application provides an electronic device, where the electronic device includes a processor, a memory, and a program or instruction stored on the memory and executable on the processor. When the program or instruction is executed by the processor, the steps of the method described in the first aspect are implemented.

[0015] In a fourth aspect, an embodiment of the present application provides a readable storage medium, where a program or instruction is stored on the readable storage medium. When the program or instruction is executed by the processor, the steps of the method described in the first aspect are implemented.

[0016] Compared with the prior art, the present invention provides a method for detecting defects of mechanical parts based on an improved YOLOv12 model. This method first collects a dataset of defective mechanical part images and performs preprocessing, including annotating the defective parts of each image and performing image data augmentation and enhancement. By providing high-quality and diverse training data and simulating complex environments, the generalization ability of the subsequent trained model and its adaptability to the complex environments of actual industrial scenarios are significantly improved. Then, the preprocessed dataset is used to train an improved YOLOv12 mechanical part defect detection model. The key to the improved model lies in its structural optimization: in the backbone network and the neck network, the C3k2_RCB feature extraction module is used to replace the original C3k2 module. This improvement enhances the model's detection ability for small targets, minute and internal defects through heterogeneous convolution kernel topology, residual feature recalibration, and the characteristics of the internally stacked RepConvBlock module, improves the feature extraction accuracy in complex industrial environments (such as oil stains and reflections), and optimizes the inference performance on edge devices. At the same time, in the neck network, the gated dynamic spatial aggregator GDSAFusion is used to replace the original Concat module. This improvement breaks through the representation bottleneck of traditional methods in detecting irregular defects (such as casting pores and machining tool marks) by combining the dynamic weight generation and context mixing mechanisms and using the dynamic kernel optimization strategy based on the non-local attention operator, significantly enhances the anti-interference ability of the model in complex industrial backgrounds, and effectively solves the problem of feature degradation under harsh working conditions such as metal surface reflections and oil stains. In addition, in the backbone network and the neck network, the A2C2f_STR module is used to replace the original A2C2f module. By introducing the visual modeling module STR that fuses cross-index temporal interaction, the model can effectively associate spatial and temporal dimension information, thereby strengthening the detection performance of continuous features of periodic defects or dynamic deformations such as cracks on the surface of rotating parts and solving the problem that the traditional A2C2f module is difficult to capture long-distance spatio-temporal dependencies due to its reliance on local convolution operations. Through the above targeted improvements to the key modules of the model, the optimized model obtained through training inherits the fast and efficient advantages of the YOLO series and greatly improves the detection accuracy, robustness, and efficiency of various complex and subtle defects of mechanical parts in complex and harsh industrial environments. Finally, the optimized model obtained through training is used to perform real-time defect detection on mechanical parts. With its high accuracy, high robustness, and high efficiency, it can timely and accurately detect and locate defects, meeting the requirements of modern manufacturing for high-quality detection. Description of the Drawings

[0017] Figure 1 It is a flowchart of a method for detecting defects of mechanical parts based on an improved YOLOv12 model provided by an embodiment of the present invention; Figure 2Schematic diagram of the model structure for the mechanical part defect detection method based on the improved YOLOv12 model; Figure 3 Schematic diagram of the C3k2_RCB module structure; Figure 4 Schematic diagram of the RepConvBlock module structure; Figure 5 Schematic diagram of the GDSAFusion module structure; Figure 6 Schematic diagram of the A2C2f_STR module structure; Figure 7 Schematic diagram of the STR module structure; Figure 8 Schematic diagram of the CITIM unit structure; Figure 9 Schematic diagram of the structure of a mechanical part defect detection system based on the improved YOLOv12 model provided by an embodiment of the present invention. Detailed implementation manners

[0018] Next, the technical solutions in the embodiments of the present application will be clearly and completely described with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0019] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects and are not used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein. In addition, "and / or" in the specification and claims means at least one of the connected objects, and the character " / " generally indicates an "or" relationship between the associated objects before and after. In the description of the present invention, "a plurality" means two or more unless otherwise specifically defined.

[0020] Embodiment 1 Please refer to Figure 1 , which is the implementation flowchart of a mechanical part defect detection method based on the improved YOLOv12 model provided by an embodiment of the present invention. Referring to Figure 1 , a mechanical part defect detection method based on the improved YOLOv12 model provided by an embodiment of the present invention includes the following steps: S1: Collect a dataset of defective mechanical part images and preprocess the dataset of mechanical part images. The preprocessing includes annotating the defective parts of each mechanical part image, as well as performing image data augmentation and enhancement.

[0021] First, collect a dataset of defective mechanical part images. Exemplarily, 3000 common mechanical part defect images are collected and photographed in a factory, including typical industrial scenario data such as internal hole defects in cast iron, surface fatigue cracks in bearings, tooth surface wear of gears, and deformation of precision shaft parts.

[0022] Second, use a computer to annotate the defective parts of each photographed photo and perform data augmentation and enhancement. Exemplarily, import the dataset of common mechanical part defect images into the X-AnyLabeling annotation tool, mark the defective parts of these 3000 images, and label them in the yolo format. The annotation file contains information such as the category number of each defective target; and use data augmentation methods such as randomly enhancing contrast, noise, flipping, and scaling to expand the image data and labels to 30000, simulating the images recognized by the camera in various extreme situations, so as to improve the generalization ability of the training model.

[0023] Finally, use a computer to divide the dataset into a validation set, a training set, and a test set, and process it into the recognition format of the YOLOv12 network model. Exemplarily, use python code to divide the mechanical part defect dataset into a training set, a validation set, and a test set according to a set ratio of 8:1:1. Among them, the training set will be used to train the model, the validation set is used for evaluation during the training process, and the test set will be used to evaluate the performance of the model, and process it into the recognition format of the YOLOv12 network model.

[0024] S2: Use the preprocessed dataset of mechanical part images to train the improved YOLOv12 mechanical part defect detection model to obtain an optimized model.

[0025] As Figure 2 shown, it is the improved YOLOv12 mechanical part defect detection model. The improved YOLOv12 mechanical part defect detection model includes: using the C3k2_RCB feature extraction module to replace the original C3k2 module in the backbone network and the neck network; using the gated dynamic spatial aggregator GDSAFusion to replace the original Concat module in the neck network; using A2C2f_STR to replace the original A2C2f module in the backbone network and the neck network.

[0026] The C3k2_RCB module realizes cross-modal feature extraction of internal defects of mechanical parts while reducing the number of model parameters through heterogeneous convolution kernel topology optimization and residual feature recalibration mechanism. Its innovative depthwise separable convolution group architecture combined with dynamic receptive field recombination technology reduces the false detection rate of traditional algorithms in micro-defect detection. At the same time, hardware-aware operator optimization makes the inference latency more stable on industrial edge devices.

[0027] As Figure 3 shown, in the step of replacing the original C3k2 module with the C3k2_RCB feature extraction module in the backbone network and the neck network, the implementation process of the C3k2_RCB feature extraction module includes: initializing the C3k2_RCB module, including receiving the number of input channels, the number of output channels, the repetition times n of the RepConvBlock module, the judgment flag, the channel scaling factor, the number of convolutional feature extraction layers, and the flag parameter indicating whether to use residual connection; inputting the feature map into the processing unit inside the C3k2_RCB module; the core processing structure of the processing unit consists of n RepConvBlock modules, specifically including: the feature map first passes through the convolutional feature extraction layer, and then every time it passes through a convolutional feature extraction layer, it uses a feature splitting operation to be divided into two feature maps. The feature map obtained by passing one of the feature maps through n RepConvBlock modules is combined with the other feature map separated by the feature splitting operation using a feature concatenation operation, and finally the features are integrated through a convolutional fusion layer to obtain the final output feature map.

[0028] In the visual inspection scenario of mechanical parts, the C3k2_RCB module plays an important role. Mechanical parts have different shapes and sizes, and some defect features are subtle. It is difficult for traditional modules to balance the detection accuracy of large and small targets. In the C3k2_RCB module, when the C3k_RCB module is enabled, it combines the channel processing ability of the original C3k2 module and the feature extraction advantages of the introduced internal stacked dynamic convolution block RepConvBlock module, which can effectively enhance the detection ability of small target defects; and the reparameterization characteristic of the RepConvBlock module enables it to extract features more accurately in complex industrial environments such as dirt and reflection interference, reducing false detection and missed detection. At the same time, the module can flexibly adjust its internal structure according to the distribution of defect categories in the data, alleviate the problem of missed detection of rare defects, and improve the robustness and detection accuracy of the model in mechanical part detection, meeting the requirements of actual industrial production for high-quality detection.

[0029] As Figure 4As shown, the RepConvBlock module's implementation process includes: performing local feature extraction on the input feature map through a 3×3 depthwise separable convolution to obtain a local feature map; processing the local feature map through a multi-branch projection module, which sequentially includes a normalization layer, a reparameterizable dilated convolution, a batch normalization layer, a channel attention module, a 1×1 convolution layer, a GELU activation function, a 3×3 depthwise separable convolution, global response normalization, and a 1×1 convolution layer; if residual scaling is enabled, scaling the input feature map and adding it to the output of the projection module to obtain the output feature; otherwise, using a residual connection and directly taking the output of the projection module as the output feature.

[0030] It can be understood that the input feature map first undergoes local feature extraction through a 3×3 depthwise separable convolution (dwconv) to capture spatial local information. Subsequently, the feature map passes through a multi-branch projection module (proj), which sequentially includes the following operations: first, standardizing the feature map through a normalization layer; then, enhancing the receptive field and fusing multi-scale information through a reparameterizable dilated convolution; further stabilizing the feature distribution through a batch normalization layer; then, adaptively weighting the channels of the feature map through a channel attention module, adaptively calibrating the channel weights, and enhancing the contribution of important features; afterwards, expanding the number of channels through a 1×1 convolution layer; introducing non-linearity by applying the GELU activation function; and then further extracting features through a 3×3 depthwise separable convolution (dwconv) to obtain the output feature ; enhancing the distinctiveness of the features through global response normalization; and finally, through a convolution layer to restore the number of channels to the original dimension. Finally, if residual scaling (ls) is enabled, the input feature map is scaled and added to the output of the projection module X proj to obtain X final , otherwise, directly using the residual connection to output X final . Its forward propagation mathematical expression is: where, represents stochastic depth.

[0031] Furthermore, the depthwise separable convolution's implementation process includes: passing the input feature map through a 3×3 depthwise separable convolution with the number of groups equal to the number of input channels. This design reduces the number of parameters from that of a standard convolution from to ( is the number of channels, (for the nuclear size), improving the computational efficiency; meanwhile, in the forward propagation, the output of the depthwise separable convolution DepthwiseConv( ) is directly added to the input feature map to form a residual connection, obtaining the output feature . Through this design, both the original feature information is retained, and the problem of gradient disappearance is alleviated through the skip connection. The mathematical expression of its forward propagation is: As Figure 5 shown, in the step of using the gated dynamic spatial aggregator GDSAFusion to replace the original Concat module in the neck network, the implementation process of the gated dynamic spatial aggregator GDSAFusion includes: concatenating the input feature map with the context feature along the channel dimension to form the fused feature ; extracting local features from the fused feature through a 3×3 depthwise separable convolution (Dwconv), and then stabilizing the feature distribution through a normalization layer (Norm) to obtain ; calculating the spatial attention weights through a query-key mechanism, where the query-key mechanism uses a relative position bias (rpb) to enhance the position perception ability; capturing multi-scale context information from the fused feature through a reparameterizable dilated convolution (lepe) to obtain , and then performing channel adaptive calibration through a channel attention module to obtain ; controlling the information flow through a gating mechanism to obtain ; performing weighted fusion of the feature processed by the gating mechanism (gate) and the feature of the residual path; using a two-level scaling (ls1 and ls2) combined with a stochastic depth dropout (drop_path) to process the weighted fused feature, obtaining an enhanced feature map, which significantly improves the training stability while maintaining the feature expression ability. The mathematical expression of its forward propagation is: Among them, represents the dimension concatenation of vectors, represents the query Q weight, represents the key K weight, represents the attention weight, represents an activation function that can normalize the input into a probability distribution vector, d represents the dimension of the key K, represents the attention channel layer.

[0032] The GDSA Fusion module combines dynamic weight generation and context mixing mechanisms. By constructing a multi-modal feature interaction matrix, it breaks through the representation bottleneck of traditional vision algorithms in the detection of irregular defects such as casting pores and machining tool marks. Its dynamic kernel optimization strategy based on non-local attention operators greatly improves the signal-to-noise ratio in complex industrial backgrounds, solving the problem of feature degradation of traditional methods under harsh working conditions such as metal surface reflection and oil stain interference.

[0033] As Figure 6 shown, the implementation process of A2C2f_STR is as follows: The input feature tensor X first accesses the initial convolutional layer constructed according to the parent class initialization mechanism. This initial convolutional layer compresses the channel dimension of the input feature tensor through low-rank mapping. By means of the sliding window convolution operation of the convolutional kernel in the feature map spatial domain, it performs the point-by-point multiplication and accumulation operation of the weights and the corresponding position elements, completing the dimension conversion process from the input channel to the hidden channel, and then extracting the feature information with basic representation ability. The feature map after the initial convolution processing will be temporarily stored in a specific list data structure. Subsequently, the feature processing flow enters the loop processing link composed of a visual modeling module (STR module) that fuses cross-index temporal interaction. In this link, taking the element at the end of the list as the input, it is sequentially passed into each STR module for processing. The feature vector X2 generated after being processed by the STR module will be fused with the original input feature tensor X through the residual mapping mechanism, and finally the fused feature Y is output. The mathematical expression of its forward propagation process is: As Figure 7 shown, the STR module (Spatio-Temporal Reformer) is a visual modeling module that fuses cross-index temporal interaction, mainly composed of double-layer normalization, adaptive residual connection, a cross-index temporal interaction modeling component (CITIM unit (Cross-Index Temporal Interaction Module)), and a feed-forward convolutional block. This module first performs the first layer of normalization (LayerNorm) processing on the input feature to obtain the normalized feature ; through the learnable parameters and it realizes the adaptive weight fusion of the normalized feature and the original input feature map (the feature map after the initial convolution processing); the fused adaptive weight fusion feature is input into the cross-index temporal interaction modeling component (CITIM unit), and the scan index and inverse index (IDs) generated by the cross-index temporal interaction modeling component are used for cross-space dimension temporal feature recombination and modeling; then it is processed by the second layer of normalization to obtain , and enter the feed-forward convolutional block Mlp to complete the non-linear transformation and output the feed-forward output ; Finally, the adaptive weight fusion feature is obtained through the residual connection and the feed-forward output are fused to obtain the final output feature vector X2. The mathematical expression of its forward propagation is as follows: Among them, IDs is an index pair composed of a scan index and an inverse scan index; the scan index is a two-dimensional feature map traversal order index generated according to a preset scan length and offset strategy, which is used to convert spatial features into a one-dimensional sequence for temporal modeling; the inverse scan index is the inverse mapping of the scan index, which is used to restore the one-dimensional sequence after temporal modeling to the original two-dimensional feature map structure to ensure the accurate restoration of the feature spatial position. The two together constitute the core index mechanism for feature cross-dimensional interaction.

[0034] As Figure 8 shown, the cross-index temporal interaction modeling component (CITIM unit) is mainly composed of an index generation module, a feature rearrangement layer, a temporal parameter projection layer, a selective scan engine, an inverse index restoration module, and an adaptive gating unit. The CITIM unit first outputs the index pair IDs for the input feature through the index generation module; uses the scan index to rearrange into a one-dimensional temporal feature ; projects into dynamic parameters (temporal decay factor), (input mapping matrix), (output mapping matrix) through the temporal parameter projection layer; the selective scan engine based on the state space model (SSM) performs temporal modeling on and outputs the one-dimensional sequence feature ; then restores to the two-dimensional spatial feature through the inverse scan index; finally, the adaptive gating unit generates weights and multiplies them element-wise with to output the finally processed feature by the cross-index temporal interaction modeling component. The mathematical expression of its forward propagation is as follows: Among them, index_scan is the index scan; represents the inverse index restoration function, which is responsible for restoring the one-dimensional temporal feature to the two-dimensional spatial feature; represents the skip connection parameter; Project is the parameter projection layer; inverse_ids represents the inverse scan index; Gating is the gating function; It is an operation to constrain the value range of learnable parameters; A_logs are the learnable parameters in the model.

[0035] Furthermore, input the training set and the validation set into the mechanical part defect detection model of the improved YOLOv12, and set the number of training times. As the number of training times increases, the loss function curve of the model gradually converges. When the loss function curve converges and stabilizes, the mechanical part defect detection model is trained to the optimal state, and its optimal model weight file is saved. Input the images to be detected in the test set into the trained mechanical part defect detection model, and output the detected images of the die parts, where the type of each detection target is included in the detection image, and the position of each target in the object detection image is marked. Download the optimal weight file and save it on the computer for subsequent use.

[0036] S3: Use the optimized model to perform real-time defect detection on mechanical parts.

[0037] Exemplarily, deploy the trained model weight file to the production environment, capture the images of mechanical parts in real time through a camera or an image acquisition device, input them into the model for analysis, and the model will quickly identify and mark the defect position and type.

[0038] In summary, in this embodiment, the improved YOLOv12 model proposed by the present invention can be applied to more complex mechanical part defect recognition application scenarios. The A2C2f_STR module strengthens the detection performance of periodic defects or dynamic deformations such as the continuous characteristics of cracks on the surface of rotating parts by introducing the visual modeling module STR that fuses cross-index temporal interaction, which associates the cross-index temporal interaction modeling space and time dimensions, and solves the problem that traditional A2C2f relies on local convolution operations and is difficult to capture long-distance spatio-temporal dependencies; the GDSAFusion module combines the dynamic weight generation and context mixing mechanisms, breaks through the representation bottleneck of traditional vision algorithms in the detection of irregular defects such as casting pores and machining tool marks by constructing a multi-modal feature interaction matrix, and its dynamic kernel optimization strategy based on non-local attention operators greatly improves the signal-to-noise ratio in complex industrial backgrounds, solving the problem of feature degradation of traditional methods under harsh working conditions such as metal surface reflection and oil stain interference; the C3k2_RCB module realizes cross-modal feature extraction of internal defects of mechanical parts while reducing the number of model parameters through heterogeneous convolution kernel topology optimization and residual feature recalibration mechanisms. Its innovative depthwise separable convolution group architecture combined with dynamic receptive field recombination technology reduces the false detection rate of traditional algorithms in the detection of tiny defects, and at the same time makes the inference latency more stable on industrial-grade edge devices through hardware-aware operator optimization.

[0039] Embodiment 2 Please refer to Figure 9, as shown in the following is a schematic structural diagram of a mechanical part defect detection system based on an improved YOLOv12 model proposed in the second embodiment of the present application. This system includes the following key modules: An image acquisition module 100, which is used to collect a defective mechanical part image data set, and preprocess the mechanical part image data set. The preprocessing includes annotating the defective parts of each mechanical part image, as well as performing image data augmentation and enhancement; A model training module 200, which is used to train an improved YOLOv12 mechanical part defect detection model using the preprocessed mechanical part image data set to obtain an optimized model. The improved YOLOv12 mechanical part defect detection model includes: Using a C3k2_RCB feature extraction module to replace the original C3k2 module in the backbone network and the neck network; Using a gated dynamic spatial aggregator GDSAFusion to replace the original Concat module in the neck network; Using an A2C2f_STR to replace the original A2C2f module in the backbone network and the neck network; A defect detection module 300, which is used to perform real-time defect detection on mechanical parts using the optimized model.

[0040] A mechanical part defect detection system based on an improved YOLOv12 model in an embodiment of the present application can be a device, or a component, an integrated circuit, or a chip in a terminal. This device can be a mobile electronic device or a non-mobile electronic device. Exemplarily, the mobile electronic device can be a mobile phone, a tablet computer, a laptop computer, a handheld computer, a vehicle-mounted electronic device, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc. The non-mobile electronic device can be a server, a Network Attached Storage (NAS), a personal computer (PC), etc. The embodiments of the present application do not make specific limitations.

[0041] A mechanical part defect detection system based on an improved YOLOv12 model in an embodiment of the present application can be a device with an operating system. This operating system can be an Android operating system, an iOS operating system, or other possible operating systems. The embodiments of the present application do not make specific limitations.

[0042] A mechanical part defect detection system based on an improved YOLOv12 model provided in an embodiment of the present application can achieveFigure 1 For the sake of avoiding repetition, the processes implemented in the method embodiments of a mechanical part defect detection method based on an improved YOLOv12 model will not be elaborated here.

[0043] Optionally, an embodiment of the present application further provides an electronic device, including a processor, a memory, and a program or instruction stored on the memory and executable on the processor. When the program or instruction is executed by the processor, it implements the processes of the above method embodiments of a mechanical part defect detection method based on an improved YOLOv12 model, and can achieve the same technical effects. For the sake of avoiding repetition, they will not be elaborated here.

[0044] An embodiment of the present application further provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by the processor, it implements the processes of the above method embodiments of a mechanical part defect detection method based on an improved YOLOv12 model, and can achieve the same technical effects. For the sake of avoiding repetition, they will not be elaborated here.

[0045] Among them, the processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs, etc.

[0046] It should be noted that in this article, the term "including", "comprising", or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article, or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article, or device. Without more limitations, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, article, or device including that element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in a reverse order according to the functions involved. For example, the described methods may be executed in an order different from that described, and various steps may be added, omitted, or combined. Additionally, the features described with reference to certain examples may be combined in other examples.

[0047] Through the description of the above embodiments, those skilled in the art can clearly understand that the above method of the embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of the present application.

[0048] The embodiments of the present application have been described above in conjunction with the accompanying drawings. However, the present application is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present application, those of ordinary skill in the art can also make many forms without departing from the purpose of the present application and the scope protected by the claims, and all of them belong to the protection scope of the present application.

Claims

1. A mechanical part defect detection method based on an improved YOLOv12 model, characterized in that The steps include: Collect a dataset of defective mechanical part images, and preprocess the mechanical part image dataset. The preprocessing includes annotating the defective parts of each mechanical part image, as well as performing image data augmentation and enhancement; Use the preprocessed mechanical part image dataset to train an improved YOLOv12 mechanical part defect detection model to obtain an optimized model. The improved YOLOv12 mechanical part defect detection model includes: Replace the original C3k2 module with the C3k2_RCB feature extraction module in the backbone network and the neck network; Replace the original Concat module with the gated dynamic spatial aggregator GDSAFusion in the neck network; Replace the original A2C2f module with A2C2f_STR in the backbone network and the neck network; Use the optimized model to perform real-time defect detection on mechanical parts.

2. A mechanical part defect detection method based on an improved YOLOv12 model according to claim 1, characterized in that, In the step of replacing the original C3k2 module with the C3k2_RCB feature extraction module in the backbone network and the neck network, the implementation process of the C3k2_RCB feature extraction module includes: Initialize the C3k2_RCB module, including receiving input channel number, output channel number, repetition times n of the RepConvBlock module, judgment flag, channel scaling factor, number of convolutional feature extraction layers, and flag parameter indicating whether to use residual connection; Input the feature map into the processing unit inside the C3k2_RCB module. The processing unit dynamically constructs a core processing structure according to the judgment flag. The core processing structure is composed of n RepConvBlock modules, specifically including: Pass the feature map through the convolutional feature extraction layer. Every time it passes through a convolutional feature extraction layer, a feature splitting operation is used to divide it into two feature maps. Combine the feature map obtained by passing one of the feature maps through n RepConvBlock modules with the other feature map split by the feature splitting operation using a feature splicing operation, and finally integrate the features through a convolutional fusion layer to obtain the final output feature map.

3. A mechanical part defect detection method based on an improved YOLOv12 model according to claim 2, characterized in that, The implementation process of the RepConvBlock module includes: Extract local features from the input feature map through the first 3×3 depthwise separable convolution to obtain a local feature map; Process the local feature map through a multi-branch projection module. The projection module sequentially includes a normalization layer, a reparameterizable dilated convolution, a batch normalization layer, a channel attention module, a first 1×1 convolutional layer, a GELU activation function, a second 3×3 depthwise separable convolution, global response normalization, and a second 1×1 convolutional layer; If residual scaling is enabled, scale the input feature map and add it to the output of the projection module to obtain the output feature; otherwise, use residual connection and directly use the output of the projection module as the output feature.

4. A mechanical part defect detection method based on an improved YOLOv12 model according to claim 1, characterized in that, In the step of replacing the original Concat module with the gated dynamic spatial aggregator GDSAFusion in the neck network, the implementation process of the gated dynamic spatial aggregator GDSAFusion includes: Concatenate the input feature map and the context feature in the channel dimension to form a fused feature; Extract local features from the fused feature through 3×3 depthwise separable convolution, and then stabilize the feature distribution through a normalization layer; Calculate the spatial attention weights through a query-key mechanism, and the query-key mechanism uses relative position biases; Capture multi-scale context information from the fused feature through reparameterizable dilated convolution, and then perform channel-wise adaptive calibration through a channel attention module; Control the information flow through a gating mechanism; Perform weighted fusion of the feature processed by the gating mechanism and the feature of the residual path; Process the weighted fused feature with two-level scaling and stochastic depth dropout to obtain an enhanced feature map.

5. A mechanical part defect detection method based on an improved YOLOv12 model according to claim 1, characterized in that, In the step of using the A2C2f_STR module to replace the original A2C2f module in the backbone network and the neck network, the implementation process of the A2C2f_STR module includes: Connect the input feature tensor to an initial convolutional layer, and the initial convolutional layer compresses the channel dimension of the input feature tensor through low-rank mapping and extracts feature information with basic representational capabilities to obtain a feature map after initial convolution processing; Temporarily store the feature map after initial convolution processing in a list data structure, and enter a loop processing link composed of visual modeling modules for fused cross-index temporal interaction. In the loop processing link, use the tail element of the list data structure as the input and sequentially pass it into each visual modeling module for fused cross-index temporal interaction for processing to obtain a feature vector generated after being processed by the visual modeling module for fused cross-index temporal interaction; Fuse the feature vector generated after being processed by the visual modeling module for fused cross-index temporal interaction and the original input feature tensor through a residual mapping mechanism to obtain the finally output fused feature.

6. A method for detecting mechanical part defects based on an improved YOLOv12 model according to claim 5, characterized in that, The implementation process of the visual modeling module for fused cross-index temporal interaction includes: Perform the first normalization processing on the feature map after initial convolution processing to obtain a first normalized feature; Realize the adaptive weight fusion of the first normalized feature and the feature map after initial convolution processing through learnable parameters to obtain an adaptive weight fusion feature; Input the adaptive weight fusion feature into a cross-index temporal interaction modeling component, and the cross-index temporal interaction modeling component uses scanning indexes and inverse scanning indexes to perform temporal feature recombination and modeling across spatial dimensions to obtain a feature processed by the cross-index temporal interaction modeling component; Perform the second normalization processing on the feature processed by the cross-index temporal interaction modeling component to obtain a second normalized feature; Input the second normalized feature into a feed-forward convolutional block to complete non-linear transformation to obtain a feed-forward output; Fuse the feature processed by the cross-index temporal interaction modeling component and the feed-forward output through a residual connection to obtain a feature vector generated after being processed by the visual modeling module for fused cross-index temporal interaction.

7. A mechanical part defect detection method based on an improved YOLOv12 model according to claim 6, characterized in that, The implementation process of the cross-index temporal interaction modeling component includes: Generate index pairs for the adaptive weight fusion feature input to the cross-index temporal interaction modeling component through an index generation module, and the index pairs include scanning indexes and inverse scanning indexes; Feature rearrangement is performed on the adaptive weight fusion features input to the cross-index temporal interaction modeling component by using the said scanning index to obtain rearranged one-dimensional temporal features; The rearranged one-dimensional temporal features are mapped into dynamic parameters through a temporal parameter projection layer; Temporal modeling is performed on the dynamic parameters by a selective scanning engine based on a state space model, and one-dimensional sequence features after temporal modeling are output; The one-dimensional sequence features after temporal modeling are restored to two-dimensional spatial features through the inverse scanning index; Gating weights are generated through an adaptive gating unit, and the gating weights are multiplied element-wise with the two-dimensional spatial features to output the finally processed features by the cross-index temporal interaction modeling component.

8. A mechanical part defect detection system based on an improved YOLOv12 model, characterized in that, It includes: An image acquisition module, configured to collect a defective mechanical part image dataset and perform preprocessing on the mechanical part image dataset, where the preprocessing includes annotating the defective parts of each mechanical part image, as well as performing image data augmentation and enhancement; A model training module, configured to train an improved YOLOv12 mechanical part defect detection model using the preprocessed mechanical part image dataset to obtain an optimized model, and the improved YOLOv12 mechanical part defect detection model includes: Using a C3k2_RCB feature extraction module to replace the original C3k2 module in the backbone network and the neck network; Using a gating dynamic space aggregator GDSAFusion to replace the original Concat module in the neck network; Using an A2C2f_STR module to replace the original A2C2f module in the backbone network and the neck network; A defect detection module, configured to perform real-time defect detection on mechanical parts using the optimized model.

9. An electronic device, characterized in that, It includes a processor, a memory, and a program or instruction stored on the memory and executable on the processor. When the program or instruction is executed by the processor, the steps of a mechanical part defect detection method based on an improved YOLOv12 model as described in any one of claims 1-7 are implemented.

10. A readable storage medium, characterized in that, A program or instruction is stored on the readable storage medium. When the program or instruction is executed by a processor, the steps of a mechanical part defect detection method based on an improved YOLOv12 model as described in any one of claims 1-7 are implemented.

Citation Information

Patent Citations

  • Industrial part defect detection method based on improved YOLOv8 model

    CN118967602A

  • Obstacle detection method of visual impaired group guiding waistcoat system based on improved YOLOv11

    CN119600523A

  • PPY-YOLO-based steel surface defect detection method and system

    CN119672031A

  • Industrial pure iron metallographic specimen defect detection optimization method based on YOLOv8obb

    CN119831964A

  • Steel surface defect detection method and system

    CN119941724A

Cited By

  • Aerial photography target detection method and device, computer equipment and storage medium

    CN120612493A

  • Wind power generation blade surface defect detection method and system based on TFPN-YOLO model

    CN121544619A

  • Multi-category electronic component surface defect detection method based on Mamba framework

    CN121660981A

  • Transformer equipment oil leakage target detection method and related device

    CN121787479A