Metallographic defect intelligent detection system and method based on improved YOLOv11

By improving the YOLOv11 model, introducing a rotating target detection head and a feature extraction module, and combining it with an attention mechanism, the problems of insufficient accuracy and real-time performance in metallographic defect detection were solved, and efficient and automated metal material defect detection was achieved.

CN120931651BActive Publication Date: 2026-02-10ZHEJIANG UNIV OF SCI & TECH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511461261.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-14
Publication Date
2026-02-10
Estimated Expiration
2045-10-14

AI Technical Summary

Technical Problem

Existing metallographic defect detection technologies are insufficient in terms of accuracy, real-time performance, and adaptability. They are difficult to effectively identify minute defects and adapt to the metallographic structure of different metal materials. Furthermore, the traditional YOLO model has a large number of parameters, making it difficult to deploy on industrial edge devices.

Method used

An improved YOLOv11 metallographic defect intelligent detection system was adopted. By introducing a rotating target inspection head (OBB), combined with the CReToNeXt feature extraction module and the SE-Net/CBAM attention mechanism, data augmentation and model training were performed to achieve accurate localization of tilted defects and feature extraction of small targets.

Benefits of technology

It significantly improves the detection accuracy and real-time performance of small targets such as microcracks, enhances the model's anti-interference ability and generalization ability, meets the needs of industrial quality inspection, and realizes efficient and automated detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120931651B_ABST
    Figure CN120931651B_ABST
Patent Text Reader

Abstract

The present application relates to a kind of metallographic defect intelligent detection system and method based on improved YOLOv11, by integrating metallographic image acquisition subsystem, rotating frame labeling and enhancement subsystem, improved neural network processing subsystem, realize the full process automation from metal material micrograph image acquisition to defect accurate identification;By introducing the rotating target detection mechanism adaptation inclined defect morphology, design progressive data enhancement strategy to strengthen small target feature learning, using replacement feature extraction module and embedding attention mechanism optimize YOLOv11 network structure, finally, while guaranteeing industrial real-time, improve metallographic defect detection precision, meet the high-precision quality inspection demand.The present application not only solves the key technical bottlenecks such as the difficulty of positioning in the existing metallographic defect detection, the difficulty of detecting small targets, model deployment, weak anti-interference ability, etc., but also realizes the efficient and automated metallographic defect detection of industrial grade, has wide application prospect and popularization value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of metal material microstructure quality inspection technology, specifically relating to an intelligent metallographic defect detection system and method based on an improved YOLOv11. Background Technology

[0002] The microstructure of metallic materials is a core factor determining their mechanical properties and reliability. Metallographic defects (such as cracks, inclusions, and porosity), as microscopic anomalies within the material, directly weaken its strength, corrosion resistance, and service life, posing a significant risk of failure in industrial components. Therefore, metallographic defect detection is a crucial aspect of metallic material quality control and is of great importance to ensuring industrial production safety. Traditional detection methods rely primarily on manual visual inspection, with technicians interpreting images using metallographic microscopes. While these methods can identify obvious defects, they suffer from low efficiency (only 5-10 images can be processed per hour), high subjectivity (false positive rate as high as 15%-20%), and insufficient accuracy (over 30% of microcracks smaller than 5μm are missed).

[0003] With the development of industrial automation technology, machine vision-based inspection methods are gradually being applied to metallographic defect detection. Automated analysis is achieved through image acquisition equipment and algorithms, improving efficiency to some extent. However, existing machine vision technology still faces significant challenges:

[0004] (1) Metallographic defects vary greatly in size (5-100μm), and the feature information of small defects (such as microcracks) is weak. Traditional image processing algorithms (such as edge detection and threshold segmentation) are not capable of extracting them, and the false negative rate exceeds 40%.

[0005] (2) Metallographic images have complex backgrounds (including grain boundaries, corrosion textures, etc.), and the contrast between defects and background is low. Traditional algorithms have difficulty accurately segmenting targets and are easily affected by texture interference, leading to misjudgment.

[0006] (3) Industrial production lines have strict requirements for detection speed. Existing machine vision methods are difficult to balance between accuracy and real-time performance, and are difficult to meet the needs of production lines.

[0007] Currently, deep learning technology has made groundbreaking progress in the field of computer vision. Object detection algorithms based on convolutional neural networks (CNNs) are widely used in industrial inspection due to their end-to-end detection capabilities, with the YOLO series algorithms becoming a mainstream tool due to their high efficiency and real-time performance. However, the standard YOLO model still has significant shortcomings in metallographic defect detection:

[0008] (1) It has limited ability to extract features of small-scale targets (such as 5-10 μm microcracks). Due to the low pixel ratio of small defects, it is easily affected by grain boundary textures, resulting in low detection accuracy.

[0009] (2) The model has insufficient generalization ability, poor adaptability to the metallographic structure of different metal materials (such as aluminum alloys and titanium alloys), and uneven detection effect of different defect types (cracks, inclusions);

[0010] (3) The feature fusion and multi-scale detection capabilities are limited. The traditional YOLO feature pyramid structure is difficult to take into account both the details of small defects and the global features of macro defects. It has low positioning accuracy for densely distributed tilted defects (such as diagonal cracks) and the intersection-over-union ratio (IoU) is less than 0.5.

[0011] (4) The model is not lightweight enough. Existing YOLO models (such as the original version of YOLOv11) have a large number of parameters, slow inference speed, and are difficult to deploy on industrial edge devices.

[0012] Therefore, given the characteristics of metallographic defects such as "microscopic scale, tilted shape, and complex background", there is an urgent need for a detection method that can accurately locate tilted defects, enhance the extraction of small target features, and also has real-time performance, in order to solve the above-mentioned problems of existing technologies. Summary of the Invention

[0013] To address the problems of redundant tilt defect localization (horizontal bounding boxes containing too much grain boundary background), low small target detection rate (indistinct microcrack features), weak generalization ability due to grain boundary texture interference, and large model parameter count making deployment difficult in metallographic defect detection (such as microcracks and tilted inclusions) of existing YOLO models, this invention proposes an intelligent metallographic defect detection system and method based on an improved YOLOv11. This system integrates a metallographic image acquisition subsystem (adapted to microscopic imaging), a rotating bounding box annotation and enhancement subsystem (for tilted defects), and an improved neural network processing subsystem (resistant to grain boundary interference), achieving full automation from microscopic image acquisition of metallic materials to accurate defect identification. This invention introduces a rotating target detection (OBB) mechanism to adapt to tilted defect morphology, designs a progressive data augmentation strategy to enhance small target feature learning, and uses a CReToNeXt feature extraction module and SE-Net / CBAM attention mechanism to optimize the network structure. Ultimately, while ensuring industrial real-time performance, it improves the accuracy of metallographic defect detection, meeting the requirements of high-precision quality inspection.

[0014] A metallographic defect intelligent detection system based on an improved YOLOv11 mainly includes three major processes: image acquisition, data annotation and enhancement, and neural network processing. The processes are coordinated through data interfaces to achieve a closed loop from metal sample to defect detection report. The system includes a rotating frame annotation and enhancement subsystem, a neural network processing subsystem, and a detection result output and visualization module.

[0015] Metallographic image acquisition subsystem: used for image acquisition of microscopic defects in different metallic materials;

[0016] Rotating bounding box annotation and enhancement subsystem: Based on a professional annotation workstation, it uses a rotating target bounding box annotation tool to annotate data, supporting accurate annotation of defects in multiple directions and at arbitrary angles; after standardization, the annotated data is input into the data enhancement module to perform geometric transformations, color perturbations, and progressive Mosaic stitching, which greatly enriches the diversity and representativeness of the training samples;

[0017] Neural network processing subsystem: By introducing a rotating target detection head (OBB), a replacement feature extraction module, and an embedded replacement attention mechanism, an improved YOLOv11 network structure is formed for model training and inference, enabling automatic defect localization and classification; this subsystem supports model training and inference on high-performance servers or edge computing terminals to achieve automatic defect localization and classification.

[0018] The detection result output and visualization module performs post-processing on the detection results, outputs the rotation box coordinates, category, and confidence level of the defects, and generates a visualization report, which is convenient for manual review and integration into the industrial quality inspection process.

[0019] Preferably, the replaced feature modules include the CReToNeXt feature extraction module, the GhostNet feature extraction module, and the MobileNet feature extraction module; the embedded attention mechanisms include the SE-Net / CBAM attention mechanism, the ECA attention mechanism, and the BAM attention mechanism.

[0020] Preferably, the metallographic image acquisition subsystem consists of a metallographic microscope, an industrial camera, and an image acquisition card; the geometric transformations include rotation, scaling, and flipping.

[0021] As a preferred option, the rotating target detection head outputs the coordinates of the four vertices of the rotating bounding box, expands the number of output channels to 8, and uses rotation IoU as the loss function.

[0022] CReToNeXt feature extraction module: adopts a multi-branch RepConv structure, integrates multi-scale convolutional kernels and batch normalization; CReToNeXt feature extraction module is deployed in both the backbone network and the neck network;

[0023] The embedded attention mechanism is either the SE-Net attention mechanism or the CBAM attention mechanism (SE-Net / CBAM attention mechanism), which is embedded in the front end of the SPPF layer.

[0024] Metallographic Image Acquisition Subsystem

[0025] Hardware components: high-magnification metallurgical microscope, industrial camera, image acquisition card;

[0026] Function: Automatically acquires microscopic images of metallic materials and outputs high-resolution images that retain defects and grain boundary details;

[0027] Composed of a high-magnification metallurgical microscope (500×~1000×), a high-resolution industrial camera, and an image acquisition card, it can achieve high-fidelity acquisition of microscopic defects in different metal materials (such as aluminum alloys, titanium alloys, etc.). During the acquisition process, automatic exposure and light source intensity adjustment suppress reflections and noise, ensuring clear separation of defects and grain boundary details.

[0028] Rotating box annotation and enhancement subsystem

[0029] Hardware components: Labeling workstation;

[0030] Software modules: X-AnyLabeling annotation tool, data augmentation engine;

[0031] Function: Completes defect rotation box annotation and sample expansion, and outputs a standardized dataset adapted to microscopic tilt defects.

[0032] Neural network processing subsystem

[0033] Hardware components: AI training server, edge computing terminal;

[0034] Software modules: Improved YOLOv11 model (including rotating object detection head (OBB Head), CReToNeXt feature extraction module, SE-Net / CBAM attention mechanism), model deployment tools;

[0035] Function: Enables feature extraction, localization, and classification of metallographic defects, supports real-time reasoning, and outputs test results that meet the needs of industrial quality inspection.

[0036] This invention also provides a metallographic defect intelligent detection method based on the improved YOLOv11, which uses the above system and includes the following steps:

[0037] Step S1: Construct a metallographic defect image acquisition subsystem to acquire images of the microscopic characteristics of metallic materials;

[0038] Step S2: Complete defect annotation, data cleaning and enhancement through the rotating box annotation and enhancement subsystem to generate a standardized dataset;

[0039] Step S3: Improve the YOLOv11 architecture in the neural network processing subsystem by introducing a rotating target detection head, replacing the feature extraction module, and embedding a replacement attention mechanism to adapt to metallographic defect characteristics;

[0040] Step S4: Train the model using a phased training strategy and dynamically optimize the parameters using the validation set;

[0041] Step S5: Evaluate the model performance and output a visualization report of the required alloy phase detection.

[0042] As a preferred option, step S1 specifically involves:

[0043] Step S11: Set up an image acquisition subsystem. The image acquisition subsystem consists of a high-magnification metallurgical microscope with a magnification of 500× to 1000×, an industrial camera with a resolution of 2048×2048 and a frame rate of 30fps that supports high-resolution imaging, and an image acquisition card. During acquisition, the light source intensity is controlled to avoid reflections interfering with the details of the metallographic structure.

[0044] Step S12: For different metallic materials (such as aluminum alloys and titanium alloys), acquire metallographic images containing typical defects, including microcracks distributed along grain boundaries, irregularly shaped inclusions, and circular pores, and save them in JPEG format, preserving the grayscale difference between defects and grain boundaries.

[0045] As a preferred option, step S2 specifically involves:

[0046] Step S21: Configure the annotation workstation, use the general rotating box annotation tool to annotate defects, and use shortcut keys to assist in annotating defects in different directions;

[0047] Step S22: Define annotation rules:

[0048] ① The defect category is uniformly marked as 0.defect;

[0049] ② The rotation box corner points are arranged in the following order: top right, bottom right, bottom left, top left, and the coordinates are standardized according to the image size.

[0050] Step S23: Data annotation: All data annotation is completed by the annotator, supporting defect annotation in multiple directions and forms, and generating JSON format annotation files;

[0051] Step S24: The data augmentation module is started, which specifically includes:

[0052] ① Geometric enhancement: Randomly rotate, horizontally flip, scale, and translate the original image; disable vertical flipping and limit the intensity of perspective transformation;

[0053] ②Color enhancement: HSV channel perturbation is used to simulate complex lighting conditions in industrial environments;

[0054] ③ Advanced Mixup Enhancement: Includes 10% probability Mixup, random erase, and RandAugment automatic enhancement strategies;

[0055] ④ Progressive Mosaic stitching: Enabled in the first 20 epochs of training and gradually disabled in the last 20 epochs, expanding the sample size to 6 times the original size;

[0056] Step S25: Divide the training set and the test set, and generate a YOLOv11-OBB adapted annotation file.

[0057] As a preferred option, the overall network structure adopts YOLOv11, and step S3 is as follows:

[0058] Step S31: Integrate the rotating target detection head, expand the horizontal box output to the eight-point coordinate output of the rotated box, and introduce the rotation IoU index into the loss function;

[0059] Step S32: Optimize the feature extraction module by replacing the C3k2 module in the original YOLOv11 with the CReToNeXt feature extraction module;

[0060] Step S33: Attention mechanism embedding. Embed the SE-Net attention mechanism or CBAM attention mechanism at the front end of the SPPF layer of the network.

[0061] Step S34: Verify effectiveness through multiple rounds of ablation experiments.

[0062] As a preferred option, step S4 specifically involves:

[0063] Step S41: Hardware environment configuration: CPU, GPU, SSD; software environment: Python + PyTorch + CUDA.

[0064] Step S42: Training parameter settings: epochs=300, imgsz=1024, batch=16, using AdamW optimizer;

[0065] Step S43: Phased training strategy: Mosaic enhancement is enabled for the first 20 epochs to strengthen the learning of defect features at multiple scales and angles; Mosaic enhancement is gradually turned off for the last 20 epochs to focus on fitting the defect distribution of the original metallographic image.

[0066] Step S44: Dynamic optimization of the validation set: Calculate mAP50-95 on the validation set every 5 epochs, focusing on microcrack recall and tilted inclusion IoU. When there is no improvement for 10 consecutive epochs, trigger the early stopping mechanism and save the optimal model.

[0067] As a preferred option, step S5 specifically involves:

[0068] Step S51: Test set evaluation: Calculate precision, recall, mAP50-95, skew defect IoU, and inference speed on the test set, and compare and verify the improvement effect;

[0069] Step S52: Output of detection results: The output includes information such as defect location, defect type, and confidence level. It supports the visualization of detection results and supports the deployment of the model on edge computing terminals.

[0070] The innovative principle of this invention's rotating target detection head (OBB Head):

[0071] Traditional YOLO series uses horizontal bounding box (HBB) output, which easily leads to inaccurate positioning of tilted defects and large background redundancy; this invention innovatively introduces a rotating target detection head (OBB Head), which directly outputs the coordinates of the four vertices of the rotating bounding box, and with the rotation IoU loss function, the model can accurately fit the actual defect morphology, especially suitable for tilted cracks and irregular inclusions distributed along grain boundaries;

[0072] Structurally, the number of output channels of the detection head has been expanded from the original 5 (x, y, w, h, θ) to 8 (x1, y1, x2, y2, x3, y3, x4, y4). The loss function adopts rotation IoU metric to optimize the overlap between the detection box and the real defect contour.

[0073] Innovative principle of CReToNeXt feature extraction module:

[0074] To address the problem that small targets such as microcracks in metallographic images are difficult to capture by traditional convolution, this invention adopts a multi-branch RepConv structure, which integrates multi-scale convolution kernels and batch normalization to improve the feature representation ability of micro-defects; the inference stage is simplified to reduce the number of model parameters, balancing efficiency and detection accuracy.

[0075] This module is deployed in both the backbone and neck networks to enhance the network's ability to perceive defects of different scales and shapes.

[0076] The innovative principle of attention mechanism:

[0077] The complex grain boundary textures in metallographic images can easily interfere with defect detection. Therefore, this invention introduces the SE-Net / CBAM attention mechanism to finely filter the high-order features extracted by the backbone network.

[0078] The SE-Net attention mechanism module models the dependencies between channels of the feature map to achieve adaptive weighting of the channel dimension, strengthen the representation of features related to defects, and suppress the response of irrelevant channels.

[0079] The CBAM attention mechanism combines channel attention and spatial attention. It first weights the channel dimension of the feature map and then weights the spatial dimension to further enhance the model's ability to focus on defective regions.

[0080] This module is embedded in the front end of the SPPF layer to improve the model's resistance to interference and its generalization ability.

[0081] Progressive Mosaic data augmentation process innovation:

[0082] In the early stages of training (the first 20 epochs), Mosaic enhancement is enabled, which randomly stitches together 4 images to greatly expand the diversity of data and enhance the model's ability to learn small targets and defects from multiple angles.

[0083] As the Mosaic is gradually turned off in the later stages of training (the last 20 epochs), the model is allowed to better fit the distribution of real metallographic images, thus improving the final inference performance.

[0084] Highly efficient collaboration is achieved among the various innovation modules:

[0085] Rotated bounding box annotation and augmentation provide high-quality, realistically distributed training samples for neural networks;

[0086] The rotating target inspection head (OBB Head) and the CReToNeXt feature module work together to improve the ability to locate and identify tilted and minute defects;

[0087] The SE-Net / CBAM attention mechanism further optimizes feature representation and enhances model robustness;

[0088] The entire process is automated, supporting closed-loop deployment from data collection, annotation, training to on-site inference.

[0089] The system can ultimately output the detection results such as the location, type, and confidence level of the defect's rotating frame, and supports visualization and manual review, meeting the needs of industrial quality inspection.

[0090] The beneficial effects of this invention are as follows:

[0091] (1) Improved accuracy in locating tilted defects;

[0092] By introducing the Rotating Target Detection (OBB) mechanism, the positioning accuracy of defects such as microcracks and tilted inclusions distributed along grain boundaries is significantly improved, solving the problem of positioning redundancy caused by excessive background in the traditional horizontal frame method, and the intersection-over-union (IoU) ratio is significantly improved.

[0093] (2) Enhanced detection capability for small targets;

[0094] Progressive Mosaic data augmentation combined with the CReToNeXt feature extraction module enhances the feature learning capability for small targets such as microcracks, improving the model's mean precision (mAP) and small target recall rate.

[0095] (3) Balancing model lightweighting with real-time performance;

[0096] The CReToNeXt feature extraction module adopts a multi-branch RepConv structure, which reduces the computational complexity of the model, improves the inference speed, and meets the needs of edge device deployment and real-time detection in industrial production lines.

[0097] (4) Optimization of anti-interference and generalization capabilities;

[0098] The SE-Net / CBAM attention mechanism effectively suppresses interference from complex backgrounds such as grain boundary textures, improving the model's adaptability and generalization ability to different metal materials (with large variations in metallographic structure). Among them, the SE-Net attention mechanism enhances the expression of useful features through channel attention, while the CBAM attention mechanism further improves the model's ability to focus on and discriminate defect regions through dual channel and spatial attention mechanisms.

[0099] (5) Full-process automation and high consistency

[0100] The system automates the entire process of metallographic image acquisition, defect annotation, data augmentation, model training, and detection result output, greatly improving detection efficiency and result consistency, and reducing human subjective error.

[0101] (6) Visualization and review support

[0102] The test results support visual display and manual verification, facilitating rapid decision-making and quality traceability in industrial settings;

[0103] (7) This invention not only solves the key technical bottlenecks in existing metallographic defect detection, such as difficulty in locating tilted defects, difficulty in detecting small targets, difficulty in model deployment, and weak anti-interference ability, but also realizes industrial-grade efficient and automated metallographic defect detection, which has broad application prospects and promotion value. Attached Figure Description

[0104] Figure 1 This is a flowchart illustrating the overall process of the metallographic defect detection method of the present invention.

[0105] Figure 2 Annotation diagram for metallographic defect detection;

[0106] Figure 3 A formatted diagram for the rotated target bounding box;

[0107] Figure 4 This is a schematic diagram of the original YOLOv11 network structure;

[0108] Figure 5 This is a diagram of the improved YOLOv11 network structure of this invention;

[0109] Figure 6 Here is a structural diagram of the CReToNeXt feature extraction module;

[0110] Figure 7 This is a structural diagram of the basic unit of ConvBNAct;

[0111] Figure 8 Here is a diagram of the Convs module structure;

[0112] Figure 9 Here is a structural diagram of the RCBS module;

[0113] Figure 10 This is a structural diagram of the RepConv multi-branch structure module. Detailed Implementation

[0114] To better understand the above technical solutions, the following will provide a detailed explanation of the technical solutions in conjunction with the accompanying drawings and specific embodiments.

[0115] Reference Figure 1 A metallographic defect intelligent detection method based on improved YOLOv11 includes the following steps:

[0116] Step S1: Construct a metallographic defect image acquisition subsystem to acquire images of the microscopic characteristics of metallic materials;

[0117] Step S2: Complete defect annotation (including fine processing of micro-defects), data cleaning and enhancement through the rotating box annotation and enhancement subsystem to generate a standardized dataset;

[0118] Step S3: Improve the YOLOv11 architecture in the neural network processing subsystem (introduce the OBB mechanism, replace the feature extraction module, and embed the attention mechanism) to adapt to the characteristics of metallographic defects;

[0119] Step S4: Train the model using a phased training strategy and dynamically optimize the parameters using the validation set;

[0120] Step S5: Evaluate the model performance and output a visualization report of the required alloy phase detection.

[0121] Specifically, step S1: Metallographic image acquisition

[0122] Step S11: Build an image acquisition subsystem, consisting of a high-magnification metallurgical microscope (500× magnification, suitable for observing microscopic defects), an industrial camera that supports high-resolution imaging (2048×2048 resolution, 30fps frame rate), and an image acquisition card. Control the light source intensity during acquisition to avoid reflections interfering with the details of the metallographic structure.

[0123] Step S12: Using aluminum alloy as the inspection object, collect a total of 120 metallographic images containing typical defects, including microcracks, inclusions and pores, and store them in 24-bit JPEG format (preserving the grayscale difference between defects and grain boundaries).

[0124] Step S2: Rotate frame annotation and enhancement subsystem operation

[0125] Step S21: Configure the annotation workstation, and annotate metallographic defects as follows. Figure 2 As shown, the general rotating box annotation tool (X-AnyLabeling) is used for defect annotation, and keyboard shortcuts are used to assist in annotating defects in different directions;

[0126] Step S22: Definition of annotation rules: ① The defect category is uniformly marked as "0.defect"; ② The rotation box corner point order is top right → bottom right → bottom left → top left, and the coordinates are standardized by image size to ensure a close fit to the microscopic defect contour. The rotation target box annotation format is as follows: Figure 3 As shown;

[0127] Step S23: Data annotation and organization: All data annotation is completed by the annotator, supporting defect annotation in multiple directions and forms, which significantly improves the efficiency and accuracy of dataset construction, generates JSON format annotation files, and is subsequently converted into txt format files adapted to YOLOv11-OBB;

[0128] Step S24: The data augmentation module is started, which specifically includes:

[0129] ① Geometric Enhancement: The original image is randomly rotated (angle range ±45°, degrees=45.0), horizontally flipped (probability 50%, fliplr=0.5), scaled (scale=0.5), and translated (translate=0.1) to enhance the model's robustness to defects at multiple angles and different locations; to avoid geometric distortion of the rotated box, vertical flipping is disabled and the intensity of perspective transformation is limited (perspective=0.0005).

[0130] ② Color enhancement: HSV channel perturbation (hsv_h=0.1, hsv_s=0.7, hsv_v=0.4) is used to simulate complex lighting conditions in industrial environments;

[0131] ③ Advanced Mixup Enhancement: Includes 10% probability Mixup (mixup=0.1), random erasing (erasing=0.4) and RandAugment automatic enhancement strategies to improve the model's robustness to occlusion, noise and other interference.

[0132] ④ Progressive Mosaic stitching: Mosaic data augmentation is enabled in the first 20 epochs of training and gradually disabled in the last 20 epochs, expanding the sample size to 6 times the original size;

[0133] Step S25: Divide the data into training and testing sets, and generate annotation files adapted for YOLOv11-OBB.

[0134] Step S3: Model Improvement of the Neural Network Processing Subsystem

[0135] The network architecture generally adopts YOLOv11, and the original structure is as follows: Figure 4 As shown, the improved structure is as follows Figure 5 As shown;

[0136] Step S31: Integrate the rotating target detection head (OBB Head), which expands the traditional horizontal frame output to a rotating frame eight-point coordinate output (x1,y1,x2,y2,x3,y3,x4,y4), and introduces the rotation IoU index into the loss function to adapt to the tilt shape and directional characteristics of metallographic defects, significantly improving the positioning accuracy of complex defects.

[0137] Step S32: Feature extraction module optimization. The C3k2 module in the original YOLOv11 is replaced with the CReToNeXt feature extraction module. The number of channels in the CReToNeXt module of the backbone network is set to 64, 128, and 256 respectively. The structure of the CReToNeXt feature extraction module is as follows: Figure 6 As shown. The core structural blocks of this module include Conv1, Conv2, Convs, and Conv3; Conv1, Conv2, and Conv3 are composed of Conv2d, BatchNorm2d, and Swish. The module is named ConvBNAct, and the basic unit structure of ConvBNAct is as follows. Figure 7 As shown. Convs consists of 3 RCBS, and the Convs module structure is as follows. Figure 8 As shown, RCBS consists of RepConv and a ConvBNAct. The RCBS module structure is as follows: Figure 9 As shown. The structural details of RepConv are as follows. Figure 10 As shown, during the training of the YOLOv11-OBB network, when the number of input and output channels is the same, RepConv consists of a 3×3 convolution with a BN layer, a 1×1 convolution with a BN layer, and a single BN layer convolution; it consists of only a 3×3 convolutional layer and a BN layer. This reduces one Conv2d module and two BatchNorm2d modules, thereby reducing computational complexity and improving the detection speed of the maintenance component.

[0138] Step S33: Embedding the attention mechanism

[0139] An SE-Net attention mechanism module is embedded at the front end of the network's SPPF layer. The goal of the SE-Net attention mechanism is to improve the network's representational power by explicitly modeling the interdependencies between channels of the network's convolutional features. After adding this attention mechanism, the network can perform feature recalibration. In this way, it can learn to use global information to selectively emphasize useful features and suppress less useful features.

[0140] Alternatively, a CBAM attention mechanism can be embedded at the front end of the SPPF layer of the network. The CBAM attention mechanism enhances the ability to represent key features through adaptive features and weighting. This module can improve the model's feature discrimination ability in complex backgrounds with only a small increase in the number of parameters. This computational efficiency makes it particularly suitable for industrial embedded devices and real-time detection scenarios, providing a cost-effective performance optimization solution for deep neural networks.

[0141] Step S34: Verify the effectiveness of the above module through multiple rounds of ablation experiments.

[0142] Step S4: Model Training

[0143] Step S41: Hardware environment configuration: CPU, GPU (such as NVIDIA series), SSD, software environment is Python + PyTorch + CUDA;

[0144] Step S42: Training parameter settings: epochs=300, imgsz=1024, batch=16, using AdamW optimizer, learning rate 0.001;

[0145] Step S43: Phased training strategy: Mosaic enhancement is enabled for the first 20 epochs to strengthen the learning of defect features at multiple scales and angles; Mosaic enhancement is gradually turned off for the last 20 epochs to focus on fitting the defect distribution of the original metallographic image.

[0146] Step S44: Dynamic optimization of the validation set: Calculate mAP50-95 on the validation set every 5 epochs, focusing on the recall rate of microcracks (small targets) and the skewed IoU. When there is no improvement for 10 consecutive epochs, trigger the early stopping mechanism and save the optimal model.

[0147] Step S5: Model Evaluation and Testing

[0148] Step S51: Test set evaluation: Calculate precision, recall, mAP50-95, skew defect IoU, and inference speed on the test set, and compare them with the improved YOLOv11 to verify the improvement effect;

[0149] Step S52: Output of detection results: The system can output information such as defect location (rotated box coordinates), defect type, and confidence level. It supports the visualization of detection results, which is convenient for manual review. It also supports the deployment of the model on the edge computing terminal to meet the real-time detection needs of industrial production lines.

[0150] Ablation experiments compared the impact of different data augmentation strategies on the performance of the YOLOv11-OBB model. The results of the data augmentation ablation experiments are shown in Table 1. The experimental environment was as follows: CPU Intel(R) Xeon(R) Platinum 8352V @2.10GHz, GPU NVIDIA GeForce RTX 4090 (24GB VRAM), based on Python 3.12, PyTorch 2.4.0 framework and CUDA 12.1 computing architecture. The specific settings are as follows:

[0151] YOLOv11+OBB: Basic model (no enhancements);

[0152] +degrees+scale: Adds geometric enhancements with ±45° random rotation and 0.5x scaling;

[0153] +mosaic: Enable progressive Mosaic (enable four-image rotation frame sensing and stitching for the first 20 rounds, and gradually disable it for the next 20 rounds).

[0154] +mixup: Adds image blending enhancement with a 10% probability;

[0155] +mosaic+mixup: Composite enhancement group (simultaneously enables the above Mosaic and Mixup strategies).

[0156] All experiments used the same parameters: Epochs=300, Imgsz=1024, batch=16, close_mosaic=20, AdamW optimizer, with mean average accuracy (mAP50-95) and model efficiency (FPS) as evaluation metrics.

[0157] Table 1

[0158] ;

[0159] As shown in Table 1, after adopting composite enhancement strategies such as mosaic+mixup, the model's mAP50-95 increased from 0.67 to 0.80 (+19.4%), significantly improving the target detection performance.

[0160] Under the same evaluation standard (COCO mAP50-95), the optimized YOLOv11-OBB model achieved a score of 0.80, significantly higher than the Shanghai Jiao Tong University team's result using YOLOv8 on 414 metallographic images for metal additive manufacturing defect detection, which was mAP50-95=0.656. This difference stems from two main factors:

[0161] Differences in task complexity: The microscopic scale (≤50μm) and morphological randomness of metallographic defects (such as tilted cracks) place higher demands on detection algorithms;

[0162] Effectiveness of the technical strategy: The progressive Mosaic and Mixup composite enhancement strategy brings an absolute improvement of 19.4% in mAP50-95, which specifically solves the problems of weak small target features and strong background interference in metallographic images.

[0163] The results of this invention not only verify the uniqueness of the method in metallographic scenarios, but also set a new performance benchmark in the field of microscopic defect detection.

[0164] The ablation experiment results for the structural improvements are shown in Table 2. The experiment was based on a self-constructed metallographic defect dataset (containing 120 24-bit JPEG images, covering defects such as cracks and inclusions, with rotated bounding boxes labeled using the X-AnyLabeling tool). The aim was to verify the impact of the CReToNeXt module and the attention mechanism on the performance of the YOLOv11+OBB model. The experimental group settings are as follows:

[0165] YOLOv11+OBB: The basic control group uses the YOLOv11 model and introduces the Rotated Object Detection (OBB) mechanism. It solves the problem of inaccurate enclosure of tilted defects by polygon coordinate annotation, without making other structural improvements.

[0166] +CReToNeXt: In the base model, the C3k2 module of the original network is replaced with the CReToNeXt module; this module reduces redundant computation and improves feature extraction efficiency by optimizing the convolutional structure (containing 3 RCBS components, each RCBS consisting of RepConv and ConvBNAct);

[0167] +SE: The SE-Net attention mechanism is embedded in the base model, which enhances the feature selection capability of the channel dimension through the "squeeze-excitation" process (Squeeze operation aggregates global information, and Excitation operation strengthens useful channel features);

[0168] +CBAM: Convolutional Block Attention Module (CBAM) is embedded in the base model. It strengthens the attention to features in both channel and spatial dimensions through dual-path modeling of channel attention (focusing on key feature channels) and spatial attention (focusing on the location of defect regions).

[0169] +CReToNeXt+SE: In the base model, C3k2 is replaced with the CReToNeXt module and the SE attention mechanism is embedded to verify the synergistic effect of feature extraction efficiency optimization and channel attention enhancement.

[0170] +CReToNeXt+CBAM: In the base model, C3k2 is replaced with the CReToNeXt module and CBAM is embedded to verify the synergistic effect of feature extraction efficiency optimization and dual-path attention enhancement.

[0171] All experiments used the same parameters: Epochs=300, Imgsz=1024, batch=16, close_mosaic=20, AdamW optimizer, and compared the performance differences of different structural improvements using accuracy (P) and mean average precision (mAP) as evaluation metrics.

[0172] Table 2

[0173] ;

[0174] Experimental results show that the model's mean average accuracy (mAP) and precision are significantly improved, meeting the high precision and robustness requirements of industrial metallographic defect detection. Integrating the rotating target head (OBB Head), the CReToNeXt feature extraction module, and the CBAM attention mechanism demonstrates performance gains. Compared to the basic model, introducing any one module alone effectively improves detection performance; however, the synergistic effect of the two modules achieves a simultaneous leap in precision and mean average accuracy (mAP). The core indicator mAP increases from 0.802 to 0.845 or 0.842, fully validating the technical superiority of this improved scheme in metallographic defect identification tasks.

[0175] This invention is not limited to the specific structures and parameters described above. For example, the CReToNeXt feature extraction module can be replaced with other lightweight multi-branch structures such as GhostNet and MobileNet; the SE-Net / CBAM attention mechanism can be replaced with channel or spatial attention structures such as ECA and BAM; the rotating object detection head (OBB Head) can adopt arbitrary rotation box parameterization methods such as four-point coordinates, center point + width and height + angle. Data augmentation methods are not limited to Mosaic and Mixup, and other augmentation methods such as CutMix and GridMask can be used.

[0176] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A metallographic defect intelligent detection system based on an improved YOLOv11, characterized in that, The system includes a rotating bounding box annotation and enhancement subsystem, a neural network processing subsystem, and a detection result output and visualization module; Metallographic image acquisition subsystem: used for image acquisition of microscopic defects in different metal materials. The image acquisition subsystem consists of a high-magnification metallographic microscope with a magnification of 500× to 1000×, an industrial camera with a resolution of 2048×2048 and a frame rate of 30fps that supports high-resolution imaging, and an image acquisition card. Rotated bounding box annotation and enhancement subsystem: Based on a professional annotation workstation, a rotated bounding box annotation tool is used for data annotation; after standardization, the annotated data is input into the data augmentation module for geometric transformation, color perturbation, and progressive Mosaic stitching. Specifically, the progressive Mosaic stitching is enabled in the first 20 epochs of training and gradually disabled in the last 20 epochs, focusing on fitting the defect distribution of the original metallographic image and expanding the sample to 6 times the original size; Neural network processing subsystem: By introducing a rotating target detection head, replacing the feature extraction module, and embedding an attention mechanism, an improved YOLOv11 network structure is formed to perform model training and inference, thereby achieving automatic defect localization and classification; among which, the replaced feature extraction module is a CReToNeXt feature extraction module, a GhostNet feature extraction module, or a MobileNet feature extraction module; The detection result output and visualization module performs post-processing on the detection results, outputs the rotation box coordinates, category, and confidence level of the defects, and generates a visualization report.

2. The intelligent metallographic defect detection system based on the improved YOLOv11 as described in claim 1, characterized in that, The embedded attention mechanism is either SE-Ne attention mechanism, CBAM attention mechanism, ECA attention mechanism, or BAM attention mechanism.

3. The intelligent metallographic defect detection system based on the improved YOLOv11 as described in claim 1, characterized in that, The metallographic image acquisition subsystem consists of a metallographic microscope, an industrial camera, and an image acquisition card; the geometric transformations include rotation, scaling, and flipping.

4. The intelligent metallographic defect detection system based on the improved YOLOv11 according to claim 2, characterized in that, Rotating target detection head: Outputs the coordinates of the four vertices of the rotating bounding box, expands the number of output channels to 8, and uses rotation IoU as the loss function; CReToNeXt feature extraction module: adopts a multi-branch RepConv structure, integrates multi-scale convolutional kernels and batch normalization; CReToNeXt feature extraction module is deployed in both the backbone network and the neck network; The embedded attention mechanism is either the SE-Net attention mechanism or the CBAM attention mechanism, which is embedded in the front end of the SPPF layer.

5. A metallographic defect intelligent detection method based on an improved YOLOv11, characterized in that, The system according to any one of claims 1-4 includes the following steps: Step S1: Construct a metallographic defect image acquisition subsystem to acquire images of the microscopic characteristics of metallic materials; Step S2: Complete defect annotation, data cleaning and enhancement through the rotating box annotation and enhancement subsystem to generate a standardized dataset; Step S3: Improve the YOLOv11 architecture in the neural network processing subsystem by introducing a rotating target detection head, replacing the feature extraction module, and embedding an attention mechanism to adapt to the characteristics of metallographic defects; Step S4: Train the model using a phased training strategy and dynamically optimize the parameters using the validation set; Step S5: Evaluate the model performance and output a visualization report of the required alloy phase detection.

6. The intelligent metallographic defect detection method based on improved YOLOv11 according to claim 5, characterized in that, Step S1 is as follows: Step S11: Set up an image acquisition subsystem. The image acquisition subsystem consists of a high-magnification metallurgical microscope with a magnification of 500× to 1000×, an industrial camera with a resolution of 2048×2048 and a frame rate of 30fps that supports high-resolution imaging, and an image acquisition card. During acquisition, the light source intensity is controlled to avoid reflections interfering with the details of the metallographic structure. Step S12: For different metallic materials, acquire metallographic images containing typical defects, including microcracks distributed along grain boundaries, irregularly shaped inclusions, and circular pores, and save them in JPEG format, preserving the grayscale differences between defects and grain boundaries.

7. The intelligent metallographic defect detection method based on the improved YOLOv11 according to claim 5, characterized in that, Step S2 is as follows: Step S21: Configure the annotation workstation, use the general rotating box annotation tool to annotate defects, and use shortcut keys to assist in annotating defects in different directions; Step S22: Define annotation rules: ① The defect category is uniformly marked as 0.defect; ② The rotation box corner points are arranged in the following order: top right, bottom right, bottom left, top left. The coordinates are standardized according to the image size. Step S23: Data annotation: All data annotation is completed by the annotator, supporting defect annotation in multiple directions and forms, and generating JSON format annotation files; Step S24: The data augmentation module is started, which specifically includes: ① Geometric enhancement: Randomly rotate, horizontally flip, scale, and translate the original image; disable vertical flipping and limit the intensity of perspective transformation; ②Color enhancement: HSV channel perturbation is used to simulate complex lighting conditions in industrial environments; ③ Advanced Mixup Enhancement: Includes 10% probability Mixup, random erase, and RandAugment automatic enhancement strategies; ④ Progressive Mosaic stitching: Enabled in the first 20 epochs of training and gradually disabled in the last 20 epochs, expanding the sample size to 6 times the original size; Step S25: Divide the training set and the test set, and generate a YOLOv11-OBB adapted annotation file.

8. The intelligent metallographic defect detection method based on the improved YOLOv11 according to claim 5, characterized in that, The network structure generally adopts YOLOv11, and step S3 is as follows: Step S31: Integrate the rotating target detection head, expand the horizontal box output to the eight-point coordinate output of the rotated box, and introduce the rotation IoU index into the loss function; Step S32: Optimize the feature extraction module by replacing the C3k2 module in the original YOLOv11 with the CReToNeXt feature extraction module; Step S33: Attention mechanism embedding. Embed the SE-Net attention mechanism or CBAM attention mechanism at the front end of the SPPF layer of the network. Step S34: Verify effectiveness through multiple rounds of ablation experiments.

9. The intelligent metallographic defect detection method based on the improved YOLOv11 according to claim 5, characterized in that, Step S4 is as follows: Step S41: Hardware environment configuration: CPU, GPU, SSD; software environment: Python + PyTorch + CUDA. Step S42: Training parameter settings: epochs=300, imgsz=1024, batch=16, using AdamW optimizer; Step S43: Phased training strategy: Mosaic enhancement is enabled for the first 20 epochs to strengthen the learning of defect features at multiple scales and from multiple perspectives; In the last 20 epochs, Mosaic enhancement is gradually turned off to focus on fitting the defect distribution of the original metallographic image; Step S44: Dynamic optimization of the validation set: Calculate mAP50-95 on the validation set every 5 epochs, focusing on microcrack recall and tilted inclusion IoU. When there is no improvement for 10 consecutive epochs, trigger the early stopping mechanism and save the optimal model.

10. The intelligent metallographic defect detection method based on the improved YOLOv11 according to claim 5, characterized in that, Step S5 is as follows: Step S51: Test set evaluation: Calculate precision, recall, mAP50-95, skew defect IoU, and inference speed on the test set, and compare and verify the improvement effect; Step S52: Output of detection results: The output includes information such as defect location, defect type, and confidence level. It supports the visualization of detection results and supports the deployment of the model on edge computing terminals.

Citation Information

Patent Citations

  • Method for detecting direction bounding box of maintenance part in real time based on YOLOv5-OBB

    CN117593509A

  • Metal surface defect detection method based on improved YOLOv5

    CN119067919A

  • UMFNet-YOLO-based joint detection algorithm under low light condition

    CN120431448A