Track surface defect detection system made of YOLOv8-based steel
By integrating the deformable convolution structure and scale attention mechanism into the YOLOv8 detection system, combined with the multi-scale feature fusion unit and training enhancement module, the problems of target diversity and complex detection environment in rail steel manufacturing are solved, and efficient and accurate rail surface defect detection is achieved.
Patent Information
- Application Number
- CN202510811943.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-18
- Publication Date
- 2025-09-19
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing defect detection systems based on the YOLO series in the rail steel manufacturing field have problems such as strong target diversity, complex detection environment and limited deployment performance, making it difficult to achieve efficient and accurate rail surface defect detection.
A steel rail surface defect detection system based on YOLOv8 is developed. It integrates a deformable convolutional structure, a scale attention mechanism, and a multi-scale feature fusion unit. It is trained with a training enhancement module to enhance the adaptability to complex backgrounds and deformed targets. Efficient deployment is achieved through a lightweight CSPDarknet backbone network.
It achieves high recall rate and high-precision positioning of various morphological defects such as long cracks and small-sized pitting, reduces missed detection and false detection rates, meets the online detection needs of high-speed production lines, and is easy to deploy on edge devices.
Smart Images

Figure CN120672724A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of track defect detection, and more specifically, to a track surface defect detection system made of steel based on YOLOv8. Background Art
[0002] With the advancement of steel manufacturing technology and the rapid construction of high-speed rail and urban rail transit, the quality and stability of rail steel have become crucial for ensuring traffic safety and transportation efficiency. During the production process, rail steel surfaces are susceptible to surface defects such as cracks, spalling, indentations, and scratches due to various processes, including casting, rolling, cooling, and transportation. If these surface defects are not discovered and repaired promptly, they can seriously impact the service life of the track and the safety of train operations.
[0003] Currently, the main methods for detecting rail steel surface defects include visual inspection, magnetic particle testing, and eddy current testing. Visual inspection is highly subjective, inefficient, and prone to missed defects, making it unsuitable for high-volume, continuous production lines. While magnetic particle testing and eddy current testing offer some automation capabilities, they require high equipment requirements, are costly, and are often limited to localized inspections, making them difficult to meet the practical needs of rapid, large-scale inspections of rail steel.
[0004] In recent years, with the continuous advancement of computer vision technology, deep learning-based image recognition methods have been widely used in the field of industrial surface defect detection. In particular, object detection algorithms such as YOLO (You Only Look Once) have become a major research hotspot in the field of automated defect detection due to their fast detection speed and high accuracy.
[0005] However, existing YOLO-based defect detection systems still face the following technical bottlenecks in rail steel manufacturing: Strong target diversity: Track surface defects vary in shape and size. Traditional convolutional structures have poor adaptability to deformable targets (such as curved surface cracks and surface indentations), making them prone to missed detection or false detection. Complex detection environment: The steel production line has a complex background, uneven lighting, and severe surface reflection, which affects the accuracy of feature extraction of defect areas; Limited deployment performance: Existing deep learning models have a large number of parameters, making them difficult to deploy efficiently in edge devices or online detection systems, which restricts the implementation of industrial applications. Summary of the Invention
[0006] The purpose of the present invention is to provide a rail surface defect detection system for steel manufacturing based on YOLOv8 to solve the problems raised in the above background technology.
[0007] A rail surface defect detection system for steel products based on YOLOv8, comprising an image acquisition device and a server, which also includes a preprocessing module, a model inference module, an evaluation feedback module, and an alarm linkage module; The image acquisition device, preprocessing module, model reasoning module, evaluation feedback module and alarm linkage module are respectively connected to the server; The image acquisition device is arranged above the rail surface detection area and is used to acquire a continuous image sequence of the rail surface; The pre-processing module performs illumination equalization, noise suppression and scale normalization on the image sequence to enhance the characteristics of the defect area; The model inference module performs defect recognition and location on the processed image sequence based on the pre-trained YOLOv8 detection model, and outputs the defect category, bounding box and confidence level; The evaluation feedback module evaluates the defect level according to the recognition result and generates a feedback signal according to the evaluation result; The alarm linkage module receives the feedback signal, triggers an audible and visual alarm, or sends a control instruction to the production control system.
[0008] Preferably, the image acquisition device includes an industrial camera, a linear array light source, and an image cache unit. The industrial camera is connected to the image cache unit, and the linear array light source is used to enhance the texture features of the track surface.
[0009] Preferably, the pre-processing module includes an image enhancement unit, an affine transformation unit, and a background modeling unit. Wherein, the image enhancement unit is used to perform operations such as histogram equalization and gamma correction on the image; The affine transformation unit is used to normalize the image to the model training size; The background modeling unit is used to suppress repeated texture interference and improve defect conspicuity.
[0010] Preferably, the model inference module includes a pre-trained YOLOv8 detection model, a scale attention mechanism unit, and a feature fusion unit; The scale attention mechanism unit is used to improve the detection accuracy of long strip defects (such as cracks); The feature fusion unit fuses multi-scale feature maps to improve the recognition rate of small defects (such as pitting).
[0011] Preferably, the pre-trained YOLOv8 detection model integrated in the model inference module includes the following essential technical features: the YOLOv8 detection model uses CSPDarknet as the backbone feature extraction network, and performs gradient diversion through the CSP neural network structure to reduce redundant calculations and achieve model lightweighting; a deformable convolution structure DCN is embedded in the detection head of YOLOv8 to dynamically sample the input feature map, thereby enhancing the model's adaptability to unstructured defects on the track surface; wherein the deformable convolution structure adjusts the spatial sampling points of the standard convolution operation by introducing a learnable offset, thereby improving the detection robustness of complex backgrounds and deformed targets. Specifically, the following calculation is performed at each convolution position: ; in, is the input feature map, is the output feature map, is the convolution kernel weight, is the current center position, is the standard convolution sampling offset, is a learnable dynamic offset.
[0012] Preferably, the evaluation feedback module includes a defect statistics unit and a grade evaluation unit; The defect statistics unit performs statistical analysis on the frequency of defects within a specified time period; The level evaluation unit calculates the defect level S based on the defect area, defect location and confidence level; ; in, is the defect area, is the position index of the defect from the edge of the rail, is the confidence level, is the weight coefficient, set by the system or user.
[0013] Preferably, if the assessed defect level S exceeds the set alarm threshold, the alarm linkage module triggers the alarm controller to output an audible and visual alarm signal, and uploads the image and location information of the defect to the cloud server or pushes it to the production management system.
[0014] Preferably, the model inference module also includes a training enhancement module, which is used to enhance the defect image during the model training stage. The enhancement methods include simulated crack texture superposition, peeling-like defect synthesis, strong light disturbance simulation, and defect random interpolation expansion to improve the generalization ability of the YOLOv8 model in defect recognition.
[0015] Preferably, the system further includes a display interaction module, which is used to display the defect detection results output by the model inference module in real time and support the operator to zoom, mark or correct the defect image through touch or gesture operations.
[0016] Compared with the prior art, the advantages of the present invention are: (1) By integrating a deformable convolutional structure into the YOLOv8 detection head, and introducing a scale attention mechanism unit and a multi-scale feature fusion unit, this solution can achieve high recall rate and high-precision positioning for various morphological defects such as long cracks and small-sized pitting, and the defect classification accuracy is improved by more than 10%.
[0017] (2) The training enhancement module of the present invention adopts multi-dimensional enhancement methods such as simulated crack superposition, quasi-peeling synthesis, strong light disturbance simulation and random interpolation expansion to effectively simulate various complex scenes in real production lines, greatly reducing missed detections and false detections caused by sample scarcity, lighting changes or background interference.
[0018] (3) The present invention adopts a lightweight CSPDarknet backbone network and a multi-threaded image stream processing design, which enables the overall model inference speed to reach over 60FPS, meeting the online detection requirements of high-speed production lines; at the same time, the algorithm can be deployed on edge devices or industrial PCs, and is easy to connect with existing MES / PLC systems. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 This is an overall block diagram of a rail surface defect detection system for steel manufacturing based on YOLOv8 of the present invention. DETAILED DESCRIPTION
[0020] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0021] A rail surface defect detection system for steel products based on YOLOv8, comprising an image acquisition device and a server, which also includes a preprocessing module, a model inference module, an evaluation feedback module, and an alarm linkage module; The image acquisition device, preprocessing module, model reasoning module, evaluation feedback module and alarm linkage module are respectively connected to the server; The image acquisition device is arranged above the rail surface detection area and is used to acquire a continuous image sequence of the rail surface; The pre-processing module performs illumination equalization, noise suppression and scale normalization on the image sequence to enhance the characteristics of the defect area; The model inference module performs defect recognition and location on the processed image sequence based on the pre-trained YOLOv8 detection model, and outputs the defect category, bounding box and confidence level; The evaluation feedback module evaluates the defect level according to the recognition result and generates a feedback signal according to the evaluation result; The alarm linkage module receives the feedback signal, triggers an audible and visual alarm, or sends a control instruction to the production control system.
[0022] The image acquisition device includes an industrial camera, a linear array light source, and an image buffer unit. The industrial camera is connected to the image buffer unit. The linear array light source is used to enhance the texture features of the track surface.
[0023] The pre-processing module includes an image enhancement unit, an affine transformation unit and a background modeling unit. Wherein, the image enhancement unit is used to perform operations such as histogram equalization and gamma correction on the image; The affine transformation unit is used to normalize the image to the model training size; The background modeling unit is used to suppress repeated texture interference and improve defect conspicuity.
[0024] The model inference module includes a pre-trained YOLOv8 detection model, a scale attention mechanism unit, and a feature fusion unit; The scale attention mechanism unit is used to improve the detection accuracy of long strip defects (such as cracks); The feature fusion unit fuses multi-scale feature maps to improve the recognition rate of small defects (such as pitting).
[0025] A scale attention mechanism unit is designed to target the slender structural characteristics of long strip defects (such as cracks) on the rail surface. It adaptively allocates attention weights between feature maps of different scales to improve the model's responsiveness to slender deformation defects.
[0026] Structural composition Input features: multi-scale feature maps from backbone networks or feature pyramids ,in Indicates different downsampling multiples (such as 8×, 16×, 32×).
[0027] Scale-ChannelAttention: For each First, apply the channel attention mechanism and calculate the channel weights: ; in, is the global average pooling, MLP is a two-layer fully connected network, Sigmoid activation; channel weighting .
[0028] Scale-SpatialAttention: Further apply spatial attention to generate a spatial weight map: ; Among them, AvgPool and MaxPool are pooling in channel dimension, Conv It is a 7×7 convolution; the final output is weighted features .
[0029] Scale fusion weight: for all scales Perform channel dimension splicing, and then pass a layer of 1×1 convolution and Softmax normalization to generate the fusion coefficient of each scale :
[0030] Weighted output: Features at each scale are weighted by The weighted sum is taken as the final output of the scale attention unit:
[0031] Dynamically emphasize the scale branches that respond most strongly to the slender shape of the crack; suppress irrelevant scale features that are very different from the crack morphology (such as large spalling); and improve the crack detection accuracy and recall rate at multiple scales.
[0032] The feature fusion unit improves the recognition rate of small-sized defects such as tiny pitting and pits, and takes into account both semantic information and spatial resolution through up / down sampling and fusion of multi-level features.
[0033] Structural composition Bottom-up feature extraction: Extract feature maps of different resolutions sequentially from the lightweight CSPDarknet backbone , corresponding to downsampling ratios of 8×, 16×, and 32×.
[0034] Top-Down Feature Pyramid (Top-DownPath): The top First perform 1×1 convolution to reduce the channel, and get ; right Upsample to Same space dimensions and Do channel splicing and then 3×3 convolution fusion to get ; Similarly, upsampling and Fusion, get .
[0035] Bottom-Up Path Augmentation: right Do 3×3 convolution downsampling and get ,and After concatenation, convolution is performed ; Similarly, Downsampling fusion to , output .
[0036] Final fusion output: multi-level After splicing in the channel dimension and integrating the channels through 1×1 convolution, a fusion feature map is formed. .
[0037] Taking into account both high-level semantics and low-level details: the spatial information of small defects (such as pitting) is The semantic context is retained in the high level Supplementary information: Bidirectional path enhancement: The bottom-up enhancement path enables mid-level features to obtain dual information from the finest-grained and most semantic features; significantly improving the detection rate of small-sized and weak-contrast defects.
[0038] The pre-trained YOLOv8 detection model integrated into the model inference module includes the following essential technical features: The YOLOv8 detection model uses CSPDarknet as the backbone feature extraction network and performs gradient shunting through the CSP neural network structure to reduce redundant computation and achieve model lightweighting. A deformable convolutional network (DCN) is embedded in the YOLOv8 detection head to dynamically sample input feature maps, thereby enhancing the model's adaptability to unstructured defects on the track surface. The deformable convolutional network introduces a learnable offset to adjust the spatial sampling points of the standard convolution operation, improving the detection robustness against complex backgrounds and deformed targets. Specifically, the following calculation is performed at each convolution position: ; in, is the input feature map, is the output feature map, is the convolution kernel weight, is the current center position, is the standard convolution sampling offset, is a learnable dynamic offset.
[0039] The evaluation feedback module includes a defect statistics unit and a grade evaluation unit; The defect statistics unit performs statistical analysis on the frequency of defects within a specified time period; The level evaluation unit calculates the defect level S based on the defect area, defect location and confidence level; ; in, is the defect area, is the position index of the defect from the edge of the rail, is the confidence level, is the weight coefficient, set by the system or user.
[0040] If the assessed defect level S exceeds the set alarm threshold, the alarm linkage module triggers the alarm controller to output an audible and visual alarm signal, and uploads the image and location information of the defect to the cloud server or pushes it to the production management system.
[0041] The model inference module also includes a training enhancement module, which is used to enhance the defect image during the model training stage. The enhancement methods include simulated crack texture superposition, peeling-like defect synthesis, strong light perturbation simulation, and defect random interpolation expansion to improve the generalization ability of the YOLOv8 model in defect recognition.
[0042] Specific implementation of training enhancement module 1. Simulated crack texture superposition Objective: To address the scarcity of real crack samples, a simulation algorithm is used to generate diverse crack textures to increase the robustness of the model to different crack morphologies.
[0043] Implementation: Extract the crack skeleton shape and width distribution information from the finite real crack image; Generate pseudo crack paths based on the trunk shape using random fractal noise (such as PerlinNoise) or Bézier curves; Random diffusion along the normal direction at the center of the path generates crack width variation; The simulated crack texture is superimposed on the original defect-free rail image according to α mixing (α∈[0.3,0.7]) to form a diversified crack sample.
[0044] 2. Synthesis of exfoliation-like defects Purpose: To simulate the irregular shape and edge characteristics of the spalling area on the rail surface and supplement the diversity of real spalling samples.
[0045] Implementation: Extract typical edge curves based on the actual peeling area contour; Use random jitter algorithms (such as Gaussian perturbation or ChamferDistance deformation) to fine-tune the edge curve to generate a new peeling profile; Fill the inside of the outline with the base color of the rail and add fine grain noise to simulate the residual material at the bottom of the pit; Blend the synthetic peeled area with the original image and add a slight feathering process at the edges to ensure a smooth transition.
[0046] 3. Strong light disturbance simulation Purpose: To reproduce the saturated highlight areas caused by high reflectivity of metal surfaces in the production line, and reduce false detections and missed detections.
[0047] Implementation: Randomly sample several highlight center points on the image; For each center point, generate the halo radius and brightness attenuation curve according to the parabolic model or Gaussian distribution; A brightness enhancement is applied to the original image, bringing the center up to 200%–300% of the original pixel value, and a slight chromatic aberration is applied on top to simulate lens flare. Optionally add a randomly oriented strip reflective texture to simulate the effect of a direct line array light source.
[0048] 4. Defect random interpolation expansion Purpose: To generate intermediate samples between a small number of real defect samples through interpolation and transformation to smooth the sample distribution.
[0049] Implementation: Randomly select two images and their annotation boxes from the same category defect sample set; Perform linear interpolation (MixUp) or region tiling (CutMix) operations at the pixel or feature level: MixUp: By Proportion Linearly blend the two images and synchronously interpolate the confidence weights of the annotation boxes; CutMix: Randomly cut out a defect area and paste it to the corresponding position of another image, and update the annotation box and category; The interpolated image is randomly rotated (±10°), scaled (0.9–1.1×), and sheared (Shear±5°), and the coordinates of the annotation box are recalculated.
[0050] The system also includes a display interaction module, which is used to display the defect detection results output by the model reasoning module in real time and support the operator to zoom, mark or correct the defect image through touch or gesture operations.
[0051] The basic principles, main features, and advantages of the present invention are shown and described above. It should be understood by those skilled in the art that the present invention is not limited to the above-described embodiments. The above-described embodiments and descriptions are merely preferred examples of the present invention and are not intended to limit the present invention. Various changes and modifications may be made to the present invention without departing from the spirit and scope of the present invention, and such changes and modifications fall within the scope of the invention claimed.
Claims
1. A rail surface defect detection system for steel products based on YOLOv8, comprising an image acquisition device and a server, characterized in that: It also includes a pre-processing module, a model reasoning module, an evaluation feedback module, and an alarm linkage module; The image acquisition device, preprocessing module, model reasoning module, evaluation feedback module and alarm linkage module are respectively connected to the server; The image acquisition device is arranged above the rail surface detection area and is used to acquire a continuous image sequence of the rail surface; The pre-processing module performs illumination equalization, noise suppression and scale normalization on the image sequence to enhance the characteristics of the defect area; The model inference module performs defect recognition and location on the processed image sequence based on the pre-trained YOLOv8 detection model, and outputs the defect category, bounding box and confidence level; The evaluation feedback module evaluates the defect level according to the recognition result and generates a feedback signal according to the evaluation result; The alarm linkage module receives the feedback signal, triggers an audible and visual alarm, or sends a control instruction to the production control system.
2. A rail surface defect detection system for steel products based on YOLOv8 according to claim 1, characterized in that: The image acquisition device includes an industrial camera, a linear array light source, and an image buffer unit. The industrial camera is connected to the image buffer unit. The linear array light source is used to enhance the texture features of the track surface.
3. The rail surface defect detection system for steel products based on YOLOv8 according to claim 1, characterized in that: The pre-processing module includes an image enhancement unit, an affine transformation unit, and a background modeling unit. Wherein, the image enhancement unit is used to perform operations such as histogram equalization and gamma correction on the image; The affine transformation unit is used to normalize the image to the model training size; The background modeling unit is used to suppress repeated texture interference and improve defect conspicuity.
4. The rail surface defect detection system for steel products based on YOLOv8 according to claim 1, characterized in that: The model inference module includes a pre-trained YOLOv8 detection model, a scale attention mechanism unit, and a feature fusion unit; The scale attention mechanism unit is used to improve the detection accuracy of long strip defects (such as cracks); The feature fusion unit fuses multi-scale feature maps to improve the recognition rate of small defects (such as pitting).
5. The rail surface defect detection system for steel products based on YOLOv8 according to claim 4, characterized in that: The pre-trained YOLOv8 detection model integrated into the model inference module includes the following essential technical features: The YOLOv8 detection model uses CSPDarknet as the backbone feature extraction network and performs gradient shunting through the CSP neural network structure to reduce redundant computation and achieve model lightweighting. A deformable convolutional network (DCN) is embedded in the YOLOv8 detection head to dynamically sample input feature maps, thereby enhancing the model's adaptability to unstructured defects on the track surface. The deformable convolutional network introduces a learnable offset to adjust the spatial sampling points of the standard convolution operation, improving the detection robustness against complex backgrounds and deformed targets. Specifically, the following calculation is performed at each convolution position: ; in, is the input feature map, is the output feature map, is the convolution kernel weight, is the current center position, is the standard convolution sampling offset, is a learnable dynamic offset.
6. The rail surface defect detection system for steel products based on YOLOv8 according to claim 1, characterized in that: The evaluation feedback module includes a defect statistics unit and a grade evaluation unit; The defect statistics unit performs statistical analysis on the frequency of defects within a specified time period; The level evaluation unit calculates the defect level S based on the defect area, defect location and confidence level; ; in, is the defect area, is the position index of the defect from the edge of the rail, is the confidence level, is the weight coefficient, set by the system or user.
7. The rail surface defect detection system for steel products based on YOLOv8 according to claim 6, characterized in that: If the assessed defect level S exceeds the set alarm threshold, the alarm linkage module triggers the alarm controller to output an audible and visual alarm signal, and uploads the image and location information of the defect to the cloud server or pushes it to the production management system.
8. The rail surface defect detection system for steel products based on YOLOv8 according to claim 1, characterized in that: The model inference module also includes a training enhancement module, which is used to enhance the defect image during the model training stage. The enhancement methods include simulated crack texture superposition, peeling-like defect synthesis, strong light disturbance simulation, and defect random interpolation expansion.
9. The rail surface defect detection system for steel products based on YOLOv8 according to claim 1, characterized in that: The system further comprises a display interaction module, which is used to display the defect detection results output by the model reasoning module in real time.
Citation Information
Cited By
Steel rail surface state detection method and system
CN120997205A
Defect detection method, device and equipment based on multi-mode cooperation and medium
CN121033026A
Method and system for identifying multiple types of defects on rolled surface of silicon steel sheet based on computer vision
CN121302897A
Silicon steel sheet rolling surface multi-type defect recognition method and system based on computer vision
CN121302897B
Railway wagon floor damage fault detection method and device and electronic equipment
CN121661059A