Railway turnout defect detection method based on SCG-DETR detection model

Through the SCG-DETR detection model, the C2f-SMP module and the CGFM module are used to optimize feature extraction and fusion, combined with IoU-aware query and Inner-GIoU loss function, the problems of low efficiency and insufficient accuracy in railway cross detection are solved, and high-precision and real-time defect detection are achieved.

CN120355657AInactive Publication Date: 2025-07-22ZHEJIANG SCI-TECH UNIV
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202510372705.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-07-22
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing railway fork detection methods rely on manual inspection and simple instruments, and have low detection efficiency and are susceptible to human factors. The deep learning methods lack accuracy in the absence of labeled data, making it difficult to effectively identify railway fork defects in complex environments.

Method used

Using the SCG-DETR detection model, the feature extraction is optimized through the C2f-SMP module, the CGFM module enhances context information fusion, and the bounding box regression of IoU-aware query is optimized, combining data enhancement and Inner-GIoU loss function to improve the detection accuracy and speed of the model in complex environments.

Benefits of technology

It realizes high-precision detection of railway fork defects in complex environments, improves detection accuracy and speed, meets real-time requirements, and reduces dependence on hardware resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120355657A_ABST
    Figure CN120355657A_ABST
Patent Text Reader

Abstract

The invention discloses a railway turnout defect detection method based on an SCG-DETR detection model. Firstly, a C2f-SMP feature extraction module is designed based on SMPConv and C2f modules, local and global information aggregation of features is optimized through continuous convolution and a gating mechanism, and the receptive field and feature extraction capability is enhanced. Secondly, a CGFM module is provided, and by introducing a CAA attention mechanism, weight distribution of a feature map is optimized, and fine-grained extraction and context information fusion are enhanced. And finally, replacing an original loss function with an Inner-GIoU based on an auxiliary frame, and improving the precision of the model while accelerating the convergence of the model. Experimental results show that the score of mAP50 of the improved model reaches 71.0% and is increased by about 4.6%, the detection speed reaches 80.5 FPS, and the requirement for real-time performance in engineering is met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer vision and railway safety detection, and in particular to a railway turnout defect detection method based on an SCG-DETR detection model. Background Art

[0002] The frog is an important part of a railway turnout, responsible for guiding the train to switch between different tracks. It must be able to ensure that the train switches tracks smoothly and seamlessly under high speed and heavy load conditions. However, with the increase in train transportation mileage, the frog is easily affected by long-term heavy loads and high-speed passage during use, resulting in defects such as block loss, cracks, and wear. If these defects are not discovered and repaired in time, they will seriously threaten the safety of train operation and may cause major accidents such as derailment, causing huge losses to passengers and railway operators. Traditional railway frog detection methods mainly rely on manual inspections and simple instrument measurements, which are not only inefficient but also easily affected by human factors.

[0003] In the field of railway defect detection, common technologies include ultrasonic testing, acoustic emission technology, electromagnetic detection technology, edge detection, Fourier transform and Canny edge detection. Zhang et al. proposed a high-speed phased array ultrasonic detection method, which significantly improved the detection speed and coverage by generating multi-angle beams and receiving defect echo signals from each channel. Ultrasonic testing is mainly used to identify defects inside materials, such as cracks and delamination, and can penetrate materials for deep detection. However, the limitation of this method is that the sensor needs to be in contact with the detection surface and is sensitive to surface roughness. Acoustic emission technology (AE) can monitor the dynamic changes of internal defects of rails in real time by detecting transient elastic waves generated when rail defects expand. This method is particularly suitable for online continuous monitoring, which can evaluate the dynamic characteristics of defects and effectively identify rail head, internal, welding and surface defects. However, its main challenge is that it is easily interfered by external noise, such as noise in train operation or natural environment. Electromagnetic detection technology includes magnetic flux leakage detection (MFL) and eddy current detection (ECI), which are mainly used for the detection of surface and internal defects of rails. During high-speed nondestructive testing, the relative movement between the rail and the testing equipment generates motion-induced eddy currents (MIEC). By analyzing the changes in the magnetic field inside and on the surface of the rail, potential defects can be identified. Magnetic flux leakage testing is usually used to detect defects such as cracks and pits on the surface of the rail. When defects occur, the magnetic lines of force inside the rail will bend and form a leakage magnetic field on the surface. Eddy current testing is mainly used to identify surface defects of conductive materials. Although it performs well in detecting surface cracks, it has limited detection capabilities for deep defects and is easily interfered by external electromagnetic signals.

[0004] With the improvement of computing power, detection methods based on deep learning have gradually become mainstream. However, deep learning methods usually rely on a large amount of labeled data and high computing resources. Deep learning technology has significantly improved detection performance by automatically learning multi-level features, especially showing excellent performance in complex scenarios. However, the success of such methods depends on a large amount of labeled data and computing resources. In practical applications, collecting and annotating large-scale datasets is often time-consuming and laborious. Especially in the small-sample problem of railway frog defects, data scarcity has become the main limitation. Summary of the Invention

[0005] The purpose of the present invention is to propose a railway turnout defect detection method based on the SCG-DETR detection model for the characteristics of railway frog defects, which have arbitrary shapes and complex geometric structures. The size of the defects is usually small, the pixel ratio in the image is low, and the provided feature information is limited, resulting in the problem of easy feature loss during detection.

[0006] The present invention is realized through the following technical solutions:

[0007] A railway turnout defect detection method based on the SCG-DETR detection model includes the following steps:

[0008] Step 1: Construct a railway turnout defect dataset. The data in the dataset are railway frog defects; the labels of railway frog defect images are defect types, and the corresponding defect types are spalling, flash, and abnormal light band; the dataset includes turnout defect images under different lighting conditions and weather conditions, and the images are normalized to 640×640 pixels in size, regularized, and data augmentation operations are performed.

[0009] Step 2: Construct an SCG-DETR detection model, and the model includes the following modules:

[0010] The model sequentially includes a feature extraction backbone (BackBone), an efficient hybrid encoder, an IoU-aware query selection, and a decoder (Decoder);

[0011] C2f-SMP module: The feature extraction backbone adopts the SMPNet network. By replacing the Bottleneck in the C2f module of the original CSPDarknet53 with the SMPC module, a convolution kernel is generated through an adaptive point movement mechanism, and the feature extraction is optimized by combining a gating mechanism, thereby significantly improving the feature extraction ability and computing efficiency. Among them, the SMPC module is composed of the SMPConv and ConvGLU modules.

[0012] CGFM module: Based on the CAA attention mechanism, adaptively weighted fusion of multi-scale feature maps is performed, and context information interaction is enhanced through global average pooling and pointwise convolution; Inner-GIoU loss function: Auxiliary bounding boxes are generated by introducing a scale factor, the bounding box regression loss is dynamically adjusted, and the model convergence is accelerated.

[0013] Specifically, the improved feature extraction backbone network consists of four stages. At the beginning of each stage, there is a 3×3 convolution (stride 2) for spatial downsampling and channel adjustment. The SMPConv in the SMPC module in C2f-SMP adopts an adaptive point movement mechanism, which can better capture long-term dependencies in long sequence data and image data, thereby improving the flexibility and accuracy of feature extraction. This method not only achieves efficient computation and parameter reduction but also maintains the high expressive power of the convolutional kernel, and is particularly suitable for processing images with rich details.

[0014] During the feature extraction process, each self-moving point in the SMPC module is updated independently, which enables the model to represent high-frequency components more effectively; compared with traditional convolution methods, the SMPConv convolution in SMPC has more advantages in feature extraction. Subsequently, features are further extracted through the Conv GLU module; this module combines convolution and gating mechanisms, and can selectively pass information channels to improve the effectiveness and flexibility of feature extraction; Conv GLU is similar to the encoder structure of Transformer, uses DWConv for feature extraction and performs non-linear transformation to further enhance the feature extraction ability. Through the dynamic gating mechanism, the gating signal of each Token comes from itself, avoiding the signal sharing problem that may be introduced by global average pooling;

[0015] Finally, to prevent the neural network from overfitting, a dropout layer is introduced at the end stage of the SMPC module, and the input features are reused through shortcut connections, further improving the stability and generalization ability of the model. Specifically, the Dropout layer randomly discards neurons, forcing the network not to rely too much on certain specific neuron combinations, enhancing the generalization ability, and at the same time preventing complex interdependencies from forming between neurons by reducing the co-adaptability of neurons; The railway frog defect detection model is trained using the dataset in step 1, and finally a trained defect detection model is obtained.

[0016] Step 3: Input the feature map obtained in Step 2, which is the original railway frog defect image, i.e., the feature map processed by the feature extraction backbone model, into the efficient hybrid encoder. This encoder consists of two parts: an attention-based intra-scale feature interaction module (AIFI) and a CNN-based cross-scale feature fusion module (CCFF). AIFI only performs interaction between scales on the S5 feature map, while the CCFF module constructs a Fusion block through convolutional layers and enters the fusion path. The role of the Fusion block is to fuse adjacent features into new features, which contains N RepBlocks, and the outputs of the two paths are fused through element-wise addition.

[0017] In the Fusion block, the present invention proposes a ContextGuideFusionModule (CGFM) module, which introduces a context attention mechanism to achieve adaptive fusion of input features. Considering the uneven feature distribution of defects in the spatial dimension in the railway frog dataset, the traditional spatial dimension-based attention mechanism has poor effects. Therefore, in the CGFM module, a context attention mechanism is adopted. This method can not only retain the original feature information but also combine context information to capture global features, thereby improving the performance of the model. The specific structures of the Fusion-IM module and the CGFM module in CCFF are as Figure 3 and Figure 4 shown, including global attention CAA, two pointwise convolutions, a batch normalization layer (BN), and a Sigmoid layer.

[0018] Step 4: Obtain the railway turnout image to be detected, input it into the trained SCG-DETR model, output the defect location and category results, and generate a heat map through Grad-CAM++ to visualize the high-confidence defect regions.

[0019] Specifically, input the feature map encoded by the efficient hybrid encoder into the IoU-aware query in the SCG-DETR model. The IoU-aware query improves the initialization of object queries and expands them into content queries and location queries (anchors). In the traditional query selection scheme, a certain number of features are usually extracted from the feature sequence output by the encoder as query objects, then passed through the decoder, and finally converted into classification scores and IoU scores by the prediction head. Most DETR variants rely on classification scores to select matching boxes, which may result in the selection of some prediction boxes with high classification scores but low IoU scores, while boxes with low classification scores but high IoU scores are ignored, thus affecting the accuracy of the model.

[0020] The IoU-aware query constrains the detection graph, enabling features with high IoU scores to obtain higher classification scores during training, while features with low IoU scores are given lower classification scores. In this way, the model can select prediction boxes with both high classification scores and high IoU scores during classification, thereby improving detection accuracy.

[0021] Furthermore, data augmentation includes random scaling with a scaling ratio of 0.5 - 1.5; horizontal / vertical flipping, brightness adjustment of ±20%, grayscale transformation, Gaussian blur with a kernel size of 3×3, and noise addition with a noise density of 0.05. The augmented dataset contains 2500 images, which are divided into a training set, a validation set, and a test set in a ratio of 7:2:1.

[0022] Furthermore, in step two, the construction of the C2f-SMP module:

[0023] The mathematical expression of SMPConv is:

[0024]

[0025] Where is a set of trainable parameters, where represents N p learnable support point positions; corresponds to the weight of each support point; corresponds to the radius parameter of each support point; |N(x)| represents the number of points in the neighborhood of point x; g(x, p i , r i ) is used to describe the relationship between point x and support point p i , where r i is the influence radius parameter; w i is the weight parameter associated with support point p i .

[0026]

[0027] g(x, p i , r i ) describes the kernel function relationship between point x and support point p i , ||x - p i || represents the Euclidean distance between point x and support point p i , r i is the radius parameter associated with support point p i . This function calculates a distance-based linear attenuation function, and this distance-based linear attenuation function is used in SMPConv to weigh the influence degree of different support points on the target point. The closer the support point, the greater the contribution. Points beyond a specific radius r i may no longer have an impact.

[0028] The ConvGLU module selectively transmits feature information through depthwise separable convolution and a gating mechanism, and its expression is:

[0029]

[0030] where σ is the Sigmoid activation function, is element-wise multiplication, where Conv represents the ordinary convolution operation; the four stages of the C2f-SMP module contain 3, 4, 6, and 3 SMPConv modules respectively. Before the start of each stage, spatial downsampling and channel adjustment (the number of channels are 64, 128, 256, and 512 respectively) are performed through 3×3 convolution (stride 2).

[0031] Furthermore, the implementation of the CGFM module includes:

[0032] Concatenate the input feature maps F i and F j , and then perform CAA attention processing to generate the attention weight F att :

[0033] F s = CAA(Cat(F i , F j ))

[0034] This means that by concatenating (Cat) the features F i and F j , and then applying a Channel Attention Adaptation (CAA) module, the fused feature F s

[0035] F att = σ(BN(PWC(BN(PWC(AVG(F s ))))

[0036] First, perform global average pooling (AVG) on F s , then apply pointwise convolution (PWC, 1×1 convolution kernel), then perform batch normalization (BN), apply pointwise convolution again, perform batch normalization again, and finally obtain the attention weight F att through the Sigmoid activation function (σ).

[0037] Fuse the features through element-wise product and cross-addition operations, and the output expression is:

[0038] F i,weight = F i × F att

[0039] F j,weight = F j × F att

[0040] F out = Cat((F i,weight + F j ),(F j,weight + F i ))

[0041] First, multiply the attention weight F att by F i to obtain the attention weight F i of the input feature map F i,weight , then add F j to get the corresponding product sum; multiply the attention feature F att by F j to obtain the attention weight F j of the input feature map F j,weight , then add F i to get the product sum; concatenate (Cat) the above two product sums to obtain the final output feature map F out .

[0042] The CAA attention mechanism captures long - range dependencies through horizontal and vertical depth - strip convolutions with kernel sizes of 7×1 and 1×7, and uses global average pooling to compress the feature dimension to 1 / 16 of the original number of channels.

[0043] The mathematical expression of IoU is as follows:

[0044]

[0045] Among them, A represents the area size of the target box, B represents the area size of the detection box, |A∩B| represents the absolute value of the area of the overlapping region between the target box and the detection box, and |A∪B| represents the absolute value of the total area sum of the target box and the detection box.

[0046] Furthermore, GIoU solves the problem that the gradient is zero when two targets have no intersection by introducing the minimum bounding rectangle of the predicted box and the ground - truth box, and its mathematical expression is:

[0047]

[0048] Among them, C is the area of the minimum bounding rectangle of the two boxes.

[0049] Next, by introducing auxiliary bounding boxes, different sizes of auxiliary bounding boxes are generated using a scaling factor to calculate the loss. Here, Inner represents the concept of auxiliary bounding boxes. By combining the auxiliary bounding boxes with GIoU, Inner-GIoU is proposed, and its mathematical expression is:

[0050] Inner-GIoU = GIoU + IoU - IoU inner

[0051] Among them, IoU inner Dynamically calculate the intersection over union through the scale factor ratio (value range 0.5 - 1.5) of the auxiliary bounding box. The specific calculation process is as follows:

[0052]

[0053] union = (w g * h g ) * (ratio) 2 + (w * h) * (ratio) 2 - inter

[0054]

[0055] Further explanation, where x g , y g represent the center coordinates of the target box, w g , h g represent the width and height of the target box, x, y represent the center coordinates of the detection box, w, h represent the width and height of the detection box, ratio represents the scaling factor of the auxiliary bounding box, b l g , b r g , b t g , b b g represent the Inner boundaries (left, right, top, bottom) of the target box respectively, b l , b r , b t , b b represent the Inner boundaries (left, right, top, bottom) of the detection box respectively, inter represents the intersection area between Inner boxes, union represents the union area between Inner boxes, and IoU inner represents the final Inner-IoU.

[0056] Furthermore, the evaluation indicators of the model SCG-DETR satisfy:

[0057] mAP50 ≥ 71.0%, AP (Average Precision) ≥ 78.8%, AR (Average Recall) ≥ 68.4%; Detection speed ≥ 80.5 FPS, number of parameters ≤ 15.5M, computational complexity ≤ 68.0 GFLOPs;

[0058] Among them, the mAP50 metric represents the average of the prediction boxes with an intersection over union greater than 50%. AP is usually an indicator to measure the detection accuracy for a certain class (or in some cases, it also refers to the average of all classes). AR is usually used to measure the average recall rate of the model under different conditions (such as different confidence thresholds, different maximum detection numbers). FPS represents the number of image frames that the model can process per second and is a commonly used indicator to measure the inference speed. The number of parameters refers to the trainable parameters (TrainableParameters) of the model, and the computational complexity represents the number of floating-point operations required for one forward inference of the model.

[0059] The effective receptive field coverage ratio (T = 20%) ≥ 0.35%, which is better than ResNet-18 (0.23%) and CSPDarknet53 (0.25%).

[0060] Furthermore, the anti-interference ability of the model SCG-DETR is achieved in the following ways:

[0061] The context information fusion of the CGFM module suppresses background noise;

[0062] The CAA attention mechanism enhances the feature weights of the defect regions;

[0063] The Inner-GIoU loss function optimizes the regression accuracy of the small object bounding boxes;

[0064] Under strong light, rain, snow, and night conditions, the fluctuation range of mAP50 ≤ ±2.5%.

[0065] Furthermore, in step 3, the parameters for training the SCG-DETR model include: the number of iterations 100, batch size 8, the optimizer is AdamW (weight decay 0.05, momentum 0.9), learning rate 0.0001, input image size 640×640, until the cross-entropy loss function converges to the desired error threshold (ε = 1e-3).

[0066] Compared with the prior art, the advantages of the present invention are:

[0067] In view of the problems of low accuracy and difficulty in deployment in complex field environments existing in the existing railway frog defect detection algorithms, the present invention proposes a railway turnout defect detection method based on an improved RT-DETR algorithm. This algorithm extracts image features by using a CNN convolutional neural network and further extracts and encodes the image features by using an efficient hybrid encoder, so as to realize the accurate detection and recognition of railway frog defects, achieving the goal of improving the detection accuracy and solving the deployment problem. BRIEF DESCRIPTION OF THE DRAWINGS

[0068] The present invention will be further described below with reference to the accompanying drawings.

[0069] Figure 1 It is a diagram of the SCG-DETR algorithm model constructed by the present invention;

[0070] Figure 2 It is the SMPNet network structure in the present invention;

[0071] Figure 3 It is the structure of the Fusion-IM module in the present invention;

[0072] Figure 4 It is the structure of the CGFM module in the present invention;

[0073] Figure 5 It is a comparison diagram of the detection of the present invention and other algorithms. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0074] The present invention will be further described below with reference to the accompanying drawings.

[0075] Embodiment 1

[0076] A railway turnout defect detection method based on an SCG-DETR detection model includes the following steps:

[0077] Step 1: Construct a railway turnout defect data set, which includes turnout defect images under different lighting conditions and weather conditions, and perform size normalization to 640×640 pixels, regularization processing and data augmentation operations on the images; 1. Data augmentation includes random scaling, with a scaling ratio of 0.5 to 1.5; horizontal / vertical flipping, brightness adjustment of ±20%, grayscale transformation, Gaussian blur, with a kernel size of 3×3 and noise addition, with a noise density of 0.05. After augmentation, the data set contains 2,500 images, which are divided into a training set, a validation set and a test set according to a ratio of 7:2:1.

[0078] Step 2: Construct an SCG-DETR detection model, which sequentially includes a feature extraction backbone (BackBone), an efficient hybrid encoder, an IoU-aware query selection and a decoder;

[0079] The feature extraction backbone network consists of four stages. At the beginning of each stage, there is a 3×3 convolution for spatial downsampling and channel adjustment. The SMPConv in the SMPC module of the C2f-SMP module adopts an adaptive point movement mechanism. During the feature extraction process, each self-moving point in the SMPC module is updated independently. Subsequently, features are further extracted through the Conv GLU module.

[0080] C2f-SMP module: The feature extraction backbone adopts the SMPNet network. By replacing the Bottleneck in the original C2f module with the SMPC module, the SMPC module includes the SMPConv and ConvGLU modules, generates convolutional kernels through the adaptive point movement mechanism, and combines the gating mechanism to optimize feature extraction.

[0081] , Construction of the C2f-SMP module:

[0082] The mathematical expression of SMPConv is:

[0083]

[0084] Among them, is the set of trainable parameters, where represents N p locations of learnable support points; corresponds to the weight of each support point; corresponds to the radius parameter of each support point; |N(x)| represents the number of points in the neighborhood of point x; g(x,p i ,r i ) is used to describe the relationship between point x and support point p i , where r i is the influence radius parameter; w i is the weight parameter associated with support point p i .

[0085] g(x,p i ,r i ) describes the kernel function relationship between point x and support point p i , ||x - p i || represents the Euclidean distance between point x and support point p i , r i is the radius parameter associated with support point p i . This function calculates a distance-based linear attenuation function. This distance-based linear attenuation function is used in SMPConv to weigh the influence degree of different support points on the target point. The closer the support point, the greater the contribution. Points beyond a specific radius r i may no longer have an impact.

[0086] The ConvGLU module selectively transmits feature information through depthwise separable convolution and a gating mechanism, and its expression is as follows:

[0087]

[0088] where σ is the Sigmoid activation function, is element-wise multiplication, where Conv represents the ordinary convolution operation; the four stages of the C2f-SMP module contain 3, 4, 6, and 3 SMPConv modules respectively. Before the start of each stage, spatial downsampling and channel adjustment (the number of channels are 64, 128, 256, and 512 respectively) are performed through 3×3 convolution (stride 2).

[0089] The efficient hybrid encoder consists of two parts: the attention-based intra-scale feature interaction module AIFI and the CNN-based cross-scale feature fusion module CCFF. AIFI only performs interaction between scales on the S5 feature map, while the CCFF module constructs a Fusion block through convolutional layers. In the Fusion block, the ContextGuideFusionModule, i.e., the CGFM module, is introduced;

[0090] CGFM module: Based on the CAA attention mechanism, it adaptively weights and fuses multi-scale feature maps, and enhances context information interaction through global average pooling and pointwise convolution;

[0091] The implementation of the CGFM module includes:

[0092] F s = CAA(Cat(F i ,F j ))

[0093] F att = σ(BN(PWC(BN(PWC(AVG(F s ))))

[0094] By concatenating (Cat) the features F i and F j , and then applying a channel attention adaptation (CAA) module, the fused feature F s is obtained. First, global average pooling (AVG) is performed on F s , then pointwise convolution (PWC, 1×1 convolution kernel) is applied, followed by batch normalization (BN), then pointwise convolution is applied again, batch normalization is performed, and finally the attention weight F att is obtained through the Sigmoid activation function (σ).

[0095] Fuse features through element-wise product and cross-addition operations, and the output expression is:

[0096] F i,weight = F i × F att

[0097] F j,weight = F j × F att

[0098] F out = Cat((F i × F att + F j ),(F j × F att + F i ))

[0099] Secondly, multiply the attention weights F att with F i , F j respectively to obtain F i,weight , F j,weight . Finally, perform diagonal addition and concatenation (Cat) operations respectively to obtain F out .

[0100] The CAA attention mechanism captures long-range dependencies through horizontal and vertical depth strip convolutions with kernel sizes of 7×1 and 1×7, and uses global average pooling to compress the feature dimension to 1 / 16 of the original number of channels.

[0101] For the IoU-aware query, the feature map encoded by the efficient hybrid encoder is input into the IoU-aware query. The IoU-aware query improves the initialization of the object query and expands it into a content query and a location query; the IoU-aware query constrains the detection map, so that during training, features with high IoU scores can obtain higher classification scores, while features with low IoU scores are given lower classification scores;

[0102] Inner-GIoU loss function: Generate auxiliary bounding boxes by introducing a scale factor, dynamically adjust the bounding box regression loss, and accelerate model convergence;

[0103] The mathematical derivation expression of the Inner-GIoU loss function is as follows:

[0104] The mathematical expression of IoU is:

[0105]

[0106] Among them, A represents the area size of the target box, B represents the area size of the detection box, |A∩B| represents the absolute value of the area of the overlapping region between the target box and the detection box, and |A∪B| represents the absolute value of the total area of the target box and the detection box.

[0107] GIoU solves the problem that the gradient is zero when two targets have no intersection by introducing the minimum bounding rectangle of the predicted box and the ground truth box to obtain the proportion of the predicted box and the ground truth box in the closure region. Its mathematical expression is:

[0108]

[0109] Among them, C is the area of the minimum bounding rectangle of the two boxes.

[0110] Next, by introducing an auxiliary bounding box, use a scaling factor to generate auxiliary bounding boxes of different sizes to calculate the loss. Its mathematical expression is:

[0111] Inner-GIoU = GIoU + IoU - IoU inner

[0112] Among them, IoU inner Dynamically calculate the intersection over union through the scale factor ratio (value range 0.5 - 1.5) of the auxiliary bounding box. The specific calculation process is:

[0113]

[0114] union = (w g * h g ) * (ratio) 2 + (w * h) * (ratio) 2 - inter

[0115]

[0116] Further explanation, where x g , y g represent the center coordinates of the target box, w g , h g represent the width and height of the target box, x, y represent the center coordinates of the detection box, w, h represent the width and height of the detection box, ratio represents the scaling factor of the auxiliary bounding box, b l g , b r g , b t g , b b g represent the Inner boundaries (left, right, top, bottom) of the target box respectively, b l,b r ,b t ,b b respectively represent the Inner boundaries (left, right, top, bottom) of the detection box, inter represents the intersection area between Inner boxes, union represents the union area between Inner boxes, and IoU inner represents the final Inner-IoU.

[0117] Step 3: Input the preprocessed dataset in Step 1 into the SCG-DETR model built in Step 2, and respectively pass through the feature extraction backbone (BackBone), efficient hybrid encoder, IoU-aware query selection, and decoder. After iterating 100 times (epoch), train to obtain a model with the highest accuracy;

[0118] The evaluation metrics of SCG-DETR satisfy:

[0119] mAP50 ≥ 71.0%, AP (average precision) ≥ 78.8%, AR (average recall rate) ≥ 68.4%;

[0120] Detection speed ≥ 80.5 FPS, number of parameters ≤ 15.5M, computational complexity ≤ 68.0 GFLOPs;

[0121] Among them, the mAP50 metric represents the average of the prediction boxes with an intersection over union greater than 50%. AP is usually an indicator to measure the detection accuracy for a certain category (or in some cases, it also refers to the average of all categories). AR is usually used to measure the average recall rate of the model under different conditions (such as different confidence thresholds, different maximum detection numbers). FPS represents the number of image frames that the model can process per second, which is a commonly used indicator to measure the inference speed. The number of parameters refers to the trainable parameters (TrainableParameters) of the model, and the computational complexity represents the number of floating-point operations required for one forward inference of the model.

[0122] The effective receptive field coverage area ratio (T = 20%) ≥ 0.35%, which is better than ResNet-18 (0.23%) and CSPDarknet53 (0.25%).

[0123] The parameters for training the SCG-DETR model include: the number of iterations is 100, the batch size is 8, the optimizer is AdamW (weight decay 0.05, momentum 0.9), the learning rate is 0.0001, the input image size is 640×640, until the cross-entropy loss function converges to the desired error threshold (ε = 1e-3).

[0124] Step 4: Obtain the railway turnout image to be detected, input it into the trained SCG-DETR model, output the defect location and category results, and generate a heat map through Grad-CAM++ to visualize the high-confidence defect areas.

[0125] In this embodiment, the anti-interference ability of the model SCG-DETR is achieved in the following ways:

[0126] The context information fusion of the CGFM module suppresses background noise;

[0127] The CAA attention mechanism enhances the feature weights of the defect areas;

[0128] The Inner-GIoU loss function optimizes the regression accuracy of small object bounding boxes;

[0129] Under strong light, rain, snow, and night conditions, the fluctuation range of mAP50 is ≤ ±2.5%.

[0130] 7. The detection method according to claim 1, wherein

[0131] Embodiment 2

[0132] As Figure 1 shown, the purpose of the present invention is to address the problems of low accuracy of existing railway frog defect algorithms and high requirements for hardware devices. Therefore, a railway frog defect detection algorithm based on the SCG-DETR detection model is proposed, and the improved algorithm is used for railway frog defect detection.

[0133] The present invention includes the following steps:

[0134] Step 1:

[0135] 1.1 Prepare the dataset and perform preprocessing

[0136] In this paper, a railway turnout defect dataset is constructed. The dataset includes turnout defect images under different lighting conditions (strong light, low light, backlight) and weather conditions (sunny, rainy, snowy, night), and the images are normalized to 640×640 pixels in size, regularized (mapped to a normal distribution), and data augmentation operations are performed. The data augmentation includes random scaling (scaling ratio 0.5 - 1.5), horizontal / vertical flipping, brightness adjustment (±20%), grayscale transformation, Gaussian blur (kernel size 3×3), and noise addition (noise density 0.05). After expansion, the dataset contains 2500 images, which are divided into a training set (1750 images), a validation set (500 images), and a test set (250 images) according to the ratio of 7:2:1;

[0137] Step 2: Construction of the SCG-DETR algorithm

[0138] In the improved RT-DETR model, namely the SCG-DETR model, the model structure is as follows. The input image first passes through SMPNet, and the SMP-CGLU module that introduces continuous convolution and a gating mechanism is used to reduce the calculation of redundant features and capture local features. Then, the AIFI module of the efficient hybrid encoder performs an attention operation on the S5 feature map output by the backbone network. After that, the vector is recombined into a two-dimensional feature, denoted as F5, as shown in Equation (2): Figure 1 As shown, the input image first passes through SMPNet, and the SMP-CGLU module that introduces continuous convolution and a gating mechanism is used to reduce the calculation of redundant features and capture local features. Then, the AIFI module of the efficient hybrid encoder performs an attention operation on the S5 feature map output by the backbone network. After completion, the vector is recombined into a two-dimensional feature, denoted as F5, as shown in Equation (2):

[0139] Q = K = V = Flatten(S5)

[0140]

[0141] F5 = Reshape(Attention(Q, K, V))

[0142] Among them, Q (query), K (key), and V (value) respectively represent the three basic tensors in the attention mechanism. The query is used to match with the key to calculate the attention distribution. The key cooperates with the query to measure the similarity between the two. The value is the feature information that will ultimately be weighted and summed. Flatten(·) means flattening the high-dimensional feature map (usually expanding the two-dimensional or three-dimensional space into a form of batch × feature dimension). d is the feature dimension or scaling factor. Reshape(·) means reshaping the tensor into the required shape after the calculation (such as restoring from batch × sequence dimension × number of channels to the original width and height spatial structure).

[0143] Then, through the F5 feature map obtained by the AIFI module, the S4 and S3 feature maps of different sizes output by the backbone network are sent to Fusion-IM in CCFF of the efficient hybrid encoder for feature fusion operation. Introducing the CAA attention mechanism can more accurately identify and locate the regions of interest. This process provides richer feature information for small object detection. The features refined by the CCFF feature pyramid network flow to the detection head. In the detection head part, the Inner-GIoU bounding box loss function can improve the convergence speed and the prediction accuracy.

[0144] Step 3: Construct the SMPNet network. The backbone network, as the core component in the algorithm model, its design and selection directly affect the final performance of the model. The backbone part of the original RT-DETR model adopts a design based on the CNN network structure, similar to ResNet, HGNetV2, and CSPDarknet53. Among them, the C2f module performs well in many object detection tasks. In the present invention, by replacing the original Bottleneck in the C2f module with SMPC, the C2f-SMP module has been significantly improved in terms of feature extraction ability and computational efficiency. The improved backbone part is as Figure 2 shown. The improved backbone network mainly consists of four stages. Before the start of each stage, there is a 3×3 convolution with a stride of 2 for spatial downsampling and channel adjustment. The SMPC module in C2f-SMP contains SMPConv and ConvGLU modules. SMPConv, by using the adaptive point movement mechanism, can better capture long-term dependencies for tasks such as long sequence data and image data, improving the flexibility and accuracy of feature extraction.

[0145] Traditional continuous convolutions usually generate convolution kernels through neural networks (such as MLP), which is computationally expensive and complex in parameter tuning. SMPConv directly generates convolution kernels through self-moving point representation (SMP), avoiding the dependence on complex neural networks. SMPConv introduces a continuous function SMP to express the convolution kernel, as shown in the following mathematical formula:

[0146]

[0147] where, is the set of trainable parameters, where represents the positions of N p learnable support points; corresponds to the weights of each support point; corresponds to the radius parameter of each support point; |N(x)| represents the number of points in the neighborhood of point x; g(x,p i ,r i ) is used to describe the relationship between point x and support point p i , where r i is the influence radius parameter; w i is the weight parameter associated with support point p i .

[0148]

[0149] g(x,p i ,r i ) describes the kernel function relationship between point x and support point p i , ||x - p i|| represents the Euclidean distance between the point x and the support point p i and r i is the radius parameter associated with the support point p i This method brings about efficient computation and parameter reduction while maintaining the high representational ability of the convolutional kernel. Each self-moving point is updated independently, enabling better representation of high-frequency components and being more suitable for processing images with rich details compared to traditional convolutional methods.

[0150] Then, feature extraction is further carried out through the ConvGLU module. The ConvGLU module combines convolution and gating mechanisms, which can selectively pass information channels, improving the effectiveness and flexibility of feature extraction. This Transformer-like encoder structure uses DWConv for feature extraction, and can further perform non-linear transformation and strengthen feature extraction. By means of the dynamic gating mechanism, it ensures that the gating signal of each Token originates from itself, avoiding the signal sharing problem that may be introduced by global average pooling. Finally, at the end of SMPC, a dropout layer is introduced to prevent the neural network from overfitting, and shortcut connections are used to reuse input features.

[0151] Step 4: Feed the feature map obtained by processing through the network in Step 3 into the efficient hybrid encoder of SCG-DETR. The efficient hybrid encoder consists of two parts, namely the AIFI module and the CCFF module. The AIFI module, full name Inner Scale Feature Interaction Module, uses the multi-head self-attention mechanism in Transformer to explore the deep correlations of the S5 feature map. With the help of Transformer, the deep correlations of the feature map can achieve the aggregation of global information and effectively avoid interference caused by complex backgrounds. The CCFF module, full name Cross-Scale Feature Fusion Mechanism, adopts the method of fusing and adding feature maps of different sizes, which can obtain feature map information of different dimensions, enabling the model to obtain more deep correlations and multi-scale information. And in this invention, by further transforming the CCFF, a feature fusion module based on the attention mechanism - CGFM (Context Guide Fusion Module), full name Context Guide Fusion Module, is proposed. After experiments and analysis on the railway frog dataset, it is found that the defect distribution has uneven feature distribution in the spatial dimension, and the use of spatial dimension-based attention has relatively poor effects. Therefore, in the Context Guide Fusion Module (CGFM), the idea of context attention mechanism is introduced, so that the fused features not only contain the information of the original features, but also can relate to the context information and capture global features, thereby improving the performance of the model. The specific structures of the Fusion-IM module and the CGFM module are as Figure 3 and Figure 4As shown, it includes global attention CAA, two pointwise convolutions, a batch normalization layer (BN), and a Sigmoid layer.

[0152] The design purpose of the CGFM module is to optimize the feature fusion and feature extraction capabilities of the network. Introducing the corresponding attention mechanism in this module can effectively improve the feature extraction ability of the module. After the feature map passed through the CAA attention is subjected to average pooling and pointwise convolution, it is then subjected to an element-wise product operation with the feature map tensor of the original input to obtain a weighted feature map tensor. In order to obtain the context information between the feature maps of each layer, the weighted feature map and the original feature map are then element-wise cross-added to strengthen the circulation between the feature maps, and then a cat operation is performed to output the final feature map. The mathematical description of the CGFM module is as follows:

[0153] F s = CAA(Cat(F i ,F j ))

[0154] F att = σ(BN(PWC(BN(PWC(AVG(F s )))))

[0155] F i,weight = F i × F att

[0156] F j,weight = F j × F att

[0157] F out = Cat((F j,weight + F j ),(F i,weight + F i )

[0158] By concatenating (Cat) the input feature maps F i and F j , and then applying a Channel Attention Adaptation (CAA) module, the fused feature map F s is obtained. First, global average pooling (AVG) is performed on the fused feature map F s , then pointwise convolution (PWC, 1×1 convolution kernel) is applied, followed by batch normalization (BN), pointwise convolution is applied again, batch normalization is performed, and finally the attention weight F att is obtained through the Sigmoid activation function (σ). Secondly, the attention weight F att is respectively multiplied with the input feature map Fi and F j are multiplied to obtain the input feature map F i of attention weights F i,weight and the input feature map F j of attention weights F j,weight Finally, the diagonal addition and concatenation (Cat) operations are performed respectively to obtain the final output feature map F out .

[0159] Step 5: Send the feature maps that have passed through the feature extraction backbone (BackBone) and the efficient hybrid encoder in Step 4 into the IoU-aware queryer. According to the above process, the corresponding defect annotation boxes are obtained to realize the detection process of railway frog defects. Under the same dataset and the same experimental conditions, this invention conducted comparative experiments using several popular object detection models. According to the experimental results in Table VI, the two-stage object algorithm Faster-RCNN generates a large number of redundant boxes during the process of generating candidate boxes because it needs to generate candidate boxes, perform object classification, and perform position regression separately during the object detection process, increasing the computational burden and failing to reach a satisfactory level in terms of accuracy and detection speed. The SSD algorithm reaches 38.2 in terms of detection speed. In an industrial scenario, an FPS of 30 can meet its actual deployment requirements, but SSD performs poorly in terms of accuracy and recall rate and cannot meet the requirements for railway detection accuracy.

[0160] In the industrial field, the relatively popular YOLO and its variants perform well in terms of detection speed and accuracy, and the former is mostly adopted in the actual application field. The performance of YOLOv5s and YOLOv8s in terms of the key indicator mAP is inferior to that of YOLOv9, and the latter has improved by 3.6% and 2.8% respectively compared with the former. However, in terms of the number of parameters and inference speed, they perform poorly and are inferior to YOLOv5s and YOLOv8s. And YOLOv10s performs even better compared with YOLOv9, with mAP50 reaching 70.0%, and AP and AR reaching 66.8% and 77.5% respectively, and there is also a great improvement in real-time performance compared with the latter. The improved model of this invention has mAP50 of 71.0%, AP of 78.8%, AR of 68.4%, the parameter size is 15.5M, and FPS is 80.5. The experimental results are shown in Table 1, and the experimental result comparison chart is as Figure 5 shown.

[0161] Table 1 Comparison table of detection accuracy and parameters of different algorithms

[0162]

[0163] According to the experiments, it can be concluded that the defect detection performance of the model proposed by the present invention is generally superior to other algorithm models, mainly manifested in the maximum, minimum, and average accuracy rates of the algorithm, and the model has a smaller number of parameters. Therefore, the stability of the model is also superior to most other algorithm models. In summary, the model proposed by the present invention has excellent defect detection performance.

[0164] As described above, only some specific embodiments of the present invention are provided, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered within the protection scope of the present invention.

[0165] Example 3

[0166] CSPDarknet53 Network, Attention Mechanism, and IoU-Aware Query Generator

[0167] 1. CSPDarknet53 Network

[0168] CSPDarknet53 is a deep convolutional neural network based on Darknet53, especially used for object detection tasks, and is used as the backbone network in models such as YOLOv4 and YOLOv5. Its design inspiration comes from CSPNet (Cross-Stage Partial Networks), aiming to improve the computational efficiency and accuracy of the network. Compared with traditional convolutional neural networks, CSPDarknet53 significantly improves the computational efficiency while ensuring high accuracy.

[0169] The core idea of CSPNet (Cross-Stage Partial Networks) is to divide the feature map into two parts, process them separately, and finally merge them. This strategy enables the model to reduce the computational amount of each layer while maintaining a deeper network structure. CSPNet enhances the diversity of features and the full utilization of information through cross-stage information flow. By combining partial shared computing and local computing, the redundant computing of the network is reduced, and the efficiency of the network is improved. Especially in deeper networks, it can significantly improve the training speed and inference performance.

[0170] Darknet53 is the backbone network of YOLOv4, which adopts multiple convolutional layers and residual connections, making it an efficient network architecture, especially suitable for object detection tasks. To further enhance the representation ability and efficiency of the network, CSPDarknet53 is optimized based on Darknet53 and introduces CSPNet technology. CSPNet improves the computational efficiency and accuracy by distributing the computing to different stages.

[0171] In addition, CSPDarknet53 continues the residual connection design of Darknet53, which helps solve the vanishing gradient problem in deep networks and ensures the effective propagation of information in the network. By using residual blocks, CSPDarknet53 can build deeper networks while avoiding overfitting and enhancing the model's expressiveness.

[0172] 2. Attention Mechanism

[0173] The attention mechanism is a commonly used technique in the process of optimizing convolutional neural networks. It often marks the more important data features in the image information by means of parameter values. Through learning and training, it changes the parameter path of the neural network and transfers the attention of the deep network to the most important areas, thus being called attention. In the field of convolutional neural networks, the attention mechanism can usually be divided into two main types: (i) Channel Attention Mechanism (SE); (ii) Spatial Attention Mechanism (SAM). The hybrid-domain attention mechanism is mainly generated by the combination of the SE mechanism and the SAM mechanism. Currently, the most typical hybrid-domain attention mechanisms are the DA attention mechanism and the CBAM attention mechanism.

[0174] Add the convolutional neural network attention mechanism module to the CGFM module to optimize the Fusion module of the network and form feature fusion optimized based on the attention mechanism. The specific structure is shown in Figure 4 The SE (Squeeze-and-Excitation) channel attention mechanism is a module used to enhance the representational ability of convolutional neural networks (CNNs), aiming to more effectively utilize the information of feature maps by adaptively adjusting the weights of channels. The core idea of the SE module is to perform adaptive recalibration of features in the channel dimension to improve the response of important features and suppress irrelevant features. The CAA (Context Anchor Attention) attention mechanism is a method to strengthen the global attention of context connection. It obtains local region features by using global average pooling and 1×1 convolution. The average pooling operation helps capture the global context information of the input feature map, and then the 1×1 convolution is used to compress the feature dimension and reduce the computational complexity. Subsequently, the CAA module adopts two deep strip convolutions, along the horizontal and vertical directions respectively. This design can be regarded as an approximation of large-kernel depth convolution, effectively expanding the receptive field of the convolution while minimizing the number of parameters and complexity as much as possible, thereby capturing the dependencies between remote pixels, and obtaining the attention weights of features with different dimensions through convolution and activation functions to ensure the improvement of model performance.

[0175] 3. IoU-aware Query Generator

[0176] In the RT-DETR model, IoU-Aware Query Selection is a strategy for initializing the decoder's object queries, aiming to improve the accuracy of object detection. In traditional object detection models, the decoder's object queries are usually randomly initialized, lacking clear physical meaning. This may make it difficult for the decoder to effectively learn features related to actual objects. To solve this problem, RT-DETR introduces the IoU-Aware Query Selection strategy. This strategy provides IoU (Intersection over Union) constraints during training to guide the model to generate higher-quality initial object queries. Specifically, during the training process, the model encourages the decoder to generate prediction boxes with a high overlap degree with the real objects, thereby improving the detection accuracy.

[0177] The following are the specific model training steps.

[0178] (2.1) Selection of the CGFM attention mechanism

[0179] When not considering the improvement of different attention mechanisms on the CGFM module, the present invention finally selects the CAA attention mechanism as the tool for feature extraction in the CGFM module through experiments. CBAM is an attention module based on channel and spatial dimensions, which is widely used in the field of object detection due to its significant effect and small computational cost. However, in this experiment, CBAM shows mediocre performance in improving accuracy. SE and CA explicitly model the interdependence between channels and enhance spatial feature representation while keeping the number of parameters and GFLOPs at a relatively low level, but their accuracy improvement in this experiment is not enough to meet the high-precision requirements of railway frog defect detection. The EMA attention mechanism significantly improves the performance of the model by preserving the association between spatial and channel information and providing more stable feature representation. Compared with traditional attention mechanisms, the accuracy improvement is relatively significant. The CAA (Context Anchor Attention) attention mechanism significantly enhances the model's understanding of global context and long-range dependencies in images by automatically focusing on important features and suppressing noise. This mechanism improves the correlation between features by weighting the information of different channels and optimizes the information flow. In terms of local-level fusion, CAA strengthens the interaction between channels. Through average pooling operations, pixel-level fusion can further process and integrate the feature maps obtained from the global and local fusion stages. While retaining the necessary information, this process effectively reduces feature redundancy and noise.

[0180] Table 1 Influence of attention mechanisms on detection accuracy

[0181] Influence

[0182] The loss function plays a crucial role in the training process of machine learning and deep learning models. It is used to measure the difference between the predicted values and the true values of the model, guiding the optimization of model parameters. The choice of the loss function directly affects the training effect, convergence speed, and final performance of the model. In this paper, Inner-GIoU is used to conduct comparative experiments with the currently mainstream loss functions on the railway frog defect dataset, with the remaining settings being the same. The experimental data is shown in Table 4. The experiments show that in the railway frog defect scenario, when using Inner-GIoU as the bounding box loss, there are significant improvements in both the precision and recall of the improved model, and the mAP metric also increases by 2.3%. In contrast, the corresponding CIoU, SIoU, DIoU, and the corresponding losses improved with Inner perform mediocrely. Under the premise of the same parameters and computational complexity, they perform worse than Inner-GIoU, demonstrating the effectiveness of introducing Inner-GIoU. The experimental results are shown in Table 2.

[0183] Table 2 Influence of Loss Functions on the Algorithm

[0184]

Claims

1. A railway turnout defect detection method based on the SCG-DETR detection model, characterized in that, It includes the following steps: Step 1: Construct a railway turnout defect dataset. The dataset includes turnout defect images under different lighting conditions and weather conditions, and the images are normalized to 640×640 pixels in size, regularized, and data augmentation operations are performed; Step 2: Construct an SCG-DETR detection model, which sequentially includes a feature extraction backbone BackBone, an efficient hybrid encoder, an IoU-aware query selection, and a decoder; The feature extraction backbone network consists of four stages. At the beginning of each stage, there is a 3×3 convolution for spatial downsampling and channel adjustment. The SMPConv in the SMPC module of the C2f-SMP module adopts an adaptive point movement mechanism; during the feature extraction process, each self-moving point of the SMPC module is updated independently. Subsequently, features are further extracted through the Conv GLU module; C2f-SMP module: The feature extraction backbone adopts the SMPNet network. By replacing the BottleNeck in the original C2f module with the SMPC module, the SMPC module includes the SMPConv and ConvGLU modules, generates convolutional kernels through the adaptive point movement mechanism, and optimizes feature extraction by combining the gating mechanism; The efficient hybrid encoder consists of two parts: an attention-based inner-scale feature interaction module AIFI and a CNN-based cross-scale feature fusion module CCFF. AIFI only performs inter-scale interaction on the S5 feature map, while the CCFF module constructs a Fusion block through convolutional layers. In the Fusion block, the ContextGuideFusionModule, i.e., the CGFM module, is introduced; CGFM module: Based on the CAA attention mechanism, it adaptively weights and fuses multi-scale feature maps, and enhances context information interaction through global average pooling and pointwise convolution; For the IoU-aware query, the feature map encoded by the efficient hybrid encoder is input into the IoU-aware query. The IoU-aware query improves the initialization of object queries and expands them into content queries and location queries; the IoU-aware query constrains the detection map, so that during the training process, features with high IoU scores can obtain higher classification scores, while features with low IoU scores are given lower classification scores; Inner-GIoU loss function: Generate auxiliary bounding boxes by introducing a scale factor, dynamically adjust the bounding box regression loss, and accelerate model convergence; Step 3: Input the preprocessed dataset in Step 1 into the SCG-DETR model built in Step 2, and pass through the feature extraction backbone BackBone, the efficient hybrid encoder, the IoU-aware query selection, and the decoder respectively. Through 100 epochs of iteration, a model with the highest accuracy is trained; Step 4: Obtain the railway turnout image to be detected, input it into the trained SCG-DETR model, output the defect location and category results, and generate a heat map through Grad-CAM++ to visualize the high-confidence defect areas.

2. The detection method according to claim 1, wherein Data augmentation includes random scaling with a scaling ratio of 0.5 to 1.5; horizontal / vertical flipping, brightness adjustment of ±20%, grayscale transformation, Gaussian blur with a kernel size of 3×3, and noise addition with a noise density of 0.

05. The augmented dataset contains 2500 images, which are divided into a training set, a validation set, and a test set in a ratio of 7:2:

1.

3. The detection method according to claim 1, wherein In step two, the construction of the C2f-SMP module: The mathematical expression of SMPConv is: Among them, is a set of trainable parameters, where represents N p learnable support point positions; the weight corresponding to each support point; the radius parameter corresponding to each support point; |N(x)| represents the number of points in the neighborhood of point x; g(x, p i , r i ) is used to describe the relationship between point x and support point p i , where r i is the influence radius parameter; w i Support point p i Associated weight parameter g(x, p i , r i ) describes the kernel function relationship between point x and support point p i . ||x - p i || represents the Euclidean distance between point x and support point p i . r i is the radius parameter related to support point p i . The function calculates a distance-based linear attenuation function. This distance-based linear attenuation function is used in SMPConv to weigh the influence degree of different support points on the target point. The closer the support point is, the greater its contribution. Points beyond a specific radius r i may no longer have an impact. The ConvGLU module selectively transmits feature information through depthwise separable convolution and a gating mechanism, and its expression is: where σ is the Sigmoid activation function, is the element-wise multiplication, where Conv represents the ordinary convolution operation; the four stages of the C2f-SMP module contain 3, 4, 6, and 3 SMPConv modules respectively, and spatial downsampling and channel adjustment are performed through 3×3 convolution before the start of each stage.

4. The detection method according to claim 3, wherein The ConvGLU module has a 3×3 convolution with a stride of 2, and the number of channels is 64, 128, 256, and 512 respectively.

5. The detection method according to claim 1, wherein The implementation of the CGFM module includes: F s = CAA(Cat(F i ,F j )) F att = σ(BN(PWC(BN(PWC(AVG(F s )))) By splicing Cat feature F i and F j , then applying a channel attention adaptation ChannelAttentionAdaptation, CAA module, to obtain the fused feature F s . First, perform global average pooling - AVG on F s , then apply pointwise convolution, followed by batch normalization - BN, apply pointwise convolution again, perform batch normalization, and finally obtain the attention weight F through the Sigmoid activation function σ att ; Fusing features through element-wise multiplication and cross addition, and the output expression is: F i,weight = F i × F att F j,weight = F j × F att F out = Cat((F i × F att + F j ),(F j × F att + F i )) Secondly, when the attention weight F att is multiplied by F i , F j respectively to obtain F i,weight , F j,weight , and finally, the diagonal elements are added respectively and the concatenation - Cat operation is performed to obtain F out . The CAA attention mechanism captures long-range dependencies through horizontal and vertical depth strip convolutions with kernel sizes of 7×1 and 1×7, and uses global average pooling to compress the feature dimension to 1 / 16 of the original number of channels.

6. The railway turnout defect detection method according to claim 1, characterized in that, The mathematical derivation expression of the Inner-GIoU loss function is as follows: The mathematical expression of IoU is: Among them, A represents the area size of the target box, B represents the area size of the detection box, |A∩B| represents the absolute value of the area of the overlapping region between the target box and the detection box, and |A∪B| represents the absolute value of the total area of the target box and the detection box. Furthermore, GIoU solves the problem of zero gradient when two targets have no intersection by introducing the minimum bounding rectangle of the predicted box and the ground truth box to obtain the proportion of the predicted box and the ground truth box in the closure region, and its mathematical expression is: Among them, C is the area of the minimum bounding rectangle of the two boxes. Next, by introducing an auxiliary bounding box, using a scaling factor to generate auxiliary bounding boxes of different sizes to calculate the loss, and its mathematical expression is: Inner-GIoU = GIoU + IoU - IoU inner Among them, IoU inner Dynamically calculate the intersection over union (IoU) through the scale factor ratio of the auxiliary bounding box. The specific calculation process is as follows: union=(w g *h g )*(ratio) 2 +(w*h)*(ratio) 2 -inter where x g , y g represent the center coordinates of the target box, w g , h g represent the width and height of the target box, x, y represent the center coordinates of the detection box, w, h represent the width and height of the detection box, ratio represents the auxiliary bounding box scaling factor, respectively represent the Inner boundaries of the target box including left, right, top, bottom, b l , b r , b t , b b respectively represent the Inner boundaries of the detection box including left, right, top, bottom, inter represents the intersection area between the Inner boxes, union represents the union area between the Inner boxes, IoU inner represents the final Inner - IoU.

7. The detection method according to claim 6, wherein The scaling factor ratio ranges from 0.5 to 1.

5.

8. The detection method according to claim 1, characterized in that, The evaluation metrics of the model SCG-DETR satisfy: mAP50≥71.0%, AP average precision≥78.8%, AR average recall≥68.4%; The detection speed≥80.5FPS, the number of parameters≤15.5M, and the computational complexity≤68.0GFLOPs; Among them, the mAP50 metric represents the average value of the predicted boxes with an intersection over union greater than 50%. AP is usually an indicator to measure the detection accuracy for a certain category. AR is usually used to measure the average recall rate of the model under different conditions. FPS represents the number of images that the model can process per second, which is a common indicator to measure the inference speed. The number of parameters refers to the trainable parameters of the model, and the computational complexity represents the number of floating-point operations required for one forward inference of the model; The effective receptive field coverage area ratio (T = 20%)≥0.35%, which is better than ResNet-18 (0.23%) and CSPDarknet53 (0.25%).

9. The detection method according to claim 1, characterized in that In step 3, the parameters for training the SCG-DETR model include: the number of iterations is 100, the batch size is 8, the optimizer is AdamW, the learning rate is 0.0001, the input image size is 640×640, until the cross-entropy loss function converges to the desired error threshold (ε = 1e-3).

10. The detection method according to claim 9, characterized in that, The AdamW weight decay is 0.05, the momentum is 0.9, and the desired error threshold is ε = 1e-3.

Citation Information

Cited By

  • Industrial surface defect detection method based on improved real-time target detection model

    CN120563514A

  • Industrial surface defect detection method based on improved real-time target detection model

    CN120563514B

  • Intelligent quantitative characterization method for wheeltrack fatigue crack based on microscopic analysis image

    CN120953294A

  • Bottle body packaging defect detection system based on machine vision

    CN121073965A

  • Method and device for quantitative representation of blood smear region under microscope

    CN121921550A